Skip to content

rustc_monomorphize: shard volatile generic CGU buckets to reduce incremental invalidation scope - #162692

Draft
Trigodil wants to merge 6 commits into
rust-lang:mainfrom
Trigodil:fine-grained-generic-cgus
Draft

Trigodil wants to merge 6 commits into
rust-lang:mainfrom
Trigodil:fine-grained-generic-cgus

Conversation

@Trigodil

@Trigodil Trigodil commented Sep 12, 2026 •

Copy link
Copy Markdown

View all comments

Summary

The current CGU partitioner assigns every monomorphized instance of a generic
function from the same module into a single "volatile" bucket, keyed on
(module_def_id, volatile). This means editing the body of any instantiation
invalidates the entire bucket -- every other instantiation in that module gets
re-emitted on the next incremental build, even ones that share no code with the
change.

This PR adds -Z fine-grained-generic-cgus, which shards each volatile bucket
into a bounded number of sub-buckets by hashing the instantiation's fully-mangled
symbol name. An edit to one instantiation's code path now only invalidates the
~1/shard_count of instantiations that hash into the same shard.

Approach

Shard count is derived from -C codegen-units, floored at the host's available
parallelism and clamped to [4, 128]. This keeps the total CGU count in the same
ballpark as a normal build while ensuring the codegen backend always has enough
independent units to keep every core busy.

Unbounded splitting (one CGU per instantiation) was tested first and regresses
heavily on real crates: merge_codegen_units is intentionally skipped in
incremental mode, so hundreds of tiny object files pile up with no recombination,
and link cost dominates. Bounded sharding avoids this entirely.

Benchmarks (polars-core, ChunkedArray)

Configuration Cold Incremental
baseline 1.000x 1.000x
unbounded (1 CGU/instantiation) ~3.06x ~20.95x (regresses)
this patch (hash-sharded) 0.988x 0.961x
hash-sharded + -Z threads=4 0.657x 0.852x
hash-sharded + threads=4 + cranelift (cu=16) nil 0.521x (best)

Tested on polars-core which has high monomorphization volume
(ChunkedArray<T> instantiations survive merge_codegen_units and make the
volatile bucket expensive to invalidate). Methodology uses a body-edit approach

  • a trailing comment does not trigger CGU invalidation since rustc fingerprints
    by item content, not file bytes.

Notes

  • Little to no effect on non-incremental builds
  • Topology-aware clustering was also prototyped and benchmarked - It lost to hash-sharding due to lack of callees to cluster by

Disclosure

Claude Sonnet 4.6 was used to speed up compiler research and to rule out dead ends in the methodology used.

@rustbot rustbot added S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue. labels Sep 12, 2026
@rust-log-analyzer

This comment has been minimized.

@Trigodil
Trigodil force-pushed the fine-grained-generic-cgus branch from b2d8f24 to cfca780 Compare September 12, 2026 16:15
@rust-log-analyzer

This comment has been minimized.

@panstromek

Copy link
Copy Markdown
Contributor

Let's try this out.

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 12, 2026
@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Sep 12, 2026
rustc_monomorphize: shard volatile generic CGU buckets to reduce incremental invalidation scope
Comment thread compiler/rustc_monomorphize/src/partitioning.rs
@panstromek

Copy link
Copy Markdown
Contributor

@bors try cancel

@rust-bors

rust-bors Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Try build cancelled. Cancelled workflows:

Hint: if you want to run another try build, you do not need to manually cancel the previous one. Just run @bors try and bors will cancel the previous build automatically.

Got it

Co-authored-by: MatyΓ‘Ε‘ Racek <panstromek@seznam.cz>
@panstromek

Copy link
Copy Markdown
Contributor

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Sep 12, 2026
rustc_monomorphize: shard volatile generic CGU buckets to reduce incremental invalidation scope
@rust-log-analyzer

This comment has been minimized.

@panstromek

Copy link
Copy Markdown
Contributor

You can ignore tidy for experiments like this, it won't fail the try build.

@Trigodil

Copy link
Copy Markdown
Author

You can ignore tidy for experiments like this, it won't fail the try build.

Gotcha, its just really annoying for me when one fails

@rust-log-analyzer

This comment has been minimized.

@Trigodil

Copy link
Copy Markdown
Author

It seems that forcing this on unconditionally breaks the tests, since they assert exact CGU names (e.g. local_generic.volatile) and now get a .shard0XX suffix appended everywhere

@rust-bors

rust-bors Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

β˜€οΈ Try build successful (CI)
Build commit: 6c4e7b9 (6c4e7b9b2b3bc7bc0abc4f3ee5402b3ecedd355e)
Base parent: 9c99d05 (9c99d05505bccb67912d68e05fe7fc7c58afcb41)

@rust-timer

This comment has been minimized.

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (6c4e7b9): comparison URL.

Overall result: βŒβœ… regressions and improvements - please read:

Benchmarking means the PR may be perf-sensitive. It's automatically marked not fit for rolling up. Overriding is possible but disadvised: it risks changing compiler perf.

Next, please: If you can, justify the regressions found in this try perf run in writing along with @rustbot label: +perf-regression-triaged. If not, fix the regressions and do another perf run. Neutral or positive results will clear the label automatically.

@bors rollup=never rustc-perf
@rustbot label: -S-waiting-on-perf +perf-regression

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
70.7% [0.1%, 1726.1%] 100
Regressions ❌
(secondary)
52.5% [0.1%, 1727.0%] 69
Improvements βœ…
(primary)
-0.3% [-0.3%, -0.2%] 4
Improvements βœ…
(secondary)
-6.9% [-66.1%, -0.1%] 11
All βŒβœ… (primary) 68.0% [-0.3%, 1726.1%] 104

Max RSS (memory usage)

Results (primary 2.4%, secondary 2.5%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
9.3% [0.5%, 44.3%] 24
Regressions ❌
(secondary)
9.0% [2.3%, 43.1%] 12
Improvements βœ…
(primary)
-7.5% [-24.4%, -0.4%] 17
Improvements βœ…
(secondary)
-8.7% [-17.4%, -4.3%] 7
All βŒβœ… (primary) 2.4% [-24.4%, 44.3%] 41

Cycles

Results (primary 82.8%, secondary 63.6%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
86.9% [0.4%, 1858.4%] 84
Regressions ❌
(secondary)
72.4% [2.2%, 1857.6%] 53
Improvements βœ…
(primary)
-1.4% [-2.7%, -0.4%] 4
Improvements βœ…
(secondary)
-14.3% [-65.8%, -2.0%] 6
All βŒβœ… (primary) 82.8% [-2.7%, 1858.4%] 88

Binary size

Results (primary 32.7%, secondary 42.8%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
32.7% [1.9%, 81.9%] 102
Regressions ❌
(secondary)
42.8% [0.1%, 118.5%] 48
Improvements βœ…
(primary)
- - 0
Improvements βœ…
(secondary)
- - 0
All βŒβœ… (primary) 32.7% [1.9%, 81.9%] 102

Bootstrap: 497.802s -> 494.017s (-0.76%)
Artifact size: 406.93 MiB -> 406.67 MiB (-0.06%)

@rustbot rustbot added the perf-regression Performance regression. label Sep 12, 2026
@rust-log-analyzer

This comment has been minimized.

@panstromek

Copy link
Copy Markdown
Contributor

The results are quite negative, so now the question is how much time you want to sink into figuring this out. There might be something to gain here, because some results are green, but I suspect it's gonna be a pretty uphill battle. This area is a bit of a tarpit, so be warned. I can help with interpreting the results if you have some questions.

If you decide to keep working on this to the point of landing something, keep in mind that this PR will probably interact with the LLM policy: https://forge.rust-lang.org/policies/llm-usage.html, so you should familiarize yourself with it and make sure you follow that when interacting with people here.

@Trigodil

Copy link
Copy Markdown
Author

I have implemented a fix, so hopefully it works, and yes, I am aware of the LLM policy of rust.

@Trigodil
Trigodil marked this pull request as ready for review September 13, 2026 15:23
@rustbot rustbot added the S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. label Sep 13, 2026
@rustbot rustbot removed the S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. label Sep 13, 2026
@rustbot

rustbot commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator

Thanks for the pull request, and welcome! The Rust Project has assigned @folkertdev (or someone else) to review your changes, you should hear from them (or someone else) within the next two weeks.

Please see the contribution instructions and our LLM policy for more information.

Why was this reviewer chosen?

The reviewer was selected based on:

  • Owners of files modified in this PR: compiler
  • compiler expanded to 76 candidates
  • Random selection from 19 candidates

@panstromek

Copy link
Copy Markdown
Contributor

To be clear, since this is not a draft anymore - I don't think we should land this atm. We can re-measure this but I kinda doubt the regression goes away and adding new unstable flag for this doesn't sound sufficiently motivated to me (maybe it also needs MCP, but I'm not sure atm).

@Trigodil

Copy link
Copy Markdown
Author

Yeah I think we can re-measure this one more time, if the regression isn't fixed, I think something else is wrong. Either way if it still regresses I think it is better off closing this pull request

@folkertdev

Copy link
Copy Markdown
Contributor

So, I guess

@bors try @rust-timer queue

@rust-timer

This comment has been minimized.

@rustbot rustbot added the S-waiting-on-perf Status: Waiting on a perf run to be completed. label Sep 13, 2026
@rust-bors

This comment has been minimized.

rust-bors Bot pushed a commit that referenced this pull request Sep 13, 2026
rustc_monomorphize: shard volatile generic CGU buckets to reduce incremental invalidation scope
@rust-bors

rust-bors Bot commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

β˜€οΈ Try build successful (CI)
Build commit: 70399b2 (70399b2e8229dcdd5f2aec0ec7edf8c9692d1e54)
Base parent: 24d4720 (24d472027454741e74f8e913755fbc7e03f02af5)

@rust-timer

This comment has been minimized.

@panstromek

panstromek commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Wait, it's still disabled by default, isn't it?

@rust-timer

Copy link
Copy Markdown
Collaborator

Finished benchmarking commit (70399b2): comparison URL.

Overall result: βœ… improvements - no action needed

Benchmarking means the PR may be perf-sensitive. Consider adding rollup=never if this change is not fit for rolling up.

@rustbot label: -S-waiting-on-perf -perf-regression

Instruction count

Our most reliable metric. Used to determine the overall result above. However, even this metric can be noisy.

mean range count
Regressions ❌
(primary)
- - 0
Regressions ❌
(secondary)
- - 0
Improvements βœ…
(primary)
- - 0
Improvements βœ…
(secondary)
-0.3% [-0.3%, -0.2%] 2
All βŒβœ… (primary) - - 0

Max RSS (memory usage)

Results (primary 1.0%, secondary -0.9%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
1.0% [1.0%, 1.0%] 1
Regressions ❌
(secondary)
3.6% [2.0%, 5.8%] 6
Improvements βœ…
(primary)
- - 0
Improvements βœ…
(secondary)
-4.2% [-5.3%, -3.1%] 8
All βŒβœ… (primary) 1.0% [1.0%, 1.0%] 1

Cycles

Results (primary 2.2%, secondary 4.6%)

A less reliable metric. May be of interest, but not used to determine the overall result above.

mean range count
Regressions ❌
(primary)
2.2% [2.2%, 2.2%] 1
Regressions ❌
(secondary)
8.1% [2.9%, 16.8%] 6
Improvements βœ…
(primary)
- - 0
Improvements βœ…
(secondary)
-5.9% [-7.7%, -4.1%] 2
All βŒβœ… (primary) 2.2% [2.2%, 2.2%] 1

Binary size

This perf run didn't have relevant results for this metric.

Bootstrap: 494.003s -> 495.931s (0.39%)
Artifact size: 406.93 MiB -> 407.03 MiB (0.02%)

@rustbot rustbot removed S-waiting-on-perf Status: Waiting on a perf run to be completed. perf-regression Performance regression. labels Sep 13, 2026
@Trigodil

Trigodil commented Sep 14, 2026 •

Copy link
Copy Markdown
Author

Hmmm, the improvement seems to be lower than my benchmarks and system, thoughts? I think the gains only kick in when you have a crate with massive generic monomorphization volume like polars

@Trigodil
Trigodil marked this pull request as draft September 14, 2026 04:58
@rustbot rustbot added S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. and removed S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. labels Sep 14, 2026
@Trigodil

Trigodil commented Sep 14, 2026 •

Copy link
Copy Markdown
Author

Let me see if I can optimize this even more
@panstromek may I ask what the CI uses as tests for benchmarks? i.e are incremental builds also tested? I would like to tune the code accordingly with the CI

@panstromek

Copy link
Copy Markdown
Contributor

The last benchmark run doesn't have fine_grained_generic_cgus option enabled, so it is more or less no-op. The improvement is probably just noise, those two benchmarks don't even run codegen (it's Doc and Check build).

@panstromek may I ask what the CI uses as tests for benchmarks? i.e are incremental builds also tested?

If you go to comparison url, you can click every benchmark and it will give you some detailed information about it and link to the source code. We run incremental, those are the incr-* scenarios. All of these are also explained if you hover over ? in the filters section.

In your case, you want to pay attention incr-* scenarios in Debug and Opt profile. incr-full is baseline we shouldn't regress, incr-unchanged is a good measure of incremental overhead when nothing changes. incr-patched is your target, but we don't have that many of those to be fair. ATM I think you should focus on making sure the existing incr-* benchmarks don't regress massively, especially the incr-full ones. Without that, we can't land this even if it improves some specific crate/scenario.

I would like to tune the code accordingly with the CI

In general, I recommend reading the profiling chapters in the dev guide, running the suite through x.py is a good start: https://rustc-dev-guide.rust-lang.org/profiling/with-rustc-perf.html, it also links to the rustc-perf docs, which has more detailed info if you need. Once you have this set up, you can also relatively easily add new benchmarks to the suite to measure them locally, there are instructions to do that in rustc-perf repo (it's a subtree in the main repo), it is mostly just adding a folder in compile-benchmarks directory.

@joshtriplett

Copy link
Copy Markdown
Member

If you want to benchmark the updated version, you may want to temporarily re-add the commit that force-enabled this (deleting the conditional), solely to ensure that it's used during the perf run.

I'd be interested in seeing how much impact this has. Whether or not it gets used in this exact form, it'd be interesting to know whether the technique is a win.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants