A benchmark set turns mostly green

Rust compiler performance improved across most tracked workloads between late July and late September 2026, according to an analysis of the project’s benchmark data by compiler contributor Nicholas Nethercote. The mean reduction in wall-clock time was 4.57% over the two-month period. Of 629 benchmark measurements, 555 improved and 74 regressed.

The result came from many changes rather than one dominant optimization. Upgrading the compiler to LLVM 23 reduced mean wall time by 1.2% across the complete benchmark collection. Enabling profile-guided optimization for Clippy, Rust’s linting tool, improved most Clippy workloads and cut wall time by as much as 18% in the strongest case.

Other patches targeted repeated allocations, specialization-graph construction and incremental compilation data. One change to implementation handling produced a mean cycle-count reduction of 1.58% across all benchmarks, while an optimization in loading incremental data reduced instruction counts by up to 6% on affected tests. Replacing dynamic dispatch in a hot trait-solver selection path also delivered smaller gains across multiple workloads.

New systems bring new costs

The gains arrived while Rust’s Nightly compiler enabled two substantial components that remain slower for some code. Polonius Alpha, a more precise borrow checker, can accept valid programs rejected by the older checker but adds measurable work in a minority of cases. Making some liveness calculations lazy reduced instruction counts for the widely used serde crate by roughly 3% to 5%, with smaller effects elsewhere.

A new trait solver likewise has performance outliers. Work aimed at those cases produced large reductions for particular crates, including reported improvements of 50%, 25% and 15% in separate examples. These focused wins do not represent average compiler performance, but they show how pathological workloads can respond to changes in algorithms and data structures.

One striking case involved a function in the cranelift-codegen crate containing more than 18,000 basic blocks. Revising the control-flow traversal used by compiler dataflow analysis cut calls to one effect-processing routine from about 1.5 million to 90,000. That change reduced wall time for a check build of the crate by about 30%. Avoiding unnecessary projection data in the same analysis reduced instruction counts by 17% on a stress benchmark.

The volume of successful work also affected the project’s integration process. Rust normally merges performance-sensitive patches separately so their effects can be measured cleanly. During this period, the queue became large enough for one rollup to combine ten performance-improving changes, followed later by another containing four. Individual effects could still be checked after merging.

The two-month picture is therefore broad rather than uniform: a few new compiler systems still impose costs on certain projects, but improvements elsewhere more than offset them across the benchmark suite. Continued profiling of those outliers will determine how much of the remaining gap can be removed without sacrificing the additional analysis those systems provide.