Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,9 +23,15 @@ Seven real bugs from rustc's history were then replayed by reverting their
fixes. mirth catches five; three of those needed a fixture addition or a
new step, made knowing the bug.

Pointed at the unmodified compiler with a wider fixture, it found two
incremental bugs that look new, where a rebuild after an edit encodes
different metadata from a clean build, and one known parallel-front-end
bug.

- [`docs/report.md`](docs/report.md): the experiment, for readers new to it
- [`docs/results.md`](docs/results.md): each edit and the output that caught it
- [`docs/regressions.md`](docs/regressions.md): the replayed bugs, caught and missed
- [`docs/hunt.md`](docs/hunt.md): bugs found in the unmodified compiler
- [`docs/plan.md`](docs/plan.md): the plan the work followed, with the properties

## An instrumented compiler
Expand All @@ -39,6 +45,7 @@ rustc/build.sh # build stage 1 through mirth-watch, then
rustc/check.sh chain # check fixtures/chain
rustc/edits.sh chain # apply, check and revert each edit in rustc/edits
EDITS=regressions rustc/edits.sh chain # the same for the past bugs in rustc/regressions
rustc/hunt.sh wide # repeated threaded builds, and P6 for each of fixtures/wide/edits
```

The build takes about an hour on 16 cores. `rustc/rmeta.toml` says what is
Expand Down
113 changes: 113 additions & 0 deletions docs/hunt.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,113 @@
# Looking for new bugs

The edits and the replayed regressions test whether mirth catches bugs that
someone put in. This tests whether it finds bugs nobody knew were there: the
unmodified pinned compiler (`ea137335b`), a wider fixture, and more
incremental edits.

`fixtures/wide` has a proc-macro crate (a derive and an attribute macro), a
library with a build script that generates code, built-in derives, const
generics, a generic associated type, `impl Trait` and `async fn` in traits,
and an exported `macro_rules!`, used by a second library and a binary.
`fixtures/wide/edits` has ten changes: a private body, a new enum variant,
a doc comment, spans only, two items swapped, the derive's output, the
build script's output, a const generic argument, a macro's expansion, and
a new private function.

`rustc/hunt.sh <fixture>` builds the fixture eight times with
`-Zthreads=8` and compares the `.rmeta` files (P5). Then, for each edit, it
compares an incremental rebuild after the edit with a clean build of the
edited source (P6), with `-Zincremental-verify-ich`, single-threaded and
with `-Zthreads=8`. `rustc/check.sh wide` runs the ordinary checks.

## Findings

| | what | status |
|---|---|---|
| 1 | incremental rebuilds encode `Generics::param_def_id_to_index` in a different order from clean builds | **looks new**; root cause found; fix and regression test written |
| 2 | incremental rebuilds encode a string literal twice where clean builds encode it once | **looks new**; root cause found, regression from #116707 (1.90); fix and regression test written |
| 3 | with `-Zthreads=8`, two traits with `-> impl Trait` methods give different metadata from run to run | known: [#162202](https://github.com/rust-lang/rust/issues/162202) |

Findings 1 and 2 are single-threaded: an ordinary `cargo build`, an edit, another
`cargo build`, and the metadata differs from a clean build of the edited source. Both come
from the same kind of mistake: a value that does not survive a round trip through the
incremental cache unchanged. All three reproduce with the official `nightly-2026-10-06`,
without mirth: `docs/hunt/repro.sh` runs them. None was searched for: P5 and P6 reported
them on the first run of the new fixture.

Draft bug reports for 1 and 2, written to be filed upstream, are
[`hunt/issue-generics-order.md`](hunt/issue-generics-order.md) and
[`hunt/issue-literal-dedup.md`](hunt/issue-literal-dedup.md). Each has a candidate fix
(`hunt/*.patch`) and a regression test in the style of rustc's `tests/run-make`
(`hunt/tests/`), which fails on the pinned compiler and passes with the fix.

With both fixes applied, P6 holds for all ten edits single-threaded, and these rustc tests
still pass: `tests/incremental`, the metadata-related UI and run-make tests,
`tests/ui/{consts,statics,const-generics}` and `tests/codegen-llvm`.

Everything else held: P1, P2, P4 and P7 on the recorded clean build; the touch-only
rebuild reused every crate's metadata; and `-Zincremental-verify-ich` found no unstable
fingerprint in 20 incremental rebuilds.

### 1. `param_def_id_to_index` order

```rust
pub struct Grid<A, B, C>(A, B, C);
impl<A, B, C> Grid<A, B, C> { pub const AREA: usize = 1; }
```

`Generics::param_def_id_to_index` is an `FxHashMap`. `HashMap`'s `Encodable` writes it in
iteration order, and `Decodable` collects it back in that order. In this example all three
keys hash to the same home bucket of a 4-bucket table, so iteration order is insertion order
rotated by one. Each time `generics_of` goes through the incremental cache, the map is
rotated again, and the metadata encoder writes it in the rotated order. The encoded order
cycles with period 3 over successive incremental rebuilds, which is exactly what a
standalone program with `std` `HashMap` and `rustc_hash` 2.1.1 predicts.

Adding a comment to the top of `lib.rs` changes the metadata of an incremental rebuild
relative to a clean one for `either`, `smallvec`, `memchr` and `arrayvec`, and making the
field an `FxIndexMap` removes the difference in all four. The same field is the cause given
in [#163878](https://github.com/rust-lang/rust/issues/163878) under `-Zthreads`.

### 2. String literals encoded twice

```rust
#[inline] pub fn a() -> &'static str { "literal" }
#[inline] pub fn b() -> &'static str { "literal" } // then edited to `{ let s = "literal"; s }`
```

String literals get their `AllocId` from `allocate_bytes_dedup`, so both bodies share one,
and a clean build encodes the allocation once. In the rebuild, `a`'s MIR comes from the
incremental cache, and decoding a memory allocation always reserves a fresh `AllocId`
(`reserve_and_set_memory_alloc`). The metadata then encodes two allocations with identical
bytes. It started with [#116707](https://github.com/rust-lang/rust/pull/116707), which gave
`ConstValue::Slice` an `AllocId`: `nightly-2025-07-24` is not affected, `nightly-2025-07-26`
is. No effect on generated code was found.

In the fixture, the two literals were the type name in `#[derive(Debug)]`'s `fmt` (which is
`#[inline]`, so its MIR is encoded) and the same name in a `const` from the fixture's own
derive. That is why removing either derive made it disappear. The first write-up of this
finding blamed hygiene data, because the rebuild resolved foreign expansions while encoding;
that was a side effect of decoding cached MIR, not the cause.

### 3. `impl Trait` in traits under `-Zthreads`

```rust
pub trait Render { fn render(&self) -> impl Sized; }
pub trait Draw { fn draw(&self) -> impl Sized; }
```

Eight builds with `-Zthreads=8` give three or four different `.rmeta` files.
A comment on [#162202](https://github.com/rust-lang/rust/issues/162202)
reports return-position `impl Trait` in traits as not reproducible, so this
is known. The fix for finding 1 does not change it. Both fixtures have such
traits, so P5 under `-Zthreads=8` fails on `wide` every time, and on
`chain` occasionally (1 build in 24).

## Not done

- Neither new finding has been reported upstream. The drafts are ready; check #163878 again
before filing the first, since it touches the same field.
- `fixtures/wide` fails `check.sh` (P5 under `-Zthreads=8`, finding 3) and P6 for most
edits (findings 1 and 2) until those are fixed. Its lists are blessed against the
unmodified compiler.
19 changes: 19 additions & 0 deletions docs/hunt/alloc-dedup-on-decode.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
--- a/compiler/rustc_middle/src/mir/interpret/mod.rs
+++ b/compiler/rustc_middle/src/mir/interpret/mod.rs
@@ -214,7 +214,15 @@
trace!("creating memory alloc ID");
let alloc = <ConstAllocation<'tcx> as Decodable<_>>::decode(decoder);
trace!("decoded alloc {:?}", alloc);
- decoder.interner().reserve_and_set_memory_alloc(alloc)
+ // Immutable memory is deduplicated when it is created (string literals, for
+ // instance), so deduplicate it here too. Otherwise one allocation decoded
+ // from the incremental cache and the same one created afresh get two
+ // `AllocId`s, and metadata depends on which results came from the cache.
+ if alloc.inner().mutability.is_not() {
+ decoder.interner().reserve_and_set_memory_dedup(alloc, CTFE_ALLOC_SALT)
+ } else {
+ decoder.interner().reserve_and_set_memory_alloc(alloc)
+ }
}
AllocDiscriminant::Fn => {
trace!("creating fn alloc ID");
20 changes: 20 additions & 0 deletions docs/hunt/generics-index-map.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
--- a/compiler/rustc_middle/src/ty/generics.rs
+++ b/compiler/rustc_middle/src/ty/generics.rs
@@ -1,7 +1,7 @@
use std::ops::ControlFlow;

use rustc_ast as ast;
-use rustc_data_structures::fx::FxHashMap;
+use rustc_data_structures::fx::FxIndexMap;
use rustc_hir::def_id::DefId;
use rustc_macros::{StableHash, TyDecodable, TyEncodable};
use rustc_span::{Span, Symbol, bug, kw};
@@ -124,7 +124,7 @@

/// Reverse map to the `index` field of each `GenericParamDef`.
#[stable_hash(ignore)]
- pub param_def_id_to_index: FxHashMap<DefId, u32>,
+ pub param_def_id_to_index: FxIndexMap<DefId, u32>,

pub has_self: bool,
pub has_late_bound_regions: Option<Span>,
8 changes: 8 additions & 0 deletions docs/hunt/hashmap-roundtrip/Cargo.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
[workspace]

[package]
name = "hashmap-roundtrip"
version = "0.0.0"
edition = "2024"
[dependencies]
rustc-hash = "=2.1.1"
60 changes: 60 additions & 0 deletions docs/hunt/hashmap-roundtrip/src/main.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
// What rustc does with Generics::param_def_id_to_index across sessions.
use rustc_hash::FxHashMap;

// DefId's Hash impl on 64-bit targets: (krate << 32) | index, as one u64.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
struct DefId {
index: u32,
krate: u32,
}
impl std::hash::Hash for DefId {
fn hash<H: std::hash::Hasher>(&self, h: &mut H) {
(((self.krate as u64) << 32) | self.index as u64).hash(h)
}
}

fn order(m: &FxHashMap<DefId, u32>) -> Vec<(u32, u32)> {
m.iter().map(|(k, v)| (k.index, *v)).collect()
}

fn main() {
if std::env::args().nth(1).as_deref() == Some("buckets") {
return buckets();
}
let first: Vec<u32> = std::env::args()
.skip(1)
.map(|a| a.parse().unwrap())
.collect();
// generics_of: own_params.iter().map(|p| (p.def_id, p.index)).collect()
let built: FxHashMap<DefId, u32> = first
.iter()
.enumerate()
.map(|(i, &d)| (DefId { index: d, krate: 0 }, i as u32))
.collect();
println!("clean session, built and encoded: {:?}", order(&built));
// Encodable writes iteration order; Decodable collects in that order.
let mut map = built;
for session in 2..=4 {
let decoded: FxHashMap<DefId, u32> = order(&map)
.into_iter()
.map(|(d, i)| (DefId { index: d, krate: 0 }, i))
.collect();
println!(
"session {session}, decoded from the cache: {:?}",
order(&decoded)
);
map = decoded;
}
}

#[allow(dead_code)]
pub fn buckets() {
use std::hash::BuildHasher;
for i in [12u32, 13, 14] {
let h = rustc_hash::FxBuildHasher.hash_one(DefId { index: i, krate: 0 });
println!(
"DefIndex {i}: hash {h:#018x}, bucket in a 4-bucket table {}",
h & 3
);
}
}
129 changes: 129 additions & 0 deletions docs/hunt/issue-generics-order.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
# Incremental rebuilds encode `generics_of` with `param_def_id_to_index` in a different order from clean builds

<!-- Draft issue for rust-lang/rust. Related: #163878 (same field, parallel front end). -->

An incremental rebuild produces different `.rmeta` bytes from a clean build of the same
source, single-threaded. The cause is that `Generics::param_def_id_to_index`, an
`FxHashMap`, does not survive a round trip through the incremental cache in the same order.

### Reproduction

```rust
// lib.rs
pub struct Grid<A, B, C>(A, B, C);
impl<A, B, C> Grid<A, B, C> { pub const AREA: usize = 1; }
```

```sh
rustc --edition 2024 --crate-type lib --emit=metadata -C incremental=incr lib.rs -o first.rmeta
sed -i '1i // a comment' lib.rs
rustc --edition 2024 --crate-type lib --emit=metadata -C incremental=incr lib.rs -o rebuilt.rmeta
rustc --edition 2024 --crate-type lib --emit=metadata -C incremental=clean lib.rs -o clean.rmeta
cmp rebuilt.rmeta clean.rmeta # differ
```

I expected `rebuilt.rmeta` and `clean.rmeta` to be identical: the same source, the same
compiler, the same flags. Instead they differ in 21 bytes: the 16-byte hash in the header
and the order of three `(DefIndex, u32)` pairs.

It is not specific to this example. A comment added at the top of `lib.rs` gives an
incremental rebuild different metadata from a clean build for
[`either`](https://crates.io/crates/either) 1.13.0, `smallvec` 1.13.2, `memchr` 2.7.4 and
`arrayvec` 0.7.6. With the change suggested below, all four match.

### Root cause

`generics_of` is `cache_on_disk`. `Generics::param_def_id_to_index` is an
`FxHashMap<DefId, u32>`, and the `Encodable` impl for `HashMap` writes entries in iteration
order. `Decodable` collects them back in that order:

- In the first session the map is built by
`own_params.iter().map(|param| (param.def_id, param.index)).collect()`, inserting in
parameter order, then written to the incremental cache in its iteration order.
- In the next session `generics_of` is green and decoded from the cache, so the map is
rebuilt by inserting in the first map's iteration order.

Iteration order of a hashbrown table is bucket order, and which bucket a key lands in
depends on insertion order when keys collide. In the example all three keys hash to
the same home bucket of a 4-bucket table (FxHash of `DefId`s 12, 13 and 14, as
`(krate << 32) | index`). The first key inserted takes bucket 3 and the others wrap to
buckets 0 and 1, so iteration order is the insertion order rotated by one. Decoding inserts
in that rotated order and rotates it again. The metadata encoder then writes the decoded
map in its own iteration order:

| session | encoded order (`DefIndex` → index) |
|---|---|
| clean build | 13→1, 14→2, 12→0 |
| incremental rebuild 1 | 14→2, 12→0, 13→1 |
| incremental rebuild 2 | 12→0, 13→1, 14→2 |
| incremental rebuild 3 | 13→1, 14→2, 12→0 (matches the clean build again) |

The table is from the `.rmeta` files rustc wrote (from byte 1355, after `03`
for the length; each entry is `CrateNum`, `DefIndex`, index). A standalone program using `std::collections::HashMap` with
`rustc_hash::FxBuildHasher` 2.1.1, building and round-tripping the map the same way, prints
exactly the same sequence (attached as `hashmap-roundtrip`; `cargo run -- 12 13 14`).

That is also why the reproduction is sensitive to unrelated items: adding or removing
items changes the `DefIndex`es, and so whether keys collide. Two colliding keys swap on
every round trip; keys that don't collide are stable.

### Impact

- An incremental rebuild's `.rmeta`, and the crate hash computed from it (#154724), depend
on how many times each `generics_of` result has been round-tripped through the cache, not
only on the source. So does the `.rlib` that embeds it, and the metadata of every crate
that depends on it, since those record the crate hash.
- I found no effect on generated code or diagnostics: the map is only used for lookups.
- It hides other incremental bugs from anyone comparing incremental and clean outputs. I
found it while doing that, and it masked a second issue (#…).

### Suggested fix

Keep the map's order deterministic across encoding and decoding. Making the field an
`FxIndexMap<DefId, u32>` does that: an `IndexMap` iterates in insertion order, and decoding
inserts in the encoded order, so a round trip is the identity. It is a two-line change
(`compiler/rustc_middle/src/ty/generics.rs`), and every construction site uses `.collect()`
and every use is `get` or indexing. Another option is not to encode the map at all and
rebuild it from `own_params` when decoding.

With the `FxIndexMap` change, the reproduction, the four crates above, and nine of ten edits
to a larger test fixture give identical metadata after an incremental rebuild. (The tenth is
the other issue, #….) With this change and the one suggested there applied together,
`tests/incremental` (180), the metadata-related UI
(532) and run-make (46) tests, `tests/ui/{consts,statics,const-generics}` (1844) and
`tests/codegen-llvm` (1122) still pass; the full test suite was not run. The `Generics`
struct is the only one I found that derives `TyEncodable`/`Encodable`, reaches metadata and
has a `HashMap` field; the others with such fields (`TypeckResults::used_trait_imports`,
`CrateInfo`, the on-disk cache footer) do not reach `.rmeta`.

#163878 reports the same field as a source of non-reproducibility under `-Zthreads`. The
change here would fix the single-threaded case. Whether it also fixes the parallel one depends
on whether those builds differ only in this map's order. The suggested change was not tested
against #163878's reproduction.

A regression test in the style of `tests/run-make` is attached
(`incr-metadata-generics-order/rmake.rs`). It fails on the current nightly and passes with
the change. It tries the example after
0 to 7 unrelated items, since collisions depend on the `DefIndex`es.

### How it was found

[mirth](https://github.com/PowderworksCode/mirth) checks properties of rustc's metadata
handling across Cargo builds. One property is that an incremental rebuild after an edit
encodes the same metadata as a clean build of the edited source.

### Meta

`rustc --version --verbose`:
```
rustc 1.101.0-nightly (ea137335b 2026-10-05)
binary: rustc
commit-hash: ea137335b78829b4514bf1b4c16302f74fab8581
host: x86_64-unknown-linux-gnu
```

The example also reproduces with 1.95.0 and 1.98.1, where only the 6 bytes of the
reordered pairs differ (those releases do not yet derive the header hash from the
metadata). Releases 1.78 to 1.82 show much larger
differences between incremental and clean metadata for this example, from some other cause,
so I could not tell when this one started.
Loading
Loading