diff --git a/.gitignore b/.gitignore index 2f7896d..89cc600 100644 --- a/.gitignore +++ b/.gitignore @@ -1 +1,2 @@ target/ +rustc/__pycache__/ diff --git a/README.md b/README.md index 1042db7..6de05a5 100644 --- a/README.md +++ b/README.md @@ -23,15 +23,20 @@ Seven real bugs from rustc's history were then replayed by reverting their fixes. mirth catches five; three of those needed a fixture addition or a new step, made knowing the bug. -Pointed at the unmodified compiler with a wider fixture, it found two -incremental bugs that look new, where a rebuild after an edit encodes -different metadata from a clean build, and one known parallel-front-end -bug. +Pointed at the unmodified compiler, it found three incremental bugs that +look new, where a rebuild after an edit publishes different metadata from a +clean build, and one known parallel-front-end bug. Two came from a wider +fixture, the third from fuzzing edits and replaying ten crates' git histories. - [`docs/report.md`](docs/report.md): the experiment, for readers new to it - [`docs/results.md`](docs/results.md): each edit and the output that caught it - [`docs/regressions.md`](docs/regressions.md): the replayed bugs, caught and missed - [`docs/hunt.md`](docs/hunt.md): bugs found in the unmodified compiler +- [`docs/scale.md`](docs/scale.md): replaying crates' histories and fuzzing edits at scale +- [`docs/properties.md`](docs/properties.md): checkable properties surveyed from 1,000 rustc bugs +- [`docs/motivating.md`](docs/motivating.md): the real rustc bugs behind each property, each reproduced before and after its fix +- [`docs/ur-queries.md`](docs/ur-queries.md): the bugs' patterns, and closed bugs' patterns, as Ur queries over rustc's source +- [`docs/shadow-mode.md`](docs/shadow-mode.md): checking reuse inside rustc, and what exists today - [`docs/plan.md`](docs/plan.md): the plan the work followed, with the properties ## An instrumented compiler @@ -46,6 +51,9 @@ rustc/check.sh chain # check fixtures/chain rustc/edits.sh chain # apply, check and revert each edit in rustc/edits EDITS=regressions rustc/edits.sh chain # the same for the past bugs in rustc/regressions rustc/hunt.sh wide # repeated threaded builds, and P6 for each of fixtures/wide/edits +rustc/fuzz.py --rustc --fixture fixtures/sink --work # random edits, P6 on each +rustc/replay.py --rustc --repo --work # a crate's history, P6 per commit +rustc/audit-options.py --rustc --source --crate fixtures/audit/lib.rs # untracked options ``` The build takes about an hour on 16 cores. `rustc/rmeta.toml` says what is diff --git a/docs/hunt.md b/docs/hunt.md index 20b656a..d0d01f5 100644 --- a/docs/hunt.md +++ b/docs/hunt.md @@ -27,6 +27,8 @@ with `-Zthreads=8`. `rustc/check.sh wide` runs the ordinary checks. | 1 | incremental rebuilds encode `Generics::param_def_id_to_index` in a different order from clean builds | **looks new**; root cause found; fix and regression test written | | 2 | incremental rebuilds encode a string literal twice where clean builds encode it once | **looks new**; root cause found, regression from #116707 (1.90); fix and regression test written | | 3 | with `-Zthreads=8`, two traits with `-> impl Trait` methods give different metadata from run to run | known: [#162202](https://github.com/rust-lang/rust/issues/162202) | +| 4 | incremental rebuilds republish the previous session's metadata when an edit moves no span, so its source map describes old files | **looks new**; found later by the fuzzer and the history replay ([`scale.md`](scale.md)); root cause found, regression from #114669 (1.90); fix and regression test written | +| 5 | `-Zemit-stack-sizes`, `-Zcodegen-source-order` and `-Zbuild-sdylib-interface` are untracked but change output that incremental compilation reuses | **looks new**; found by a query written from a closed bug and an option audit ([`ur-queries.md`](ur-queries.md)); report drafted | Findings 1 and 2 are single-threaded: an ordinary `cargo build`, an edit, another `cargo build`, and the metadata differs from a clean build of the edited source. Both come @@ -35,9 +37,12 @@ incremental cache unchanged. All three reproduce with the official `nightly-2026 without mirth: `docs/hunt/repro.sh` runs them. None was searched for: P5 and P6 reported them on the first run of the new fixture. -Draft bug reports for 1 and 2, written to be filed upstream, are -[`hunt/issue-generics-order.md`](hunt/issue-generics-order.md) and -[`hunt/issue-literal-dedup.md`](hunt/issue-literal-dedup.md). Each has a candidate fix +Draft bug reports for 1, 2, 4 and 5, written to be filed upstream, are +[`hunt/issue-generics-order.md`](hunt/issue-generics-order.md), +[`hunt/issue-literal-dedup.md`](hunt/issue-literal-dedup.md) and +[`hunt/issue-stale-metadata-reuse.md`](hunt/issue-stale-metadata-reuse.md) and +[`hunt/issue-untracked-options.md`](hunt/issue-untracked-options.md), the last a comment for +rust-lang/rust#84232. Each has a candidate fix (`hunt/*.patch`) and a regression test in the style of rustc's `tests/run-make` (`hunt/tests/`), which fails on the pinned compiler and passes with the fix. diff --git a/docs/hunt/alloc-dedup-on-decode.patch b/docs/hunt/alloc-dedup-on-decode.patch index 8959c8a..9d6916d 100644 --- a/docs/hunt/alloc-dedup-on-decode.patch +++ b/docs/hunt/alloc-dedup-on-decode.patch @@ -1,19 +1,41 @@ --- a/compiler/rustc_middle/src/mir/interpret/mod.rs +++ b/compiler/rustc_middle/src/mir/interpret/mod.rs -@@ -214,7 +214,15 @@ - trace!("creating memory alloc ID"); - let alloc = as Decodable<_>>::decode(decoder); +@@ -99,6 +99,10 @@ + VTable, + Static, + Type, ++ /// Memory that was deduplicated when it was created (string literals, for instance), and ++ /// must be deduplicated again when decoded: otherwise the same allocation decoded from the ++ /// incremental cache and created afresh would get two `AllocId`s. ++ DedupAlloc, + } + + pub fn specialized_encode_alloc_id<'tcx, E: TyEncoder<'tcx>>( +@@ -109,7 +113,13 @@ + match tcx.global_alloc(alloc_id) { + GlobalAlloc::Memory(alloc) => { + trace!("encoding {:?} with {:#?}", alloc_id, alloc); +- AllocDiscriminant::Alloc.encode(encoder); ++ let deduplicated = tcx.alloc_map.dedup.lock().get(&(GlobalAlloc::Memory(alloc), CTFE_ALLOC_SALT)) ++ == Some(&alloc_id); ++ if deduplicated { ++ AllocDiscriminant::DedupAlloc.encode(encoder); ++ } else { ++ AllocDiscriminant::Alloc.encode(encoder); ++ } + alloc.encode(encoder); + } + GlobalAlloc::Function { instance } => { +@@ -216,6 +226,12 @@ trace!("decoded alloc {:?}", alloc); -- decoder.interner().reserve_and_set_memory_alloc(alloc) -+ // Immutable memory is deduplicated when it is created (string literals, for -+ // instance), so deduplicate it here too. Otherwise one allocation decoded -+ // from the incremental cache and the same one created afresh get two -+ // `AllocId`s, and metadata depends on which results came from the cache. -+ if alloc.inner().mutability.is_not() { -+ decoder.interner().reserve_and_set_memory_dedup(alloc, CTFE_ALLOC_SALT) -+ } else { -+ decoder.interner().reserve_and_set_memory_alloc(alloc) -+ } + decoder.interner().reserve_and_set_memory_alloc(alloc) } ++ AllocDiscriminant::DedupAlloc => { ++ trace!("creating deduplicated memory alloc ID"); ++ let alloc = as Decodable<_>>::decode(decoder); ++ trace!("decoded alloc {:?}", alloc); ++ decoder.interner().reserve_and_set_memory_dedup(alloc, CTFE_ALLOC_SALT) ++ } AllocDiscriminant::Fn => { trace!("creating fn alloc ID"); + let instance = ty::Instance::decode(decoder); diff --git a/docs/hunt/generics-index-map.patch b/docs/hunt/generics-index-map.patch index b6fdf09..ea913e6 100644 --- a/docs/hunt/generics-index-map.patch +++ b/docs/hunt/generics-index-map.patch @@ -18,3 +18,12 @@ pub has_self: bool, pub has_late_bound_regions: Option, +@@ -132,8 +132,6 @@ + + impl std::fmt::Debug for Generics { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> Result<(), std::fmt::Error> { +- // ironically, we get this warning because of what we're trying to fix. +- #[expect(rustc::potential_query_instability)] + let mut stabilized_hashmap = self.param_def_id_to_index.iter().collect::>(); + stabilized_hashmap.sort_by_key(|(_, v)| **v); + f.debug_struct("Generics") diff --git a/docs/hunt/issue-generics-order.md b/docs/hunt/issue-generics-order.md index 4ed51b3..d353e97 100644 --- a/docs/hunt/issue-generics-order.md +++ b/docs/hunt/issue-generics-order.md @@ -77,21 +77,31 @@ every round trip; keys that don't collide are stable. - It hides other incremental bugs from anyone comparing incremental and clean outputs. I found it while doing that, and it masked a second issue (#…). +The instability was known in one place: `impl Debug for Generics` collects and sorts the map +before printing it, under `#[expect(rustc::potential_query_instability)]` and the comment +"ironically, we get this warning because of what we're trying to fix". The encoding was +not given the same treatment. + ### Suggested fix Keep the map's order deterministic across encoding and decoding. Making the field an `FxIndexMap` does that: an `IndexMap` iterates in insertion order, and decoding -inserts in the encoded order, so a round trip is the identity. It is a two-line change -(`compiler/rustc_middle/src/ty/generics.rs`), and every construction site uses `.collect()` +inserts in the encoded order, so a round trip is the identity. It is a small change +(`compiler/rustc_middle/src/ty/generics.rs`; the `#[expect]` in the `Debug` impl then has +nothing to expect and goes too), and every construction site uses `.collect()` and every use is `get` or indexing. Another option is not to encode the map at all and rebuild it from `own_params` when decoding. With the `FxIndexMap` change, the reproduction, the four crates above, and nine of ten edits to a larger test fixture give identical metadata after an incremental rebuild. (The tenth is -the other issue, #….) With this change and the one suggested there applied together, -`tests/incremental` (180), the metadata-related UI -(532) and run-make (46) tests, `tests/ui/{consts,statics,const-generics}` (1844) and -`tests/codegen-llvm` (1122) still pass; the full test suite was not run. The `Generics` +the issue about string literals, #….) + +With all three changes proposed in this series applied (this one and the two in #… and +#…), each attached regression test passes, and each fails when only its own change is +removed. These rustc tests still pass: `tests/incremental` (180), the UI tests in +`tests/ui/{deprecation,crate-loading,rmeta,extern,cross-crate}` (532), 46 metadata-related +`tests/run-make` tests, `tests/ui/{consts,statics,const-generics}` (1844) and +`tests/codegen-llvm` (1122). The full test suite was not run. A fuzzer making random edits to a 1,200-line test workspace ran 10,717 edits on the patched compiler, 8,936 of which built and were compared with a clean build, and a replay of ten crates' git histories compared 4,862 commits; neither found a difference. The `Generics` struct is the only one I found that derives `TyEncodable`/`Encodable`, reaches metadata and has a `HashMap` field; the others with such fields (`TypeckResults::used_trait_imports`, `CrateInfo`, the on-disk cache footer) do not reach `.rmeta`. diff --git a/docs/hunt/issue-literal-dedup.md b/docs/hunt/issue-literal-dedup.md index d4716a4..2980a5c 100644 --- a/docs/hunt/issue-literal-dedup.md +++ b/docs/hunt/issue-literal-dedup.md @@ -79,32 +79,48 @@ two and the clean build one. ### Suggested fix -Decode immutable memory allocations the way they were created, deduplicated: +Decode an allocation the way it was created. When encoding a memory allocation, check +whether the deduplication map holds exactly this `AllocId` for it, under `CTFE_ALLOC_SALT`. +If it does, encode it with a new discriminant, and decode that one through +`reserve_and_set_memory_dedup` (attached, `alloc-dedup-on-decode.patch`): ```diff - AllocDiscriminant::Alloc => { - let alloc = as Decodable<_>>::decode(decoder); -- decoder.interner().reserve_and_set_memory_alloc(alloc) -+ if alloc.inner().mutability.is_not() { -+ decoder.interner().reserve_and_set_memory_dedup(alloc, CTFE_ALLOC_SALT) -+ } else { -+ decoder.interner().reserve_and_set_memory_alloc(alloc) -+ } - } + GlobalAlloc::Memory(alloc) => { +- AllocDiscriminant::Alloc.encode(encoder); ++ let deduplicated = tcx.alloc_map.dedup.lock().get(&(GlobalAlloc::Memory(alloc), CTFE_ALLOC_SALT)) ++ == Some(&alloc_id); ++ if deduplicated { ++ AllocDiscriminant::DedupAlloc.encode(encoder); ++ } else { ++ AllocDiscriminant::Alloc.encode(encoder); ++ } + alloc.encode(encoder); + } + ... ++ AllocDiscriminant::DedupAlloc => { ++ let alloc = as Decodable<_>>::decode(decoder); ++ decoder.interner().reserve_and_set_memory_dedup(alloc, CTFE_ALLOC_SALT) ++ } ``` -With it, together with the fix for #… (the other issue), the reproduction and all ten -single-threaded incremental edits to a larger test fixture give identical metadata. These all -still pass: `tests/incremental` (180), `tests/ui/{consts,statics,const-generics}` (1844), -`tests/codegen-llvm` (1122), and the 532 UI and 46 run-make tests about metadata and crate -loading that I run. The full test suite was not run. - -It does more than strictly needed: it would also merge an immutable allocation decoded from a -dependency's metadata with an identical local one, and two immutable allocations that were -distinct when created (results of different constants, say). As far as I know, neither has -a guaranteed unique address, but someone who knows the const-eval memory model should -confirm. A narrower fix records, when encoding, whether an allocation was created -through deduplication and with which salt, and repeats exactly that when decoding. +A simpler change, deduplicating every immutable allocation on decode, is wrong. I tried it +first, and it fixes this reproduction but breaks the same property elsewhere. It merges +allocations that a clean build keeps apart, because they were never deduplicated when +created. On serde at commit `2f58a20` ("Inline is_human_readable", 2017), an incremental +rebuild then encoded 59 bytes fewer of `interpret-alloc-index` than a clean build +(`-Zmeta-stats`). With the narrower change above, that commit and its neighbours match. + +Two things for a reviewer. Only `CTFE_ALLOC_SALT` is handled: allocations deduplicated +under other salts (Miri's) are encoded as before, which only matters if those reach an +incremental cache. And the change adds a discriminant to the encoding of allocations in +metadata and in the incremental cache, so it may want a metadata version bump. + +With all three changes proposed in this series applied (this one and the two in #… and +#…), each attached regression test passes, and each fails when only its own change is +removed. These rustc tests still pass: `tests/incremental` (180), the UI tests in +`tests/ui/{deprecation,crate-loading,rmeta,extern,cross-crate}` (532), 46 metadata-related +`tests/run-make` tests, `tests/ui/{consts,statics,const-generics}` (1844) and +`tests/codegen-llvm` (1122). The full test suite was not run. A fuzzer making random edits to a 1,200-line test workspace ran 10,717 edits on the patched compiler, 8,936 of which built and were compared with a clean build, and a replay of ten crates' git histories compared 4,862 commits; neither found a difference. A regression test in the style of `tests/run-make`, which fails before the change and passes after, is attached (`incr-metadata-literal-dedup/rmake.rs`). diff --git a/docs/hunt/issue-stale-metadata-reuse.md b/docs/hunt/issue-stale-metadata-reuse.md new file mode 100644 index 0000000..e40414d --- /dev/null +++ b/docs/hunt/issue-stale-metadata-reuse.md @@ -0,0 +1,118 @@ +# Incremental rebuilds reuse stale metadata when an edit moves no span + + + +When an edit changes a source file without changing any query result (a comment added at +the end of a file, for instance), an incremental rebuild republishes the previous +session's `.rmeta` unchanged. That metadata describes the old file: its source map has the +old content hash and length. A clean build of the same source encodes the new ones. +Dependents then fail to find the source, so diagnostics that point into the dependency lose +their snippet. + +### Reproduction + +```sh +printf 'pub fn f(a: u32) -> u32 { a }\n' > dep.rs +rustc --edition 2024 --crate-type lib --crate-name dep --emit=metadata,link -C incremental=incr --out-dir rebuilt dep.rs +cp rebuilt/libdep.rmeta before.rmeta +printf '// a comment at the end\n' >> dep.rs +rustc --edition 2024 --crate-type lib --crate-name dep --emit=metadata,link -C incremental=incr --out-dir rebuilt dep.rs +rustc --edition 2024 --crate-type lib --crate-name dep --emit=metadata,link -C incremental=clean --out-dir clean dep.rs +cmp before.rmeta rebuilt/libdep.rmeta # identical: the old metadata was reused +cmp rebuilt/libdep.rmeta clean/libdep.rmeta # differ +``` + +I expected the rebuilt metadata to equal the clean build's. Instead it is the previous +session's, byte for byte. The clean build differs in the source map entry for `dep.rs`: +its length, line table and content hash. + +What a dependent sees: + +```rust +// user.rs +pub fn use_it() -> u32 { dep::f() } +``` + +```sh +rustc --edition 2024 --crate-type lib --crate-name user --extern dep=rebuilt/libdep.rlib user.rs +``` + +Built against the incrementally rebuilt `dep`, the note has no snippet and a different +column: + +```text +note: function defined here + --> dep.rs:1:7 +``` + +Built against the clean `dep`: + +```text +note: function defined here + --> dep.rs:1:8 + | +1 | pub fn f(a: u32) -> u32 { a } + | ^ +``` + +### Root cause + +Since #114669 ("Make metadata a workproduct and reuse it", merged 2025-07-04), metadata is +encoded inside a dep-graph task, and a later session reuses the saved file when that node +can be marked green (`encode_metadata`, `rustc_metadata/src/rmeta/encoder.rs`). The node's +dependencies are the queries read while encoding. But `encode_source_map` reads the +`SourceFile`s straight from `tcx.sess.source_map()`, which is not a query, so the encoded +names, lengths, line tables and content hashes are not dependencies. An edit that changes +no query result, such as a comment after the last item or text inside a comment that keeps +every span where it was, leaves the node green. The old file is reused and describes source +that no longer exists. + +Bisection: `nightly-2025-07-03` (`667787527`) rebuilds the metadata; +`nightly-2025-07-06` (`5adb489a8`) reuses it. 1.88.0 and 1.89.0 are not affected; +1.90.0 to the current nightly are. + +### It happens in real histories + +Replaying git histories with an incremental rebuild per commit, compared with a clean build +each time, hit this on ordinary commits. The rebuilt `.rmeta` differs from the clean one only +in the header hash and the content hash of the edited file: + +- `memchr`, 5 consecutive commits in 2018 (e.g. `49ab2ef`, "fallback: fix variable name in + docstrings": the edited text is in a module compiled out on x86_64); +- `smallvec`, 2 commits (e.g. `01354f1`, three lines changed in `lib.rs`); +- `hashbrown`, 1 commit (`957b590`, one line in `src/raw/mod.rs`). + +### Suggested fix + +Make the metadata task depend on the source files it encodes. The attached patch adds an +`eval_always` query, `local_source_files_fingerprint`, which hashes every local +`SourceFile`'s name, length and content hash. It also reads that query at the start of the +metadata task. An `eval_always` query runs again in every session, but a node depending on +it stays green when its result is unchanged, so metadata is still reused when the files are +truly unchanged. With all three changes proposed in this series applied (this one and the two in #… and +#…), each attached regression test passes, and each fails when only its own change is +removed. These rustc tests still pass: `tests/incremental` (180), the UI tests in +`tests/ui/{deprecation,crate-loading,rmeta,extern,cross-crate}` (532), 46 metadata-related +`tests/run-make` tests, `tests/ui/{consts,statics,const-generics}` (1844) and +`tests/codegen-llvm` (1122). The full test suite was not run. A fuzzer making random edits to a 1,200-line test workspace ran 10,717 edits on the patched compiler, 8,936 of which built and were compared with a clean build, and a replay of ten crates' git histories compared 4,862 commits; neither found a difference. + +A regression test in the style of `tests/run-make` is attached +(`incr-metadata-stale-source/rmake.rs`). It fails on the current nightly and passes with +the change. + +### How it was found + +[mirth](https://github.com/PowderworksCode/mirth) checks properties of rustc's metadata +handling across Cargo builds. One is that an incremental rebuild after an edit encodes the +same metadata as a clean build of the edited source. A fuzzer making random edits to a test +workspace and a replay of crates' git histories both reported it. + +### Meta + +`rustc --version --verbose`: +``` +rustc 1.101.0-nightly (ea137335b 2026-10-05) +binary: rustc +commit-hash: ea137335b78829b4514bf1b4c16302f74fab8581 +host: x86_64-unknown-linux-gnu +``` diff --git a/docs/hunt/issue-untracked-options.md b/docs/hunt/issue-untracked-options.md new file mode 100644 index 0000000..c67f832 --- /dev/null +++ b/docs/hunt/issue-untracked-options.md @@ -0,0 +1,59 @@ +# `-Zemit-stack-sizes`, `-Zcodegen-source-order` and `-Zbuild-sdylib-interface` are untracked but change output that incremental compilation reuses + + + +All three are marked `[UNTRACKED]` in `compiler/rustc_session/src/options.rs`, so they're +left out of the dependency-tracking hash. Adding one between two incremental sessions +leaves the second session's results as the first's, and the option has no effect. A clean +build with the option produces something different. + +### Reproduction + +`lib.rs` can be any library with non-generic code of its own; the one used here is +`fixtures/audit/lib.rs` in mirth. + +```sh +rustc --edition 2021 --crate-type lib -C incremental=incr --out-dir out lib.rs +rustc --edition 2021 --crate-type lib -C incremental=incr --out-dir out lib.rs -Zemit-stack-sizes +rustc --edition 2021 --crate-type lib -C incremental=clean --out-dir clean-out lib.rs -Zemit-stack-sizes +# count .stack_sizes sections in each rlib's objects +for d in out clean-out; do (mkdir $d/x && cd $d/x && ar x ../*.rlib && for f in *.o; do readelf -S $f; done | grep -c stack_sizes); done +``` + +On `nightly-2026-10-06` the incremental rebuild has 0 `.stack_sizes` sections and the clean +build 276. With `-Zbuild-sdylib-interface`, which compiles an interface without function +bodies, the clean build prints 27 warnings (unused variables in the bodies it skipped) and +writes different metadata and objects. The incremental rebuild prints the first session's 5 +warnings and keeps its full build. + +A tool that does this for every untracked boolean option, comparing metadata, each rlib +member (object names have their incremental session suffix removed, since the objects are +otherwise identical), diagnostics and files written, is +[`rustc/audit-options.py`](../../rustc/audit-options.py) in mirth: + +```text +(control: no option) same +-Csave-temps=yes 128 files only a clean build writes +-Zbuild-sdylib-interface=yes metadata; object code (44 rlib members differ or exist on one side); diagnostics (5 lines incrementally, 27 clean) +-Zcodegen-source-order=yes object code (1 rlib members differ or exist on one side) +-Zemit-stack-sizes=yes object code (42 rlib members differ or exist on one side) +``` + +The other 45 boolean untracked options gave the same output incrementally as clean. +(`-Zdump-dep-graph` and `-Zno-parallel-backend` failed to build this crate either way.) +`-Csave-temps` writing no temporaries for reused codegen units is probably fine for a +debugging option. + +### Suggested fix + +Mark `emit_stack_sizes`, `codegen_source_order` and `build_sdylib_interface` `[TRACKED]`. +An audit like the one above could run in CI over every untracked option, so a new option +marked `[UNTRACKED]` that changes reused output is caught when it is added; that is the +long tail this issue's discussion worries about. + +### How it was found + +[mirth](https://github.com/PowderworksCode/mirth) wrote the pattern behind #66955 +(`--remap-path-prefix` untracked) as a query over rustc's source, reads of options, then +joined it with the options marked `[UNTRACKED]`. The audit is the differential check this +issue's third comment suggests. diff --git a/docs/hunt/metadata-source-files.patch b/docs/hunt/metadata-source-files.patch new file mode 100644 index 0000000..9e2a56c --- /dev/null +++ b/docs/hunt/metadata-source-files.patch @@ -0,0 +1,56 @@ +--- a/compiler/rustc_middle/src/queries.rs ++++ b/compiler/rustc_middle/src/queries.rs +@@ -190,6 +190,15 @@ + desc { "get the value of an environment variable" } + } + ++ /// A fingerprint of every local source file's name, length and contents. Metadata ++ /// encodes all three in its source map, so reusing metadata from the incremental ++ /// cache has to depend on them; nothing else that metadata reads does. ++ query local_source_files_fingerprint(_: ()) -> Svh { ++ // The source map is global state ++ eval_always ++ desc { "fingerprinting the local source files" } ++ } ++ + query resolutions(_: ()) -> &'tcx ResolverGlobalCtxt { + desc { "getting the resolver outputs" } + } +--- a/compiler/rustc_metadata/src/rmeta/encoder.rs ++++ b/compiler/rustc_metadata/src/rmeta/encoder.rs +@@ -15,6 +15,7 @@ + use rustc_data_structures::owned_slice::slice_owned; + use rustc_data_structures::stable_hash::{StableHash, StableHasher}; + use rustc_data_structures::sync::{par_for_each_in, par_join}; ++use rustc_data_structures::svh::Svh; + use rustc_data_structures::temp_dir::MaybeTempDir; + use rustc_data_structures::thousands::usize_with_underscores; + use rustc_hir as hir; +@@ -2598,6 +2599,8 @@ + dep_node, + tcx, + || { ++ // Make the metadata depend on the source files it describes. ++ let _ = tcx.local_source_files_fingerprint(()); + with_encode_metadata_header(tcx, path, |ecx| { + // Encode all the entries and extra information in the crate, + // culminating in the `CrateRoot` which points to all of it. +@@ -2768,6 +2771,18 @@ + + pub(crate) fn provide(providers: &mut Providers) { + *providers = Providers { ++ local_source_files_fingerprint: |tcx, ()| { ++ let mut hasher = StableHasher::new(); ++ for file in tcx.sess.source_map().files().iter() { ++ if file.is_imported() { ++ continue; ++ } ++ std::hash::Hash::hash(&file.name, &mut hasher); ++ std::hash::Hash::hash(&file.unnormalized_source_len, &mut hasher); ++ std::hash::Hash::hash(&file.src_hash, &mut hasher); ++ } ++ Svh::new(hasher.finish()) ++ }, + doc_link_resolutions: |tcx, def_id| { + tcx.resolutions(()) + .doc_link_resolutions diff --git a/docs/hunt/repro.sh b/docs/hunt/repro.sh index 78de6e1..51b6bb3 100755 --- a/docs/hunt/repro.sh +++ b/docs/hunt/repro.sh @@ -30,3 +30,15 @@ cp "$here/p6-literals/after.rs" "$d/lib.rs" rc --crate-type lib --emit=metadata,link -Cincremental="$d/i3" --out-dir "$d/o1" "$d/lib.rs" rc --crate-type lib --emit=metadata,link -Cincremental="$d/i4" --out-dir "$d/o2" "$d/lib.rs" cmp -s "$d/o1/liblib.rmeta" "$d/o2/liblib.rmeta" && echo same || echo DIFFER + +echo -n "stale-source, incremental rebuild after a comment at the end vs a clean build: " +mkdir -p "$d/s1" "$d/s2" +printf 'pub fn f(a: u32) -> u32 { a }\n' > "$d/dep.rs" +rc --crate-type lib --crate-name dep --emit=metadata,link -Cincremental="$d/i5" --out-dir "$d/s1" "$d/dep.rs" +cp "$d/s1/libdep.rmeta" "$d/before.rmeta" +printf '// a comment at the end\n' >> "$d/dep.rs" +rc --crate-type lib --crate-name dep --emit=metadata,link -Cincremental="$d/i5" --out-dir "$d/s1" "$d/dep.rs" +rc --crate-type lib --crate-name dep --emit=metadata,link -Cincremental="$d/i6" --out-dir "$d/s2" "$d/dep.rs" +if cmp -s "$d/s1/libdep.rmeta" "$d/s2/libdep.rmeta"; then echo same +elif cmp -s "$d/s1/libdep.rmeta" "$d/before.rmeta"; then echo "DIFFER (the previous session's metadata, republished)" +else echo DIFFER; fi diff --git a/docs/hunt/tests/incr-metadata-stale-source/rmake.rs b/docs/hunt/tests/incr-metadata-stale-source/rmake.rs new file mode 100644 index 0000000..30f2611 --- /dev/null +++ b/docs/hunt/tests/incr-metadata-stale-source/rmake.rs @@ -0,0 +1,40 @@ +//@ needs-target-std +// +// An incremental rebuild must encode the same metadata as a clean build of the same source. +// Metadata is reused from the incremental cache when its dep-node is green, but its source +// map records each file's length and content hash, which no query tracked. So after an edit +// that changes no query result, a comment at the end of the file, the rebuild republished the +// previous session's metadata, describing a file that no longer exists. + +use run_make_support::{rfs, rustc}; + +fn build(incremental: &str, out_dir: &str) -> Vec { + rustc() + .input("lib.rs") + .crate_name("foo") + .crate_type("lib") + .emit("metadata,link") + .incremental(incremental) + .out_dir(out_dir) + .run(); + rfs::read(format!("{out_dir}/libfoo.rmeta")) +} + +fn main() { + rfs::write("lib.rs", "pub fn f(a: u32) -> u32 { a }\n"); + let before = build("incr", "out"); + rfs::write( + "lib.rs", + "pub fn f(a: u32) -> u32 { a }\n// a comment at the end\n", + ); + let incremental = build("incr", "out"); + let clean = build("clean-incr", "clean-out"); + assert!( + incremental != before, + "the rebuild republished the previous session's metadata" + ); + assert!( + incremental == clean, + "the incremental rebuild's metadata differs from a clean build's" + ); +} diff --git a/docs/motivating.md b/docs/motivating.md new file mode 100644 index 0000000..93905e3 --- /dev/null +++ b/docs/motivating.md @@ -0,0 +1,312 @@ +# The bugs behind each property + +mirth's properties were designed from first principles: what must hold for rustc's +handling of metadata and incremental compilation to be correct. This document records the +real rust-lang/rust bugs that violated each one, so a maintainer can see why a check exists. +Each bug was **reproduced on a toolchain from before its fix and shown fixed on one after**, +with no mirth involved: the original bug, as users met it. + +The bugs were found by searching rust-lang/rust for each property; every issue and PR was +read, and every reproduction run on this machine (Linux x86_64). Reproduce one with +`docs/motivating/run.py `; the outputs used here are in `docs/motivating/out/`. + +Some bugs appear under two properties. Three found by mirth itself +([`hunt/`](hunt)) are listed under P6 at the end. + +## P1: An `.rmeta` reaches its final path only by a rename of a fully written file + +A reader that opens a half-written or truncated `.rmeta` misreads it, usually as an ICE in the decoder. rustc writes the metadata to a temporary directory and renames it into place; P1 checks that protocol on every process. The first bug below is where the protocol came from; the other two published a truncated file through the rename. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#45841](https://github.com/rust-lang/rust/issues/45841) ICE: index out of bounds in libserialize/leb128.rs | [#45899](https://github.com/rust-lang/rust/pull/45899) | 2017-11-18 | `nightly-2017-11-17` → `nightly-2017-11-20` | reproduced | — | +| [#117254](https://github.com/rust-lang/rust/issues/117254) FileEncoder delayed error reporting is still broken | [#117301](https://github.com/rust-lang/rust/pull/117301) | 2023-11-26 | `nightly-2023-11-26` → `nightly-2023-11-28` | reproduced | — | +| [#119456](https://github.com/rust-lang/rust/issues/119456) ICE 'range start index ... out of range for slice of length 16384': rmeta I/O er | [#119510](https://github.com/rust-lang/rust/pull/119510) | 2024-01-03 | `nightly-2024-01-01` → `nightly-2024-01-05` | reproduced | — | + +**#45841.** Before the fix, rustc_trans::back::link::emit_metadata did File::create(out_filename) + write_all straight into the final lib.rmeta path (O_WRONLY|O_CREAT|O_TRUNC). Another rustc that was searching the same -L directory could open the file while it was truncated or only partly written, and then panicked decoding it. The issue thread (arielb1: "the compiler writing metadata in parts, so that another instance of the compiler can read metadata while it is being partially written to. Should be fixable by doing an atomic rename") and PR #45899 ("atomically write .rmeta outputs to avoid races ... write a temporary file and then rename it") record exactly this. The fix added the rmeta* tempdir inside the output directory plus fs::rename, and the comment "To avoid races with another rustc process scanning the output directory..." is still in compiler/rustc_metadata/src/fs.rs today. This is the bug that P1 itself comes from. + +*This run:* before: the .rmeta is rewritten in place (same inode), and 2 of 1,704 concurrent readers hit the issue's ICE in leb128.rs. After: replaced by rename, 0 of 1,712. Outputs: [`before`](motivating/out/45841/before.txt), [`after`](motivating/out/45841/after.txt). + +*Expected, from the issue and the research:* nightly-2017-11-17 (rustc 1.23.0-nightly d0f8e2913 2017-11-16): run.sh prints the same inode before and after the rebuild, then "BUG: liba.rmeta rewritten in place" and "BUG: old hard link sees new bytes". strace shows openat("st/liba.rmeta", O_WRONLY|O_CREAT|O_TRUNC) with no rename. nightly-2017-11-20 (5041b3bb3 2017-11-19): the inode changes, then "OK: liba.rmeta replaced by rename" and "OK: old hard link still holds previous metadata". strace shows the write going to st/rmeta.XXXX/rust.metadata.bin, followed by rename(... , "st/liba.rmeta"). race.sh on 2017-11-17 gave 1 failure in 1078 reader runs in one 60s attempt and 0 in another 90s attempt. The failure was the exact ICE from the issue: "thread 'rustc' panicked at 'index out of bounds: the len is 2432992 but the index is 6805837', /checkout/src/libserialize/leb128.rs:59:20". On 2017-11-20: 0 failures in 1719 runs. + +**#117254.** rmeta encoding never called FileEncoder::finish, so an I/O error while writing the temp file (ENOSPC in crater, EFBIG here) was swallowed. rustc then renamed the short temp file to the final lib.rmeta and exited 0. Downstream crates ICE'd in MemDecoder ("We're going ahead to decode a result which was not completely written out", issue text). The rename itself is atomic, but the file it publishes is not fully written, which breaks the second half of P1. Introduced by the delayed error scheme of #94732 and not fixed by #115542. #117301 added the finish() check, but only with emit_err; see the next entry for the remaining hole. + +*This run:* before: a write error is swallowed (exit 0), a 1 MiB truncated .rmeta is published, and a dependent ICEs. After: an error and exit 1, but the truncated file is still published. Outputs: [`before`](motivating/out/117254/before.txt), [`after`](motivating/out/117254/after.txt). + +*Expected, from the issue and the research:* nightly-2023-11-26 (1.76.0-nightly f5dc2653f 2023-11-25): no diagnostic, "rustc exit=0", and out/libbig.rmeta is exactly 1048576 bytes, a truncated file at the final path. The downstream rustc then ICEs in rustc_serialize/src/opaque.rs ("range start index ... out of range for slice of length 1048576"). nightly-2023-11-28 (49b3924bd 2023-11-27): rustc now prints "error: failed to write to `.../rmetaXXXX/lib.rmeta`: File too large (os error 27)" and exits 1. However, the 1048576-byte out/libbig.rmeta is still published, and the downstream rustc still panics with "range start index 14979595 out of range for slice of length 1048576" (fixed fully by #119510 below). + +**#119456.** After #117301, rmeta write errors were reported with emit_err, which does not stop compilation. rustc kept going and renamed the incomplete temp file over lib.rmeta, and cargo could start a dependent build that read it (PR #119510: "there is a window of time between the call to emit_err and the full error reporting where rustc believes it has emitted a valid rmeta file and will permit Cargo to launch a build for a dependent crate"). Switching to emit_fatal aborts before the rename, so nothing partial reaches the final path. The reporter hit this on 1.75.0 with typst as a dependency (disk-full); the PR author reproduced it with an LD_PRELOAD write() that randomly returns ENOSPC. + +*This run:* before: an error and exit 1, yet the truncated .rmeta is published and a dependent ICEs. After: nothing is published, and the dependent gets `can't find crate`. Outputs: [`before`](motivating/out/119456/before.txt), [`after`](motivating/out/119456/after.txt). + +*Expected, from the issue and the research:* nightly-2024-01-01 (1.77.0-nightly e51e98dde 2023-12-31): "error: failed to write to `.../rmetaXXXX/lib.rmeta`: File too large (os error 27)" and exit 1, yet `ls -l out` shows libbig.rmeta at 1048576 bytes. The reader then panics: "thread 'rustc' panicked at .../compiler/rustc_serialize/src/opaque.rs:262:42: range start index 12750977 out of range for slice of length 1048576". nightly-2024-01-05 (f688dd684 2024-01-04): same error and exit 1, but out/ is empty (the truncated file is never renamed into place), and the reader gets a clean "error[E0463]: can't find crate for `big`". Current stable 1.97.1 behaves the same as 2024-01-05 (the temp file is now named full.rmeta). + +## P2: No process opens a dependency's `.rmeta` before it has been renamed into place + +With pipelining, a dependent starts as soon as its dependency's metadata exists, so the order of writes and opens across processes matters. P2 compares the timestamps of opens in readers with renames in writers. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#45841](https://github.com/rust-lang/rust/issues/45841) ICE: index out of bounds in libserialize/leb128.rs | [#45899](https://github.com/rust-lang/rust/pull/45899) | 2017-11-18 | `nightly-2017-11-17` → `nightly-2017-11-20` | reproduced | — | +| [#68149](https://github.com/rust-lang/rust/issues/68149) Spurious rebuilds under pipelining: a consumer opened | [#68298](https://github.com/rust-lang/rust/pull/68298) | 2020-01-23 | `nightly-2020-01-22` → `nightly-2020-01-24` | reproduced | — | + +**#45841.** Before the fix, rustc_trans::back::link::emit_metadata did File::create(out_filename) + write_all straight into the final lib.rmeta path (O_WRONLY|O_CREAT|O_TRUNC). Another rustc that was searching the same -L directory could open the file while it was truncated or only partly written, and then panicked decoding it. The issue thread (arielb1: "the compiler writing metadata in parts, so that another instance of the compiler can read metadata while it is being partially written to. Should be fixable by doing an atomic rename") and PR #45899 ("atomically write .rmeta outputs to avoid races ... write a temporary file and then rename it") record exactly this. The fix added the rmeta* tempdir inside the output directory plus fs::rename, and the comment "To avoid races with another rustc process scanning the output directory..." is still in compiler/rustc_metadata/src/fs.rs today. This is the bug that P1 itself comes from. + +*This run:* before: the .rmeta is rewritten in place (same inode), and 2 of 1,704 concurrent readers hit the issue's ICE in leb128.rs. After: replaced by rename, 0 of 1,712. Outputs: [`before`](motivating/out/45841/before.txt), [`after`](motivating/out/45841/after.txt). + +*Expected, from the issue and the research:* nightly-2017-11-17 (rustc 1.23.0-nightly d0f8e2913 2017-11-16): run.sh prints the same inode before and after the rebuild, then "BUG: liba.rmeta rewritten in place" and "BUG: old hard link sees new bytes". strace shows openat("st/liba.rmeta", O_WRONLY|O_CREAT|O_TRUNC) with no rename. nightly-2017-11-20 (5041b3bb3 2017-11-19): the inode changes, then "OK: liba.rmeta replaced by rename" and "OK: old hard link still holds previous metadata". strace shows the write going to st/rmeta.XXXX/rust.metadata.bin, followed by rename(... , "st/liba.rmeta"). race.sh on 2017-11-17 gave 1 failure in 1078 reader runs in one 60s attempt and 0 in another 90s attempt. The failure was the exact ICE from the issue: "thread 'rustc' panicked at 'index out of bounds: the len is 2432992 but the index is 6805837', /checkout/src/libserialize/leb128.rs:59:20". On 2017-11-20: 0 failures in 1719 runs. + +**#68149.** Under cargo pipelining, a consumer starts as soon as the dependency's .rmeta is published, while the dependency's rustc is still producing the .rlib. When the locator resolved a transitive dependency by searching -L dependency=, it opened and remembered the .rlib if it was present, even though an rlib-only build needs just the .rmeta. ehuss traced the timeline: T1 libcore.rmeta emitted; T2 the backtrace build starts; T3 libcore.rlib emitted; T4 backtrace loads libcore.rlib; T5 its dep-info lists both. Cargo backdates the dep-info to T2, so the next build sees rlib mtime T3 > T2 and spuriously rebuilds std crates (reported by petrochenkov with -Zbinary-dep-depinfo in x.py). This is a P2-adjacent violation: a pipelined consumer reaches into a dependency artifact (.rlib) that is not yet finished or published from its point of view. PR #68298 ('Avoid declaring a fake dependency edge') stopped storing rlib/dylib paths when only producing an rlib. Caveat: the file opened is the .rlib, not the .rmeta. + +*This run:* before: the consumer opens and records the dependency's in-flight .rlib. After: only the .rmeta files. Outputs: [`before`](motivating/out/68149/before.txt), [`after`](motivating/out/68149/after.txt). + +*Expected, from the issue and the research:* Run on this VM. nightly-2020-01-22 (rustc 5e8897b7b 2020-01-21, which does not contain merge be663bf85): c's dep-info lists D/liba-x.rlib as well as liba-x.rmeta and libb-x.rmeta, and strace shows c opening D/liba-x.rlib. nightly-2020-01-24 (rustc 41f41b235 2020-01-23, which contains it): only liba-x.rmeta and libb-x.rmeta are listed and opened. This deterministic check shows the cause: the consumer opening the in-flight rlib. The user-visible symptom, a spurious cargo rebuild, needs the rlib to appear in the small window between the consumer starting and loading crates (ehuss: 'the time between T2 and T4 is extremely small'), plus -Zbinary-dep-depinfo, so the cargo-level symptom is not cheap to reproduce. + +## P3: A dependent reads only table entries the dependency wrote + +An entry the writer never wrote reads as a default, without complaint, so a table that stops being written shows up as wrong behaviour far away: a missing bound, a missing attribute, or an ICE in the reader. The `written` column of the blessed list records it. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#122859](https://github.com/rust-lang/rust/issues/122859) Implied bound not implied across crates: associated-type bounds in supertrait po | [#122891](https://github.com/rust-lang/rust/pull/122891) | 2024-03-24 | `nightly-2024-03-24` → `nightly-2024-03-26` | reproduced | caught when the fix (#122891) is reverted: the list shows the table no longer written | +| [#130201](https://github.com/rust-lang/rust/issues/130201) ICE: `coroutine_by_move_body_def_id` unsupported by its crate when calling a for | [#130201](https://github.com/rust-lang/rust/pull/130201) | 2024-09-17 | `nightly-2024-09-17` → `nightly-2024-09-19` | reproduced | caught when the fix is reverted: the build ICEs | +| [#144004](https://github.com/rust-lang/rust/issues/144004) rustdoc drops #[no_mangle] / #[link_section] from inlined cross-crate re-exports | [#144050](https://github.com/rust-lang/rust/pull/144050) | 2025-07-19 | `nightly-2025-07-06` → `nightly-2025-07-24` | reproduced | missed when the fix (#144050) is reverted: only rustdoc reads these attributes | + +**#122859.** For traits, the metadata encoder assumed `implied_predicates_of` was equal to `super_predicates_of` and only wrote the latter. So the implied predicates that come from associated type bounds (`trait Bar: Super`) were never written. The dependent crate read a smaller predicate list without noticing, and lost the bound. That shows up as a wrong E0277 that only happens across crates. The PR says: 'The assumption that they didn't differ was hard-coded in #107614, so in cross-crate positions this means that we forget the implied predicates from associated type bounds.' The same code compiles when crate_b is a local module. + +*This run:* before: E0277, the implied bound lost across crates. After: compiles. Outputs: [`before`](motivating/out/122859/before.txt), [`after`](motivating/out/122859/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2024-03-24, main.rs fails with `error[E0277]: the trait bound `<::FooAssoc as Super>::SuperAssoc: Unsatisfied` is not satisfied`. On nightly-2024-03-26 it compiles (there is only a dead-code warning). Pasting crate_b's contents into main.rs as `mod crate_b` compiles on both, which shows that only the cross-crate (metadata) path is affected. + +**#130201.** The dependency crate never wrote the `coroutine_by_move_body_def_id` table entry, and never wrote optimized_mir for the synthetic by-move body. When the dependent crate asked for that entry, the lookup found nothing and fell through to a missing provider, which ICEs. The PR text says: 'We weren't encoding this query in the metadata though, nor were we properly recording that synthetic MIR in `mir_keys`, so the `optimized_mir` wasn't getting encoded either!' There is no separate issue; PR #130201 itself is the reference, and its regression test is tests/ui/async-await/async-closures/foreign.rs. + +*This run:* before: the dependent ICEs (`coroutine_by_move_body_def_id` unsupported by its crate). After: compiles. Outputs: [`before`](motivating/out/130201/before.txt), [`after`](motivating/out/130201/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2024-09-17, compiling main.rs ICEs with: `error: internal compiler error: compiler/rustc_middle/src/query/plumbing.rs:664:5: `tcx.coroutine_by_move_body_def_id(DefId(20:6 ~ foreign[87b3]::closure::{closure#0}::{closure#0}))` unsupported by its crate; perhaps the `coroutine_by_move_body_def_id` query was never assigned a provider function`. On nightly-2024-09-19 it compiles cleanly and produces the `main` binary. + +**#144004.** `no_mangle` and `link_section` were moved out of the generic encoded attribute list, and nothing else wrote them to the attribute table. A dependent crate that reads the dependency's attributes from metadata (here rustdoc inlining a `pub use a::*` re-export) silently gets no such attribute, with no error. The PR title is 'Fix encoding of link_section and no_mangle cross crate', and it fixes it by always encoding them. This is the 'missing attributes cross-crate' kind of P3 violation. The result is silent information loss, not an ICE. + +*This run:* before: rustdoc shows neither attribute on the re-exports. After: `no_mangle` and `link_section` shown. Outputs: [`before`](motivating/out/144004/before.txt), [`after`](motivating/out/144004/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2025-07-06 all four pages (doc/b/fn.f0.html, fn.f1.html, static.S0.html, static.S1.html) show no attribute, so every grep prints nothing. On nightly-2025-07-24 they show `no_mangle`, `link_section = ".here"`, `no_mangle` and `link_section = ".there"`. The issue adds that 1.88.0 stable also lacked both attributes, and 1.89 beta showed no_mangle but not link_section. The regression came and went with attribute-parsing refactors. + +## P4: Encoding reads no untracked state + +Environment variables, the clock, randomly seeded maps and files that are not declared inputs all reach outputs without incremental compilation or Cargo knowing. P4 records such reads while encoding. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#40364](https://github.com/rust-lang/rust/issues/40364) env!/option_env! values were baked into the output, but the env vars were not re | [#71858](https://github.com/rust-lang/rust/pull/71858) | 2020-06-26 | `1.45.0` → `1.46.0` | reproduced | — | +| [#66955](https://github.com/rust-lang/rust/issues/66955) --remap-path-prefix was UNTRACKED: changing it under incremental reused stale co | [#84233](https://github.com/rust-lang/rust/pull/84233) | 2021-04-29 | `nightly-2021-04-29` → `nightly-2021-05-01` | reproduced | — | +| [#111227](https://github.com/rust-lang/rust/issues/111227) debugger_visualizer files | [#111641](https://github.com/rust-lang/rust/pull/111641) | 2023-05-19 | `nightly-2023-05-18` → `nightly-2023-05-21` | partly reproduced | — | +| [#138678](https://github.com/rust-lang/rust/issues/138678) Randomly seeded HashMap | [#138678](https://github.com/rust-lang/rust/pull/138678) | 2025-03-28 | `nightly-2025-03-28` → `nightly-2025-03-30` | reproduced | caught when the fix is reverted: P5, P5 with threads, the touch rebuild and P6 | + +**#40364.** P4: env!() reads the process environment while compiling, and the value ends up in the output. rustc did not record that the variable was read, so build tools treated the crate as fresh after the variable changed. PR #71858 added '# env-dep:KEY=VALUE' lines to dep-info, closing #40364, #44074 and #70517. Cargo then read those lines and rebuilt on a change (rust-lang/cargo PR #8421, merged 2020-06-30). Both parts are needed for the repro, so it uses stable releases with the matching cargo: 1.45.0 has neither part, 1.46.0 has both. rustc nightlies after 2020-06-26 have the dep-info lines; cargo's rebuild arrived when the cargo submodule was next updated in early July 2020. + +*This run:* before: changing the variable leaves the binary printing the old value, and dep-info has no env-dep. After: the new value, and `# env-dep:MIRTH_DEMO=two`. Outputs: [`before`](motivating/out/40364/before.txt), [`after`](motivating/out/40364/after.txt). + +*Expected, from the issue and the research:* Verified locally. With 1.45.0 the two runs print 'one' then 'one': the stale binary is reused after MIRTH_DEMO changed. With 1.46.0 they print 'one' then 'two'. On 1.46.0 the dep-info file also has a '# env-dep:MIRTH_DEMO=two' line; 1.45.0 has none. + +**#66955.** P4, read broadly as untracked state that reaches the output: #48162 made --remap-path-prefix UNTRACKED so that the crate hash stayed the same. Remapped paths still went into the outputs, so after the flag changed, an incremental rebuild reused cached codegen units holding the old prefix. The result was a mix of old and new remapped paths. The fix (#84233) added TRACKED_NO_CRATE_HASH: the option invalidates the incremental cache but stays out of the crate hash. Note: in this era metadata was always re-encoded, so the rmeta itself picked up the new path. The stale data is in the object code and debuginfo inside the rlib. + +*This run:* before: after changing the remap, the rlib still holds 2 copies of the old path. After: only the new path. Outputs: [`before`](motivating/out/66955/before.txt), [`after`](motivating/out/66955/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2021-04-29, after the second build (remap to /BBBB_second) the rlib still contains '2 /AAAA_first' and only '1 /BBBB_second'; the stale object code and debuginfo were reused. On nightly-2021-05-01 it contains '3 /BBBB_second' and no /AAAA_first. + +**#111227.** P4: the contents of files named by #![debugger_visualizer(natvis_file / gdb_script_file)] were read and encoded into crate metadata (the debugger_visualizers query), but those files were not tracked inputs. They were missing from dep-info (#111226), so cargo never rebuilt after a change. The incremental system also did not see them: changing the natvis file gave 'internal compiler error: encountered incremental compilation error with debugger_visualizers' (#111227), and changed GDB scripts were not picked up (#111295). The fix (#111641) made debugger_visualizers an eval_always query computed from the AST and added the files to dep-info. Its tests include run-make/incremental-debugger-visualizer, which greps the rmeta for the file contents. + +*Status:* partly reproduced: `foo.py` is missing from the dep-info before the fix and present after it, but the stale metadata and ICE the issue describes did not appear here. + +*This run:* before: the visualizer file is missing from dep-info. After: listed. The stale metadata did not appear here. Outputs: [`before`](motivating/out/111227/before.txt), [`after`](motivating/out/111227/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2023-05-18 the script prints 'foo.py NOT in dep-info'. The second, incremental compile prints 'error: internal compiler error: encountered incremental compilation error with debugger_visualizers(foo[47df])', and libfoo.rmeta still contains 'Natvis v1', so the metadata is stale. On nightly-2023-05-21 foo.py is listed in foo.d, the rebuild succeeds, and the rmeta contains 'Natvis v2'. + +**#138678.** P4: metadata encoding read a randomly seeded std HashMap. rustc_resolve::rustdoc::parse_links walks pulldown-cmark's reference_definitions(), a HashMap with RandomState, and pushes the links in that iteration order. The resulting list of doc-link candidates is encoded into crate metadata, so two runs of rustc on the same input wrote different .rmeta bytes. The bug came in with #136363 (merged 2025-02-16). The fix (#138678) sorts the links by label. #138678 is the PR; it has no separate issue, and its description says the nondeterminism was found in Bazel lib.rmeta outputs. + +*This run:* before: 8 builds, 8 different .rmeta files. After: 8 identical. Outputs: [`before`](motivating/out/138678/before.txt), [`after`](motivating/out/138678/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2025-03-28 (rustc 1.87.0-nightly 3f5502370 2025-03-27), the 8 identical compilations gave 8 different liblib.rmeta hashes, each counted once. On nightly-2025-03-30 (1.88.0-nightly 1799887bb 2025-03-29), all 8 gave the same hash (count 8). + +## P5: Two clean builds give the same bytes + +Nondeterminism in metadata breaks reproducible builds and makes crate hashes, and so everything downstream, depend on chance. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#34902](https://github.com/rust-lang/rust/issues/34902) Metadata xrefs encoded in pointer-address | [#35984](https://github.com/rust-lang/rust/pull/35984) | 2016-08-28 | `nightly-2016-08-27` → `nightly-2016-08-30` | reproduced | — | +| [#65036](https://github.com/rust-lang/rust/issues/65036) Module re-exports serialized into metadata in FxHashMap order keyed by | [#65043](https://github.com/rust-lang/rust/pull/65043) | 2019-10-06 | `nightly-2019-10-05` → `nightly-2019-10-08` | reproduced | — | +| [#159677](https://github.com/rust-lang/rust/issues/159677) .rmeta contents depend on unrelated files in the library search path | [#159718](https://github.com/rust-lang/rust/pull/159718) | 2026-07-24 | `nightly-2026-07-20` → `nightly-2026-09-25` | reproduced | not replayed (needs a decoy crate in the search path) | +| [#129094](https://github.com/rust-lang/rust/issues/129094) Parallel frontend: derives make metadata irreproducible | [#161450](https://github.com/rust-lang/rust/pull/161450) | 2026-09-16 | `nightly-2026-07-20` → `nightly-2026-09-25` | reproduced | not replayed (the revert does not apply cleanly) | + +**#34902.** encode_xrefs iterated an FnvHashMap, u32>. Its keys are interned ty::Predicate pointers, so the hash and the iteration order depend on heap addresses, which ASLR changes on every run. The commit 'Make metadata encoding deterministic' in PR #35984 ('Steps towards reproducible builds', cc tracking issue #34902) sorts the xrefs by their ID before encoding. This is the classic single-threaded pointer-address nondeterminism: the same command run twice in the same directory gives different rust.metadata.bin bytes, while the object file stays identical. + +*This run:* before: 5 builds, 5 different rlibs (only the metadata member differs); identical with ASLR off. After: 5 identical. Outputs: [`before`](motivating/out/34902/before.txt), [`after`](motivating/out/34902/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2016-08-27 (rustc 1.13.0-nightly 198713106) 10 runs gave 10 different liblib.rlib hashes. Only the rust.metadata.bin member differs; lib.0.o is identical. Under 'setarch -R' (ASLR off) the runs are identical, which confirms that pointer addresses are the cause. On nightly-2016-08-30 (77d2cd28f) all 10 runs give the same hash. Very old toolchain: install it with 'rustup toolchain install nightly-2016-08-27 --profile minimal'. + +**#65036.** The PR text says: 're-exports end up getting serialized into crate metadata, which means that metadata generation was non-deterministic'. A module's resolutions were an FxHashMap<(Ident, Namespace), ...>, and Ident hashes by Symbol interner index. Anything that changes interning order changes the order of re-exports in .rmeta. The fix changes Resolutions to an FxIndexMap. The reported symptom (#65036) was flaky diagnostics, std::mem::transmute vs std::intrinsics::transmute. The reproduction below uses the same unrelated-file-in-search-path trigger as #159677. A proc macro interns the item names after the extern crate lookup has interned 'foo_bar', so the indices shift. + +*This run:* before: two builds differ at byte 5163, re-exports in a different order. After: identical. Outputs: [`before`](motivating/out/65036/before.txt), [`after`](motivating/out/65036/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2019-10-05 (rustc 1.40.0-nightly 2e7244807 2019-10-04) cmp reports 'differ: byte 5157', and the strings diff shows the item_N re-export names in a different order (item_21, item_1, item_24, item_11, ...). On nightly-2019-10-08 (f3c9cece7 2019-10-07) the script prints IDENTICAL. + +**#159677.** Building the same client.rs with the same flags twice gives different .rmeta bytes if an unrelated rlib whose name starts with the dependency's name (libfoo_bar.rlib next to libfoo.rlib) is in the -L directory. The crate locator opens libfoo_bar.rlib to read its crate name, which interns the extra Symbol `foo_bar`. That shifts the interner indices of symbols interned later, such as the doc-link strings. DocLinkResMap was an UnordMap (FxHashMap) keyed by (Symbol, Namespace) and hashed by interner index, and it was encoded in hash-iteration order. The fix makes it an FxIndexMap, so entries are encoded in insertion order. The PR adds the regression test tests/run-make/rmeta-unrelated-search-path-files. + +*This run:* before: adding an unrelated library to the search path changes the .rmeta (differs at byte 1256). After: identical. Outputs: [`before`](motivating/out/159677/before.txt), [`after`](motivating/out/159677/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2026-07-20 (and on stable 1.97.1) cmp prints 'client1.rmeta client2.rmeta differ: byte 1256' (1263 on 1.97.1); the differing bytes are the reordered doc-link entry 'crate::Client'. On nightly-2026-09-25 the files are identical and the script prints IDENTICAL. The fix merged 2026-07-24T09:12Z, so nightly-2026-07-26 and later should be fixed. + +**#129094.** With -Zthreads=N, the order in which SyntaxContexts and expansions are reached during metadata encoding depends on thread scheduling, so a tiny derive-heavy crate compiled repeatedly with the same inputs gives different rlibs. The fix ('Fix non-deterministic encoding of syntax contexts', which reiterates #157409) adds deterministic encoding indices in rustc_metadata/rmeta/encoder.rs and rustc_span/hygiene.rs. It also adds derives-issue-129094.rs to tests/run-make/parallel-reproducible-build. Unlike the other entries this is not single-threaded: it needs the parallel frontend. + +*This run:* before: 20 threaded builds give 4 different rlibs. After: 20 identical. Outputs: [`before`](motivating/out/129094/before.txt), [`after`](motivating/out/129094/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2026-07-20, 20 runs gave 3 distinct rlib hashes (15/4/1). On nightly-2026-09-25 all 20 runs give one hash. Race-dependent: it needs -Zthreads>1 and several runs, and how often it shows up depends on core count and scheduling. It reproduced readily on this VM. + +## P5t: Two clean builds with `-Zthreads` give the same bytes + +The parallel front end adds a new source of nondeterminism: the order in which threads create and intern things. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#140413](https://github.com/rust-lang/rust/issues/140413) parallel rustc: static mut refs not reproducible | [#144722](https://github.com/rust-lang/rust/pull/144722) | 2025-08-13 | `nightly-2025-07-26` → `nightly-2025-12-10` | reproduced | — | +| [#150451](https://github.com/rust-lang/rust/issues/150451) parallel compiler: thread::spawn-ing loop not reproducible | [#160197](https://github.com/rust-lang/rust/pull/160197) | 2026-09-07 | `nightly-2026-07-20` → `nightly-2026-09-25` | reproduced | — | +| [#129094](https://github.com/rust-lang/rust/issues/129094) Parallel frontend: derives make metadata irreproducible | [#161450](https://github.com/rust-lang/rust/pull/161450) | 2026-09-16 | `nightly-2026-07-20` → `nightly-2026-09-25` | reproduced | not replayed (the revert does not apply cleanly) | + +**#140413.** With -Zthreads=50, building the same binary repeatedly gives different bytes. Mono items were sorted by (DefId, SymbolName) to pick their order in the output. DefId indices are allocated in access order, which is nondeterministic under the parallel front end. PR #144722 ('Fix parallel rustc not being reproducible due to unstable sorts of items') stops sorting by DefId. Two later facts confirm the fix: PR #161353 (merged 2026-09-01; it closed the issue and added tests/run-make/parallel-reproducible-build) found by bisection that the regression flips at nightly-2025-08-14, at commit #144722. The same PR also fixes the async-closure case #140425 (closed 2025-08-13). + +*This run:* before: 15 threaded builds give 2 different binaries. After: 15 identical. Outputs: [`before`](motivating/out/140413/before.txt), [`after`](motivating/out/140413/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2025-07-26 there were 2 distinct md5s of the binary over 15 runs. On nightly-2025-12-10 there was 1. The bisected boundary is nightly-2025-08-13 (bad) to nightly-2025-08-14 (good). Caveat: building this file as --crate-type=lib --emit=metadata is still nondeterministic, even on nightly-2026-10-06 (7-8 distinct rmeta md5s out of 10). That case is the still-open follow-up #162203, so use this repro for the binary only. + +**#150451.** With -Zthreads=3, compiling a tiny lib that uses thread::spawn gives a different .rlib each time. The output differs when bitcode is embedded or LTO is used. The cause is that the srcloc 'cookies' attached to LLVM inline asm were assigned nondeterministically by the parallel front end. PR #160197 ('Restrict LLVM inline asm location cookie usage. Fixes #150451') landed in rollup #162434 on 2026-09-07 and added tests/run-make/parallel-reproducible-inline-asm-cookie. The bug is in object code and bitcode, not in rmeta, but it breaks reproducibility of the published artifact. + +*This run:* before: 15 threaded builds give 10 different rlibs. After: 15 identical. Outputs: [`before`](motivating/out/150451/before.txt), [`after`](motivating/out/150451/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2026-07-20 there were 7 distinct rlib md5s over 15 runs. On nightly-2026-09-25 there was 1. + +**#129094.** With -Zthreads=N, the order in which SyntaxContexts and expansions are reached during metadata encoding depends on thread scheduling, so a tiny derive-heavy crate compiled repeatedly with the same inputs gives different rlibs. The fix ('Fix non-deterministic encoding of syntax contexts', which reiterates #157409) adds deterministic encoding indices in rustc_metadata/rmeta/encoder.rs and rustc_span/hygiene.rs. It also adds derives-issue-129094.rs to tests/run-make/parallel-reproducible-build. Unlike the other entries this is not single-threaded: it needs the parallel frontend. + +*This run:* before: 20 threaded builds give 4 different rlibs. After: 20 identical. Outputs: [`before`](motivating/out/129094/before.txt), [`after`](motivating/out/129094/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2026-07-20, 20 runs gave 3 distinct rlib hashes (15/4/1). On nightly-2026-09-25 all 20 runs give one hash. Race-dependent: it needs -Zthreads>1 and several runs, and how often it shows up depends on core count and scheduling. It reproduced readily on this VM. + +## P6: An incremental rebuild gives the same results as a clean build + +Incremental compilation is only correct if what it reuses is what it would have computed. When it is not, the result is a stale output, a miscompilation or a wrong diagnostic, which `cargo clean` fixes. These bugs are why P6 compares an incremental rebuild with a clean build of the same source. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#82920](https://github.com/rust-lang/rust/issues/82920) Miscompilation with incr. comp. | [#83074](https://github.com/rust-lang/rust/pull/83074) | 2021-03-15 | `nightly-2021-03-13` → `nightly-2021-03-17` | reproduced | — | +| [#89598](https://github.com/rust-lang/rust/issues/89598) VTable-related miscompilation with incremental compilation | [#89619](https://github.com/rust-lang/rust/pull/89619) | 2021-10-08 | `nightly-2021-10-07` → `nightly-2021-10-10` | reproduced | — | +| [#135514](https://github.com/rust-lang/rust/issues/135514) Rust 1.84 sometimes allows overlapping impls in incremental re-builds | [#133828](https://github.com/rust-lang/rust/pull/133828) | 2024-12-05 | `1.84.0` → `1.85.0` | reproduced | — | +| [#139407](https://github.com/rust-lang/rust/issues/139407) Instructions missing from | [#139453](https://github.com/rust-lang/rust/pull/139453) | 2025-04-11 | `nightly-2025-04-10` → `nightly-2025-04-13` | reproduced | — | +| [#162901](https://github.com/rust-lang/rust/issues/162901) Diagnostic deduplication breaks with incr comp: incremental builds print 4 copie | [#163461](https://github.com/rust-lang/rust/pull/163461) | 2026-10-02 | `nightly-2026-10-01` → `nightly-2026-10-03` | reproduced | — | + +**#82920.** Bounds and predicates were sorted by DefId, which is not stable across sessions. When two trait declarations swap places, the query result changes but is still treated as green and reused, so the vtable layout and the method call sites disagree. The incremental binary calls the wrong trait method, while a clean build is correct. The original report was a rust-analyzer test suite failing with out-of-bounds panics and segfaults until `cargo clean`. Labelled I-unsound. + +*This run:* before: the incremental pass-2 binary panics (`left: 2, right: 1`) while a clean build of the same source passes. After: both pass. Outputs: [`before`](motivating/out/82920/before.txt), [`after`](motivating/out/82920/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2021-03-13, the incremental pass 2 binary panics with "assertion failed: `(left == right)` left: `2`, right: `1`" (method_two was called for method_one), while a clean rpass2 build prints 'clean ok'. On nightly-2021-03-17 both passes print ok. Regression test: tests/incremental/issue-82920-predicate-order-miscompile.rs. + +**#89598.** #86475 added an untracked global vtable cache to tcx. After trait methods are reordered, the object file for the main CGU is reused from the incremental cache with the old vtable layout, while mod1 is recompiled for the new layout. The incremental binary then calls method2 where a clean build calls method1, so the binaries differ (a miscompile). P-critical, I-unsound, regression-from-stable-to-beta. + +*This run:* before: the incremental pass-2 binary calls the wrong method (`left: 42, right: 17`); the clean build passes. After: both pass. Outputs: [`before`](motivating/out/89598/before.txt), [`after`](motivating/out/89598/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2021-10-07, pass 1 prints 'pass1 ok'; the incremental pass 2 binary panics with "assertion failed: `(left == right)` left: `42`, right: `17`" because method2 was called through the stale vtable. The clean build of rpass2 runs fine. On nightly-2021-10-10 both passes print ok. Regression test: tests/incremental/reorder_vtable.rs. + +**#135514.** A clean build of the edited source fails with E0119 (conflicting implementations). The incremental rebuild after the same edit accepts it, because coherence results computed through the new solver's cache had no dependency edges. The accepted program is a safe transmute: Vec becomes String and prints ABC. The diagnostics and the success or failure differ from a clean build. Labelled I-unsound, regression-from-stable-to-stable. #135522 added tests/incremental/overlapping-impls-in-new-solver-issue-135514.rs. The fix was already on master before the issue was filed, so 1.85 is fixed and 1.84.x is affected. + +*This run:* before: the incremental rebuild accepts overlapping impls and runs, while a clean build gives E0119. After: both give E0119. Outputs: [`before`](motivating/out/135514/before.txt), [`after`](motivating/out/135514/after.txt). + +*Expected, from the issue and the research:* Verified locally. On 1.84.0, pass 1 prints 'pass1'; the incremental pass 2 compiles without error and prints 'ABC' then 'incremental pass2 built and ran'; the clean build of the same source fails with error[E0119]: conflicting implementations of trait `Other` for type `S`. On 1.85.0 the incremental pass 2 also fails with E0119. On 1.83.0 the old solver rejects pass 1 itself, so use 1.84.0. For nightlies, try nightly-2024-12-04 (bug) and nightly-2024-12-07 (fixed); these were not run. + +**#139407.** Object files were hard-linked from fixed temp paths into the incremental session directory. A session that fails in the assembler (left unfinalized) overwrites a temp file that a previous, finalized session still hard-links. On the next successful build the reused object holds code from the failed session's source, so the binary differs from a clean build. The fix gives temp files a per-invocation random prefix. Labelled I-unsound. The run-make regression test is tests/run-make/dirty-incr-due-to-hard-link. The reporter hit it with plain `cargo run` (no -Csave-temps) on aarch64-apple-darwin. The test, used here, needs -Csave-temps to keep the temp files on x86_64 Linux. It needs a failed build in between (an asm error), but it is deterministic. + +*This run:* before: after a failed build is fixed back, the rebuilt binary still contains the failed session's code and panics. After: it passes. Outputs: [`before`](motivating/out/139407/before.txt), [`after`](motivating/out/139407/after.txt). + +*Expected, from the issue and the research:* Verified locally on x86_64 Linux. On nightly-2025-04-10 (and on stable 1.70.0, 1.86.0 and 1.87.0), pass 1 prints 'pass1 ok' and pass 2 fails with "error: invalid instruction mnemonic 'missing'". Pass 3 has the same source as pass 1, yet its binary panics with "assertion `left == right` failed left: 1 right: 0": a() from the failed cfail2 session leaked into the reused object. On nightly-2025-04-13 and on 1.88.0, pass 3 prints 'pass3 ok'. + +**#162901.** In incremental mode, spans carry a parent (incremental-relative spans). The diagnostic dedup hash included Span::parent, so identical diagnostics hashed differently and were all emitted. The diagnostics from an incremental compile differ from a non-incremental compile of the same source. The fix PR also fixes #106571 (regression from #84762, Jan 2023), where `cargo check --message-format=json` with incremental on prints identical compiler-message lines twice for a proc-macro-generated error, and CARGO_INCREMENTAL=0 does not. No edit is needed: the difference is already there between an incremental and a non-incremental build. A P6 check that compares against a clean incremental build would not see it. It shows only when the reference build is non-incremental (CARGO_INCREMENTAL=0) or when diagnostic counts are compared. + +*This run:* before: the incremental build prints the error 4 times, a clean build once. After: once each. Outputs: [`before`](motivating/out/162901/before.txt), [`after`](motivating/out/162901/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2026-10-01 (also 1.85.0, 1.92.0, 1.96.1), the non-incremental build prints 1 E0277 error and the incremental build prints 4 identical copies ('aborting due to 4 previous errors'). On nightly-2026-10-03 both print 1. Stable 1.70.0 prints 1 in both modes, so this form of the bug arrived later; #106571's proc-macro form dates from nightly-2023-01-03. Regression test: tests/ui/diagnostic-flags/deduplicate-diagnostics-incr.rs. + +## P7: Nothing is left behind in the output directory + +Leftover temporary files waste space and can be picked up by later builds. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#107001](https://github.com/rust-lang/rust/issues/107001) rustc leaves *.rcgu.o object files in the output directory when a post-monomorph | [#110107](https://github.com/rust-lang/rust/pull/110107) | 2023-04-21 | `nightly-2023-04-21` → `nightly-2023-04-22` | reproduced | — | +| [#139899](https://github.com/rust-lang/rust/issues/139899) rustdoc --test leaves rustdoctest* temporary directories behind whenever a docte | [#140706](https://github.com/rust-lang/rust/pull/140706) | 2025-05-08 | `nightly-2025-05-07` → `nightly-2025-05-09` | reproduced | — | + +**#107001.** The const-prop lints (unconditional_panic, arithmetic_overflow) ran in mir_drops_elaborated_and_const_checked, which was only forced lazily, so their errors could fire after codegen had already written the CGU object files. rustc then aborted without deleting them, leaving `..rcgu.o` / `..-cgu.N.rcgu.o` in the output directory (target/debug/deps under cargo). PR #110107 ('Ensure mir_drops_elaborated_and_const_checked when requiring codegen') makes sure that query runs before codegen. Its description says: 'may emit errors while codegen has started, and the compiler would exit leaving object code files around. Found by @cuviper in #109731' (cuviper's comment there: 'each time I try one of these failing tests, it's leaving temporary *.rcgu.o files around'). Issue #107001 is the standalone report with this exact repro. It is still marked open, but its repro stops leaking at this PR. I bisected the nightlies myself and the boundary is exactly this merge: nightly-2023-04-21 (8bdcc62cb) leaks and nightly-2023-04-22 (fec9adcdb) is clean. The commit range between them contains #110107. + +*This run:* before: a failed compile leaves `.rcgu.o` files in the output directory. After: none. Outputs: [`before`](motivating/out/107001/before.txt), [`after`](motivating/out/107001/after.txt). + +*Expected, from the issue and the research:* Both toolchains fail the same way: 'error: this operation will panic at runtime ... index out of bounds: the length is 5 but the index is 9' with `#[deny(unconditional_panic)]`, exit=1. On nightly-2023-04-21 (also stable 1.70.0), `ls -A out` lists 6 leftover object files, e.g. `code.1tgaf0fuackrygys.rcgu.o code.code.56d798bc-cgu.0.rcgu.o ...`. On nightly-2023-04-22 (also stable 1.71.0 and everything since, through nightly-2026-10-06), `out` is empty. Verified locally. + +**#139899.** When any doctest failed, rustdoc (or libtest) called process::exit, so the TempDir destructor never ran and the `rustdoctestXXXXXX` directory was never removed. Issue #139899 reports 6197 of them piling up in /tmp. PR #140706 ('[rustdoc] Ensure that temporary doctest folder is correctly removed even if doctests failed') adds a libtest hook that runs after all tests and cleans the folder up. It also adds the regression test tests/run-make/rustdoc/doctest/tempdir-removal, whose two input files are used verbatim below. The leftovers go to TMPDIR, not to target/, so this hits P7 only when TMPDIR points inside the build tree. It is still a real, fixed 'temp dir left behind on the failure path' bug in the toolchain that `cargo test --doc` runs. Note: this is rustdoc, not rustc. + +*This run:* before: each failing doctest leaves a `rustdoctest*` directory (4 left). After: none. Outputs: [`before`](motivating/out/139899/before.txt), [`after`](motivating/out/139899/after.txt). + +*Expected, from the issue and the research:* Every run exits 101 (failed doctest) on both toolchains. On nightly-2025-05-07, `ls -A tmp` shows one leftover directory per run (4 in total), e.g. `rustdoctestLP6B1A rustdoctesteGUqc8 rustdoctest0kc32h rustdoctestqVdzVs`. On nightly-2025-05-09, `tmp` is empty. Verified locally for both editions (2018 per-test and 2024 merged doctests). + +## tracked: Cross-crate reads are tracked, and metadata is reused when nothing changed + +Every query that reads another crate's metadata must record a dependency on that crate, or a later session reuses a stale result. The `tracked` column of the blessed list records it per query; the touch-only rebuild records whether metadata was reused. + +| issue | fixed by | merged | before → after | status | mirth, with the fix reverted | +|---|---|---|---|---|---| +| [#82920](https://github.com/rust-lang/rust/issues/82920) Miscompilation with incr. comp. | [#83074](https://github.com/rust-lang/rust/pull/83074) | 2021-03-15 | `nightly-2021-03-13` → `nightly-2021-03-17` | reproduced | — | +| [#84252](https://github.com/rust-lang/rust/issues/84252) ICE: found unstable fingerprints for has_global_allocator(): the query read untr | [#84260](https://github.com/rust-lang/rust/pull/84260) | 2021-04-17 | `nightly-2021-04-16` → `nightly-2021-04-19` | reproduced | — | +| [#89598](https://github.com/rust-lang/rust/issues/89598) VTable-related miscompilation with incremental compilation | [#89619](https://github.com/rust-lang/rust/pull/89619) | 2021-10-08 | `nightly-2021-10-07` → `nightly-2021-10-10` | reproduced | — | +| [#111295](https://github.com/rust-lang/rust/issues/111295) debugger_visualizer: edits to the visualizer script are not picked up under incr | [#111641](https://github.com/rust-lang/rust/pull/111641) | 2023-05-19 | `nightly-2023-05-18` → `nightly-2023-05-21` | reproduced | — | +| [#114669](https://github.com/rust-lang/rust/issues/114669) Metadata was never reused across incremental sessions: always re-encoded | [#114669](https://github.com/rust-lang/rust/pull/114669) | 2025-07-04 | `nightly-2025-07-03` → `nightly-2025-07-06` | reproduced | with #143247 reverted (metadata depending on a node that is never green), the touch rebuild's record catches it | + +**#82920.** Bounds and predicates were sorted by DefId, which is not stable across sessions. When two trait declarations swap places, the query result changes but is still treated as green and reused, so the vtable layout and the method call sites disagree. The incremental binary calls the wrong trait method, while a clean build is correct. The original report was a rust-analyzer test suite failing with out-of-bounds panics and segfaults until `cargo clean`. Labelled I-unsound. + +*This run:* before: the incremental pass-2 binary panics (`left: 2, right: 1`) while a clean build of the same source passes. After: both pass. Outputs: [`before`](motivating/out/82920/before.txt), [`after`](motivating/out/82920/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2021-03-13, the incremental pass 2 binary panics with "assertion failed: `(left == right)` left: `2`, right: `1`" (method_two was called for method_one), while a clean rpass2 build prints 'clean ok'. On nightly-2021-03-17 both passes print ok. Regression test: tests/incremental/issue-82920-predicate-order-miscompile.rs. + +**#84252.** has_global_allocator was answered from the CStore (the crate loader's per-crate metadata state) without recording any dependency. PR #84260: 'This query reads from untracked global state in `CStore`'. It was fixed by making it eval_always. The repro is the regression test tests/incremental/issue-84252-global-alloc.rs: removing #[global_allocator] between sessions leaves the cached result stale, and recomputing it gives a different value. The sibling fix #83153 (extern_mod_stmt_cnum made eval_always for the same reason, issue #83126) is a related example. #83126 is still open on GitHub, so I left it out as a separate entry. + +*This run:* before: the second incremental build ICEs with unstable fingerprints for `has_global_allocator`. After: both builds succeed. Outputs: [`before`](motivating/out/84252/before.txt), [`after`](motivating/out/84252/after.txt). + +*Expected, from the issue and the research:* Before the fix (nightly-2021-04-16): the second build panics with "found unstable fingerprints for has_global_allocator(lib[8787]): false" (assertion left/right Fingerprint mismatch in rustc_query_system plumbing.rs) and exits non-zero. After the fix (nightly-2021-04-19): rc1=0 and rc2=0. + +**#89598.** #86475 added an untracked global vtable cache to tcx. After trait methods are reordered, the object file for the main CGU is reused from the incremental cache with the old vtable layout, while mod1 is recompiled for the new layout. The incremental binary then calls method2 where a clean build calls method1, so the binaries differ (a miscompile). P-critical, I-unsound, regression-from-stable-to-beta. + +*This run:* before: the incremental pass-2 binary calls the wrong method (`left: 42, right: 17`); the clean build passes. After: both pass. Outputs: [`before`](motivating/out/89598/before.txt), [`after`](motivating/out/89598/after.txt). + +*Expected, from the issue and the research:* Verified locally. On nightly-2021-10-07, pass 1 prints 'pass1 ok'; the incremental pass 2 binary panics with "assertion failed: `(left == right)` left: `42`, right: `17`" because method2 was called through the stale vtable. The clean build of rpass2 runs fine. On nightly-2021-10-10 both passes print ok. Regression test: tests/incremental/reorder_vtable.rs. + +**#111295.** The debugger_visualizers query result (the script contents, which are also encoded into crate metadata so downstream binaries embed upstream visualizers) did not depend on the script file. After the .py file was edited, an incremental rebuild reused the stale contents. #111641 ('Fix dependency tracking for debugger visualizers') made the query eval_always, and the fix also hashes the visualizer contents into crate_hash. The code comment says 'that content is exported into crate metadata, so any changes to it need to be reflected in the crate hash', so downstream crates that read it from metadata also see the change. The same PR fixes #111227 (an ICE from changing a natvis file with --crate-type=rlib) and #111226 (dep-info). The repro is single-crate, as in the issue. No gdb is needed: read the .debug_gdb_scripts section directly. + +*This run:* before: the rebuilt binary still embeds the old visualizer script. After: the new one. Outputs: [`before`](motivating/out/111295/before.txt), [`after`](motivating/out/111295/after.txt). + +*Expected, from the issue and the research:* Before the fix (nightly-2023-05-18): the rebuilt binary's .debug_gdb_scripts still contains print('hello!'), the stale script. After the fix (nightly-2023-05-21): it contains print('hello world'). Linux/ELF only; needs binutils objcopy. + +**#114669.** A performance violation of 'metadata is reused (not re-encoded) when nothing it depends on changed'. Before July 2025, metadata encoding depended on the forever-red DepNode (iter_local_def_id / def_path_table read DepNodeIndex::FOREVER_RED_NODE) and was not a dep-graph task, so every incremental session re-encoded the .rmeta from scratch. #143247 (merged 2025-07-04, split out of #114669 'for perf') removed the forever-red read by depending on `analysis` instead. #114669 (merged 2025-07-04) wraps encoding in a Metadata dep-node task, saves the rmeta as a work product ('metadata') in the incremental dir, and when the node is green it hardlinks or copies the saved file instead of encoding ('can yield substantial gains (~10%)... if all the changes are in upstream crates and have no effect on it'). Observable without logs: the reused rmeta is a hardlink of the work product, so its inode stays the same across no-op sessions. Verified on both toolchains. This is a numbered PR, not an issue: no separate GitHub issue exists, so the issue field holds the PR number. + +*This run:* before: an unchanged rebuild re-encodes the metadata (new inode). After: reused (same inode, hard-linked to the work product). Outputs: [`before`](motivating/out/114669/before.txt), [`after`](motivating/out/114669/after.txt). + +*Expected, from the issue and the research:* Before (nightly-2025-07-03): run1 and run2 report different inodes and links=1, so the rmeta is freshly encoded on the unchanged rebuild. After (nightly-2025-07-06): links=2 (the file is hardlinked to the 'metadata' work product in the incremental dir) and run2 has the same inode as run1, so the metadata was reused rather than re-encoded. Note: after the fix, editing even a private fn body still changed the inode in my test (the Metadata node went red), so the no-op rebuild is the clean demonstration. Needs a filesystem with hardlinks; on one without them link_or_copy falls back to copying, and the inode check does not apply. + +## Found by mirth (P6) + +| bug | reproduce | cause | since | +|---|---|---|---| +| `Generics::param_def_id_to_index` order changes on each round trip through the incremental cache | `docs/hunt/repro.sh` (p6-generics) | an `FxHashMap` encoded in iteration order | at least 1.95 | +| a string literal encoded twice after an incremental rebuild | `docs/hunt/repro.sh` (p6-literals) | literals deduplicated when created but not when decoded | 1.90 (#116707) | +| the previous session's metadata republished after an edit that moves no span | `docs/hunt/repro.sh` (stale-source) | source map file hashes and lengths not tracked | 1.90 (#114669) | + +Each reproduces on the official nightly and on stable 1.98.1, and has a draft report, a +candidate fix and a regression test in [`hunt/`](hunt). + +## Not covered here + +The 29 properties in [`properties.md`](properties.md) cite the bugs that suggested them, +checked against the issue text, but those bugs have not been reproduced this way. diff --git a/docs/motivating/bugs.json b/docs/motivating/bugs.json new file mode 100644 index 0000000..0076df0 --- /dev/null +++ b/docs/motivating/bugs.json @@ -0,0 +1,649 @@ +[ + { + "issue": 34902, + "title": "Metadata xrefs encoded in pointer-address (ASLR) order: rlib differs on every run", + "fix_pr": 35984, + "merged": "2016-08-28", + "how_it_violates": "encode_xrefs iterated an FnvHashMap, u32>. Its keys are interned ty::Predicate pointers, so the hash and the iteration order depend on heap addresses, which ASLR changes on every run. The commit 'Make metadata encoding deterministic' in PR #35984 ('Steps towards reproducible builds', cc tracking issue #34902) sorts the xrefs by their ID before encoding. This is the classic single-threaded pointer-address nondeterminism: the same command run twice in the same directory gives different rust.metadata.bin bytes, while the object file stays identical.", + "before_toolchain": "nightly-2016-08-27", + "after_toolchain": "nightly-2016-08-30", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.sh", + "content": "{\necho 'pub trait Tr {}'\nfor i in $(seq 1 30); do echo \"pub struct S$i;\"; echo \"pub fn f$i + Clone + Into>(_t: T) {}\"; done\n} > lib.rs\n" + } + ], + "commands": "sh gen.sh\nfor i in 1 2 3 4 5; do rm -rf o$i; mkdir o$i; rustc +TOOLCHAIN lib.rs --crate-type rlib --crate-name lib --out-dir o$i; done\nmd5sum o*/liblib.rlib\n(cd o1 && ar x liblib.rlib); (cd o2 && ar x liblib.rlib); md5sum o1/rust.metadata.bin o2/rust.metadata.bin o1/lib.0.o o2/lib.0.o\n# control: with ASLR disabled the old compiler is deterministic\nfor i in 1 2; do rm -rf r$i; mkdir r$i; setarch -R rustc +TOOLCHAIN lib.rs --crate-type rlib --crate-name lib --out-dir r$i; done; md5sum r*/liblib.rlib", + "observe": "Verified locally. On nightly-2016-08-27 (rustc 1.13.0-nightly 198713106) 10 runs gave 10 different liblib.rlib hashes. Only the rust.metadata.bin member differs; lib.0.o is identical. Under 'setarch -R' (ASLR off) the runs are identical, which confirms that pointer addresses are the cause. On nightly-2016-08-30 (77d2cd28f) all 10 runs give the same hash. Very old toolchain: install it with 'rustup toolchain install nightly-2016-08-27 --profile minimal'.", + "properties": [ + "P5" + ] + }, + { + "issue": 45841, + "title": "ICE: index out of bounds in libserialize/leb128.rs (rustc wrote .rmeta in place; a concurrent rustc read the half-written file)", + "fix_pr": 45899, + "merged": "2017-11-18", + "how_it_violates": "Before the fix, rustc_trans::back::link::emit_metadata did File::create(out_filename) + write_all straight into the final lib.rmeta path (O_WRONLY|O_CREAT|O_TRUNC). Another rustc that was searching the same -L directory could open the file while it was truncated or only partly written, and then panicked decoding it. The issue thread (arielb1: \"the compiler writing metadata in parts, so that another instance of the compiler can read metadata while it is being partially written to. Should be fixable by doing an atomic rename\") and PR #45899 (\"atomically write .rmeta outputs to avoid races ... write a temporary file and then rename it\") record exactly this. The fix added the rmeta* tempdir inside the output directory plus fs::rename, and the comment \"To avoid races with another rustc process scanning the output directory...\" is still in compiler/rustc_metadata/src/fs.rs today. This is the bug that P1 itself comes from.", + "before_toolchain": "nightly-2017-11-17", + "after_toolchain": "nightly-2017-11-20", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "#![crate_type = \"lib\"]\npub fn f() -> u32 { 1 }\n" + }, + { + "path": "run.sh", + "content": "#!/bin/sh\n# Deterministic check: is liba.rmeta overwritten in place (same inode) or replaced by rename?\n# usage: sh run.sh TOOLCHAIN (run in a fresh dir containing a.rs; appends to a.rs)\nset -e\nTC=$1; D=out-$TC; mkdir $D\nrustc +$TC --crate-type lib --emit=metadata a.rs --out-dir $D\nln $D/liba.rmeta $D/snapshot.rmeta\nino_before=$(stat -c %i $D/liba.rmeta)\nprintf 'pub fn g() {}\\n' >> a.rs\nrustc +$TC --crate-type lib --emit=metadata a.rs --out-dir $D\nino_after=$(stat -c %i $D/liba.rmeta)\necho \"inode before=$ino_before after=$ino_after\"\nif [ \"$ino_before\" = \"$ino_after\" ]; then echo \"BUG: liba.rmeta rewritten in place\"; else echo \"OK: liba.rmeta replaced by rename\"; fi\nif cmp -s $D/liba.rmeta $D/snapshot.rmeta; then echo \"BUG: old hard link sees new bytes\"; else echo \"OK: old hard link still holds previous metadata\"; fi\n" + }, + { + "path": "race.sh", + "content": "#!/bin/bash\n# Optional: shows the actual symptom (torn read -> ICE). A rare race (about 1 in 1000-3000 reader runs here).\n# usage: bash race.sh TOOLCHAIN [SECONDS] [READERS]\nTC=$1; T=${2:-60}; R=${3:-12}; D=race-$TC; mkdir -p $D\npython3 -c 'print(\"#![crate_type=\\\"lib\\\"]\"); [print(f\"pub fn f{i}(x: u64) -> u64 {{ x.wrapping_mul({i}) }}\\npub struct S{i} {{ pub a: u64, pub b: String }}\") for i in range(20000)]' > big.rs\necho 'extern crate big;' > user.rs\nrustc +$TC --crate-type lib --emit=metadata big.rs --out-dir $D\nend=$((SECONDS+T))\n( while [ $SECONDS -lt $end ]; do rustc +$TC --emit=metadata big.rs --out-dir $D 2>/dev/null; done ) &\nfor r in $(seq $R); do\n ( mkdir -p $D/u$r; n=0; while [ $SECONDS -lt $end ]; do n=$((n+1));\n rustc +$TC --crate-type lib --emit=metadata user.rs -L $D --out-dir $D/u$r > $D/u$r/err 2>&1 || cp $D/u$r/err $D/fail-$r-$n.err; done; echo $n > $D/u$r/count ) &\ndone\nwait\ntotal=$(cat $D/u*/count | paste -sd+ | bc); fails=$(ls $D | grep -c '^fail-')\necho \"$TC: reader runs=$total failed=$fails\"\nfor f in $(ls $D | grep '^fail-' | head -3); do echo \"--- $f\"; grep -m3 -E 'error|panicked' $D/$f; done\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\nsh run.sh TOOLCHAIN\n# syscall-level view (needs strace):\nstrace -f -e trace=openat,rename rustc +TOOLCHAIN --crate-type lib --emit=metadata a.rs --crate-name a --out-dir st 2>&1 | grep rmeta\n# optional, probabilistic symptom reproduction (run more than once):\nbash race.sh TOOLCHAIN 90", + "observe": "nightly-2017-11-17 (rustc 1.23.0-nightly d0f8e2913 2017-11-16): run.sh prints the same inode before and after the rebuild, then \"BUG: liba.rmeta rewritten in place\" and \"BUG: old hard link sees new bytes\". strace shows openat(\"st/liba.rmeta\", O_WRONLY|O_CREAT|O_TRUNC) with no rename. nightly-2017-11-20 (5041b3bb3 2017-11-19): the inode changes, then \"OK: liba.rmeta replaced by rename\" and \"OK: old hard link still holds previous metadata\". strace shows the write going to st/rmeta.XXXX/rust.metadata.bin, followed by rename(... , \"st/liba.rmeta\"). race.sh on 2017-11-17 gave 1 failure in 1078 reader runs in one 60s attempt and 0 in another 90s attempt. The failure was the exact ICE from the issue: \"thread 'rustc' panicked at 'index out of bounds: the len is 2432992 but the index is 6805837', /checkout/src/libserialize/leb128.rs:59:20\". On 2017-11-20: 0 failures in 1719 runs.", + "properties": [ + "P1", + "P2" + ] + }, + { + "issue": 65036, + "title": "Module re-exports serialized into metadata in FxHashMap order keyed by (Ident, Namespace), which is Symbol-index dependent", + "fix_pr": 65043, + "merged": "2019-10-06", + "how_it_violates": "The PR text says: 're-exports end up getting serialized into crate metadata, which means that metadata generation was non-deterministic'. A module's resolutions were an FxHashMap<(Ident, Namespace), ...>, and Ident hashes by Symbol interner index. Anything that changes interning order changes the order of re-exports in .rmeta. The fix changes Resolutions to an FxIndexMap. The reported symptom (#65036) was flaky diagnostics, std::mem::transmute vs std::intrinsics::transmute. The reproduction below uses the same unrelated-file-in-search-path trigger as #159677. A proc macro interns the item names after the extern crate lookup has interned 'foo_bar', so the indices shift.", + "before_toolchain": "nightly-2019-10-05", + "after_toolchain": "nightly-2019-10-08", + "reproducible_cheaply": true, + "files": [ + { + "path": "foo.rs", + "content": "pub struct Foo;\n" + }, + { + "path": "foo_bar.rs", + "content": "pub struct FooBar;\n" + }, + { + "path": "gen.rs", + "content": "extern crate proc_macro;\nuse proc_macro::TokenStream;\n#[proc_macro]\npub fn make(_: TokenStream) -> TokenStream {\n (1..=40).map(|i| format!(\"pub fn item_{}() {{}}\\n\", i)).collect::().parse().unwrap()\n}\n#[proc_macro]\npub fn reexports(_: TokenStream) -> TokenStream {\n (1..=40).map(|i| format!(\"pub use inner::item_{};\\n\", i)).collect::().parse().unwrap()\n}\n" + }, + { + "path": "client.rs", + "content": "extern crate foo;\nextern crate gen;\nmod inner { gen::make!(); }\ngen::reexports!();\n" + } + ], + "commands": "rustc +TOOLCHAIN gen.rs --crate-type proc-macro\nrustc +TOOLCHAIN foo.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client1.rmeta\nrustc +TOOLCHAIN foo_bar.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client2.rmeta\ncmp client1.rmeta client2.rmeta && echo IDENTICAL\ndiff <(strings -n4 client1.rmeta) <(strings -n4 client2.rmeta) | head", + "observe": "Verified locally. On nightly-2019-10-05 (rustc 1.40.0-nightly 2e7244807 2019-10-04) cmp reports 'differ: byte 5157', and the strings diff shows the item_N re-export names in a different order (item_21, item_1, item_24, item_11, ...). On nightly-2019-10-08 (f3c9cece7 2019-10-07) the script prints IDENTICAL.", + "properties": [ + "P5" + ] + }, + { + "issue": 68149, + "title": "Spurious rebuilds under pipelining: a consumer opened (and recorded in its dep-info) a dependency's .rlib that was still being produced, instead of only the .rmeta", + "fix_pr": 68298, + "merged": "2020-01-23", + "how_it_violates": "Under cargo pipelining, a consumer starts as soon as the dependency's .rmeta is published, while the dependency's rustc is still producing the .rlib. When the locator resolved a transitive dependency by searching -L dependency=, it opened and remembered the .rlib if it was present, even though an rlib-only build needs just the .rmeta. ehuss traced the timeline: T1 libcore.rmeta emitted; T2 the backtrace build starts; T3 libcore.rlib emitted; T4 backtrace loads libcore.rlib; T5 its dep-info lists both. Cargo backdates the dep-info to T2, so the next build sees rlib mtime T3 > T2 and spuriously rebuilds std crates (reported by petrochenkov with -Zbinary-dep-depinfo in x.py). This is a P2-adjacent violation: a pipelined consumer reaches into a dependency artifact (.rlib) that is not yet finished or published from its point of view. PR #68298 ('Avoid declaring a fake dependency edge') stopped storing rlib/dylib paths when only producing an rlib. Caveat: the file opened is the .rlib, not the .rmeta.", + "before_toolchain": "nightly-2020-01-22", + "after_toolchain": "nightly-2020-01-24", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "pub fn fa() -> u32 { 1 }\n" + }, + { + "path": "b.rs", + "content": "extern crate a;\npub fn fb() -> u32 { a::fa() }\n" + }, + { + "path": "c.rs", + "content": "extern crate b;\npub fn fc() -> u32 { b::fb() }\n" + }, + { + "path": "run.sh", + "content": "#!/bin/bash\n# usage: run.sh TOOLCHAIN -- shows which dependency files the consumer `c` opened / recorded\nT=$1; rm -rf D; mkdir D\nR=\"rustc +$T --crate-type lib -C extra-filename=-x -L dependency=D --out-dir D\"\n$R --emit=metadata,link a.rs\n$R --emit=metadata,link --extern a=D/liba-x.rmeta b.rs\n# c is compiled like cargo does under pipelining: only b's .rmeta is given; a is found by search\n$R -Zbinary-dep-depinfo --emit=dep-info,metadata,link --extern b=D/libb-x.rmeta c.rs\necho \"== $T: deps recorded by c:\"; tr ' ' '\\n' < D/c-x.d | grep -E 'lib[ab]-x\\.r[a-z]+$' | sort -u\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\nchmod +x run.sh\n./run.sh TOOLCHAIN\n# which dependency files c actually opens:\nstrace -f -e trace=openat -o st.txt rustc +TOOLCHAIN --crate-type lib -C extra-filename=-x -L dependency=D --out-dir D --emit=metadata,link --extern b=D/libb-x.rmeta c.rs; grep -o 'D/lib[ab]-x\\.r[a-z]*' st.txt | sort | uniq -c", + "observe": "Run on this VM. nightly-2020-01-22 (rustc 5e8897b7b 2020-01-21, which does not contain merge be663bf85): c's dep-info lists D/liba-x.rlib as well as liba-x.rmeta and libb-x.rmeta, and strace shows c opening D/liba-x.rlib. nightly-2020-01-24 (rustc 41f41b235 2020-01-23, which contains it): only liba-x.rmeta and libb-x.rmeta are listed and opened. This deterministic check shows the cause: the consumer opening the in-flight rlib. The user-visible symptom, a spurious cargo rebuild, needs the rlib to appear in the small window between the consumer starting and loading crates (ehuss: 'the time between T2 and T4 is extremely small'), plus -Zbinary-dep-depinfo, so the cargo-level symptom is not cheap to reproduce.", + "properties": [ + "P2" + ] + }, + { + "issue": 40364, + "title": "env!/option_env! values were baked into the output, but the env vars were not reported as dependencies: cargo did not rebuild when the variable changed", + "fix_pr": 71858, + "merged": "2020-06-26", + "how_it_violates": "P4: env!() reads the process environment while compiling, and the value ends up in the output. rustc did not record that the variable was read, so build tools treated the crate as fresh after the variable changed. PR #71858 added '# env-dep:KEY=VALUE' lines to dep-info, closing #40364, #44074 and #70517. Cargo then read those lines and rebuilt on a change (rust-lang/cargo PR #8421, merged 2020-06-30). Both parts are needed for the repro, so it uses stable releases with the matching cargo: 1.45.0 has neither part, 1.46.0 has both. rustc nightlies after 2020-06-26 have the dep-info lines; cargo's rebuild arrived when the cargo submodule was next updated in early July 2020.", + "before_toolchain": "1.45.0", + "after_toolchain": "1.46.0", + "reproducible_cheaply": true, + "files": [ + { + "path": "Cargo.toml", + "content": "[package]\nname = \"envdemo\"\nversion = \"0.1.0\"\nedition = \"2018\"\n" + }, + { + "path": "src/main.rs", + "content": "fn main() {\n println!(\"{}\", env!(\"MIRTH_DEMO\"));\n}\n" + } + ], + "commands": "cargo +TOOLCHAIN clean -q; MIRTH_DEMO=one cargo +TOOLCHAIN run -q; MIRTH_DEMO=two cargo +TOOLCHAIN run -q; grep env-dep target/debug/deps/envdemo-*.d || echo 'no env-dep in dep-info'", + "observe": "Verified locally. With 1.45.0 the two runs print 'one' then 'one': the stale binary is reused after MIRTH_DEMO changed. With 1.46.0 they print 'one' then 'two'. On 1.46.0 the dep-info file also has a '# env-dep:MIRTH_DEMO=two' line; 1.45.0 has none.", + "properties": [ + "P4" + ] + }, + { + "issue": 82920, + "title": "Miscompilation with incr. comp. (supertrait predicates sorted by DefId; reordering trait declarations breaks dyn method calls)", + "fix_pr": 83074, + "merged": "2021-03-15", + "how_it_violates": "Bounds and predicates were sorted by DefId, which is not stable across sessions. When two trait declarations swap places, the query result changes but is still treated as green and reused, so the vtable layout and the method call sites disagree. The incremental binary calls the wrong trait method, while a clean build is correct. The original report was a rust-analyzer test suite failing with out-of-bounds panics and segfaults until `cargo clean`. Labelled I-unsound.", + "before_toolchain": "nightly-2021-03-13", + "after_toolchain": "nightly-2021-03-17", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "trait MyTrait: One + Two {}\nimpl One for T {\n fn method_one(&self) -> usize {\n 1\n }\n}\nimpl Two for T {\n fn method_two(&self) -> usize {\n 2\n }\n}\nimpl MyTrait for T {}\n\nfn main() {\n let a: &dyn MyTrait = &true;\n assert_eq!(a.method_one(), 1);\n assert_eq!(a.method_two(), 2);\n}\n\n// Re-order traits 'One' and 'Two' between compilation sessions\n\n#[cfg(rpass1)]\ntrait One { fn method_one(&self) -> usize; }\n\ntrait Two { fn method_two(&self) -> usize; }\n\n#[cfg(rpass2)]\ntrait One { fn method_one(&self) -> usize; }\n" + } + ], + "commands": "rm -rf inc\nrustc +TOOLCHAIN --cfg rpass1 -C incremental=inc main.rs -o a.out && ./a.out && echo pass1 ok\nrustc +TOOLCHAIN --cfg rpass2 -C incremental=inc main.rs -o a.out && ./a.out && echo pass2 ok\n# control: a clean build of the edited source passes\nrustc +TOOLCHAIN --cfg rpass2 main.rs -o clean.out && ./clean.out && echo clean ok", + "observe": "Verified locally. On nightly-2021-03-13, the incremental pass 2 binary panics with \"assertion failed: `(left == right)` left: `2`, right: `1`\" (method_two was called for method_one), while a clean rpass2 build prints 'clean ok'. On nightly-2021-03-17 both passes print ok. Regression test: tests/incremental/issue-82920-predicate-order-miscompile.rs.", + "properties": [ + "P6", + "tracked" + ] + }, + { + "issue": 84252, + "title": "ICE: found unstable fingerprints for has_global_allocator(): the query read untracked CStore state", + "fix_pr": 84260, + "merged": "2021-04-17", + "how_it_violates": "has_global_allocator was answered from the CStore (the crate loader's per-crate metadata state) without recording any dependency. PR #84260: 'This query reads from untracked global state in `CStore`'. It was fixed by making it eval_always. The repro is the regression test tests/incremental/issue-84252-global-alloc.rs: removing #[global_allocator] between sessions leaves the cached result stale, and recomputing it gives a different value. The sibling fix #83153 (extern_mod_stmt_cnum made eval_always for the same reason, issue #83126) is a related example. #83126 is still open on GitHub, so I left it out as a separate entry.", + "before_toolchain": "nightly-2021-04-16", + "after_toolchain": "nightly-2021-04-19", + "reproducible_cheaply": true, + "files": [ + { + "path": "lib.rs", + "content": "#![crate_type=\"lib\"]\n#![crate_type=\"cdylib\"]\n\n#[allow(unused_imports)]\nuse std::alloc::System;\n\n#[cfg(bpass1)]\n#[global_allocator]\nstatic ALLOC: System = System;\n" + } + ], + "commands": "rm -rf inc out; mkdir out\nrustc +TOOLCHAIN -C incremental=inc --out-dir out --cfg bpass1 lib.rs; echo rc1=$?\nrustc +TOOLCHAIN -C incremental=inc --out-dir out lib.rs; echo rc2=$?", + "observe": "Before the fix (nightly-2021-04-16): the second build panics with \"found unstable fingerprints for has_global_allocator(lib[8787]): false\" (assertion left/right Fingerprint mismatch in rustc_query_system plumbing.rs) and exits non-zero. After the fix (nightly-2021-04-19): rc1=0 and rc2=0.", + "properties": [ + "tracked" + ] + }, + { + "issue": 66955, + "title": "--remap-path-prefix was UNTRACKED: changing it under incremental reused stale codegen and debuginfo with the old path", + "fix_pr": 84233, + "merged": "2021-04-29", + "how_it_violates": "P4, read broadly as untracked state that reaches the output: #48162 made --remap-path-prefix UNTRACKED so that the crate hash stayed the same. Remapped paths still went into the outputs, so after the flag changed, an incremental rebuild reused cached codegen units holding the old prefix. The result was a mix of old and new remapped paths. The fix (#84233) added TRACKED_NO_CRATE_HASH: the option invalidates the incremental cache but stays out of the crate hash. Note: in this era metadata was always re-encoded, so the rmeta itself picked up the new path. The stale data is in the object code and debuginfo inside the rlib.", + "before_toolchain": "nightly-2021-04-29", + "after_toolchain": "nightly-2021-05-01", + "reproducible_cheaply": true, + "files": [ + { + "path": "lib.rs", + "content": "pub fn f() -> u32 { 1 }\n#[inline] pub fn g() -> &'static str { file!() }\n" + }, + { + "path": "run.sh", + "content": "#!/bin/sh\nTC=$1\nrm -rf out incr; mkdir out\nfor P in /AAAA_first /BBBB_second; do\n rustc +$TC lib.rs --crate-type=rlib --out-dir out -g -C incremental=incr --remap-path-prefix=\"$PWD=$P\"\n echo \"after remap to $P, rlib contains:\"; strings out/liblib.rlib | grep -oE '/(AAAA_first|BBBB_second)' | sort | uniq -c\ndone\n" + } + ], + "commands": "sh run.sh TOOLCHAIN", + "observe": "Verified locally. On nightly-2021-04-29, after the second build (remap to /BBBB_second) the rlib still contains '2 /AAAA_first' and only '1 /BBBB_second'; the stale object code and debuginfo were reused. On nightly-2021-05-01 it contains '3 /BBBB_second' and no /AAAA_first.", + "properties": [ + "P4" + ] + }, + { + "issue": 89598, + "title": "VTable-related miscompilation with incremental compilation (reordering trait methods calls the wrong method)", + "fix_pr": 89619, + "merged": "2021-10-08", + "how_it_violates": "#86475 added an untracked global vtable cache to tcx. After trait methods are reordered, the object file for the main CGU is reused from the incremental cache with the old vtable layout, while mod1 is recompiled for the new layout. The incremental binary then calls method2 where a clean build calls method1, so the binaries differ (a miscompile). P-critical, I-unsound, regression-from-stable-to-beta.", + "before_toolchain": "nightly-2021-10-07", + "after_toolchain": "nightly-2021-10-10", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "trait Foo {\n #[cfg(rpass1)]\n fn method1(&self) -> u32;\n\n fn method2(&self) -> u32;\n\n #[cfg(rpass2)]\n fn method1(&self) -> u32;\n}\n\nimpl Foo for u32 {\n fn method1(&self) -> u32 { 17 }\n fn method2(&self) -> u32 { 42 }\n}\n\nfn main() {\n let x: &dyn Foo = &0u32;\n assert_eq!(mod1::foo(x), 17);\n}\n\nmod mod1 {\n pub(super) fn foo(x: &dyn super::Foo) -> u32 {\n x.method1()\n }\n}\n" + } + ], + "commands": "rm -rf inc\nrustc +TOOLCHAIN --cfg rpass1 -C incremental=inc main.rs -o a.out && ./a.out && echo pass1 ok\nrustc +TOOLCHAIN --cfg rpass2 -C incremental=inc main.rs -o a.out && ./a.out && echo pass2 ok\n# control: a clean build of the edited source passes\nrustc +TOOLCHAIN --cfg rpass2 main.rs -o clean.out && ./clean.out && echo clean ok", + "observe": "Verified locally. On nightly-2021-10-07, pass 1 prints 'pass1 ok'; the incremental pass 2 binary panics with \"assertion failed: `(left == right)` left: `42`, right: `17`\" because method2 was called through the stale vtable. The clean build of rpass2 runs fine. On nightly-2021-10-10 both passes print ok. Regression test: tests/incremental/reorder_vtable.rs.", + "properties": [ + "P6", + "tracked" + ] + }, + { + "issue": 107001, + "title": "rustc leaves *.rcgu.o object files in the output directory when a post-monomorphization lint (unconditional_panic / arithmetic_overflow) fails the build after codegen has started", + "fix_pr": 110107, + "merged": "2023-04-21", + "how_it_violates": "The const-prop lints (unconditional_panic, arithmetic_overflow) ran in mir_drops_elaborated_and_const_checked, which was only forced lazily, so their errors could fire after codegen had already written the CGU object files. rustc then aborted without deleting them, leaving `..rcgu.o` / `..-cgu.N.rcgu.o` in the output directory (target/debug/deps under cargo). PR #110107 ('Ensure mir_drops_elaborated_and_const_checked when requiring codegen') makes sure that query runs before codegen. Its description says: 'may emit errors while codegen has started, and the compiler would exit leaving object code files around. Found by @cuviper in #109731' (cuviper's comment there: 'each time I try one of these failing tests, it's leaving temporary *.rcgu.o files around'). Issue #107001 is the standalone report with this exact repro. It is still marked open, but its repro stops leaking at this PR. I bisected the nightlies myself and the boundary is exactly this merge: nightly-2023-04-21 (8bdcc62cb) leaks and nightly-2023-04-22 (fec9adcdb) is clean. The commit range between them contains #110107.", + "before_toolchain": "nightly-2023-04-21", + "after_toolchain": "nightly-2023-04-22", + "reproducible_cheaply": true, + "files": [ + { + "path": "code.rs", + "content": "fn main() {\n let a = [1, 2, 3, 4, 5];\n let _x = a[9];\n}\n" + } + ], + "commands": "mkdir out\nrustc +TOOLCHAIN code.rs --out-dir out; echo \"exit=$?\"\nls -A out", + "observe": "Both toolchains fail the same way: 'error: this operation will panic at runtime ... index out of bounds: the length is 5 but the index is 9' with `#[deny(unconditional_panic)]`, exit=1. On nightly-2023-04-21 (also stable 1.70.0), `ls -A out` lists 6 leftover object files, e.g. `code.1tgaf0fuackrygys.rcgu.o code.code.56d798bc-cgu.0.rcgu.o ...`. On nightly-2023-04-22 (also stable 1.71.0 and everything since, through nightly-2026-10-06), `out` is empty. Verified locally.", + "properties": [ + "P7" + ] + }, + { + "issue": 111227, + "title": "debugger_visualizer files (#111226 / #111227 / #111295) were read into rmeta but tracked by neither dep-info nor incremental: stale rmeta and an incremental ICE", + "fix_pr": 111641, + "merged": "2023-05-19", + "how_it_violates": "P4: the contents of files named by #![debugger_visualizer(natvis_file / gdb_script_file)] were read and encoded into crate metadata (the debugger_visualizers query), but those files were not tracked inputs. They were missing from dep-info (#111226), so cargo never rebuilt after a change. The incremental system also did not see them: changing the natvis file gave 'internal compiler error: encountered incremental compilation error with debugger_visualizers' (#111227), and changed GDB scripts were not picked up (#111295). The fix (#111641) made debugger_visualizers an eval_always query computed from the AST and added the files to dep-info. Its tests include run-make/incremental-debugger-visualizer, which greps the rmeta for the file contents.", + "before_toolchain": "nightly-2023-05-18", + "after_toolchain": "nightly-2023-05-21", + "reproducible_cheaply": true, + "files": [ + { + "path": "foo.rs", + "content": "#![debugger_visualizer(natvis_file = \"./foo.natvis\")]\n#![debugger_visualizer(gdb_script_file = \"./foo.py\")]\n\npub struct Foo {\n pub x: u32,\n}\n" + }, + { + "path": "run.sh", + "content": "#!/bin/sh\nTC=$1\nrm -rf out incr; mkdir out\necho \"GDB script v1\" > foo.py; echo \"Natvis v1\" > foo.natvis\nrustc +$TC foo.rs --crate-type=rlib --emit=metadata,dep-info --out-dir out -C incremental=incr\ngrep -q foo.py out/foo.d && echo \"foo.py in dep-info\" || echo \"foo.py NOT in dep-info\"\necho \"Natvis v2\" > foo.natvis\nrustc +$TC foo.rs --crate-type=rlib --emit=metadata --out-dir out -C incremental=incr 2>&1 | head -3\ngrep -a -o \"Natvis v[12]\" out/libfoo.rmeta\n" + } + ], + "commands": "sh run.sh TOOLCHAIN", + "observe": "Verified locally. On nightly-2023-05-18 the script prints 'foo.py NOT in dep-info'. The second, incremental compile prints 'error: internal compiler error: encountered incremental compilation error with debugger_visualizers(foo[47df])', and libfoo.rmeta still contains 'Natvis v1', so the metadata is stale. On nightly-2023-05-21 foo.py is listed in foo.d, the rebuild succeeds, and the rmeta contains 'Natvis v2'.", + "properties": [ + "P4" + ] + }, + { + "issue": 111295, + "title": "debugger_visualizer: edits to the visualizer script are not picked up under incremental (stale content that is exported into metadata)", + "fix_pr": 111641, + "merged": "2023-05-19", + "how_it_violates": "The debugger_visualizers query result (the script contents, which are also encoded into crate metadata so downstream binaries embed upstream visualizers) did not depend on the script file. After the .py file was edited, an incremental rebuild reused the stale contents. #111641 ('Fix dependency tracking for debugger visualizers') made the query eval_always, and the fix also hashes the visualizer contents into crate_hash. The code comment says 'that content is exported into crate metadata, so any changes to it need to be reflected in the crate hash', so downstream crates that read it from metadata also see the change. The same PR fixes #111227 (an ICE from changing a natvis file with --crate-type=rlib) and #111226 (dep-info). The repro is single-crate, as in the issue. No gdb is needed: read the .debug_gdb_scripts section directly.", + "before_toolchain": "nightly-2023-05-18", + "after_toolchain": "nightly-2023-05-21", + "reproducible_cheaply": true, + "files": [ + { + "path": "foo.rs", + "content": "#![debugger_visualizer(gdb_script_file = \"foo.py\")]\n\nfn main() {\n let x = 1;\n println!(\"breakpoint {x}\");\n}\n" + } + ], + "commands": "rm -rf incremental foo\necho \"print('hello!')\" > foo.py\nrustc +TOOLCHAIN -Cincremental=incremental -g foo.rs\necho \"print('hello world')\" > foo.py\nrustc +TOOLCHAIN -Cincremental=incremental -g foo.rs\nobjcopy -O binary --only-section=.debug_gdb_scripts foo /dev/stdout | tr -c '[:print:]' ' ' | grep -o \"print('[^']*')\"", + "observe": "Before the fix (nightly-2023-05-18): the rebuilt binary's .debug_gdb_scripts still contains print('hello!'), the stale script. After the fix (nightly-2023-05-21): it contains print('hello world'). Linux/ELF only; needs binutils objcopy.", + "properties": [ + "tracked" + ] + }, + { + "issue": 117254, + "title": "FileEncoder delayed error reporting is still broken (rmeta write errors swallowed; truncated .rmeta published with exit 0)", + "fix_pr": 117301, + "merged": "2023-11-26", + "how_it_violates": "rmeta encoding never called FileEncoder::finish, so an I/O error while writing the temp file (ENOSPC in crater, EFBIG here) was swallowed. rustc then renamed the short temp file to the final lib.rmeta and exited 0. Downstream crates ICE'd in MemDecoder (\"We're going ahead to decode a result which was not completely written out\", issue text). The rename itself is atomic, but the file it publishes is not fully written, which breaks the second half of P1. Introduced by the delayed error scheme of #94732 and not fixed by #115542. #117301 added the finish() check, but only with emit_err; see the next entry for the remaining hole.", + "before_toolchain": "nightly-2023-11-26", + "after_toolchain": "nightly-2023-11-28", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.py", + "content": "print('#![crate_type=\"lib\"]')\nfor i in range(20000):\n print(f'pub fn f{i}(x: u64) -> u64 {{ x.wrapping_mul({i}) }}')\n print(f'pub struct S{i} {{ pub a: u64, pub b: String }}')\n" + }, + { + "path": "user.rs", + "content": "extern crate big;\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\npython3 gen.py > big.rs # ~6.8 MB of metadata\nmkdir out\n# Cap file size at 1 MiB and ignore SIGXFSZ so write() fails with EFBIG (a cheap stand-in for a full disk, no root needed)\nbash -c \"trap '' XFSZ; ulimit -f 1024; rustc +TOOLCHAIN --crate-type lib --emit=metadata big.rs --out-dir out; echo rustc exit=\\$?\"\nls -l out\nrustc +TOOLCHAIN --crate-type lib --emit=metadata user.rs -L out -o u.rmeta 2>&1 | grep -m2 -E 'panicked|range|error'", + "observe": "nightly-2023-11-26 (1.76.0-nightly f5dc2653f 2023-11-25): no diagnostic, \"rustc exit=0\", and out/libbig.rmeta is exactly 1048576 bytes, a truncated file at the final path. The downstream rustc then ICEs in rustc_serialize/src/opaque.rs (\"range start index ... out of range for slice of length 1048576\"). nightly-2023-11-28 (49b3924bd 2023-11-27): rustc now prints \"error: failed to write to `.../rmetaXXXX/lib.rmeta`: File too large (os error 27)\" and exits 1. However, the 1048576-byte out/libbig.rmeta is still published, and the downstream rustc still panics with \"range start index 14979595 out of range for slice of length 1048576\" (fixed fully by #119510 below).", + "properties": [ + "P1" + ] + }, + { + "issue": 119456, + "title": "ICE 'range start index ... out of range for slice of length 16384': rmeta I/O error reported with emit_err, so the truncated temp file is still renamed to the final path", + "fix_pr": 119510, + "merged": "2024-01-03", + "how_it_violates": "After #117301, rmeta write errors were reported with emit_err, which does not stop compilation. rustc kept going and renamed the incomplete temp file over lib.rmeta, and cargo could start a dependent build that read it (PR #119510: \"there is a window of time between the call to emit_err and the full error reporting where rustc believes it has emitted a valid rmeta file and will permit Cargo to launch a build for a dependent crate\"). Switching to emit_fatal aborts before the rename, so nothing partial reaches the final path. The reporter hit this on 1.75.0 with typst as a dependency (disk-full); the PR author reproduced it with an LD_PRELOAD write() that randomly returns ENOSPC.", + "before_toolchain": "nightly-2024-01-01", + "after_toolchain": "nightly-2024-01-05", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.py", + "content": "print('#![crate_type=\"lib\"]')\nfor i in range(20000):\n print(f'pub fn f{i}(x: u64) -> u64 {{ x.wrapping_mul({i}) }}')\n print(f'pub struct S{i} {{ pub a: u64, pub b: String }}')\n" + }, + { + "path": "user.rs", + "content": "extern crate big;\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\npython3 gen.py > big.rs\nmkdir out\nbash -c \"trap '' XFSZ; ulimit -f 1024; rustc +TOOLCHAIN --crate-type lib --emit=metadata big.rs --out-dir out; echo rustc exit=\\$?\"\nls -l out\nrustc +TOOLCHAIN --crate-type lib --emit=metadata user.rs -L out -o u.rmeta 2>&1 | grep -m2 -E 'panicked|range|error'", + "observe": "nightly-2024-01-01 (1.77.0-nightly e51e98dde 2023-12-31): \"error: failed to write to `.../rmetaXXXX/lib.rmeta`: File too large (os error 27)\" and exit 1, yet `ls -l out` shows libbig.rmeta at 1048576 bytes. The reader then panics: \"thread 'rustc' panicked at .../compiler/rustc_serialize/src/opaque.rs:262:42: range start index 12750977 out of range for slice of length 1048576\". nightly-2024-01-05 (f688dd684 2024-01-04): same error and exit 1, but out/ is empty (the truncated file is never renamed into place), and the reader gets a clean \"error[E0463]: can't find crate for `big`\". Current stable 1.97.1 behaves the same as 2024-01-05 (the temp file is now named full.rmeta).", + "properties": [ + "P1" + ] + }, + { + "issue": 122859, + "title": "Implied bound not implied across crates: associated-type bounds in supertrait position were dropped from metadata (implied_predicates encoded as super_predicates)", + "fix_pr": 122891, + "merged": "2024-03-24", + "how_it_violates": "For traits, the metadata encoder assumed `implied_predicates_of` was equal to `super_predicates_of` and only wrote the latter. So the implied predicates that come from associated type bounds (`trait Bar: Super`) were never written. The dependent crate read a smaller predicate list without noticing, and lost the bound. That shows up as a wrong E0277 that only happens across crates. The PR says: 'The assumption that they didn't differ was hard-coded in #107614, so in cross-crate positions this means that we forget the implied predicates from associated type bounds.' The same code compiles when crate_b is a local module.", + "before_toolchain": "nightly-2024-03-24", + "after_toolchain": "nightly-2024-03-26", + "reproducible_cheaply": true, + "files": [ + { + "path": "crate_b.rs", + "content": "pub trait Foo { type FooAssoc: Bar; }\npub trait Bar: Super {}\npub trait Super { type SuperAssoc; }\npub trait Bound: Unsatisfied {}\npub trait Unsatisfied {}\n" + }, + { + "path": "main.rs", + "content": "use crate_b::{Foo, Super, Unsatisfied};\nfn foo() {\n unsatisfied::<::SuperAssoc>()\n}\nfn unsatisfied() {}\nfn main() {}\n" + } + ], + "commands": "rustc +TOOLCHAIN --edition 2021 --crate-type lib crate_b.rs\nrustc +TOOLCHAIN --edition 2021 main.rs -L . --extern crate_b", + "observe": "Verified locally. On nightly-2024-03-24, main.rs fails with `error[E0277]: the trait bound `<::FooAssoc as Super>::SuperAssoc: Unsatisfied` is not satisfied`. On nightly-2024-03-26 it compiles (there is only a dead-code warning). Pasting crate_b's contents into main.rs as `mod crate_b` compiles on both, which shows that only the cross-crate (metadata) path is affected.", + "properties": [ + "P3" + ] + }, + { + "issue": 130201, + "title": "ICE: `coroutine_by_move_body_def_id` unsupported by its crate when calling a foreign crate's async closure as AsyncFnOnce (query result and by-move MIR were never encoded in rmeta)", + "fix_pr": 130201, + "merged": "2024-09-17", + "how_it_violates": "The dependency crate never wrote the `coroutine_by_move_body_def_id` table entry, and never wrote optimized_mir for the synthetic by-move body. When the dependent crate asked for that entry, the lookup found nothing and fell through to a missing provider, which ICEs. The PR text says: 'We weren't encoding this query in the metadata though, nor were we properly recording that synthetic MIR in `mir_keys`, so the `optimized_mir` wasn't getting encoded either!' There is no separate issue; PR #130201 itself is the reference, and its regression test is tests/ui/async-await/async-closures/foreign.rs.", + "before_toolchain": "nightly-2024-09-17", + "after_toolchain": "nightly-2024-09-19", + "reproducible_cheaply": true, + "files": [ + { + "path": "foreign.rs", + "content": "#![feature(async_closure)]\npub fn closure() -> impl async Fn() {\n async || {}\n}\n" + }, + { + "path": "main.rs", + "content": "#![feature(async_closure)]\nextern crate foreign;\nasync fn call_once(f: impl async FnOnce()) {\n f().await;\n}\nfn main() {\n let _ = call_once(foreign::closure());\n}\n" + } + ], + "commands": "rustc +TOOLCHAIN --edition 2021 --crate-type lib foreign.rs\nrustc +TOOLCHAIN --edition 2021 main.rs -L .", + "observe": "Verified locally. On nightly-2024-09-17, compiling main.rs ICEs with: `error: internal compiler error: compiler/rustc_middle/src/query/plumbing.rs:664:5: `tcx.coroutine_by_move_body_def_id(DefId(20:6 ~ foreign[87b3]::closure::{closure#0}::{closure#0}))` unsupported by its crate; perhaps the `coroutine_by_move_body_def_id` query was never assigned a provider function`. On nightly-2024-09-19 it compiles cleanly and produces the `main` binary.", + "properties": [ + "P3" + ] + }, + { + "issue": 135514, + "title": "Rust 1.84 sometimes allows overlapping impls in incremental re-builds (new solver did not record deps for cached tasks)", + "fix_pr": 133828, + "merged": "2024-12-05", + "how_it_violates": "A clean build of the edited source fails with E0119 (conflicting implementations). The incremental rebuild after the same edit accepts it, because coherence results computed through the new solver's cache had no dependency edges. The accepted program is a safe transmute: Vec becomes String and prints ABC. The diagnostics and the success or failure differ from a clean build. Labelled I-unsound, regression-from-stable-to-stable. #135522 added tests/incremental/overlapping-impls-in-new-solver-issue-135514.rs. The fix was already on master before the issue was filed, so 1.85 is fixed and 1.84.x is affected.", + "before_toolchain": "1.84.0", + "after_toolchain": "1.85.0", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "trait Trait {}\n\nstruct S0(T);\n\nstruct S(T);\nimpl Trait for S where S0: Trait {}\n\nstruct W;\n\ntrait Other {\n type Choose;\n}\n\n// first impl\nimpl Other for T {\n type Choose = L;\n}\n\n// second impl: overlaps the first one once `S: Trait` holds\nimpl Other for S {\n type Choose = R;\n}\n\n#[cfg(rpass1)]\nimpl Trait for W {}\n\n#[cfg(rpass1)]\npub fn transmute(_l: L) -> R {\n todo!();\n}\n\n#[cfg(rpass2)]\nimpl Trait for S {}\n\n#[cfg(rpass2)]\nfn use_first_impl(l: L) -> <::To as Other>::Choose {\n l\n}\n\n#[cfg(rpass2)]\nfn use_second_impl(l: as Other>::Choose) -> R {\n l\n}\n\ntrait TyEq {\n type To;\n}\nimpl TyEq for T {\n type To = T;\n}\n\n#[cfg(rpass2)]\nfn transmute_inner(l: L) -> R\nwhere\n T: Trait + TyEq>,\n{\n use_second_impl::(use_first_impl::(l))\n}\n\n#[cfg(rpass2)]\npub fn transmute(l: L) -> R {\n transmute_inner::, L, R>(l)\n}\n\nfn main() {\n if cfg!(rpass2) {\n let v = vec![65_u8, 66, 67];\n let s: String = transmute(v);\n println!(\"{}\", s);\n } else {\n println!(\"pass1\");\n }\n}\n" + } + ], + "commands": "rm -rf inc\nrustc +TOOLCHAIN --cfg rpass1 -A warnings -C incremental=inc main.rs -o a.out && ./a.out\nrustc +TOOLCHAIN --cfg rpass2 -A warnings -C incremental=inc main.rs -o a.out && ./a.out && echo 'incremental pass2 built and ran'\n# control: clean build of the edited source\nrustc +TOOLCHAIN --cfg rpass2 -A warnings main.rs -o clean.out", + "observe": "Verified locally. On 1.84.0, pass 1 prints 'pass1'; the incremental pass 2 compiles without error and prints 'ABC' then 'incremental pass2 built and ran'; the clean build of the same source fails with error[E0119]: conflicting implementations of trait `Other` for type `S`. On 1.85.0 the incremental pass 2 also fails with E0119. On 1.83.0 the old solver rejects pass 1 itself, so use 1.84.0. For nightlies, try nightly-2024-12-04 (bug) and nightly-2024-12-07 (fixed); these were not run.", + "properties": [ + "P6" + ] + }, + { + "issue": 138678, + "title": "Randomly seeded HashMap (pulldown-cmark reference_definitions) iterated while collecting doc links: lib.rmeta differs between identical runs", + "fix_pr": 138678, + "merged": "2025-03-28", + "how_it_violates": "P4: metadata encoding read a randomly seeded std HashMap. rustc_resolve::rustdoc::parse_links walks pulldown-cmark's reference_definitions(), a HashMap with RandomState, and pushes the links in that iteration order. The resulting list of doc-link candidates is encoded into crate metadata, so two runs of rustc on the same input wrote different .rmeta bytes. The bug came in with #136363 (merged 2025-02-16). The fix (#138678) sorts the links by label. #138678 is the PR; it has no separate issue, and its description says the nondeterminism was found in Bazel lib.rmeta outputs.", + "before_toolchain": "nightly-2025-03-28", + "after_toolchain": "nightly-2025-03-30", + "reproducible_cheaply": true, + "files": [ + { + "path": "lib.rs", + "content": "//! Crate docs.\n//!\n//! [l0]: crate::S0\n//! [l1]: crate::S1\n//! [l2]: crate::S2\n//! [l3]: crate::S3\n//! [l4]: crate::S4\n//! [l5]: crate::S5\n//! [l6]: crate::S6\n//! [l7]: crate::S7\n//! [l8]: crate::S8\n//! [l9]: crate::S9\n//! [l10]: crate::S10\n//! [l11]: crate::S11\n//! [l12]: crate::S12\n//! [l13]: crate::S13\n//! [l14]: crate::S14\n//! [l15]: crate::S15\npub struct S0;\npub struct S1;\npub struct S2;\npub struct S3;\npub struct S4;\npub struct S5;\npub struct S6;\npub struct S7;\npub struct S8;\npub struct S9;\npub struct S10;\npub struct S11;\npub struct S12;\npub struct S13;\npub struct S14;\npub struct S15;\n" + } + ], + "commands": "for i in 1 2 3 4 5 6 7 8; do rm -rf out; mkdir out; rustc +TOOLCHAIN lib.rs --crate-type=rlib --emit=metadata --out-dir out 2>/dev/null; sha256sum out/liblib.rmeta | cut -c1-16; done | sort | uniq -c", + "observe": "Verified locally. On nightly-2025-03-28 (rustc 1.87.0-nightly 3f5502370 2025-03-27), the 8 identical compilations gave 8 different liblib.rmeta hashes, each counted once. On nightly-2025-03-30 (1.88.0-nightly 1799887bb 2025-03-29), all 8 gave the same hash (count 8).", + "properties": [ + "P4" + ] + }, + { + "issue": 139407, + "title": "Instructions missing from (naked_)asm blocks after fixing an assembler error and rebuilding (hard-linked temp files corrupt the incremental cache)", + "fix_pr": 139453, + "merged": "2025-04-11", + "how_it_violates": "Object files were hard-linked from fixed temp paths into the incremental session directory. A session that fails in the assembler (left unfinalized) overwrites a temp file that a previous, finalized session still hard-links. On the next successful build the reused object holds code from the failed session's source, so the binary differs from a clean build. The fix gives temp files a per-invocation random prefix. Labelled I-unsound. The run-make regression test is tests/run-make/dirty-incr-due-to-hard-link. The reporter hit it with plain `cargo run` (no -Csave-temps) on aarch64-apple-darwin. The test, used here, needs -Csave-temps to keep the temp files on x86_64 Linux. It needs a failed build in between (an asm error), but it is deterministic.", + "before_toolchain": "nightly-2025-04-10", + "after_toolchain": "nightly-2025-04-13", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "#[inline(never)]\n#[cfg(any(rpass1, rpass3))]\nfn a() -> i32 {\n 0\n}\n\n#[cfg(any(cfail2))]\nfn a() -> i32 {\n 1\n}\n\nfn main() {\n evil::evil();\n assert_eq!(a(), 0);\n}\n\nmod evil {\n #[cfg(any(rpass1, rpass3))]\n pub fn evil() {\n unsafe {\n std::arch::asm!(\"/* */\");\n }\n }\n\n #[cfg(any(cfail2))]\n pub fn evil() {\n unsafe {\n std::arch::asm!(\"missing\");\n }\n }\n}\n" + } + ], + "commands": "rm -rf inc a.out *.o *.rcgu* \nR=\"rustc +TOOLCHAIN -C incremental=inc -C save-temps main.rs -o a.out\"\n$R --cfg rpass1 && ./a.out && echo pass1 ok\n$R --cfg cfail2; echo '(pass2 is expected to fail with an assembler error)'\n$R --cfg rpass3 && ./a.out && echo pass3 ok", + "observe": "Verified locally on x86_64 Linux. On nightly-2025-04-10 (and on stable 1.70.0, 1.86.0 and 1.87.0), pass 1 prints 'pass1 ok' and pass 2 fails with \"error: invalid instruction mnemonic 'missing'\". Pass 3 has the same source as pass 1, yet its binary panics with \"assertion `left == right` failed left: 1 right: 0\": a() from the failed cfail2 session leaked into the reused object. On nightly-2025-04-13 and on 1.88.0, pass 3 prints 'pass3 ok'.", + "properties": [ + "P6" + ] + }, + { + "issue": 139899, + "title": "rustdoc --test leaves rustdoctest* temporary directories behind whenever a doctest fails (compile error or panic)", + "fix_pr": 140706, + "merged": "2025-05-08", + "how_it_violates": "When any doctest failed, rustdoc (or libtest) called process::exit, so the TempDir destructor never ran and the `rustdoctestXXXXXX` directory was never removed. Issue #139899 reports 6197 of them piling up in /tmp. PR #140706 ('[rustdoc] Ensure that temporary doctest folder is correctly removed even if doctests failed') adds a libtest hook that runs after all tests and cleans the folder up. It also adds the regression test tests/run-make/rustdoc/doctest/tempdir-removal, whose two input files are used verbatim below. The leftovers go to TMPDIR, not to target/, so this hits P7 only when TMPDIR points inside the build tree. It is still a real, fixed 'temp dir left behind on the failure path' bug in the toolchain that `cargo test --doc` runs. Note: this is rustdoc, not rustc.", + "before_toolchain": "nightly-2025-05-07", + "after_toolchain": "nightly-2025-05-09", + "reproducible_cheaply": true, + "files": [ + { + "path": "compile-error.rs", + "content": "#![doc(test(attr(deny(warnings))))]\n\n//! ```\n//! let a = 12;\n//! ```\n" + }, + { + "path": "run-error.rs", + "content": "//! ```\n//! panic!();\n//! ```\n" + } + ], + "commands": "mkdir tmp\nfor f in compile-error.rs run-error.rs; do for ed in 2018 2024; do TMPDIR=$PWD/tmp rustdoc +TOOLCHAIN --test $f --edition $ed >/dev/null 2>&1; echo \"$f $ed exit=$?\"; done; done\nls -A tmp", + "observe": "Every run exits 101 (failed doctest) on both toolchains. On nightly-2025-05-07, `ls -A tmp` shows one leftover directory per run (4 in total), e.g. `rustdoctestLP6B1A rustdoctesteGUqc8 rustdoctest0kc32h rustdoctestqVdzVs`. On nightly-2025-05-09, `tmp` is empty. Verified locally for both editions (2018 per-test and 2024 merged doctests).", + "properties": [ + "P7" + ] + }, + { + "issue": 114669, + "title": "Metadata was never reused across incremental sessions: always re-encoded (fixed by #114669 'Make metadata a workproduct and reuse it', with prerequisite #143247 'Avoid depending on forever-red DepNode when encoding metadata')", + "fix_pr": 114669, + "merged": "2025-07-04", + "how_it_violates": "A performance violation of 'metadata is reused (not re-encoded) when nothing it depends on changed'. Before July 2025, metadata encoding depended on the forever-red DepNode (iter_local_def_id / def_path_table read DepNodeIndex::FOREVER_RED_NODE) and was not a dep-graph task, so every incremental session re-encoded the .rmeta from scratch. #143247 (merged 2025-07-04, split out of #114669 'for perf') removed the forever-red read by depending on `analysis` instead. #114669 (merged 2025-07-04) wraps encoding in a Metadata dep-node task, saves the rmeta as a work product ('metadata') in the incremental dir, and when the node is green it hardlinks or copies the saved file instead of encoding ('can yield substantial gains (~10%)... if all the changes are in upstream crates and have no effect on it'). Observable without logs: the reused rmeta is a hardlink of the work product, so its inode stays the same across no-op sessions. Verified on both toolchains. This is a numbered PR, not an issue: no separate GitHub issue exists, so the issue field holds the PR number.", + "before_toolchain": "nightly-2025-07-03", + "after_toolchain": "nightly-2025-07-06", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "pub fn public() -> u32 { private() }\nfn private() -> u32 { 1 }\n" + } + ], + "commands": "rm -rf out inc; mkdir out\nfor i in 1 2; do rustc +TOOLCHAIN --edition 2021 --crate-type lib --emit=metadata,link -C incremental=inc --out-dir out a.rs; echo \"run$i inode=$(stat -c %i out/liba.rmeta) links=$(stat -c %h out/liba.rmeta)\"; done", + "observe": "Before (nightly-2025-07-03): run1 and run2 report different inodes and links=1, so the rmeta is freshly encoded on the unchanged rebuild. After (nightly-2025-07-06): links=2 (the file is hardlinked to the 'metadata' work product in the incremental dir) and run2 has the same inode as run1, so the metadata was reused rather than re-encoded. Note: after the fix, editing even a private fn body still changed the inode in my test (the Metadata node went red), so the no-op rebuild is the clean demonstration. Needs a filesystem with hardlinks; on one without them link_or_copy falls back to copying, and the inode check does not apply.", + "properties": [ + "tracked" + ] + }, + { + "issue": 144004, + "title": "rustdoc drops #[no_mangle] / #[link_section] from inlined cross-crate re-exports: those attributes were not encoded in crate metadata", + "fix_pr": 144050, + "merged": "2025-07-19", + "how_it_violates": "`no_mangle` and `link_section` were moved out of the generic encoded attribute list, and nothing else wrote them to the attribute table. A dependent crate that reads the dependency's attributes from metadata (here rustdoc inlining a `pub use a::*` re-export) silently gets no such attribute, with no error. The PR title is 'Fix encoding of link_section and no_mangle cross crate', and it fixes it by always encoding them. This is the 'missing attributes cross-crate' kind of P3 violation. The result is silent information loss, not an ICE.", + "before_toolchain": "nightly-2025-07-06", + "after_toolchain": "nightly-2025-07-24", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "#[unsafe(no_mangle)]\npub fn f0() {}\n#[unsafe(link_section = \".here\")]\npub fn f1() {}\n#[unsafe(no_mangle)]\npub static S0: () = ();\n#[unsafe(link_section = \".there\")]\npub static S1: () = ();\n" + }, + { + "path": "b.rs", + "content": "pub use a::*;\n" + } + ], + "commands": "rustc +TOOLCHAIN a.rs --crate-type lib --edition 2024\nrustdoc +TOOLCHAIN b.rs --edition 2024 -L. --extern a\nfor f in fn.f0 fn.f1 static.S0 static.S1; do echo -n \"$f: \"; grep -o 'no_mangle\\|link_section[^<]*' doc/b/$f.html | sort -u | tr '\\n' ' '; echo; done", + "observe": "Verified locally. On nightly-2025-07-06 all four pages (doc/b/fn.f0.html, fn.f1.html, static.S0.html, static.S1.html) show no attribute, so every grep prints nothing. On nightly-2025-07-24 they show `no_mangle`, `link_section = \".here\"`, `no_mangle` and `link_section = \".there\"`. The issue adds that 1.88.0 stable also lacked both attributes, and 1.89 beta showed no_mangle but not link_section. The regression came and went with attribute-parsing refactors.", + "properties": [ + "P3" + ] + }, + { + "issue": 140413, + "title": "parallel rustc: static mut refs not reproducible", + "fix_pr": 144722, + "merged": "2025-08-13", + "how_it_violates": "With -Zthreads=50, building the same binary repeatedly gives different bytes. Mono items were sorted by (DefId, SymbolName) to pick their order in the output. DefId indices are allocated in access order, which is nondeterministic under the parallel front end. PR #144722 ('Fix parallel rustc not being reproducible due to unstable sorts of items') stops sorting by DefId. Two later facts confirm the fix: PR #161353 (merged 2026-09-01; it closed the issue and added tests/run-make/parallel-reproducible-build) found by bisection that the regression flips at nightly-2025-08-14, at commit #144722. The same PR also fixes the async-closure case #140425 (closed 2025-08-13).", + "before_toolchain": "nightly-2025-07-26", + "after_toolchain": "nightly-2025-12-10", + "reproducible_cheaply": true, + "files": [ + { + "path": "static.rs", + "content": "// Checks that mutable static items can have mutable slices and other references\n\npub static mut TEST: &'static mut [isize] = &mut [1];\npub static mut EMPTY: &'static mut [isize] = &mut [];\npub static mut INT: &'static mut isize = &mut 1;\n\n// And the same for raw pointers.\n\npub static mut TEST_RAW: *mut [isize] = &mut [1isize] as *mut _;\npub static mut EMPTY_RAW: *mut [isize] = &mut [] as *mut _;\npub static mut INT_RAW: *mut isize = &mut 1isize as *mut _;\n\npub fn main() {\n unsafe {\n TEST[0] += 1;\n assert_eq!(TEST[0], 2);\n *INT_RAW += 1;\n assert_eq!(*INT_RAW, 2);\n }\n}\n" + } + ], + "commands": "for i in $(seq 1 15); do rustc +TOOLCHAIN static.rs -Zthreads=50 -o a 2>/dev/null; md5sum < a; done | sort | uniq -c", + "observe": "Verified locally. On nightly-2025-07-26 there were 2 distinct md5s of the binary over 15 runs. On nightly-2025-12-10 there was 1. The bisected boundary is nightly-2025-08-13 (bad) to nightly-2025-08-14 (good). Caveat: building this file as --crate-type=lib --emit=metadata is still nondeterministic, even on nightly-2026-10-06 (7-8 distinct rmeta md5s out of 10). That case is the still-open follow-up #162203, so use this repro for the binary only.", + "properties": [ + "P5t" + ] + }, + { + "issue": 159677, + "title": ".rmeta contents depend on unrelated files in the library search path (doc_link_resolutions encoded in Symbol-index hash order)", + "fix_pr": 159718, + "merged": "2026-07-24", + "how_it_violates": "Building the same client.rs with the same flags twice gives different .rmeta bytes if an unrelated rlib whose name starts with the dependency's name (libfoo_bar.rlib next to libfoo.rlib) is in the -L directory. The crate locator opens libfoo_bar.rlib to read its crate name, which interns the extra Symbol `foo_bar`. That shifts the interner indices of symbols interned later, such as the doc-link strings. DocLinkResMap was an UnordMap (FxHashMap) keyed by (Symbol, Namespace) and hashed by interner index, and it was encoded in hash-iteration order. The fix makes it an FxIndexMap, so entries are encoded in insertion order. The PR adds the regression test tests/run-make/rmeta-unrelated-search-path-files.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": true, + "files": [ + { + "path": "client.rs", + "content": "//! [crate::Client]\n\nextern crate foo;\n\npub struct Client;\n" + }, + { + "path": "foo.rs", + "content": "pub struct Foo;\n" + }, + { + "path": "foo_bar.rs", + "content": "pub struct FooBar;\n" + } + ], + "commands": "rustc +TOOLCHAIN foo.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client1.rmeta\nrustc +TOOLCHAIN foo_bar.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client2.rmeta\ncmp client1.rmeta client2.rmeta && echo IDENTICAL", + "observe": "Verified locally. On nightly-2026-07-20 (and on stable 1.97.1) cmp prints 'client1.rmeta client2.rmeta differ: byte 1256' (1263 on 1.97.1); the differing bytes are the reordered doc-link entry 'crate::Client'. On nightly-2026-09-25 the files are identical and the script prints IDENTICAL. The fix merged 2026-07-24T09:12Z, so nightly-2026-07-26 and later should be fixed.", + "properties": [ + "P5" + ] + }, + { + "issue": 150451, + "title": "parallel compiler: thread::spawn-ing loop not reproducible (LLVM inline-asm location cookies)", + "fix_pr": 160197, + "merged": "2026-09-07", + "how_it_violates": "With -Zthreads=3, compiling a tiny lib that uses thread::spawn gives a different .rlib each time. The output differs when bitcode is embedded or LTO is used. The cause is that the srcloc 'cookies' attached to LLVM inline asm were assigned nondeterministically by the parallel front end. PR #160197 ('Restrict LLVM inline asm location cookie usage. Fixes #150451') landed in rollup #162434 on 2026-09-07 and added tests/run-make/parallel-reproducible-inline-asm-cookie. The bug is in object code and bitcode, not in rmeta, but it breaks reproducibility of the published artifact.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": true, + "files": [ + { + "path": "spawn.rs", + "content": "use std::thread;\n\nfn _main() {\n let _t1 = thread::spawn(|| {\n for _ in 0..100 {\n println!(\"test\");\n }\n });\n}\n" + } + ], + "commands": "for i in $(seq 1 15); do rustc +TOOLCHAIN spawn.rs --crate-type=lib -Zthreads=3 -Clink-dead-code=true -Copt-level=0 -Cembed-bitcode=true -o libs.rlib 2>/dev/null; md5sum < libs.rlib; done | sort | uniq -c", + "observe": "Verified locally. On nightly-2026-07-20 there were 7 distinct rlib md5s over 15 runs. On nightly-2026-09-25 there was 1.", + "properties": [ + "P5t" + ] + }, + { + "issue": 129094, + "title": "Parallel frontend: derives make metadata irreproducible (non-deterministic encoding of syntax contexts)", + "fix_pr": 161450, + "merged": "2026-09-16", + "how_it_violates": "With -Zthreads=N, the order in which SyntaxContexts and expansions are reached during metadata encoding depends on thread scheduling, so a tiny derive-heavy crate compiled repeatedly with the same inputs gives different rlibs. The fix ('Fix non-deterministic encoding of syntax contexts', which reiterates #157409) adds deterministic encoding indices in rustc_metadata/rmeta/encoder.rs and rustc_span/hygiene.rs. It also adds derives-issue-129094.rs to tests/run-make/parallel-reproducible-build. Unlike the other entries this is not single-threaded: it needs the parallel frontend.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": false, + "files": [ + { + "path": "derives.rs", + "content": "#![crate_type = \"lib\"]\n#[derive(Clone, Copy, Hash, PartialEq, PartialOrd)]\nstruct PackedPoint {\n x: u32,\n}\n" + } + ], + "commands": "for i in $(seq 1 20); do rm -rf o$i; mkdir o$i; rustc +TOOLCHAIN derives.rs -Zthreads=16 -Copt-level=3 --out-dir o$i 2>/dev/null; done\nmd5sum o*/*.rlib | awk '{print $1}' | sort | uniq -c", + "observe": "Verified locally. On nightly-2026-07-20, 20 runs gave 3 distinct rlib hashes (15/4/1). On nightly-2026-09-25 all 20 runs give one hash. Race-dependent: it needs -Zthreads>1 and several runs, and how often it shows up depends on core count and scheduling. It reproduced readily on this VM.", + "properties": [ + "P5", + "P5t" + ] + }, + { + "issue": 162901, + "title": "Diagnostic deduplication breaks with incr comp: incremental builds print 4 copies of an error a non-incremental build prints once (also fixes #106571, duplicate JSON messages)", + "fix_pr": 163461, + "merged": "2026-10-02", + "how_it_violates": "In incremental mode, spans carry a parent (incremental-relative spans). The diagnostic dedup hash included Span::parent, so identical diagnostics hashed differently and were all emitted. The diagnostics from an incremental compile differ from a non-incremental compile of the same source. The fix PR also fixes #106571 (regression from #84762, Jan 2023), where `cargo check --message-format=json` with incremental on prints identical compiler-message lines twice for a proc-macro-generated error, and CARGO_INCREMENTAL=0 does not. No edit is needed: the difference is already there between an incremental and a non-incremental build. A P6 check that compares against a clean incremental build would not see it. It shows only when the reference build is non-incremental (CARGO_INCREMENTAL=0) or when diagnostic counts are compared.", + "before_toolchain": "nightly-2026-10-01", + "after_toolchain": "nightly-2026-10-03", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "#[global_allocator]\nstatic A: usize = 0;\n\nfn main() {}\n" + } + ], + "commands": "rm -rf inc\necho \"non-incremental: $(rustc +TOOLCHAIN main.rs -o x 2>&1 | grep -c '^error\\[E0277\\]') E0277 errors\"\necho \"incremental: $(rustc +TOOLCHAIN -C incremental=inc main.rs -o x 2>&1 | grep -c '^error\\[E0277\\]') E0277 errors\"", + "observe": "Verified locally. On nightly-2026-10-01 (also 1.85.0, 1.92.0, 1.96.1), the non-incremental build prints 1 E0277 error and the incremental build prints 4 identical copies ('aborting due to 4 previous errors'). On nightly-2026-10-03 both print 1. Stable 1.70.0 prints 1 in both modes, so this form of the bug arrived later; #106571's proc-macro form dates from nightly-2023-01-03. Regression test: tests/ui/diagnostic-flags/deduplicate-diagnostics-incr.rs.", + "properties": [ + "P6" + ] + } +] \ No newline at end of file diff --git a/docs/motivating/make_doc.py b/docs/motivating/make_doc.py new file mode 100644 index 0000000..af9c014 --- /dev/null +++ b/docs/motivating/make_doc.py @@ -0,0 +1,141 @@ +#!/usr/bin/env python3 +"""Generate docs/motivating.md from bugs.json and the outputs in out/.""" +import json +from pathlib import Path + +here = Path(__file__).resolve().parent +bugs = json.loads((here / "bugs.json").read_text()) + +STATUS = {b["issue"]: "reproduced" for b in bugs} +STATUS[111227] = ("partly reproduced: `foo.py` is missing from the dep-info before the fix and " + "present after it, but the stale metadata and ICE the issue describes did not " + "appear here") + +# What this run showed, from out//{before,after}.txt. +THIS_RUN = { + 34902: "before: 5 builds, 5 different rlibs (only the metadata member differs); identical with ASLR off. After: 5 identical", + 45841: "before: the .rmeta is rewritten in place (same inode), and 2 of 1,704 concurrent readers hit the issue's ICE in leb128.rs. After: replaced by rename, 0 of 1,712", + 65036: "before: two builds differ at byte 5163, re-exports in a different order. After: identical", + 68149: "before: the consumer opens and records the dependency's in-flight .rlib. After: only the .rmeta files", + 40364: "before: changing the variable leaves the binary printing the old value, and dep-info has no env-dep. After: the new value, and `# env-dep:MIRTH_DEMO=two`", + 82920: "before: the incremental pass-2 binary panics (`left: 2, right: 1`) while a clean build of the same source passes. After: both pass", + 84252: "before: the second incremental build ICEs with unstable fingerprints for `has_global_allocator`. After: both builds succeed", + 66955: "before: after changing the remap, the rlib still holds 2 copies of the old path. After: only the new path", + 89598: "before: the incremental pass-2 binary calls the wrong method (`left: 42, right: 17`); the clean build passes. After: both pass", + 107001: "before: a failed compile leaves `.rcgu.o` files in the output directory. After: none", + 111227: "before: the visualizer file is missing from dep-info. After: listed. The stale metadata did not appear here", + 111295: "before: the rebuilt binary still embeds the old visualizer script. After: the new one", + 117254: "before: a write error is swallowed (exit 0), a 1 MiB truncated .rmeta is published, and a dependent ICEs. After: an error and exit 1, but the truncated file is still published", + 119456: "before: an error and exit 1, yet the truncated .rmeta is published and a dependent ICEs. After: nothing is published, and the dependent gets `can't find crate`", + 122859: "before: E0277, the implied bound lost across crates. After: compiles", + 130201: "before: the dependent ICEs (`coroutine_by_move_body_def_id` unsupported by its crate). After: compiles", + 135514: "before: the incremental rebuild accepts overlapping impls and runs, while a clean build gives E0119. After: both give E0119", + 138678: "before: 8 builds, 8 different .rmeta files. After: 8 identical", + 139407: "before: after a failed build is fixed back, the rebuilt binary still contains the failed session's code and panics. After: it passes", + 139899: "before: each failing doctest leaves a `rustdoctest*` directory (4 left). After: none", + 114669: "before: an unchanged rebuild re-encodes the metadata (new inode). After: reused (same inode, hard-linked to the work product)", + 144004: "before: rustdoc shows neither attribute on the re-exports. After: `no_mangle` and `link_section` shown", + 140413: "before: 15 threaded builds give 2 different binaries. After: 15 identical", + 159677: "before: adding an unrelated library to the search path changes the .rmeta (differs at byte 1256). After: identical", + 150451: "before: 15 threaded builds give 10 different rlibs. After: 15 identical", + 129094: "before: 20 threaded builds give 4 different rlibs. After: 20 identical", + 162901: "before: the incremental build prints the error 4 times, a clean build once. After: once each", +} + +# The fixes replayed by reverting them on the pinned compiler (docs/regressions.md). +MIRTH = { + 122859: "caught when the fix (#122891) is reverted: the list shows the table no longer written", + 130201: "caught when the fix is reverted: the build ICEs", + 138678: "caught when the fix is reverted: P5, P5 with threads, the touch rebuild and P6", + 144004: "missed when the fix (#144050) is reverted: only rustdoc reads these attributes", + 114669: "with #143247 reverted (metadata depending on a node that is never green), the touch rebuild's record catches it", + 129094: "not replayed (the revert does not apply cleanly)", + 159677: "not replayed (needs a decoy crate in the search path)", +} + +PROPS = [ + ("P1", "An `.rmeta` reaches its final path only by a rename of a fully written file", + "A reader that opens a half-written or truncated `.rmeta` misreads it, usually as an ICE " + "in the decoder. rustc writes the metadata to a temporary directory and renames it into " + "place; P1 checks that protocol on every process. The first bug below is where the " + "protocol came from; the other two published a truncated file through the rename."), + ("P2", "No process opens a dependency's `.rmeta` before it has been renamed into place", + "With pipelining, a dependent starts as soon as its dependency's metadata exists, so the " + "order of writes and opens across processes matters. P2 compares the timestamps of opens " + "in readers with renames in writers."), + ("P3", "A dependent reads only table entries the dependency wrote", + "An entry the writer never wrote reads as a default, without complaint, so a table that " + "stops being written shows up as wrong behaviour far away: a missing bound, a missing " + "attribute, or an ICE in the reader. The `written` column of the blessed list records it."), + ("P4", "Encoding reads no untracked state", + "Environment variables, the clock, randomly seeded maps and files that are not declared " + "inputs all reach outputs without incremental compilation or Cargo knowing. P4 records " + "such reads while encoding."), + ("P5", "Two clean builds give the same bytes", + "Nondeterminism in metadata breaks reproducible builds and makes crate hashes, and so " + "everything downstream, depend on chance."), + ("P5t", "Two clean builds with `-Zthreads` give the same bytes", + "The parallel front end adds a new source of nondeterminism: the order in which threads " + "create and intern things."), + ("P6", "An incremental rebuild gives the same results as a clean build", + "Incremental compilation is only correct if what it reuses is what it would have " + "computed. When it is not, the result is a stale output, a miscompilation or a wrong " + "diagnostic, which `cargo clean` fixes. These bugs are why P6 compares an incremental " + "rebuild with a clean build of the same source."), + ("P7", "Nothing is left behind in the output directory", + "Leftover temporary files waste space and can be picked up by later builds."), + ("tracked", "Cross-crate reads are tracked, and metadata is reused when nothing changed", + "Every query that reads another crate's metadata must record a dependency on that crate, " + "or a later session reuses a stale result. The `tracked` column of the blessed list " + "records it per query; the touch-only rebuild records whether metadata was reused."), +] + +out = ["""# The bugs behind each property + +mirth's properties were designed from first principles: what must hold for rustc's +handling of metadata and incremental compilation to be correct. This document records the +real rust-lang/rust bugs that violated each one, so a maintainer can see why a check exists. +Each bug was **reproduced on a toolchain from before its fix and shown fixed on one after**, +with no mirth involved: the original bug, as users met it. + +The bugs were found by searching rust-lang/rust for each property; every issue and PR was +read, and every reproduction run on this machine (Linux x86_64). Reproduce one with +`docs/motivating/run.py `; the outputs used here are in `docs/motivating/out/`. + +Some bugs appear under two properties. Three found by mirth itself +([`hunt/`](hunt)) are listed under P6 at the end. +"""] +for key, title, why in PROPS: + rows = [b for b in bugs if key in b["properties"]] + out.append(f"## {key}: {title}\n\n{why}\n") + out.append("| issue | fixed by | merged | before → after | status | mirth, with the fix reverted |\n|---|---|---|---|---|---|") + for b in rows: + out.append(f"| [#{b['issue']}](https://github.com/rust-lang/rust/issues/{b['issue']}) {b['title'].split(' (')[0][:80]} " + f"| [#{b['fix_pr']}](https://github.com/rust-lang/rust/pull/{b['fix_pr']}) | {b['merged']} " + f"| `{b['before_toolchain']}` → `{b['after_toolchain']}` | {STATUS[b['issue']].split(':')[0]} " + f"| {MIRTH.get(b['issue'], '—')} |") + out.append("") + for b in rows: + out.append(f"**#{b['issue']}.** {b['how_it_violates']}\n") + if STATUS[b["issue"]] != "reproduced": + out.append(f"*Status:* {STATUS[b['issue']]}.\n") + out.append(f"*This run:* {THIS_RUN[b['issue']]}. Outputs: [`before`](motivating/out/{b['issue']}/before.txt), [`after`](motivating/out/{b['issue']}/after.txt).\n") + out.append(f"*Expected, from the issue and the research:* {b['observe']}\n") +out.append("""## Found by mirth (P6) + +| bug | reproduce | cause | since | +|---|---|---|---| +| `Generics::param_def_id_to_index` order changes on each round trip through the incremental cache | `docs/hunt/repro.sh` (p6-generics) | an `FxHashMap` encoded in iteration order | at least 1.95 | +| a string literal encoded twice after an incremental rebuild | `docs/hunt/repro.sh` (p6-literals) | literals deduplicated when created but not when decoded | 1.90 (#116707) | +| the previous session's metadata republished after an edit that moves no span | `docs/hunt/repro.sh` (stale-source) | source map file hashes and lengths not tracked | 1.90 (#114669) | + +Each reproduces on the official nightly and on stable 1.98.1, and has a draft report, a +candidate fix and a regression test in [`hunt/`](hunt). + +## Not covered here + +The 29 properties in [`properties.md`](properties.md) cite the bugs that suggested them, +checked against the issue text, but those bugs have not been reproduced this way. +""") +(here.parent / "motivating.md").write_text("\n".join(out)) +print("written", sum(len(x) for x in out)) diff --git a/docs/motivating/out/107001/after.txt b/docs/motivating/out/107001/after.txt new file mode 100644 index 0000000..be8fb2b --- /dev/null +++ b/docs/motivating/out/107001/after.txt @@ -0,0 +1,16 @@ +$ toolchain nightly-2023-04-22 +exit=1 + +--- stderr --- +error: this operation will panic at runtime + --> code.rs:3:14 + | +3 | let _x = a[9]; + | ^^^^ index out of bounds: the length is 5 but the index is 9 + | + = note: `#[deny(unconditional_panic)]` on by default + +error: aborting due to previous error + + +exit 0 diff --git a/docs/motivating/out/107001/before.txt b/docs/motivating/out/107001/before.txt new file mode 100644 index 0000000..0f1ec64 --- /dev/null +++ b/docs/motivating/out/107001/before.txt @@ -0,0 +1,22 @@ +$ toolchain nightly-2023-04-21 +exit=1 +code.4fkn3jumrthh88oi.rcgu.o +code.code.49928fca7798b8c0-cgu.0.rcgu.o +code.code.49928fca7798b8c0-cgu.1.rcgu.o +code.code.49928fca7798b8c0-cgu.2.rcgu.o +code.code.49928fca7798b8c0-cgu.3.rcgu.o +code.code.49928fca7798b8c0-cgu.4.rcgu.o + +--- stderr --- +error: this operation will panic at runtime + --> code.rs:3:14 + | +3 | let _x = a[9]; + | ^^^^ index out of bounds: the length is 5 but the index is 9 + | + = note: `#[deny(unconditional_panic)]` on by default + +error: aborting due to previous error + + +exit 0 diff --git a/docs/motivating/out/111227/after.txt b/docs/motivating/out/111227/after.txt new file mode 100644 index 0000000..a7ddc2d --- /dev/null +++ b/docs/motivating/out/111227/after.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2023-05-21 +foo.py in dep-info +Natvis v2 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/111227/before.txt b/docs/motivating/out/111227/before.txt new file mode 100644 index 0000000..5c55066 --- /dev/null +++ b/docs/motivating/out/111227/before.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2023-05-18 +foo.py NOT in dep-info +Natvis v2 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/111295/after.txt b/docs/motivating/out/111295/after.txt new file mode 100644 index 0000000..69591f2 --- /dev/null +++ b/docs/motivating/out/111295/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2023-05-21 +print('hello world') + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/111295/before.txt b/docs/motivating/out/111295/before.txt new file mode 100644 index 0000000..10254e1 --- /dev/null +++ b/docs/motivating/out/111295/before.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2023-05-18 +print('hello!') + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/114669/after.txt b/docs/motivating/out/114669/after.txt new file mode 100644 index 0000000..84702af --- /dev/null +++ b/docs/motivating/out/114669/after.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2025-07-06 +run1 inode=68412 links=2 +run2 inode=68412 links=2 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/114669/before.txt b/docs/motivating/out/114669/before.txt new file mode 100644 index 0000000..dfcb919 --- /dev/null +++ b/docs/motivating/out/114669/before.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2025-07-03 +run1 inode=68409 links=1 +run2 inode=68417 links=1 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/117254/after.txt b/docs/motivating/out/117254/after.txt new file mode 100644 index 0000000..c372129 --- /dev/null +++ b/docs/motivating/out/117254/after.txt @@ -0,0 +1,19 @@ +$ toolchain nightly-2023-11-28 + + nightly-2023-11-28-x86_64-unknown-linux-gnu unchanged - rustc 1.76.0-nightly (49b3924bd 2023-11-27) + +rustc exit=1 +total 1024 +-rw-r--r-- 1 exedev exedev 1048576 Oct 7 09:55 libbig.rmeta +thread 'rustc' panicked at /rustc/49b3924bd4a34d3cf9c37b74120fba78d9712ab8/compiler/rustc_serialize/src/opaque.rs:262:42: +range start index 14979596 out of range for slice of length 1048576 + +--- stderr --- +info: syncing channel updates for nightly-2023-11-28-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) +error: failed to write to `/tmp/tmpmgamaf61/out/rmetaPRkzQ4/lib.rmeta`: File too large (os error 27) + +error: aborting due to 1 previous error + + +exit 0 diff --git a/docs/motivating/out/117254/before.txt b/docs/motivating/out/117254/before.txt new file mode 100644 index 0000000..f69bd54 --- /dev/null +++ b/docs/motivating/out/117254/before.txt @@ -0,0 +1,15 @@ +$ toolchain nightly-2023-11-26 + + nightly-2023-11-26-x86_64-unknown-linux-gnu unchanged - rustc 1.76.0-nightly (f5dc2653f 2023-11-25) + +rustc exit=0 +total 1024 +-rw-r--r-- 1 exedev exedev 1048576 Oct 7 09:55 libbig.rmeta +thread 'rustc' panicked at /rustc/f5dc2653fdd8b5d177b2ccbd84057954340a89fc/compiler/rustc_serialize/src/opaque.rs:233:42: +range start index 14979654 out of range for slice of length 1048576 + +--- stderr --- +info: syncing channel updates for nightly-2023-11-26-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) + +exit 0 diff --git a/docs/motivating/out/119456/after.txt b/docs/motivating/out/119456/after.txt new file mode 100644 index 0000000..6329d5e --- /dev/null +++ b/docs/motivating/out/119456/after.txt @@ -0,0 +1,18 @@ +$ toolchain nightly-2024-01-05 + + nightly-2024-01-05-x86_64-unknown-linux-gnu unchanged - rustc 1.77.0-nightly (f688dd684 2024-01-04) + +rustc exit=1 +total 0 +error[E0463]: can't find crate for `big` +error: aborting due to 1 previous error + +--- stderr --- +info: syncing channel updates for nightly-2024-01-05-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) +error: failed to write to `/tmp/tmpuriudou1/out/rmetaSwrGS7/lib.rmeta`: File too large (os error 27) + +error: aborting due to 1 previous error + + +exit 0 diff --git a/docs/motivating/out/119456/before.txt b/docs/motivating/out/119456/before.txt new file mode 100644 index 0000000..bac1862 --- /dev/null +++ b/docs/motivating/out/119456/before.txt @@ -0,0 +1,19 @@ +$ toolchain nightly-2024-01-01 + + nightly-2024-01-01-x86_64-unknown-linux-gnu unchanged - rustc 1.77.0-nightly (e51e98dde 2023-12-31) + +rustc exit=1 +total 1024 +-rw-r--r-- 1 exedev exedev 1048576 Oct 7 09:58 libbig.rmeta +thread 'rustc' panicked at /rustc/e51e98dde6a60637b6a71b8105245b629ac3fe77/compiler/rustc_serialize/src/opaque.rs:262:42: +range start index 12750978 out of range for slice of length 1048576 + +--- stderr --- +info: syncing channel updates for nightly-2024-01-01-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) +error: failed to write to `/tmp/tmp6y9blil5/out/rmeta0Ew0Gj/lib.rmeta`: File too large (os error 27) + +error: aborting due to 1 previous error + + +exit 0 diff --git a/docs/motivating/out/122859/after.txt b/docs/motivating/out/122859/after.txt new file mode 100644 index 0000000..055b834 --- /dev/null +++ b/docs/motivating/out/122859/after.txt @@ -0,0 +1,21 @@ +$ toolchain nightly-2024-03-26 + +--- stderr --- +warning: function `foo` is never used + --> main.rs:2:4 + | +2 | fn foo() { + | ^^^ + | + = note: `#[warn(dead_code)]` on by default + +warning: function `unsatisfied` is never used + --> main.rs:5:4 + | +5 | fn unsatisfied() {} + | ^^^^^^^^^^^ + +warning: 2 warnings emitted + + +exit 0 diff --git a/docs/motivating/out/122859/before.txt b/docs/motivating/out/122859/before.txt new file mode 100644 index 0000000..9c7a01b --- /dev/null +++ b/docs/motivating/out/122859/before.txt @@ -0,0 +1,24 @@ +$ toolchain nightly-2024-03-24 + +--- stderr --- +error[E0277]: the trait bound `<::FooAssoc as Super>::SuperAssoc: Unsatisfied` is not satisfied + --> main.rs:3:19 + | +3 | unsatisfied::<::SuperAssoc>() + | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ the trait `Unsatisfied` is not implemented for `<::FooAssoc as Super>::SuperAssoc` + | +note: required by a bound in `unsatisfied` + --> main.rs:5:19 + | +5 | fn unsatisfied() {} + | ^^^^^^^^^^^ required by this bound in `unsatisfied` +help: consider further restricting the associated type + | +2 | fn foo() where <::FooAssoc as Super>::SuperAssoc: Unsatisfied { + | ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ + +error: aborting due to 1 previous error + +For more information about this error, try `rustc --explain E0277`. + +exit 1 diff --git a/docs/motivating/out/129094/after.txt b/docs/motivating/out/129094/after.txt new file mode 100644 index 0000000..3fe42d0 --- /dev/null +++ b/docs/motivating/out/129094/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2026-09-25 + 20 7a33160ebe06a5d81f95097df5a15c19 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/129094/before.txt b/docs/motivating/out/129094/before.txt new file mode 100644 index 0000000..709f24d --- /dev/null +++ b/docs/motivating/out/129094/before.txt @@ -0,0 +1,9 @@ +$ toolchain nightly-2026-07-20 + 6 2767dded22f4a777db9435725b80762e + 9 3540b5d654bada0e28cc9429ec201f4e + 4 3864815b421379871b31100482f3d25e + 1 f6cfefae36a42b22d720deae4c6b0524 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/130201/after.txt b/docs/motivating/out/130201/after.txt new file mode 100644 index 0000000..844bfa0 --- /dev/null +++ b/docs/motivating/out/130201/after.txt @@ -0,0 +1,5 @@ +$ toolchain nightly-2024-09-19 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/130201/before.txt b/docs/motivating/out/130201/before.txt new file mode 100644 index 0000000..5a0a127 --- /dev/null +++ b/docs/motivating/out/130201/before.txt @@ -0,0 +1,94 @@ +$ toolchain nightly-2024-09-17 + +--- stderr --- +error: internal compiler error: compiler/rustc_middle/src/query/plumbing.rs:664:5: `tcx.coroutine_by_move_body_def_id(DefId(20:6 ~ foreign[87b3]::closure::{closure#0}::{closure#0}))` unsupported by its crate; perhaps the `coroutine_by_move_body_def_id` query was never assigned a provider function + +thread 'rustc' panicked at compiler/rustc_middle/src/query/plumbing.rs:664:5: +Box +stack backtrace: + 0: 0x7f84c127fdaa - ::fmt::he49223be0828c2da + 1: 0x7f84c1a03297 - core::fmt::write::h4fe12f6cb3392adb + 2: 0x7f84c2916433 - std::io::Write::write_fmt::hec0f9e215795ad8e + 3: 0x7f84c127fc02 - std::sys::backtrace::BacktraceLock::print::hb2fbe0e79fb51bb5 + 4: 0x7f84c1282381 - std::panicking::default_hook::{{closure}}::h5458010bed85c320 + 5: 0x7f84c12821b4 - std::panicking::default_hook::h6a18671e790f3216 + 6: 0x7f84c03832ef - std[854462dde0a0a441]::panicking::update_hook::>::{closure#0} + 7: 0x7f84c1282aa8 - std::panicking::rust_panic_with_hook::h1f6a459c24b862a1 + 8: 0x7f84c03bc921 - std[854462dde0a0a441]::panicking::begin_panic::::{closure#0} + 9: 0x7f84c03b0136 - std[854462dde0a0a441]::sys::backtrace::__rust_end_short_backtrace::::{closure#0}, !> + 10: 0x7f84c03ab8a9 - std[854462dde0a0a441]::panicking::begin_panic:: + 11: 0x7f84c03c5c41 - ::emit_producing_guarantee + 12: 0x7f84c09e36b4 - rustc_middle[3a7265ad3fb88dbe]::util::bug::opt_span_bug_fmt::::{closure#0} + 13: 0x7f84c09c987a - rustc_middle[3a7265ad3fb88dbe]::ty::context::tls::with_opt::::{closure#0}, !>::{closure#0} + 14: 0x7f84c09c972b - rustc_middle[3a7265ad3fb88dbe]::ty::context::tls::with_context_opt::::{closure#0}, !>::{closure#0}, !> + 15: 0x7f84be020380 - rustc_middle[3a7265ad3fb88dbe]::util::bug::bug_fmt + 16: 0x7f84c09e85b6 - rustc_middle[3a7265ad3fb88dbe]::query::plumbing::default_extern_query + 17: 0x7f84c09935f4 - <::default::{closure#29} as core[a99a3101f2a6dd0e]::ops::function::FnOnce<(rustc_middle[3a7265ad3fb88dbe]::ty::context::TyCtxt, rustc_span[a7f461c3849a3fed]::def_id::DefId)>>::call_once + 18: 0x7f84c0dba94b - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 19: 0x7f84c1a2f5ee - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::>, false, false, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 20: 0x7f84c0dc842e - rustc_query_impl[9cce4e1893ae50bb]::query_impl::coroutine_by_move_body_def_id::get_query_non_incr::__rust_end_short_backtrace + 21: 0x7f84c222843f - rustc_middle[3a7265ad3fb88dbe]::query::plumbing::query_get_at::>> + 22: 0x7f84c09e3363 - ::coroutine_layout + 23: 0x7f84c2291462 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of_uncached + 24: 0x7f84c228b986 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of + 25: 0x7f84c228b911 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 26: 0x7f84c228ab93 - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>, false, true, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 27: 0x7f84c228a82d - rustc_query_impl[9cce4e1893ae50bb]::query_impl::layout_of::get_query_non_incr::__rust_end_short_backtrace + 28: 0x7f84c228959d - , rustc_ty_utils[696df349ec243f2a]::layout::layout_of_uncached::{closure#13}>>, core[a99a3101f2a6dd0e]::result::Result> as core[a99a3101f2a6dd0e]::iter::traits::iterator::Iterator>::next + 29: 0x7f84c228d596 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of_uncached + 30: 0x7f84c228b986 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of + 31: 0x7f84c228b911 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 32: 0x7f84c228ab93 - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>, false, true, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 33: 0x7f84c228a82d - rustc_query_impl[9cce4e1893ae50bb]::query_impl::layout_of::get_query_non_incr::__rust_end_short_backtrace + 34: 0x7f84c2289997 - , rustc_ty_utils[696df349ec243f2a]::layout::layout_of_uncached::{closure#13}>>, core[a99a3101f2a6dd0e]::result::Result> as core[a99a3101f2a6dd0e]::iter::traits::iterator::Iterator>::next + 35: 0x7f84c228d596 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of_uncached + 36: 0x7f84c228b986 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of + 37: 0x7f84c228b911 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 38: 0x7f84c228ab93 - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>, false, true, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 39: 0x7f84c228a82d - rustc_query_impl[9cce4e1893ae50bb]::query_impl::layout_of::get_query_non_incr::__rust_end_short_backtrace + 40: 0x7f84c2288b9d - rustc_middle[3a7265ad3fb88dbe]::query::plumbing::query_get_at::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>> + 41: 0x7f84c228bd2f - rustc_ty_utils[696df349ec243f2a]::layout::layout_of + 42: 0x7f84c228b911 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 43: 0x7f84c228ab93 - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>, false, true, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 44: 0x7f84c228a82d - rustc_query_impl[9cce4e1893ae50bb]::query_impl::layout_of::get_query_non_incr::__rust_end_short_backtrace + 45: 0x7f84c2288b9d - rustc_middle[3a7265ad3fb88dbe]::query::plumbing::query_get_at::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>> + 46: 0x7f84c11fde2a - as rustc_middle[3a7265ad3fb88dbe]::ty::layout::LayoutOf>::spanned_layout_of + 47: 0x7f84c11ea1c6 - >, rustc_ty_utils[696df349ec243f2a]::layout::coroutine_layout::{closure#2}>, core[a99a3101f2a6dd0e]::iter::sources::once::Once>>, core[a99a3101f2a6dd0e]::iter::adapters::map::Map, rustc_ty_utils[696df349ec243f2a]::layout::coroutine_layout::{closure#1}>>>, core[a99a3101f2a6dd0e]::result::Result> as core[a99a3101f2a6dd0e]::iter::traits::iterator::Iterator>::next + 48: 0x7f84c2292fe0 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of_uncached + 49: 0x7f84c228b986 - rustc_ty_utils[696df349ec243f2a]::layout::layout_of + 50: 0x7f84c228b911 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 51: 0x7f84c228ab93 - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::, rustc_middle[3a7265ad3fb88dbe]::query::erase::Erased<[u8; 16usize]>>, false, true, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 52: 0x7f84c228a82d - rustc_query_impl[9cce4e1893ae50bb]::query_impl::layout_of::get_query_non_incr::__rust_end_short_backtrace + 53: 0x7f84bf1207fd - ::run_lint + 54: 0x7f84c1a0894c - rustc_mir_transform[528fb1b057dfa389]::run_analysis_to_runtime_passes + 55: 0x7f84c1ed6a92 - rustc_mir_transform[528fb1b057dfa389]::mir_drops_elaborated_and_const_checked + 56: 0x7f84c1ed63d5 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 57: 0x7f84c1ece57d - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::>, false, false, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 58: 0x7f84c1ecdf25 - rustc_query_impl[9cce4e1893ae50bb]::query_impl::mir_drops_elaborated_and_const_checked::get_query_non_incr::__rust_end_short_backtrace + 59: 0x7f84c1ec7f8c - rustc_interface[32c659705ba33f15]::passes::run_required_analyses + 60: 0x7f84c272251e - rustc_interface[32c659705ba33f15]::passes::analysis + 61: 0x7f84c27224f1 - rustc_query_impl[9cce4e1893ae50bb]::plumbing::__rust_begin_short_backtrace::> + 62: 0x7f84c28a962e - rustc_query_system[53985dc8c6f0f2ea]::query::plumbing::try_execute_query::>, false, false, false>, rustc_query_impl[9cce4e1893ae50bb]::plumbing::QueryCtxt, false> + 63: 0x7f84c28a938f - rustc_query_impl[9cce4e1893ae50bb]::query_impl::analysis::get_query_non_incr::__rust_end_short_backtrace + 64: 0x7f84c270c9fc - rustc_interface[32c659705ba33f15]::interface::run_compiler::, rustc_driver_impl[ba21f6c8d520bf6e]::run_compiler::{closure#0}>::{closure#1} + 65: 0x7f84c2852c10 - std[854462dde0a0a441]::sys::backtrace::__rust_begin_short_backtrace::, rustc_driver_impl[ba21f6c8d520bf6e]::run_compiler::{closure#0}>::{closure#1}, core[a99a3101f2a6dd0e]::result::Result<(), rustc_span[a7f461c3849a3fed]::ErrorGuaranteed>>::{closure#0}, core[a99a3101f2a6dd0e]::result::Result<(), rustc_span[a7f461c3849a3fed]::ErrorGuaranteed>>::{closure#0}::{closure#0}, core[a99a3101f2a6dd0e]::result::Result<(), rustc_span[a7f461c3849a3fed]::ErrorGuaranteed>> + 66: 0x7f84c285327a - <::spawn_unchecked_, rustc_driver_impl[ba21f6c8d520bf6e]::run_compiler::{closure#0}>::{closure#1}, core[a99a3101f2a6dd0e]::result::Result<(), rustc_span[a7f461c3849a3fed]::ErrorGuaranteed>>::{closure#0}, core[a99a3101f2a6dd0e]::result::Result<(), rustc_span[a7f461c3849a3fed]::ErrorGuaranteed>>::{closure#0}::{closure#0}, core[a99a3101f2a6dd0e]::result::Result<(), rustc_span[a7f461c3849a3fed]::ErrorGuaranteed>>::{closure#1} as core[a99a3101f2a6dd0e]::ops::function::FnOnce<()>>::call_once::{shim:vtable#0} + 67: 0x7f84c285366b - std::sys::pal::unix::thread::Thread::new::thread_start::h01ba1847dfc214af + 68: 0x7f84bcc8ab84 - + 69: 0x7f84bcd17d6c - + 70: 0x0 - + +note: we would appreciate a bug report: https://github.com/rust-lang/rust/issues/new?labels=C-bug%2C+I-ICE%2C+T-compiler&template=ice.md + +note: please make sure that you have updated to the latest nightly + +note: please attach the file at `/tmp/tmp707dxio7/rustc-ice-2026-10-07T09_55_12-2414492.txt` to your bug report + +query stack during panic: +#0 [coroutine_by_move_body_def_id] looking up the coroutine by-move body for `foreign::closure::{closure#0}::{closure#0}` +#1 [layout_of] computing layout of `{async closure body@foreign::closure::{closure#0}::{closure#0}}` +end of query stack +error: aborting due to 1 previous error + + +exit 101 diff --git a/docs/motivating/out/135514/after.txt b/docs/motivating/out/135514/after.txt new file mode 100644 index 0000000..008806d --- /dev/null +++ b/docs/motivating/out/135514/after.txt @@ -0,0 +1,30 @@ +$ toolchain 1.85.0 +pass1 + +--- stderr --- +error[E0119]: conflicting implementations of trait `Other` for type `S` + --> main.rs:20:1 + | +15 | impl Other for T { + | -------------------------- first implementation here +... +20 | impl Other for S { + | ^^^^^^^^^^^^^^^^^^^^^^ conflicting implementation for `S` + +error: aborting due to 1 previous error + +For more information about this error, try `rustc --explain E0119`. +error[E0119]: conflicting implementations of trait `Other` for type `S` + --> main.rs:20:1 + | +15 | impl Other for T { + | -------------------------- first implementation here +... +20 | impl Other for S { + | ^^^^^^^^^^^^^^^^^^^^^^ conflicting implementation for `S` + +error: aborting due to 1 previous error + +For more information about this error, try `rustc --explain E0119`. + +exit 1 diff --git a/docs/motivating/out/135514/before.txt b/docs/motivating/out/135514/before.txt new file mode 100644 index 0000000..31f9b46 --- /dev/null +++ b/docs/motivating/out/135514/before.txt @@ -0,0 +1,20 @@ +$ toolchain 1.84.0 +pass1 +ABC +incremental pass2 built and ran + +--- stderr --- +error[E0119]: conflicting implementations of trait `Other` for type `S` + --> main.rs:20:1 + | +15 | impl Other for T { + | -------------------------- first implementation here +... +20 | impl Other for S { + | ^^^^^^^^^^^^^^^^^^^^^^ conflicting implementation for `S` + +error: aborting due to 1 previous error + +For more information about this error, try `rustc --explain E0119`. + +exit 1 diff --git a/docs/motivating/out/138678/after.txt b/docs/motivating/out/138678/after.txt new file mode 100644 index 0000000..bd75fc7 --- /dev/null +++ b/docs/motivating/out/138678/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2025-03-30 + 8 c01cd55bf6f3e169 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/138678/before.txt b/docs/motivating/out/138678/before.txt new file mode 100644 index 0000000..63e1356 --- /dev/null +++ b/docs/motivating/out/138678/before.txt @@ -0,0 +1,13 @@ +$ toolchain nightly-2025-03-28 + 1 5222c28142da99a8 + 1 5717d08526356af8 + 1 675e29a6aba0e847 + 1 7bbc9ca532eb4f24 + 1 8382fc8319ec3359 + 1 b0f39be99129acc7 + 1 cbaa0f7373d94da1 + 1 e52b42802cf07a85 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/139407/after.txt b/docs/motivating/out/139407/after.txt new file mode 100644 index 0000000..9ff0b55 --- /dev/null +++ b/docs/motivating/out/139407/after.txt @@ -0,0 +1,22 @@ +$ toolchain nightly-2025-04-13 +pass1 ok +(pass2 is expected to fail with an assembler error) +pass3 ok + +--- stderr --- +error: invalid instruction mnemonic 'missing' + --> main.rs:28:30 + | +28 | std::arch::asm!("missing"); + | ^^^^^^^ + | +note: instantiated into assembly here + --> :2:2 + | +2 | missing + | ^^^^^^^ + +error: aborting due to 1 previous error + + +exit 0 diff --git a/docs/motivating/out/139407/before.txt b/docs/motivating/out/139407/before.txt new file mode 100644 index 0000000..15f6f0e --- /dev/null +++ b/docs/motivating/out/139407/before.txt @@ -0,0 +1,27 @@ +$ toolchain nightly-2025-04-10 +pass1 ok +(pass2 is expected to fail with an assembler error) + +--- stderr --- +error: invalid instruction mnemonic 'missing' + --> main.rs:28:30 + | +28 | std::arch::asm!("missing"); + | ^^^^^^^ + | +note: instantiated into assembly here + --> :2:2 + | +2 | missing + | ^^^^^^^ + +error: aborting due to 1 previous error + + +thread 'main' panicked at main.rs:14:5: +assertion `left == right` failed + left: 1 + right: 0 +note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace + +exit 101 diff --git a/docs/motivating/out/139899/after.txt b/docs/motivating/out/139899/after.txt new file mode 100644 index 0000000..867bbad --- /dev/null +++ b/docs/motivating/out/139899/after.txt @@ -0,0 +1,9 @@ +$ toolchain nightly-2025-05-09 +compile-error.rs 2018 exit=101 +compile-error.rs 2024 exit=101 +run-error.rs 2018 exit=101 +run-error.rs 2024 exit=101 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/139899/before.txt b/docs/motivating/out/139899/before.txt new file mode 100644 index 0000000..8b3187b --- /dev/null +++ b/docs/motivating/out/139899/before.txt @@ -0,0 +1,13 @@ +$ toolchain nightly-2025-05-07 +compile-error.rs 2018 exit=101 +compile-error.rs 2024 exit=101 +run-error.rs 2018 exit=101 +run-error.rs 2024 exit=101 +rustdoctest9b2TYz +rustdoctestdZtosn +rustdoctesteK4t1v +rustdoctesttT11Mx + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/140413/after.txt b/docs/motivating/out/140413/after.txt new file mode 100644 index 0000000..9c115b5 --- /dev/null +++ b/docs/motivating/out/140413/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2025-12-10 + 15 89679f9d1e5abdffc5e8140137da3a64 - + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/140413/before.txt b/docs/motivating/out/140413/before.txt new file mode 100644 index 0000000..0097de2 --- /dev/null +++ b/docs/motivating/out/140413/before.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2025-07-26 + 10 099ddaff6232a5431a603199ddd3dfab - + 5 4f4824253541548e77af66ec004bcf13 - + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/144004/after.txt b/docs/motivating/out/144004/after.txt new file mode 100644 index 0000000..68d0239 --- /dev/null +++ b/docs/motivating/out/144004/after.txt @@ -0,0 +1,9 @@ +$ toolchain nightly-2025-07-24 +fn.f0: no_mangle +fn.f1: link_section = ".here"] +static.S0: no_mangle +static.S1: link_section = ".there"] + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/144004/before.txt b/docs/motivating/out/144004/before.txt new file mode 100644 index 0000000..b76b2f8 --- /dev/null +++ b/docs/motivating/out/144004/before.txt @@ -0,0 +1,9 @@ +$ toolchain nightly-2025-07-06 +fn.f0: +fn.f1: +static.S0: +static.S1: + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/150451/after.txt b/docs/motivating/out/150451/after.txt new file mode 100644 index 0000000..5b4f29f --- /dev/null +++ b/docs/motivating/out/150451/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2026-09-25 + 15 4c12e7fe4e3294ab66037611a6747a8c - + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/150451/before.txt b/docs/motivating/out/150451/before.txt new file mode 100644 index 0000000..720f040 --- /dev/null +++ b/docs/motivating/out/150451/before.txt @@ -0,0 +1,15 @@ +$ toolchain nightly-2026-07-20 + 4 1fca635fe5eb6613ae1fa80b0b457272 - + 1 28e2f25e12c2619c1192cd4a339bf7fb - + 1 2d725850d2881f27cc8e0b48d9e5bb1b - + 1 694c765ce32f960a21f34bb6beecde6e - + 1 7c7f5d7d1cf15be16e7b98af752c6cb8 - + 1 7f14e8090a992bc283a68e9e969efb1f - + 2 8d6935e039d07693fc8e7cd93a4108ff - + 1 a0f7ad17dd6b7bed3681e8f6bda2f6b9 - + 1 b79d3cb8f329b1d875a7a63c87ccdc3a - + 2 ff88983d7dc4c62ed3f33218fab43246 - + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/159677/after.txt b/docs/motivating/out/159677/after.txt new file mode 100644 index 0000000..07941d4 --- /dev/null +++ b/docs/motivating/out/159677/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2026-09-25 +IDENTICAL + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/159677/before.txt b/docs/motivating/out/159677/before.txt new file mode 100644 index 0000000..0a0a088 --- /dev/null +++ b/docs/motivating/out/159677/before.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2026-07-20 +client1.rmeta client2.rmeta differ: byte 1256, line 7 + +--- stderr --- + +exit 1 diff --git a/docs/motivating/out/162901/after.txt b/docs/motivating/out/162901/after.txt new file mode 100644 index 0000000..df8f1e7 --- /dev/null +++ b/docs/motivating/out/162901/after.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2026-10-03 +non-incremental: 1 E0277 errors +incremental: 1 E0277 errors + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/162901/before.txt b/docs/motivating/out/162901/before.txt new file mode 100644 index 0000000..02c765e --- /dev/null +++ b/docs/motivating/out/162901/before.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2026-10-01 +non-incremental: 1 E0277 errors +incremental: 4 E0277 errors + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/34902/after.txt b/docs/motivating/out/34902/after.txt new file mode 100644 index 0000000..a285bee --- /dev/null +++ b/docs/motivating/out/34902/after.txt @@ -0,0 +1,16 @@ +$ toolchain nightly-2016-08-30 +50e58a5d263a4b4c036eb475d954c0e1 o1/liblib.rlib +50e58a5d263a4b4c036eb475d954c0e1 o2/liblib.rlib +50e58a5d263a4b4c036eb475d954c0e1 o3/liblib.rlib +50e58a5d263a4b4c036eb475d954c0e1 o4/liblib.rlib +50e58a5d263a4b4c036eb475d954c0e1 o5/liblib.rlib +d9c31955006a49def3e7d07bcd165436 o1/rust.metadata.bin +d9c31955006a49def3e7d07bcd165436 o2/rust.metadata.bin +5c51b89a636d171a66f665172438427c o1/lib.0.o +5c51b89a636d171a66f665172438427c o2/lib.0.o +50e58a5d263a4b4c036eb475d954c0e1 r1/liblib.rlib +50e58a5d263a4b4c036eb475d954c0e1 r2/liblib.rlib + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/34902/before.txt b/docs/motivating/out/34902/before.txt new file mode 100644 index 0000000..f5af88f --- /dev/null +++ b/docs/motivating/out/34902/before.txt @@ -0,0 +1,16 @@ +$ toolchain nightly-2016-08-27 +56ae7d365435e777da2d5020a7954786 o1/liblib.rlib +118877e4c5977cf74fe097d8524fee12 o2/liblib.rlib +146874bb536b97b105f605aa83541f39 o3/liblib.rlib +25ac068a53158a16267463cde8c7bf1c o4/liblib.rlib +3a2a27963189a0b74d31d32158b868e5 o5/liblib.rlib +e405b0329b0b7032c4e50a43f2914ca9 o1/rust.metadata.bin +9c2955556242bfadc34f6cce43b99b19 o2/rust.metadata.bin +5c51b89a636d171a66f665172438427c o1/lib.0.o +5c51b89a636d171a66f665172438427c o2/lib.0.o +85b98cc729c8229ec4964252bcb08688 r1/liblib.rlib +85b98cc729c8229ec4964252bcb08688 r2/liblib.rlib + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/40364/after.txt b/docs/motivating/out/40364/after.txt new file mode 100644 index 0000000..47317e5 --- /dev/null +++ b/docs/motivating/out/40364/after.txt @@ -0,0 +1,8 @@ +$ toolchain 1.46.0 +one +two +# env-dep:MIRTH_DEMO=two + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/40364/before.txt b/docs/motivating/out/40364/before.txt new file mode 100644 index 0000000..fbd58df --- /dev/null +++ b/docs/motivating/out/40364/before.txt @@ -0,0 +1,8 @@ +$ toolchain 1.45.0 +one +one +no env-dep in dep-info + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/45841/after.txt b/docs/motivating/out/45841/after.txt new file mode 100644 index 0000000..054c9d1 --- /dev/null +++ b/docs/motivating/out/45841/after.txt @@ -0,0 +1,14 @@ +$ toolchain nightly-2017-11-20 + + nightly-2017-11-20-x86_64-unknown-linux-gnu unchanged - rustc 1.23.0-nightly (5041b3bb3 2017-11-19) + +inode before=68362 after=68359 +OK: liba.rmeta replaced by rename +OK: old hard link still holds previous metadata +nightly-2017-11-20: reader runs=1712 failed=0 + +--- stderr --- +info: syncing channel updates for nightly-2017-11-20-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) + +exit 0 diff --git a/docs/motivating/out/45841/before.txt b/docs/motivating/out/45841/before.txt new file mode 100644 index 0000000..c70565d --- /dev/null +++ b/docs/motivating/out/45841/before.txt @@ -0,0 +1,22 @@ +$ toolchain nightly-2017-11-17 + + nightly-2017-11-17-x86_64-unknown-linux-gnu unchanged - rustc 1.23.0-nightly (d0f8e2913 2017-11-16) + +inode before=68388 after=68388 +BUG: liba.rmeta rewritten in place +BUG: old hard link sees new bytes +nightly-2017-11-17: reader runs=1704 failed=2 +--- fail-8-25.err +error: internal compiler error: unexpected panic +note: the compiler unexpectedly panicked. this is a bug. +thread 'rustc' panicked at 'index out of bounds: the len is 2088928 but the index is 6805836', /checkout/src/libserialize/leb128.rs:59:20 +--- fail-9-76.err +error: internal compiler error: unexpected panic +note: the compiler unexpectedly panicked. this is a bug. +thread 'rustc' panicked at 'index out of bounds: the len is 4186080 but the index is 6805836', /checkout/src/libserialize/leb128.rs:59:20 + +--- stderr --- +info: syncing channel updates for nightly-2017-11-17-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) + +exit 0 diff --git a/docs/motivating/out/65036/after.txt b/docs/motivating/out/65036/after.txt new file mode 100644 index 0000000..e1868db --- /dev/null +++ b/docs/motivating/out/65036/after.txt @@ -0,0 +1,6 @@ +$ toolchain nightly-2019-10-08 +IDENTICAL + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/65036/before.txt b/docs/motivating/out/65036/before.txt new file mode 100644 index 0000000..cc63dbb --- /dev/null +++ b/docs/motivating/out/65036/before.txt @@ -0,0 +1,16 @@ +$ toolchain nightly-2019-10-05 +client1.rmeta client2.rmeta differ: byte 5163, line 16 +107,108d106 +< item_21 +< item_1 +110,118d107 +< item_24 +< item_11 +< item_31 +< item_37 +< item_40 +< item_28 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/66955/after.txt b/docs/motivating/out/66955/after.txt new file mode 100644 index 0000000..3828fe6 --- /dev/null +++ b/docs/motivating/out/66955/after.txt @@ -0,0 +1,9 @@ +$ toolchain nightly-2021-05-01 +after remap to /AAAA_first, rlib contains: + 3 /AAAA_first +after remap to /BBBB_second, rlib contains: + 3 /BBBB_second + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/66955/before.txt b/docs/motivating/out/66955/before.txt new file mode 100644 index 0000000..30f8ae0 --- /dev/null +++ b/docs/motivating/out/66955/before.txt @@ -0,0 +1,10 @@ +$ toolchain nightly-2021-04-29 +after remap to /AAAA_first, rlib contains: + 3 /AAAA_first +after remap to /BBBB_second, rlib contains: + 2 /AAAA_first + 1 /BBBB_second + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/68149/after.txt b/docs/motivating/out/68149/after.txt new file mode 100644 index 0000000..f7e885c --- /dev/null +++ b/docs/motivating/out/68149/after.txt @@ -0,0 +1,15 @@ +$ toolchain nightly-2020-01-24 + + nightly-2020-01-24-x86_64-unknown-linux-gnu unchanged - rustc 1.42.0-nightly (41f41b235 2020-01-23) + +== nightly-2020-01-24: deps recorded by c: +/tmp/tmpy6sh50ae/D/liba-x.rmeta +/tmp/tmpy6sh50ae/D/libb-x.rmeta + 1 D/liba-x.rmeta + 1 D/libb-x.rmeta + +--- stderr --- +info: syncing channel updates for nightly-2020-01-24-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) + +exit 0 diff --git a/docs/motivating/out/68149/before.txt b/docs/motivating/out/68149/before.txt new file mode 100644 index 0000000..35044e3 --- /dev/null +++ b/docs/motivating/out/68149/before.txt @@ -0,0 +1,16 @@ +$ toolchain nightly-2020-01-22 + + nightly-2020-01-22-x86_64-unknown-linux-gnu unchanged - rustc 1.42.0-nightly (5e8897b7b 2020-01-21) + +== nightly-2020-01-22: deps recorded by c: +/tmp/tmpxseyosna/D/liba-x.rlib +/tmp/tmpxseyosna/D/liba-x.rmeta +/tmp/tmpxseyosna/D/libb-x.rmeta + 1 D/liba-x.rlib + 1 D/libb-x.rmeta + +--- stderr --- +info: syncing channel updates for nightly-2020-01-22-x86_64-unknown-linux-gnu +info: checking for self-update (current version: 1.29.1) + +exit 0 diff --git a/docs/motivating/out/82920/after.txt b/docs/motivating/out/82920/after.txt new file mode 100644 index 0000000..24b842f --- /dev/null +++ b/docs/motivating/out/82920/after.txt @@ -0,0 +1,8 @@ +$ toolchain nightly-2021-03-17 +pass1 ok +pass2 ok +clean ok + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/82920/before.txt b/docs/motivating/out/82920/before.txt new file mode 100644 index 0000000..b4078b4 --- /dev/null +++ b/docs/motivating/out/82920/before.txt @@ -0,0 +1,11 @@ +$ toolchain nightly-2021-03-13 +pass1 ok +clean ok + +--- stderr --- +thread 'main' panicked at 'assertion failed: `(left == right)` + left: `2`, + right: `1`', main.rs:16:5 +note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace + +exit 0 diff --git a/docs/motivating/out/84252/after.txt b/docs/motivating/out/84252/after.txt new file mode 100644 index 0000000..d294288 --- /dev/null +++ b/docs/motivating/out/84252/after.txt @@ -0,0 +1,7 @@ +$ toolchain nightly-2021-04-19 +rc1=0 +rc2=0 + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/84252/before.txt b/docs/motivating/out/84252/before.txt new file mode 100644 index 0000000..57781f7 --- /dev/null +++ b/docs/motivating/out/84252/before.txt @@ -0,0 +1,25 @@ +$ toolchain nightly-2021-04-16 +rc1=0 +rc2=101 + +--- stderr --- +thread 'rustc' panicked at 'assertion failed: `(left == right)` + left: `Some(Fingerprint(13307939394407243476, 123911068766761009))`, + right: `Some(Fingerprint(15782864888164328018, 3268950574376745062))`: found unstable fingerprints for has_global_allocator(lib[8787]): false', /rustc/7af1f55ae359e731c2c303f5d98e42a1a8163af0/compiler/rustc_query_system/src/query/plumbing.rs:593:5 +note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace + +error: internal compiler error: unexpected panic + +note: the compiler unexpectedly panicked. this is a bug. + +note: we would appreciate a bug report: https://github.com/rust-lang/rust/issues/new?labels=C-bug%2C+I-ICE%2C+T-compiler&template=ice.md + +note: rustc 1.53.0-nightly (7af1f55ae 2021-04-15) running on x86_64-unknown-linux-gnu + +note: compiler flags: -C incremental + +query stack during panic: +#0 [has_global_allocator] checking if the crate has_global_allocator +end of query stack + +exit 0 diff --git a/docs/motivating/out/89598/after.txt b/docs/motivating/out/89598/after.txt new file mode 100644 index 0000000..ea6fe1b --- /dev/null +++ b/docs/motivating/out/89598/after.txt @@ -0,0 +1,8 @@ +$ toolchain nightly-2021-10-10 +pass1 ok +pass2 ok +clean ok + +--- stderr --- + +exit 0 diff --git a/docs/motivating/out/89598/before.txt b/docs/motivating/out/89598/before.txt new file mode 100644 index 0000000..6d8457e --- /dev/null +++ b/docs/motivating/out/89598/before.txt @@ -0,0 +1,11 @@ +$ toolchain nightly-2021-10-07 +pass1 ok +clean ok + +--- stderr --- +thread 'main' panicked at 'assertion failed: `(left == right)` + left: `42`, + right: `17`', main.rs:18:5 +note: run with `RUST_BACKTRACE=1` environment variable to display a backtrace + +exit 0 diff --git a/docs/motivating/research.json b/docs/motivating/research.json new file mode 100644 index 0000000..51c694b --- /dev/null +++ b/docs/motivating/research.json @@ -0,0 +1,710 @@ +[ + { + "key": "P1", + "bugs": [ + { + "issue": 45841, + "title": "ICE: index out of bounds in libserialize/leb128.rs (rustc wrote .rmeta in place; a concurrent rustc read the half-written file)", + "fix_pr": 45899, + "merged": "2017-11-18", + "how_it_violates": "Before the fix, rustc_trans::back::link::emit_metadata did File::create(out_filename) + write_all straight into the final lib.rmeta path (O_WRONLY|O_CREAT|O_TRUNC). Another rustc that was searching the same -L directory could open the file while it was truncated or only partly written, and then panicked decoding it. The issue thread (arielb1: \"the compiler writing metadata in parts, so that another instance of the compiler can read metadata while it is being partially written to. Should be fixable by doing an atomic rename\") and PR #45899 (\"atomically write .rmeta outputs to avoid races ... write a temporary file and then rename it\") record exactly this. The fix added the rmeta* tempdir inside the output directory plus fs::rename, and the comment \"To avoid races with another rustc process scanning the output directory...\" is still in compiler/rustc_metadata/src/fs.rs today. This is the bug that P1 itself comes from.", + "before_toolchain": "nightly-2017-11-17", + "after_toolchain": "nightly-2017-11-20", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "#![crate_type = \"lib\"]\npub fn f() -> u32 { 1 }\n" + }, + { + "path": "run.sh", + "content": "#!/bin/sh\n# Deterministic check: is liba.rmeta overwritten in place (same inode) or replaced by rename?\n# usage: sh run.sh TOOLCHAIN (run in a fresh dir containing a.rs; appends to a.rs)\nset -e\nTC=$1; D=out-$TC; mkdir $D\nrustc +$TC --crate-type lib --emit=metadata a.rs --out-dir $D\nln $D/liba.rmeta $D/snapshot.rmeta\nino_before=$(stat -c %i $D/liba.rmeta)\nprintf 'pub fn g() {}\\n' >> a.rs\nrustc +$TC --crate-type lib --emit=metadata a.rs --out-dir $D\nino_after=$(stat -c %i $D/liba.rmeta)\necho \"inode before=$ino_before after=$ino_after\"\nif [ \"$ino_before\" = \"$ino_after\" ]; then echo \"BUG: liba.rmeta rewritten in place\"; else echo \"OK: liba.rmeta replaced by rename\"; fi\nif cmp -s $D/liba.rmeta $D/snapshot.rmeta; then echo \"BUG: old hard link sees new bytes\"; else echo \"OK: old hard link still holds previous metadata\"; fi\n" + }, + { + "path": "race.sh", + "content": "#!/bin/bash\n# Optional: shows the actual symptom (torn read -> ICE). A rare race (about 1 in 1000-3000 reader runs here).\n# usage: bash race.sh TOOLCHAIN [SECONDS] [READERS]\nTC=$1; T=${2:-60}; R=${3:-12}; D=race-$TC; mkdir -p $D\npython3 -c 'print(\"#![crate_type=\\\"lib\\\"]\"); [print(f\"pub fn f{i}(x: u64) -> u64 {{ x.wrapping_mul({i}) }}\\npub struct S{i} {{ pub a: u64, pub b: String }}\") for i in range(20000)]' > big.rs\necho 'extern crate big;' > user.rs\nrustc +$TC --crate-type lib --emit=metadata big.rs --out-dir $D\nend=$((SECONDS+T))\n( while [ $SECONDS -lt $end ]; do rustc +$TC --emit=metadata big.rs --out-dir $D 2>/dev/null; done ) &\nfor r in $(seq $R); do\n ( mkdir -p $D/u$r; n=0; while [ $SECONDS -lt $end ]; do n=$((n+1));\n rustc +$TC --crate-type lib --emit=metadata user.rs -L $D --out-dir $D/u$r > $D/u$r/err 2>&1 || cp $D/u$r/err $D/fail-$r-$n.err; done; echo $n > $D/u$r/count ) &\ndone\nwait\ntotal=$(cat $D/u*/count | paste -sd+ | bc); fails=$(ls $D | grep -c '^fail-')\necho \"$TC: reader runs=$total failed=$fails\"\nfor f in $(ls $D | grep '^fail-' | head -3); do echo \"--- $f\"; grep -m3 -E 'error|panicked' $D/$f; done\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\nsh run.sh TOOLCHAIN\n# syscall-level view (needs strace):\nstrace -f -e trace=openat,rename rustc +TOOLCHAIN --crate-type lib --emit=metadata a.rs --crate-name a --out-dir st 2>&1 | grep rmeta\n# optional, probabilistic symptom reproduction (run more than once):\nbash race.sh TOOLCHAIN 90", + "observe": "nightly-2017-11-17 (rustc 1.23.0-nightly d0f8e2913 2017-11-16): run.sh prints the same inode before and after the rebuild, then \"BUG: liba.rmeta rewritten in place\" and \"BUG: old hard link sees new bytes\". strace shows openat(\"st/liba.rmeta\", O_WRONLY|O_CREAT|O_TRUNC) with no rename. nightly-2017-11-20 (5041b3bb3 2017-11-19): the inode changes, then \"OK: liba.rmeta replaced by rename\" and \"OK: old hard link still holds previous metadata\". strace shows the write going to st/rmeta.XXXX/rust.metadata.bin, followed by rename(... , \"st/liba.rmeta\"). race.sh on 2017-11-17 gave 1 failure in 1078 reader runs in one 60s attempt and 0 in another 90s attempt. The failure was the exact ICE from the issue: \"thread 'rustc' panicked at 'index out of bounds: the len is 2432992 but the index is 6805837', /checkout/src/libserialize/leb128.rs:59:20\". On 2017-11-20: 0 failures in 1719 runs." + }, + { + "issue": 117254, + "title": "FileEncoder delayed error reporting is still broken (rmeta write errors swallowed; truncated .rmeta published with exit 0)", + "fix_pr": 117301, + "merged": "2023-11-26", + "how_it_violates": "rmeta encoding never called FileEncoder::finish, so an I/O error while writing the temp file (ENOSPC in crater, EFBIG here) was swallowed. rustc then renamed the short temp file to the final lib.rmeta and exited 0. Downstream crates ICE'd in MemDecoder (\"We're going ahead to decode a result which was not completely written out\", issue text). The rename itself is atomic, but the file it publishes is not fully written, which breaks the second half of P1. Introduced by the delayed error scheme of #94732 and not fixed by #115542. #117301 added the finish() check, but only with emit_err; see the next entry for the remaining hole.", + "before_toolchain": "nightly-2023-11-26", + "after_toolchain": "nightly-2023-11-28", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.py", + "content": "print('#![crate_type=\"lib\"]')\nfor i in range(20000):\n print(f'pub fn f{i}(x: u64) -> u64 {{ x.wrapping_mul({i}) }}')\n print(f'pub struct S{i} {{ pub a: u64, pub b: String }}')\n" + }, + { + "path": "user.rs", + "content": "extern crate big;\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\npython3 gen.py > big.rs # ~6.8 MB of metadata\nmkdir out\n# Cap file size at 1 MiB and ignore SIGXFSZ so write() fails with EFBIG (a cheap stand-in for a full disk, no root needed)\nbash -c \"trap '' XFSZ; ulimit -f 1024; rustc +TOOLCHAIN --crate-type lib --emit=metadata big.rs --out-dir out; echo rustc exit=\\$?\"\nls -l out\nrustc +TOOLCHAIN --crate-type lib --emit=metadata user.rs -L out -o u.rmeta 2>&1 | grep -m2 -E 'panicked|range|error'", + "observe": "nightly-2023-11-26 (1.76.0-nightly f5dc2653f 2023-11-25): no diagnostic, \"rustc exit=0\", and out/libbig.rmeta is exactly 1048576 bytes, a truncated file at the final path. The downstream rustc then ICEs in rustc_serialize/src/opaque.rs (\"range start index ... out of range for slice of length 1048576\"). nightly-2023-11-28 (49b3924bd 2023-11-27): rustc now prints \"error: failed to write to `.../rmetaXXXX/lib.rmeta`: File too large (os error 27)\" and exits 1. However, the 1048576-byte out/libbig.rmeta is still published, and the downstream rustc still panics with \"range start index 14979595 out of range for slice of length 1048576\" (fixed fully by #119510 below)." + }, + { + "issue": 119456, + "title": "ICE 'range start index ... out of range for slice of length 16384': rmeta I/O error reported with emit_err, so the truncated temp file is still renamed to the final path", + "fix_pr": 119510, + "merged": "2024-01-03", + "how_it_violates": "After #117301, rmeta write errors were reported with emit_err, which does not stop compilation. rustc kept going and renamed the incomplete temp file over lib.rmeta, and cargo could start a dependent build that read it (PR #119510: \"there is a window of time between the call to emit_err and the full error reporting where rustc believes it has emitted a valid rmeta file and will permit Cargo to launch a build for a dependent crate\"). Switching to emit_fatal aborts before the rename, so nothing partial reaches the final path. The reporter hit this on 1.75.0 with typst as a dependency (disk-full); the PR author reproduced it with an LD_PRELOAD write() that randomly returns ENOSPC.", + "before_toolchain": "nightly-2024-01-01", + "after_toolchain": "nightly-2024-01-05", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.py", + "content": "print('#![crate_type=\"lib\"]')\nfor i in range(20000):\n print(f'pub fn f{i}(x: u64) -> u64 {{ x.wrapping_mul({i}) }}')\n print(f'pub struct S{i} {{ pub a: u64, pub b: String }}')\n" + }, + { + "path": "user.rs", + "content": "extern crate big;\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\npython3 gen.py > big.rs\nmkdir out\nbash -c \"trap '' XFSZ; ulimit -f 1024; rustc +TOOLCHAIN --crate-type lib --emit=metadata big.rs --out-dir out; echo rustc exit=\\$?\"\nls -l out\nrustc +TOOLCHAIN --crate-type lib --emit=metadata user.rs -L out -o u.rmeta 2>&1 | grep -m2 -E 'panicked|range|error'", + "observe": "nightly-2024-01-01 (1.77.0-nightly e51e98dde 2023-12-31): \"error: failed to write to `.../rmetaXXXX/lib.rmeta`: File too large (os error 27)\" and exit 1, yet `ls -l out` shows libbig.rmeta at 1048576 bytes. The reader then panics: \"thread 'rustc' panicked at .../compiler/rustc_serialize/src/opaque.rs:262:42: range start index 12750977 out of range for slice of length 1048576\". nightly-2024-01-05 (f688dd684 2024-01-04): same error and exit 1, but out/ is empty (the truncated file is never renamed into place), and the reader gets a clean \"error[E0463]: can't find crate for `big`\". Current stable 1.97.1 behaves the same as 2024-01-05 (the temp file is now named full.rmeta)." + } + ], + "notes": "I ran all three reproductions on this machine (scratch dirs /tmp/p1repro.LXfv and /tmp/p1enc.olUX) and got the outputs quoted above. I checked every issue and PR number, quote and merge date with gh.\n\nNightly naming:\n- nightly-2024-01-02 and nightly-2024-01-03 do not exist on rustup, so I used 2024-01-01 as the \"before\" build for #119510.\n- nightly-2017-11-17 reports commit date 2017-11-16, and nightly-2017-11-20 reports 2017-11-19.\n\nBug 1 (#45841/#45899):\n- The in-place write is deterministic and cheap to show (same inode, and an extra hard link sees the new bytes; strace shows O_TRUNC on the final path). The actual torn-read ICE needs a race and is rare: 1 hit in roughly 2700 reader runs across two attempts. race.sh is included but should be treated as probabilistic.\n- Same era, probably the same cause, not confirmed: #45581 (cargo check on 1.21.0, the same leb128 ICE with \\\"the len is 32768\\\"; closed in 2023 for lack of a repro). #45581 is not a confirmed P1 bug.\n\nBugs 2 and 3 (#117254/#117301 and #119456/#119510): these are two steps of the same chain (#94732 introduced delayed FileEncoder errors). Both are reproduced cheaply with `ulimit -f` plus `trap '' XFSZ`, which turns write() into EFBIG instead of needing a full disk or root. #124686 (merged 2024-05-22) later added a footer so readers detect truncated rmeta/incremental files; that hardens the reader side and is not itself a P1 bug.\n\nChecked and not used (not P1 violations):\n- #82047 (merged 2021-03-08) is a perf change: on Linux it remove_file()s the destination before the rename (non_durable_rename, still present). That leaves a brief window where the .rmeta is missing, though never torn. I found no issue reporting harm from it.\n- #97485 made the Rust archive writer write a temp file and rename, but it did so in the same PR that replaced LLVM's writer (which already did temp+rename), so no released toolchain wrote rlibs in place.\n- #159619 (closed, not merged) concerns link_or_copy's fs::copy fallback writing in place on Haiku (#120581, #138847). The writer/reader interleaving there is unproven.\n- #89300 (open) is a disk-full problem in incremental compilation.\n- #155417 is about rejecting forged rmeta, not partial writes.\n\nThe GitHub search API was rate-limited for part of this work, so coverage of older issues may be incomplete." + }, + { + "key": "P2", + "bugs": [ + { + "issue": 45841, + "title": "ICE: index out of bounds in libserialize/leb128.rs (a rustc reads another rustc's half-written .rmeta while searching -L dependency= for a transitive crate)", + "fix_pr": 45899, + "merged": "2017-11-18", + "how_it_violates": "Before the fix, rustc wrote `--emit=metadata` output straight to its final path (open with O_CREAT|O_TRUNC, then one write() of the whole blob). A second rustc that loads a transitive dependency by scanning `-L dependency=DIR` for `lib-*.rmeta` opens every candidate and decodes its root to compare the crate hash. It can therefore open a sibling version's .rmeta while that file's writer is still filling it in, and the decoder panics on the truncated blob. This is P2 exactly: a reader opens a dependency .rmeta before its writer has finished it and published it. Issue #45841 is a Fuchsia `cargo check` build with two versions of `slab` in one deps dir. PR #45899 ('rustc_trans: atomically write .rmeta outputs to avoid races.') introduced the write-into-a-tempdir-in-the-output-dir then fs::rename scheme. The 'To avoid races with another rustc process scanning the output directory' comment in compiler/rustc_metadata/src/fs.rs still comes from that PR (commit f5952804251b).", + "before_toolchain": "nightly-2017-11-17", + "after_toolchain": "nightly-2017-11-20", + "reproducible_cheaply": false, + "files": [ + { + "path": "gen_a.py", + "content": "# generates a.rs: a crate with a large (~18 MB) .rmeta so the write window is wide enough to hit\nwith open('a.rs','w') as f:\n f.write('#![crate_name=\"a\"]\\n')\n for i in range(20000):\n f.write(f'pub struct S{i} {{ pub x{i}: u64, pub y: Vec }}\\n#[inline] pub fn f{i}(s: &S{i}) -> u64 {{ s.x{i} + s.y.len() as u64 + {i} }}\\n')\n" + }, + { + "path": "b.rs", + "content": "extern crate a;\npub fn g(s: &a::S0) -> u64 { a::f0(s) }\n" + }, + { + "path": "c.rs", + "content": "extern crate b;\npub fn h() -> u64 { b::g(unimplemented!()) }\n" + }, + { + "path": "race.sh", + "content": "#!/bin/bash\n# usage: race.sh TOOLCHAIN SECONDS WRITERS READERS\n# writers keep re-emitting version v2 of crate `a` into D; readers compile c, whose dep b\n# needs a (v1). a is found by scanning D, so readers also open liba-v2.rmeta.\nT=$1; SECS=${2:-60}; NW=${3:-3}; NR=${4:-8}\nR=\"rustc +$T\"\nrm -rf D out* stop fail.*; mkdir -p D\n$R --crate-type lib --emit=metadata -C metadata=v1 -C extra-filename=-v1 --out-dir D a.rs || exit 1\n$R --crate-type lib --emit=metadata -C metadata=b --out-dir D -L dependency=D --extern a=D/liba-v1.rmeta b.rs || exit 1\nfor w in $(seq $NW); do ( while [ ! -e stop ]; do $R --crate-type lib --emit=metadata -C metadata=v2 -C extra-filename=-v2 --out-dir D a.rs 2>/dev/null; done ) & done\nfor r in $(seq $NR); do ( mkdir -p out$r; n=0; f=0; while [ ! -e stop ]; do n=$((n+1)); $R --crate-type lib --emit=metadata --out-dir out$r -L dependency=D --extern b=D/libb.rmeta c.rs >out$r/log 2>&1 || { f=$((f+1)); cp out$r/log fail.$r.$f; }; done; echo \"$n $f\" > out$r/count ) & done\nsleep $SECS; touch stop; wait\ncat out*/count | awk '{n+=$1; f+=$2} END {print \"'$T': \" f \"/\" n \" reader runs failed\"}'\nls fail.* 2>/dev/null | head -1 | xargs -r grep -h -m3 \"panicked\\|error\"\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\npython3 gen_a.py\nchmod +x race.sh\n./race.sh TOOLCHAIN 90 4 10\n# deterministic check of the decoder half (old toolchain only): put a truncated copy in place of the sibling\n# head -c 9000000 D/liba-v1.rmeta > D/liba-v2.rmeta; rustc +TOOLCHAIN --crate-type lib --emit=metadata --out-dir out -L dependency=D --extern b=D/libb.rmeta c.rs", + "observe": "Run on this VM (16 cores, ext4). nightly-2017-11-17 (rustc d0f8e2913 2017-11-16, which does not contain merge 8752aeed3): '13/2345 reader runs failed', each an ICE with \"thread 'rustc' panicked at 'index out of bounds: the len is 958432 but the index is 18371597', /checkout/src/libserialize/leb128.rs:59:20\", the same panic as the issue. nightly-2017-11-20 (rustc 5041b3bb3 2017-11-19, which contains it): '0/2296 reader runs failed'. strace of the old writer shows openat(\"D/liba-v2.rmeta\", O_WRONLY|O_CREAT|O_TRUNC) then a single 18 MB write taking about 7 ms, which is the window. The fixed toolchain writes into an rmetaXXXX tempdir in D and renames. It is a race, so it needs several parallel readers and writers and about a minute. With one reader and one writer, 100 runs gave 0 hits. An empty sibling file is silently skipped; only a partially written one panics." + }, + { + "issue": 68149, + "title": "Spurious rebuilds under pipelining: a consumer opened (and recorded in its dep-info) a dependency's .rlib that was still being produced, instead of only the .rmeta", + "fix_pr": 68298, + "merged": "2020-01-23", + "how_it_violates": "Under cargo pipelining, a consumer starts as soon as the dependency's .rmeta is published, while the dependency's rustc is still producing the .rlib. When the locator resolved a transitive dependency by searching -L dependency=, it opened and remembered the .rlib if it was present, even though an rlib-only build needs just the .rmeta. ehuss traced the timeline: T1 libcore.rmeta emitted; T2 the backtrace build starts; T3 libcore.rlib emitted; T4 backtrace loads libcore.rlib; T5 its dep-info lists both. Cargo backdates the dep-info to T2, so the next build sees rlib mtime T3 > T2 and spuriously rebuilds std crates (reported by petrochenkov with -Zbinary-dep-depinfo in x.py). This is a P2-adjacent violation: a pipelined consumer reaches into a dependency artifact (.rlib) that is not yet finished or published from its point of view. PR #68298 ('Avoid declaring a fake dependency edge') stopped storing rlib/dylib paths when only producing an rlib. Caveat: the file opened is the .rlib, not the .rmeta.", + "before_toolchain": "nightly-2020-01-22", + "after_toolchain": "nightly-2020-01-24", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "pub fn fa() -> u32 { 1 }\n" + }, + { + "path": "b.rs", + "content": "extern crate a;\npub fn fb() -> u32 { a::fa() }\n" + }, + { + "path": "c.rs", + "content": "extern crate b;\npub fn fc() -> u32 { b::fb() }\n" + }, + { + "path": "run.sh", + "content": "#!/bin/bash\n# usage: run.sh TOOLCHAIN -- shows which dependency files the consumer `c` opened / recorded\nT=$1; rm -rf D; mkdir D\nR=\"rustc +$T --crate-type lib -C extra-filename=-x -L dependency=D --out-dir D\"\n$R --emit=metadata,link a.rs\n$R --emit=metadata,link --extern a=D/liba-x.rmeta b.rs\n# c is compiled like cargo does under pipelining: only b's .rmeta is given; a is found by search\n$R -Zbinary-dep-depinfo --emit=dep-info,metadata,link --extern b=D/libb-x.rmeta c.rs\necho \"== $T: deps recorded by c:\"; tr ' ' '\\n' < D/c-x.d | grep -E 'lib[ab]-x\\.r[a-z]+$' | sort -u\n" + } + ], + "commands": "rustup toolchain install TOOLCHAIN --profile minimal\nchmod +x run.sh\n./run.sh TOOLCHAIN\n# which dependency files c actually opens:\nstrace -f -e trace=openat -o st.txt rustc +TOOLCHAIN --crate-type lib -C extra-filename=-x -L dependency=D --out-dir D --emit=metadata,link --extern b=D/libb-x.rmeta c.rs; grep -o 'D/lib[ab]-x\\.r[a-z]*' st.txt | sort | uniq -c", + "observe": "Run on this VM. nightly-2020-01-22 (rustc 5e8897b7b 2020-01-21, which does not contain merge be663bf85): c's dep-info lists D/liba-x.rlib as well as liba-x.rmeta and libb-x.rmeta, and strace shows c opening D/liba-x.rlib. nightly-2020-01-24 (rustc 41f41b235 2020-01-23, which contains it): only liba-x.rmeta and libb-x.rmeta are listed and opened. This deterministic check shows the cause: the consumer opening the in-flight rlib. The user-visible symptom, a spurious cargo rebuild, needs the rlib to appear in the small window between the consumer starting and loading crates (ehuss: 'the time between T2 and T4 is extremely small'), plus -Zbinary-dep-depinfo, so the cargo-level symptom is not cheap to reproduce." + } + ], + "notes": "I found only two real, fixed rust-lang/rust bugs for P2, and reproduced both before and after the fix on this VM. Scratch dirs: /tmp/p2r (#45841) and /tmp/p2b (#68149).\n\n1. #45841 / PR #45899 is the founding P2 bug. It introduced the 'write to a tempdir inside the output dir, then fs::rename' code. That code is still in compiler/rustc_metadata/src/fs.rs: the \"avoid races with another rustc process scanning the output directory\" comment and the emit_wrapper_file doc. I found it by git blame: f5952804251b (2017-11-09) is the commit of PR #45899. The PR was merged 2017-11-18 as 8752aeed3. Toolchain containment was checked with the GitHub compare API.\n\n2. #68149 / PR #68298 is P2-adjacent. A pipelined consumer opened the dependency's not-yet-published .rlib. Its cause shows up deterministically in dep-info and strace output.\n\nRelated items that are unfixed or unmerged, so they are not included above:\n- rust-lang/cargo#16790 (open): 'only metadata stub found for rlib dependency' in pipelined or concurrent cargo builds. Reports so far bisect it to 1.93 vs 1.92; candidates are cargo#16230 / cargo#16177. No fix yet.\n- rust-lang/rust#159619 (closed, not merged) and cargo#17241 (closed): make the link_or_copy copy fallback atomic. Motivated by Haiku metadata-decode ICEs, rust-lang/rust#120581 and #138847; the mechanism was not confirmed.\n- #149711 (open): Miri E0460 'found possibly newer version'. This is stale sysroot or dep tracking, not a race.\n\nPR #88368 records the maintainer view that the locator deliberately tolerates partially written rlibs during pipelining (ehuss: \"there are race conditions in the pipelining system whereby rustc may attempt to load a partially written rlib, and we want to ignore that\"). That is useful context for how P2 is scoped. Note that an empty sibling .rmeta is silently skipped; only a partially written one ICEs.\n\nNone of the known-bug list (#122891, #130201, #138678, #143247, #82047, #144050, #162910, #159677/#159718, #161450/#129094, #117301, #114669) is a P2 race. #82047 only switched the metadata rename to a non-durable rename for performance.\n\nGitHub issue search was rate-limited once; searches for pipelining, rmeta race, partially written, invalid metadata and artifact notification in rust-lang/rust and rust-lang/cargo turned up no other fixed P2 bugs." + }, + { + "key": "P3", + "bugs": [ + { + "issue": 130201, + "title": "ICE: `coroutine_by_move_body_def_id` unsupported by its crate when calling a foreign crate's async closure as AsyncFnOnce (query result and by-move MIR were never encoded in rmeta)", + "fix_pr": 130201, + "merged": "2024-09-17", + "how_it_violates": "The dependency crate never wrote the `coroutine_by_move_body_def_id` table entry, and never wrote optimized_mir for the synthetic by-move body. When the dependent crate asked for that entry, the lookup found nothing and fell through to a missing provider, which ICEs. The PR text says: 'We weren't encoding this query in the metadata though, nor were we properly recording that synthetic MIR in `mir_keys`, so the `optimized_mir` wasn't getting encoded either!' There is no separate issue; PR #130201 itself is the reference, and its regression test is tests/ui/async-await/async-closures/foreign.rs.", + "before_toolchain": "nightly-2024-09-17", + "after_toolchain": "nightly-2024-09-19", + "reproducible_cheaply": true, + "files": [ + { + "path": "foreign.rs", + "content": "#![feature(async_closure)]\npub fn closure() -> impl async Fn() {\n async || {}\n}\n" + }, + { + "path": "main.rs", + "content": "#![feature(async_closure)]\nextern crate foreign;\nasync fn call_once(f: impl async FnOnce()) {\n f().await;\n}\nfn main() {\n let _ = call_once(foreign::closure());\n}\n" + } + ], + "commands": "rustc +TOOLCHAIN --edition 2021 --crate-type lib foreign.rs\nrustc +TOOLCHAIN --edition 2021 main.rs -L .", + "observe": "Verified locally. On nightly-2024-09-17, compiling main.rs ICEs with: `error: internal compiler error: compiler/rustc_middle/src/query/plumbing.rs:664:5: `tcx.coroutine_by_move_body_def_id(DefId(20:6 ~ foreign[87b3]::closure::{closure#0}::{closure#0}))` unsupported by its crate; perhaps the `coroutine_by_move_body_def_id` query was never assigned a provider function`. On nightly-2024-09-19 it compiles cleanly and produces the `main` binary." + }, + { + "issue": 122859, + "title": "Implied bound not implied across crates: associated-type bounds in supertrait position were dropped from metadata (implied_predicates encoded as super_predicates)", + "fix_pr": 122891, + "merged": "2024-03-24", + "how_it_violates": "For traits, the metadata encoder assumed `implied_predicates_of` was equal to `super_predicates_of` and only wrote the latter. So the implied predicates that come from associated type bounds (`trait Bar: Super`) were never written. The dependent crate read a smaller predicate list without noticing, and lost the bound. That shows up as a wrong E0277 that only happens across crates. The PR says: 'The assumption that they didn't differ was hard-coded in #107614, so in cross-crate positions this means that we forget the implied predicates from associated type bounds.' The same code compiles when crate_b is a local module.", + "before_toolchain": "nightly-2024-03-24", + "after_toolchain": "nightly-2024-03-26", + "reproducible_cheaply": true, + "files": [ + { + "path": "crate_b.rs", + "content": "pub trait Foo { type FooAssoc: Bar; }\npub trait Bar: Super {}\npub trait Super { type SuperAssoc; }\npub trait Bound: Unsatisfied {}\npub trait Unsatisfied {}\n" + }, + { + "path": "main.rs", + "content": "use crate_b::{Foo, Super, Unsatisfied};\nfn foo() {\n unsatisfied::<::SuperAssoc>()\n}\nfn unsatisfied() {}\nfn main() {}\n" + } + ], + "commands": "rustc +TOOLCHAIN --edition 2021 --crate-type lib crate_b.rs\nrustc +TOOLCHAIN --edition 2021 main.rs -L . --extern crate_b", + "observe": "Verified locally. On nightly-2024-03-24, main.rs fails with `error[E0277]: the trait bound `<::FooAssoc as Super>::SuperAssoc: Unsatisfied` is not satisfied`. On nightly-2024-03-26 it compiles (there is only a dead-code warning). Pasting crate_b's contents into main.rs as `mod crate_b` compiles on both, which shows that only the cross-crate (metadata) path is affected." + }, + { + "issue": 144004, + "title": "rustdoc drops #[no_mangle] / #[link_section] from inlined cross-crate re-exports: those attributes were not encoded in crate metadata", + "fix_pr": 144050, + "merged": "2025-07-19", + "how_it_violates": "`no_mangle` and `link_section` were moved out of the generic encoded attribute list, and nothing else wrote them to the attribute table. A dependent crate that reads the dependency's attributes from metadata (here rustdoc inlining a `pub use a::*` re-export) silently gets no such attribute, with no error. The PR title is 'Fix encoding of link_section and no_mangle cross crate', and it fixes it by always encoding them. This is the 'missing attributes cross-crate' kind of P3 violation. The result is silent information loss, not an ICE.", + "before_toolchain": "nightly-2025-07-06", + "after_toolchain": "nightly-2025-07-24", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "#[unsafe(no_mangle)]\npub fn f0() {}\n#[unsafe(link_section = \".here\")]\npub fn f1() {}\n#[unsafe(no_mangle)]\npub static S0: () = ();\n#[unsafe(link_section = \".there\")]\npub static S1: () = ();\n" + }, + { + "path": "b.rs", + "content": "pub use a::*;\n" + } + ], + "commands": "rustc +TOOLCHAIN a.rs --crate-type lib --edition 2024\nrustdoc +TOOLCHAIN b.rs --edition 2024 -L. --extern a\nfor f in fn.f0 fn.f1 static.S0 static.S1; do echo -n \"$f: \"; grep -o 'no_mangle\\|link_section[^<]*' doc/b/$f.html | sort -u | tr '\\n' ' '; echo; done", + "observe": "Verified locally. On nightly-2025-07-06 all four pages (doc/b/fn.f0.html, fn.f1.html, static.S0.html, static.S1.html) show no attribute, so every grep prints nothing. On nightly-2025-07-24 they show `no_mangle`, `link_section = \".here\"`, `no_mangle` and `link_section = \".there\"`. The issue adds that 1.88.0 stable also lacked both attributes, and 1.89 beta showed no_mangle but not link_section. The regression came and went with attribute-parsing refactors." + } + ], + "notes": "I ran all three reproductions with rustup on this VM. Each one fails on the \"before\" toolchain and passes on the \"after\" one exactly as described. The work dirs were /tmp/p3/a (#130201), /tmp/p3/b (#122891 / issue #122859) and /tmp/p3/c (#144050 / issue #144004).\n\nAll issue and PR numbers come from `gh issue view` / `gh pr view` text:\n- PR #122891 \"Encode implied predicates for traits\" closes issue #122859 and merged 2024-03-24T13:16Z.\n- PR #130201 \"Encode `coroutine_by_move_body_def_id` in crate metadata\" has no linked issue and merged 2024-09-17T19:35Z.\n- PR #144050 closes issue #144004 and merged 2025-07-19T08:02Z.\n\nThe other known bugs in the list are about determinism, durability or incremental compilation (#138678, #143247, #82047, #162910, #159718, #161450, #117301, #114669), not P3, so I left them out.\n\nI could not search for more candidates, such as other 'missing optimized MIR' or 'unsupported by its crate' ICEs. Every `gh search issues` call came back HTTP 403 because the API rate limit was exceeded. A later pass could look for similar PRs titled 'Encode X in metadata' for more P3 examples.\n\nThe ICE in #130201 is the clearest P3 signal: the query falls back to a provider that does not exist because the table entry was never written. #122859 and #144004 are the silent kind: the dependent crate reads an empty or default value with no error, which is exactly what P3 is meant to catch." + }, + { + "key": "P4", + "bugs": [ + { + "issue": 138678, + "title": "Randomly seeded HashMap (pulldown-cmark reference_definitions) iterated while collecting doc links: lib.rmeta differs between identical runs", + "fix_pr": 138678, + "merged": "2025-03-28", + "how_it_violates": "P4: metadata encoding read a randomly seeded std HashMap. rustc_resolve::rustdoc::parse_links walks pulldown-cmark's reference_definitions(), a HashMap with RandomState, and pushes the links in that iteration order. The resulting list of doc-link candidates is encoded into crate metadata, so two runs of rustc on the same input wrote different .rmeta bytes. The bug came in with #136363 (merged 2025-02-16). The fix (#138678) sorts the links by label. #138678 is the PR; it has no separate issue, and its description says the nondeterminism was found in Bazel lib.rmeta outputs.", + "before_toolchain": "nightly-2025-03-28", + "after_toolchain": "nightly-2025-03-30", + "reproducible_cheaply": true, + "files": [ + { + "path": "lib.rs", + "content": "//! Crate docs.\n//!\n//! [l0]: crate::S0\n//! [l1]: crate::S1\n//! [l2]: crate::S2\n//! [l3]: crate::S3\n//! [l4]: crate::S4\n//! [l5]: crate::S5\n//! [l6]: crate::S6\n//! [l7]: crate::S7\n//! [l8]: crate::S8\n//! [l9]: crate::S9\n//! [l10]: crate::S10\n//! [l11]: crate::S11\n//! [l12]: crate::S12\n//! [l13]: crate::S13\n//! [l14]: crate::S14\n//! [l15]: crate::S15\npub struct S0;\npub struct S1;\npub struct S2;\npub struct S3;\npub struct S4;\npub struct S5;\npub struct S6;\npub struct S7;\npub struct S8;\npub struct S9;\npub struct S10;\npub struct S11;\npub struct S12;\npub struct S13;\npub struct S14;\npub struct S15;\n" + } + ], + "commands": "for i in 1 2 3 4 5 6 7 8; do rm -rf out; mkdir out; rustc +TOOLCHAIN lib.rs --crate-type=rlib --emit=metadata --out-dir out 2>/dev/null; sha256sum out/liblib.rmeta | cut -c1-16; done | sort | uniq -c", + "observe": "Verified locally. On nightly-2025-03-28 (rustc 1.87.0-nightly 3f5502370 2025-03-27), the 8 identical compilations gave 8 different liblib.rmeta hashes, each counted once. On nightly-2025-03-30 (1.88.0-nightly 1799887bb 2025-03-29), all 8 gave the same hash (count 8)." + }, + { + "issue": 111227, + "title": "debugger_visualizer files (#111226 / #111227 / #111295) were read into rmeta but tracked by neither dep-info nor incremental: stale rmeta and an incremental ICE", + "fix_pr": 111641, + "merged": "2023-05-19", + "how_it_violates": "P4: the contents of files named by #![debugger_visualizer(natvis_file / gdb_script_file)] were read and encoded into crate metadata (the debugger_visualizers query), but those files were not tracked inputs. They were missing from dep-info (#111226), so cargo never rebuilt after a change. The incremental system also did not see them: changing the natvis file gave 'internal compiler error: encountered incremental compilation error with debugger_visualizers' (#111227), and changed GDB scripts were not picked up (#111295). The fix (#111641) made debugger_visualizers an eval_always query computed from the AST and added the files to dep-info. Its tests include run-make/incremental-debugger-visualizer, which greps the rmeta for the file contents.", + "before_toolchain": "nightly-2023-05-18", + "after_toolchain": "nightly-2023-05-21", + "reproducible_cheaply": true, + "files": [ + { + "path": "foo.rs", + "content": "#![debugger_visualizer(natvis_file = \"./foo.natvis\")]\n#![debugger_visualizer(gdb_script_file = \"./foo.py\")]\n\npub struct Foo {\n pub x: u32,\n}\n" + }, + { + "path": "run.sh", + "content": "#!/bin/sh\nTC=$1\nrm -rf out incr; mkdir out\necho \"GDB script v1\" > foo.py; echo \"Natvis v1\" > foo.natvis\nrustc +$TC foo.rs --crate-type=rlib --emit=metadata,dep-info --out-dir out -C incremental=incr\ngrep -q foo.py out/foo.d && echo \"foo.py in dep-info\" || echo \"foo.py NOT in dep-info\"\necho \"Natvis v2\" > foo.natvis\nrustc +$TC foo.rs --crate-type=rlib --emit=metadata --out-dir out -C incremental=incr 2>&1 | head -3\ngrep -a -o \"Natvis v[12]\" out/libfoo.rmeta\n" + } + ], + "commands": "sh run.sh TOOLCHAIN", + "observe": "Verified locally. On nightly-2023-05-18 the script prints 'foo.py NOT in dep-info'. The second, incremental compile prints 'error: internal compiler error: encountered incremental compilation error with debugger_visualizers(foo[47df])', and libfoo.rmeta still contains 'Natvis v1', so the metadata is stale. On nightly-2023-05-21 foo.py is listed in foo.d, the rebuild succeeds, and the rmeta contains 'Natvis v2'." + }, + { + "issue": 40364, + "title": "env!/option_env! values were baked into the output, but the env vars were not reported as dependencies: cargo did not rebuild when the variable changed", + "fix_pr": 71858, + "merged": "2020-06-26", + "how_it_violates": "P4: env!() reads the process environment while compiling, and the value ends up in the output. rustc did not record that the variable was read, so build tools treated the crate as fresh after the variable changed. PR #71858 added '# env-dep:KEY=VALUE' lines to dep-info, closing #40364, #44074 and #70517. Cargo then read those lines and rebuilt on a change (rust-lang/cargo PR #8421, merged 2020-06-30). Both parts are needed for the repro, so it uses stable releases with the matching cargo: 1.45.0 has neither part, 1.46.0 has both. rustc nightlies after 2020-06-26 have the dep-info lines; cargo's rebuild arrived when the cargo submodule was next updated in early July 2020.", + "before_toolchain": "1.45.0", + "after_toolchain": "1.46.0", + "reproducible_cheaply": true, + "files": [ + { + "path": "Cargo.toml", + "content": "[package]\nname = \"envdemo\"\nversion = \"0.1.0\"\nedition = \"2018\"\n" + }, + { + "path": "src/main.rs", + "content": "fn main() {\n println!(\"{}\", env!(\"MIRTH_DEMO\"));\n}\n" + } + ], + "commands": "cargo +TOOLCHAIN clean -q; MIRTH_DEMO=one cargo +TOOLCHAIN run -q; MIRTH_DEMO=two cargo +TOOLCHAIN run -q; grep env-dep target/debug/deps/envdemo-*.d || echo 'no env-dep in dep-info'", + "observe": "Verified locally. With 1.45.0 the two runs print 'one' then 'one': the stale binary is reused after MIRTH_DEMO changed. With 1.46.0 they print 'one' then 'two'. On 1.46.0 the dep-info file also has a '# env-dep:MIRTH_DEMO=two' line; 1.45.0 has none." + }, + { + "issue": 66955, + "title": "--remap-path-prefix was UNTRACKED: changing it under incremental reused stale codegen and debuginfo with the old path", + "fix_pr": 84233, + "merged": "2021-04-29", + "how_it_violates": "P4, read broadly as untracked state that reaches the output: #48162 made --remap-path-prefix UNTRACKED so that the crate hash stayed the same. Remapped paths still went into the outputs, so after the flag changed, an incremental rebuild reused cached codegen units holding the old prefix. The result was a mix of old and new remapped paths. The fix (#84233) added TRACKED_NO_CRATE_HASH: the option invalidates the incremental cache but stays out of the crate hash. Note: in this era metadata was always re-encoded, so the rmeta itself picked up the new path. The stale data is in the object code and debuginfo inside the rlib.", + "before_toolchain": "nightly-2021-04-29", + "after_toolchain": "nightly-2021-05-01", + "reproducible_cheaply": true, + "files": [ + { + "path": "lib.rs", + "content": "pub fn f() -> u32 { 1 }\n#[inline] pub fn g() -> &'static str { file!() }\n" + }, + { + "path": "run.sh", + "content": "#!/bin/sh\nTC=$1\nrm -rf out incr; mkdir out\nfor P in /AAAA_first /BBBB_second; do\n rustc +$TC lib.rs --crate-type=rlib --out-dir out -g -C incremental=incr --remap-path-prefix=\"$PWD=$P\"\n echo \"after remap to $P, rlib contains:\"; strings out/liblib.rlib | grep -oE '/(AAAA_first|BBBB_second)' | sort | uniq -c\ndone\n" + } + ], + "commands": "sh run.sh TOOLCHAIN", + "observe": "Verified locally. On nightly-2021-04-29, after the second build (remap to /BBBB_second) the rlib still contains '2 /AAAA_first' and only '1 /BBBB_second'; the stale object code and debuginfo were reused. On nightly-2021-05-01 it contains '3 /BBBB_second' and no /AAAA_first." + } + ], + "notes": "I ran all four reproductions locally on the toolchains listed, and they behave as described: the bug shows on the before toolchain and not on the after one. Every issue and PR number, and every merge date, was checked with gh.\n\nOrder and how well each fits P4:\n1. #138678: a randomly seeded HashMap leaks into rmeta bytes. This is the purest P4 case and already on your known list. The bug was introduced by #136363 (2025-02-16). #138678 is a PR with no linked issue, so I put it in the issue field too.\n2. #111226/#111227/#111295, fixed by #111641: file contents went into rmeta without being tracked. This is the best match for \"include_str!-like untracked files\": incremental hits an ICE, the rmeta goes stale, and cargo does not rebuild.\n3. #40364 (also #44074 and #70517), fixed by rustc #71858 plus cargo #8421: env! values were not tracked. It reproduces at the cargo level, so I used the stable releases 1.45.0 and 1.46.0 because both rustc and cargo had to change.\n4. #66955, fixed by #84233: an UNTRACKED CLI option; it fits only under a broad reading of P4. The stale data is in the codegen and debuginfo in the rlib, not in the rmeta. Before about 2025 rmeta was always re-encoded, so the rmeta itself was correct.\n\nConsidered and left out:\n- #118204: MACOSX_DEPLOYMENT_TARGET is not tracked. It is still OPEN. #130883 added an env_var_os query in 2025-03 as groundwork, but this is macOS-only and not yet fixed.\n- #113990: the stable crate hash depends on the host tuple. Still open, and it needs two hosts.\n- #113584: closed without an identified fix.\n- #151868 and #152150: closed as not actionable.\n- #65036/#65043: FxHashMap re-export order, 2019. Fx hashing is deterministic from run to run, so this cannot be reproduced cheaply.\n- #76496/#76515: only affected rustc's own crates.\n- #58069 (include!): closed without a fix PR.\n- #87899: a rustdoc-only problem.\n- tracked_env/tracked_path (#99515, #74690): still unstable and open.\n- Incremental env! (tests/incremental/env, added with #130883): there was no prior bug, since macro expansion reruns every session.\n\nReproduction directories are in /tmp: p4doc, p4dv, p4env and p4remap." + }, + { + "key": "P5", + "bugs": [ + { + "issue": 159677, + "title": ".rmeta contents depend on unrelated files in the library search path (doc_link_resolutions encoded in Symbol-index hash order)", + "fix_pr": 159718, + "merged": "2026-07-24", + "how_it_violates": "Building the same client.rs with the same flags twice gives different .rmeta bytes if an unrelated rlib whose name starts with the dependency's name (libfoo_bar.rlib next to libfoo.rlib) is in the -L directory. The crate locator opens libfoo_bar.rlib to read its crate name, which interns the extra Symbol `foo_bar`. That shifts the interner indices of symbols interned later, such as the doc-link strings. DocLinkResMap was an UnordMap (FxHashMap) keyed by (Symbol, Namespace) and hashed by interner index, and it was encoded in hash-iteration order. The fix makes it an FxIndexMap, so entries are encoded in insertion order. The PR adds the regression test tests/run-make/rmeta-unrelated-search-path-files.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": true, + "files": [ + { + "path": "client.rs", + "content": "//! [crate::Client]\n\nextern crate foo;\n\npub struct Client;\n" + }, + { + "path": "foo.rs", + "content": "pub struct Foo;\n" + }, + { + "path": "foo_bar.rs", + "content": "pub struct FooBar;\n" + } + ], + "commands": "rustc +TOOLCHAIN foo.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client1.rmeta\nrustc +TOOLCHAIN foo_bar.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client2.rmeta\ncmp client1.rmeta client2.rmeta && echo IDENTICAL", + "observe": "Verified locally. On nightly-2026-07-20 (and on stable 1.97.1) cmp prints 'client1.rmeta client2.rmeta differ: byte 1256' (1263 on 1.97.1); the differing bytes are the reordered doc-link entry 'crate::Client'. On nightly-2026-09-25 the files are identical and the script prints IDENTICAL. The fix merged 2026-07-24T09:12Z, so nightly-2026-07-26 and later should be fixed." + }, + { + "issue": 34902, + "title": "Metadata xrefs encoded in pointer-address (ASLR) order: rlib differs on every run", + "fix_pr": 35984, + "merged": "2016-08-28", + "how_it_violates": "encode_xrefs iterated an FnvHashMap, u32>. Its keys are interned ty::Predicate pointers, so the hash and the iteration order depend on heap addresses, which ASLR changes on every run. The commit 'Make metadata encoding deterministic' in PR #35984 ('Steps towards reproducible builds', cc tracking issue #34902) sorts the xrefs by their ID before encoding. This is the classic single-threaded pointer-address nondeterminism: the same command run twice in the same directory gives different rust.metadata.bin bytes, while the object file stays identical.", + "before_toolchain": "nightly-2016-08-27", + "after_toolchain": "nightly-2016-08-30", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.sh", + "content": "{\necho 'pub trait Tr {}'\nfor i in $(seq 1 30); do echo \"pub struct S$i;\"; echo \"pub fn f$i + Clone + Into>(_t: T) {}\"; done\n} > lib.rs\n" + } + ], + "commands": "sh gen.sh\nfor i in 1 2 3 4 5; do rm -rf o$i; mkdir o$i; rustc +TOOLCHAIN lib.rs --crate-type rlib --crate-name lib --out-dir o$i; done\nmd5sum o*/liblib.rlib\n(cd o1 && ar x liblib.rlib); (cd o2 && ar x liblib.rlib); md5sum o1/rust.metadata.bin o2/rust.metadata.bin o1/lib.0.o o2/lib.0.o\n# control: with ASLR disabled the old compiler is deterministic\nfor i in 1 2; do rm -rf r$i; mkdir r$i; setarch -R rustc +TOOLCHAIN lib.rs --crate-type rlib --crate-name lib --out-dir r$i; done; md5sum r*/liblib.rlib", + "observe": "Verified locally. On nightly-2016-08-27 (rustc 1.13.0-nightly 198713106) 10 runs gave 10 different liblib.rlib hashes. Only the rust.metadata.bin member differs; lib.0.o is identical. Under 'setarch -R' (ASLR off) the runs are identical, which confirms that pointer addresses are the cause. On nightly-2016-08-30 (77d2cd28f) all 10 runs give the same hash. Very old toolchain: install it with 'rustup toolchain install nightly-2016-08-27 --profile minimal'." + }, + { + "issue": 65036, + "title": "Module re-exports serialized into metadata in FxHashMap order keyed by (Ident, Namespace), which is Symbol-index dependent", + "fix_pr": 65043, + "merged": "2019-10-06", + "how_it_violates": "The PR text says: 're-exports end up getting serialized into crate metadata, which means that metadata generation was non-deterministic'. A module's resolutions were an FxHashMap<(Ident, Namespace), ...>, and Ident hashes by Symbol interner index. Anything that changes interning order changes the order of re-exports in .rmeta. The fix changes Resolutions to an FxIndexMap. The reported symptom (#65036) was flaky diagnostics, std::mem::transmute vs std::intrinsics::transmute. The reproduction below uses the same unrelated-file-in-search-path trigger as #159677. A proc macro interns the item names after the extern crate lookup has interned 'foo_bar', so the indices shift.", + "before_toolchain": "nightly-2019-10-05", + "after_toolchain": "nightly-2019-10-08", + "reproducible_cheaply": true, + "files": [ + { + "path": "foo.rs", + "content": "pub struct Foo;\n" + }, + { + "path": "foo_bar.rs", + "content": "pub struct FooBar;\n" + }, + { + "path": "gen.rs", + "content": "extern crate proc_macro;\nuse proc_macro::TokenStream;\n#[proc_macro]\npub fn make(_: TokenStream) -> TokenStream {\n (1..=40).map(|i| format!(\"pub fn item_{}() {{}}\\n\", i)).collect::().parse().unwrap()\n}\n#[proc_macro]\npub fn reexports(_: TokenStream) -> TokenStream {\n (1..=40).map(|i| format!(\"pub use inner::item_{};\\n\", i)).collect::().parse().unwrap()\n}\n" + }, + { + "path": "client.rs", + "content": "extern crate foo;\nextern crate gen;\nmod inner { gen::make!(); }\ngen::reexports!();\n" + } + ], + "commands": "rustc +TOOLCHAIN gen.rs --crate-type proc-macro\nrustc +TOOLCHAIN foo.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client1.rmeta\nrustc +TOOLCHAIN foo_bar.rs --crate-type rlib\nrustc +TOOLCHAIN client.rs --crate-type rlib --emit metadata -L . -o client2.rmeta\ncmp client1.rmeta client2.rmeta && echo IDENTICAL\ndiff <(strings -n4 client1.rmeta) <(strings -n4 client2.rmeta) | head", + "observe": "Verified locally. On nightly-2019-10-05 (rustc 1.40.0-nightly 2e7244807 2019-10-04) cmp reports 'differ: byte 5157', and the strings diff shows the item_N re-export names in a different order (item_21, item_1, item_24, item_11, ...). On nightly-2019-10-08 (f3c9cece7 2019-10-07) the script prints IDENTICAL." + }, + { + "issue": 129094, + "title": "Parallel frontend: derives make metadata irreproducible (non-deterministic encoding of syntax contexts)", + "fix_pr": 161450, + "merged": "2026-09-16", + "how_it_violates": "With -Zthreads=N, the order in which SyntaxContexts and expansions are reached during metadata encoding depends on thread scheduling, so a tiny derive-heavy crate compiled repeatedly with the same inputs gives different rlibs. The fix ('Fix non-deterministic encoding of syntax contexts', which reiterates #157409) adds deterministic encoding indices in rustc_metadata/rmeta/encoder.rs and rustc_span/hygiene.rs. It also adds derives-issue-129094.rs to tests/run-make/parallel-reproducible-build. Unlike the other entries this is not single-threaded: it needs the parallel frontend.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": false, + "files": [ + { + "path": "derives.rs", + "content": "#![crate_type = \"lib\"]\n#[derive(Clone, Copy, Hash, PartialEq, PartialOrd)]\nstruct PackedPoint {\n x: u32,\n}\n" + } + ], + "commands": "for i in $(seq 1 20); do rm -rf o$i; mkdir o$i; rustc +TOOLCHAIN derives.rs -Zthreads=16 -Copt-level=3 --out-dir o$i 2>/dev/null; done\nmd5sum o*/*.rlib | awk '{print $1}' | sort | uniq -c", + "observe": "Verified locally. On nightly-2026-07-20, 20 runs gave 3 distinct rlib hashes (15/4/1). On nightly-2026-09-25 all 20 runs give one hash. Race-dependent: it needs -Zthreads>1 and several runs, and how often it shows up depends on core count and scheduling. It reproduced readily on this VM." + } + ], + "notes": "I ran all four reproductions on this machine with rustup toolchains and saw the bug before the fix and identical outputs after. The scratch copies are in /tmp/p5a (#159677), /tmp/p5b (#65036), /tmp/p5c (#129094) and /tmp/p5d (#34902/#35984).\n\nFor #34902, the PR #35984 is real ('Steps towards reproducible builds', merged 2016-08-28), but #34902 is the umbrella tracking issue: the PR body says 'cc #34902' and does not close a specific bug. The metadata fix in that PR is the commit 'Make metadata encoding deterministic', which sorts encode_xrefs.\n\nFor #65036, the issue itself only reports flaky ui-test diagnostics. That PR #65043 serialized re-exports nondeterministically into metadata is stated in the PR body. The reproduction for it is mine, built on the same search-path trigger as #159677.\n\nRuled out:\n- #113584 and #119372: caused by the compiler_builtins build script iterating a HashMap, not a rustc bug.\n- #124357 and #124306: closed without a rustc fix; same symptom as #159677.\n- #140061: caused by a HashMap in a third-party proc macro.\n- #153898 (strict version hash depends on rust-src being present): still open, no fix.\n- #120825, #112586, #129117, #88982: path remapping problems or Windows/PDB-only problems.\n- #162910 (DefPathHashMap): parallel-frontend only.\n\nGitHub search rate limits cut some queries short. The A-reproducibility label listing was complete." + }, + { + "key": "P5t", + "bugs": [ + { + "issue": 129094, + "title": "derives: parallel compiler makes builds irreproducible (non-deterministic encoding of syntax contexts in crate metadata)", + "fix_pr": 161450, + "merged": "2026-09-16", + "how_it_violates": "With -Zthreads=16, compiling the same crate twice gives a different .rlib. The lib.rmeta member inside the rlib is itself different from run to run. The cause is that SyntaxContext/hygiene data was encoded into metadata (and the on-disk cache) in an order that depended on thread scheduling. PR #161450 ('Fix non-deterministic encoding of syntax contexts'; touches rustc_metadata/src/rmeta/encoder.rs and rustc_span/src/hygiene.rs) encodes them in a deterministic order and adds the regression test tests/run-make/parallel-reproducible-build/derives-issue-129094.rs.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": true, + "files": [ + { + "path": "derives.rs", + "content": "#![crate_type = \"lib\"]\n#[derive(Clone, Copy, Hash, PartialEq, PartialOrd)]\nstruct PackedPoint {\n x: u32,\n}\n" + } + ], + "commands": "for i in $(seq 1 15); do rm -rf x; mkdir x; rustc +TOOLCHAIN derives.rs -Zthreads=16 -Copt-level=3 -o x/libd.rlib 2>/dev/null; md5sum < x/libd.rlib; (cd x && ar x libd.rlib lib.rmeta && md5sum < lib.rmeta | sed 's/^/rmeta /'); done | sort | uniq -c", + "observe": "Verified locally. On nightly-2026-07-20 the rlib had 4 different md5s over 15 runs, and the lib.rmeta extracted from it had 3 different md5s over 10 runs. On nightly-2026-09-25 there was a single md5 for both. Note: on the old nightly, --emit=metadata alone was stable for this input (1 md5). The nondeterminism shows up in a full build, where codegen runs alongside metadata encoding." + }, + { + "issue": 140413, + "title": "parallel rustc: static mut refs not reproducible", + "fix_pr": 144722, + "merged": "2025-08-13", + "how_it_violates": "With -Zthreads=50, building the same binary repeatedly gives different bytes. Mono items were sorted by (DefId, SymbolName) to pick their order in the output. DefId indices are allocated in access order, which is nondeterministic under the parallel front end. PR #144722 ('Fix parallel rustc not being reproducible due to unstable sorts of items') stops sorting by DefId. Two later facts confirm the fix: PR #161353 (merged 2026-09-01; it closed the issue and added tests/run-make/parallel-reproducible-build) found by bisection that the regression flips at nightly-2025-08-14, at commit #144722. The same PR also fixes the async-closure case #140425 (closed 2025-08-13).", + "before_toolchain": "nightly-2025-07-26", + "after_toolchain": "nightly-2025-12-10", + "reproducible_cheaply": true, + "files": [ + { + "path": "static.rs", + "content": "// Checks that mutable static items can have mutable slices and other references\n\npub static mut TEST: &'static mut [isize] = &mut [1];\npub static mut EMPTY: &'static mut [isize] = &mut [];\npub static mut INT: &'static mut isize = &mut 1;\n\n// And the same for raw pointers.\n\npub static mut TEST_RAW: *mut [isize] = &mut [1isize] as *mut _;\npub static mut EMPTY_RAW: *mut [isize] = &mut [] as *mut _;\npub static mut INT_RAW: *mut isize = &mut 1isize as *mut _;\n\npub fn main() {\n unsafe {\n TEST[0] += 1;\n assert_eq!(TEST[0], 2);\n *INT_RAW += 1;\n assert_eq!(*INT_RAW, 2);\n }\n}\n" + } + ], + "commands": "for i in $(seq 1 15); do rustc +TOOLCHAIN static.rs -Zthreads=50 -o a 2>/dev/null; md5sum < a; done | sort | uniq -c", + "observe": "Verified locally. On nightly-2025-07-26 there were 2 distinct md5s of the binary over 15 runs. On nightly-2025-12-10 there was 1. The bisected boundary is nightly-2025-08-13 (bad) to nightly-2025-08-14 (good). Caveat: building this file as --crate-type=lib --emit=metadata is still nondeterministic, even on nightly-2026-10-06 (7-8 distinct rmeta md5s out of 10). That case is the still-open follow-up #162203, so use this repro for the binary only." + }, + { + "issue": 150451, + "title": "parallel compiler: thread::spawn-ing loop not reproducible (LLVM inline-asm location cookies)", + "fix_pr": 160197, + "merged": "2026-09-07", + "how_it_violates": "With -Zthreads=3, compiling a tiny lib that uses thread::spawn gives a different .rlib each time. The output differs when bitcode is embedded or LTO is used. The cause is that the srcloc 'cookies' attached to LLVM inline asm were assigned nondeterministically by the parallel front end. PR #160197 ('Restrict LLVM inline asm location cookie usage. Fixes #150451') landed in rollup #162434 on 2026-09-07 and added tests/run-make/parallel-reproducible-inline-asm-cookie. The bug is in object code and bitcode, not in rmeta, but it breaks reproducibility of the published artifact.", + "before_toolchain": "nightly-2026-07-20", + "after_toolchain": "nightly-2026-09-25", + "reproducible_cheaply": true, + "files": [ + { + "path": "spawn.rs", + "content": "use std::thread;\n\nfn _main() {\n let _t1 = thread::spawn(|| {\n for _ in 0..100 {\n println!(\"test\");\n }\n });\n}\n" + } + ], + "commands": "for i in $(seq 1 15); do rustc +TOOLCHAIN spawn.rs --crate-type=lib -Zthreads=3 -Clink-dead-code=true -Copt-level=0 -Cembed-bitcode=true -o libs.rlib 2>/dev/null; md5sum < libs.rlib; done | sort | uniq -c", + "observe": "Verified locally. On nightly-2026-07-20 there were 7 distinct rlib md5s over 15 runs. On nightly-2026-09-25 there was 1." + } + ], + "notes": "I found three fixed P5 bugs and ran every reproduction locally on the before and after nightlies. Each one shows multiple distinct md5s before the fix and a single md5 after it. Issue and PR facts come from gh: issue bodies, closing and cross-reference timeline events, and PR bodies and file lists. The test sources come from the rust-lang/rust master test files.\n\nI found the candidates by searching rust-lang/rust for issues labelled both A-parallel-compiler and A-reproducibility. Fix PRs and their merge dates come from the GitHub timeline.\n\nHow the bugs relate to rmeta:\n- #129094 is the most relevant to P5: the lib.rmeta inside the rlib changes from run to run.\n- #140413 is about the binary.\n- #150451 is about object code and bitcode.\n\n#140425 (async closures not reproducible) was closed by the same PR as #140413, #144722 (merged 2025-08-13). I did not run it separately.\n\nOpen, unfixed bugs (no fix PR, so they are listed here rather than as entries):\n- #163878, \"generics_of encoding not reproducible\". This one breaks rmeta directly. The `param_def_id_to_index` FxHashMap in Generics is encoded in iteration order. I verified it still reproduces on nightly-2026-10-06 with `rustc --edition 2024 -Zthreads=2 --crate-type=lib --emit=metadata`: 6 distinct rmeta md5s over 20 runs. The input is struct A with two methods returning `impl Iterator + '_`. The issue says the open PR #162809 (\"Remap def indices for deterministic metadata encoding\") does not fix it.\n- #162203, static mut refs in a lib. This is the follow-up to #140413. I verified that `--crate-type=lib --emit=metadata -Zthreads=50` on static.rs gives 7-8 distinct rmeta md5s per 10 runs on nightly-2025-12-10, 2026-09-25 and 2026-10-06.\n- #162202, async fns with elided or named lifetimes in a lib (`-Zthreads=30 --crate-type=lib`). I did not verify this one.\n- #154278, `--emit=mir` not reproducible because of AllocIds.\n- #142063, cycle diagnostics change from run to run under -Zthreads.\n- #146616, the query depth limit check is less stable under the parallel frontend.\n- #154314, an umbrella issue for parallel frontend test failures.\n\nClosed without a code fix:\n- #117776 (ripgrep binaries differ with -Zthreads=8) and #154259 (.llvm. symbol suffixes differ in librustc_driver with parallel-frontend-threads) were closed on 2026-09-04 and 2026-09-16. The timeline has no closing PR for either.\n- #157665 (unstable AllocIds in diagnostics) was closed 2026-07-02, also with no closing PR.\n- #157659 (`dcx.try_steal_modify_and_emit_err` race) was closed 2026-06-12. I did not look into it.\n\nAll of these bugs are races, so a single run can match by luck. Run 10-20 times and count distinct md5s, as the commands do. More threads (-Zthreads=16 to 50) make the bug show up more often." + }, + { + "key": "P6", + "bugs": [ + { + "issue": 89598, + "title": "VTable-related miscompilation with incremental compilation (reordering trait methods calls the wrong method)", + "fix_pr": 89619, + "merged": "2021-10-08", + "how_it_violates": "#86475 added an untracked global vtable cache to tcx. After trait methods are reordered, the object file for the main CGU is reused from the incremental cache with the old vtable layout, while mod1 is recompiled for the new layout. The incremental binary then calls method2 where a clean build calls method1, so the binaries differ (a miscompile). P-critical, I-unsound, regression-from-stable-to-beta.", + "before_toolchain": "nightly-2021-10-07", + "after_toolchain": "nightly-2021-10-10", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "trait Foo {\n #[cfg(rpass1)]\n fn method1(&self) -> u32;\n\n fn method2(&self) -> u32;\n\n #[cfg(rpass2)]\n fn method1(&self) -> u32;\n}\n\nimpl Foo for u32 {\n fn method1(&self) -> u32 { 17 }\n fn method2(&self) -> u32 { 42 }\n}\n\nfn main() {\n let x: &dyn Foo = &0u32;\n assert_eq!(mod1::foo(x), 17);\n}\n\nmod mod1 {\n pub(super) fn foo(x: &dyn super::Foo) -> u32 {\n x.method1()\n }\n}\n" + } + ], + "commands": "rm -rf inc\nrustc +TOOLCHAIN --cfg rpass1 -C incremental=inc main.rs -o a.out && ./a.out && echo pass1 ok\nrustc +TOOLCHAIN --cfg rpass2 -C incremental=inc main.rs -o a.out && ./a.out && echo pass2 ok\n# control: a clean build of the edited source passes\nrustc +TOOLCHAIN --cfg rpass2 main.rs -o clean.out && ./clean.out && echo clean ok", + "observe": "Verified locally. On nightly-2021-10-07, pass 1 prints 'pass1 ok'; the incremental pass 2 binary panics with \"assertion failed: `(left == right)` left: `42`, right: `17`\" because method2 was called through the stale vtable. The clean build of rpass2 runs fine. On nightly-2021-10-10 both passes print ok. Regression test: tests/incremental/reorder_vtable.rs." + }, + { + "issue": 82920, + "title": "Miscompilation with incr. comp. (supertrait predicates sorted by DefId; reordering trait declarations breaks dyn method calls)", + "fix_pr": 83074, + "merged": "2021-03-15", + "how_it_violates": "Bounds and predicates were sorted by DefId, which is not stable across sessions. When two trait declarations swap places, the query result changes but is still treated as green and reused, so the vtable layout and the method call sites disagree. The incremental binary calls the wrong trait method, while a clean build is correct. The original report was a rust-analyzer test suite failing with out-of-bounds panics and segfaults until `cargo clean`. Labelled I-unsound.", + "before_toolchain": "nightly-2021-03-13", + "after_toolchain": "nightly-2021-03-17", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "trait MyTrait: One + Two {}\nimpl One for T {\n fn method_one(&self) -> usize {\n 1\n }\n}\nimpl Two for T {\n fn method_two(&self) -> usize {\n 2\n }\n}\nimpl MyTrait for T {}\n\nfn main() {\n let a: &dyn MyTrait = &true;\n assert_eq!(a.method_one(), 1);\n assert_eq!(a.method_two(), 2);\n}\n\n// Re-order traits 'One' and 'Two' between compilation sessions\n\n#[cfg(rpass1)]\ntrait One { fn method_one(&self) -> usize; }\n\ntrait Two { fn method_two(&self) -> usize; }\n\n#[cfg(rpass2)]\ntrait One { fn method_one(&self) -> usize; }\n" + } + ], + "commands": "rm -rf inc\nrustc +TOOLCHAIN --cfg rpass1 -C incremental=inc main.rs -o a.out && ./a.out && echo pass1 ok\nrustc +TOOLCHAIN --cfg rpass2 -C incremental=inc main.rs -o a.out && ./a.out && echo pass2 ok\n# control: a clean build of the edited source passes\nrustc +TOOLCHAIN --cfg rpass2 main.rs -o clean.out && ./clean.out && echo clean ok", + "observe": "Verified locally. On nightly-2021-03-13, the incremental pass 2 binary panics with \"assertion failed: `(left == right)` left: `2`, right: `1`\" (method_two was called for method_one), while a clean rpass2 build prints 'clean ok'. On nightly-2021-03-17 both passes print ok. Regression test: tests/incremental/issue-82920-predicate-order-miscompile.rs." + }, + { + "issue": 135514, + "title": "Rust 1.84 sometimes allows overlapping impls in incremental re-builds (new solver did not record deps for cached tasks)", + "fix_pr": 133828, + "merged": "2024-12-05", + "how_it_violates": "A clean build of the edited source fails with E0119 (conflicting implementations). The incremental rebuild after the same edit accepts it, because coherence results computed through the new solver's cache had no dependency edges. The accepted program is a safe transmute: Vec becomes String and prints ABC. The diagnostics and the success or failure differ from a clean build. Labelled I-unsound, regression-from-stable-to-stable. #135522 added tests/incremental/overlapping-impls-in-new-solver-issue-135514.rs. The fix was already on master before the issue was filed, so 1.85 is fixed and 1.84.x is affected.", + "before_toolchain": "1.84.0", + "after_toolchain": "1.85.0", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "trait Trait {}\n\nstruct S0(T);\n\nstruct S(T);\nimpl Trait for S where S0: Trait {}\n\nstruct W;\n\ntrait Other {\n type Choose;\n}\n\n// first impl\nimpl Other for T {\n type Choose = L;\n}\n\n// second impl: overlaps the first one once `S: Trait` holds\nimpl Other for S {\n type Choose = R;\n}\n\n#[cfg(rpass1)]\nimpl Trait for W {}\n\n#[cfg(rpass1)]\npub fn transmute(_l: L) -> R {\n todo!();\n}\n\n#[cfg(rpass2)]\nimpl Trait for S {}\n\n#[cfg(rpass2)]\nfn use_first_impl(l: L) -> <::To as Other>::Choose {\n l\n}\n\n#[cfg(rpass2)]\nfn use_second_impl(l: as Other>::Choose) -> R {\n l\n}\n\ntrait TyEq {\n type To;\n}\nimpl TyEq for T {\n type To = T;\n}\n\n#[cfg(rpass2)]\nfn transmute_inner(l: L) -> R\nwhere\n T: Trait + TyEq>,\n{\n use_second_impl::(use_first_impl::(l))\n}\n\n#[cfg(rpass2)]\npub fn transmute(l: L) -> R {\n transmute_inner::, L, R>(l)\n}\n\nfn main() {\n if cfg!(rpass2) {\n let v = vec![65_u8, 66, 67];\n let s: String = transmute(v);\n println!(\"{}\", s);\n } else {\n println!(\"pass1\");\n }\n}\n" + } + ], + "commands": "rm -rf inc\nrustc +TOOLCHAIN --cfg rpass1 -A warnings -C incremental=inc main.rs -o a.out && ./a.out\nrustc +TOOLCHAIN --cfg rpass2 -A warnings -C incremental=inc main.rs -o a.out && ./a.out && echo 'incremental pass2 built and ran'\n# control: clean build of the edited source\nrustc +TOOLCHAIN --cfg rpass2 -A warnings main.rs -o clean.out", + "observe": "Verified locally. On 1.84.0, pass 1 prints 'pass1'; the incremental pass 2 compiles without error and prints 'ABC' then 'incremental pass2 built and ran'; the clean build of the same source fails with error[E0119]: conflicting implementations of trait `Other` for type `S`. On 1.85.0 the incremental pass 2 also fails with E0119. On 1.83.0 the old solver rejects pass 1 itself, so use 1.84.0. For nightlies, try nightly-2024-12-04 (bug) and nightly-2024-12-07 (fixed); these were not run." + }, + { + "issue": 139407, + "title": "Instructions missing from (naked_)asm blocks after fixing an assembler error and rebuilding (hard-linked temp files corrupt the incremental cache)", + "fix_pr": 139453, + "merged": "2025-04-11", + "how_it_violates": "Object files were hard-linked from fixed temp paths into the incremental session directory. A session that fails in the assembler (left unfinalized) overwrites a temp file that a previous, finalized session still hard-links. On the next successful build the reused object holds code from the failed session's source, so the binary differs from a clean build. The fix gives temp files a per-invocation random prefix. Labelled I-unsound. The run-make regression test is tests/run-make/dirty-incr-due-to-hard-link. The reporter hit it with plain `cargo run` (no -Csave-temps) on aarch64-apple-darwin. The test, used here, needs -Csave-temps to keep the temp files on x86_64 Linux. It needs a failed build in between (an asm error), but it is deterministic.", + "before_toolchain": "nightly-2025-04-10", + "after_toolchain": "nightly-2025-04-13", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "#[inline(never)]\n#[cfg(any(rpass1, rpass3))]\nfn a() -> i32 {\n 0\n}\n\n#[cfg(any(cfail2))]\nfn a() -> i32 {\n 1\n}\n\nfn main() {\n evil::evil();\n assert_eq!(a(), 0);\n}\n\nmod evil {\n #[cfg(any(rpass1, rpass3))]\n pub fn evil() {\n unsafe {\n std::arch::asm!(\"/* */\");\n }\n }\n\n #[cfg(any(cfail2))]\n pub fn evil() {\n unsafe {\n std::arch::asm!(\"missing\");\n }\n }\n}\n" + } + ], + "commands": "rm -rf inc a.out *.o *.rcgu* \nR=\"rustc +TOOLCHAIN -C incremental=inc -C save-temps main.rs -o a.out\"\n$R --cfg rpass1 && ./a.out && echo pass1 ok\n$R --cfg cfail2; echo '(pass2 is expected to fail with an assembler error)'\n$R --cfg rpass3 && ./a.out && echo pass3 ok", + "observe": "Verified locally on x86_64 Linux. On nightly-2025-04-10 (and on stable 1.70.0, 1.86.0 and 1.87.0), pass 1 prints 'pass1 ok' and pass 2 fails with \"error: invalid instruction mnemonic 'missing'\". Pass 3 has the same source as pass 1, yet its binary panics with \"assertion `left == right` failed left: 1 right: 0\": a() from the failed cfail2 session leaked into the reused object. On nightly-2025-04-13 and on 1.88.0, pass 3 prints 'pass3 ok'." + }, + { + "issue": 162901, + "title": "Diagnostic deduplication breaks with incr comp: incremental builds print 4 copies of an error a non-incremental build prints once (also fixes #106571, duplicate JSON messages)", + "fix_pr": 163461, + "merged": "2026-10-02", + "how_it_violates": "In incremental mode, spans carry a parent (incremental-relative spans). The diagnostic dedup hash included Span::parent, so identical diagnostics hashed differently and were all emitted. The diagnostics from an incremental compile differ from a non-incremental compile of the same source. The fix PR also fixes #106571 (regression from #84762, Jan 2023), where `cargo check --message-format=json` with incremental on prints identical compiler-message lines twice for a proc-macro-generated error, and CARGO_INCREMENTAL=0 does not. No edit is needed: the difference is already there between an incremental and a non-incremental build. A P6 check that compares against a clean incremental build would not see it. It shows only when the reference build is non-incremental (CARGO_INCREMENTAL=0) or when diagnostic counts are compared.", + "before_toolchain": "nightly-2026-10-01", + "after_toolchain": "nightly-2026-10-03", + "reproducible_cheaply": true, + "files": [ + { + "path": "main.rs", + "content": "#[global_allocator]\nstatic A: usize = 0;\n\nfn main() {}\n" + } + ], + "commands": "rm -rf inc\necho \"non-incremental: $(rustc +TOOLCHAIN main.rs -o x 2>&1 | grep -c '^error\\[E0277\\]') E0277 errors\"\necho \"incremental: $(rustc +TOOLCHAIN -C incremental=inc main.rs -o x 2>&1 | grep -c '^error\\[E0277\\]') E0277 errors\"", + "observe": "Verified locally. On nightly-2026-10-01 (also 1.85.0, 1.92.0, 1.96.1), the non-incremental build prints 1 E0277 error and the incremental build prints 4 identical copies ('aborting due to 4 previous errors'). On nightly-2026-10-03 both print 1. Stable 1.70.0 prints 1 in both modes, so this form of the bug arrived later; #106571's proc-macro form dates from nightly-2023-01-03. Regression test: tests/ui/diagnostic-flags/deduplicate-diagnostics-incr.rs." + } + ], + "notes": "All five reproductions were run locally: the old toolchain shows the bug and the newer one does not. Scripts and outputs are in /tmp/p6//. Each issue and PR number, its merge date and its links were checked with gh (issue timeline cross-references and the PR body saying 'Fixes #N'). Merge times: #89619 2021-10-08T11:44Z, #83074 2021-03-15T08:49Z, #133828 2024-12-05T07:07Z, #139453 2025-04-11T16:59Z, #163461 2026-10-02T05:00Z.\n\nNone of the user-supplied known bugs (#122891 \u2026 #114669) is a P6 incremental-vs-clean bug except possibly #114669 (a perf workproduct change, not a bug), so none is included.\n\nThe first two (#89598, #82920) are the cleanest P6 miscompiles: a few lines of Rust, a single rustc invocation per step, and an assertion that fails only in the incremental binary.\n\n#135514 differs in success or failure and in diagnostics (a clean build gives E0119, the incremental build compiles UB). Pin it with stable 1.84.0 vs 1.85.0, because 1.83.0 rejects revision 1 itself.\n\n#139407 needs an intermediate failed session; with -Csave-temps it is deterministic on x86_64 Linux.\n\n#162901/#106571 is a diagnostics-only difference that does not depend on an edit. It shows only when the reference build is non-incremental.\n\nConsidered and dropped:\n- #45469 (2017, stale panic span): a simple repro on nightly-2017-10-30 did not reproduce. The bug needs specific nested-body contexts.\n- #123695 (segfault): this was an LLVM bug, needing an LLVM update.\n- #108216 (macOS unpacked debuginfo only).\n- #84955 (marker traits) and #65401 (delayed bugs): not tried.\n\nOther A-incr-comp + I-unsound candidates for later: #110140, #84955, #65401.\n\nGitHub search hit its rate limit (30/min); space out search calls." + }, + { + "key": "P7", + "bugs": [ + { + "issue": 107001, + "title": "rustc leaves *.rcgu.o object files in the output directory when a post-monomorphization lint (unconditional_panic / arithmetic_overflow) fails the build after codegen has started", + "fix_pr": 110107, + "merged": "2023-04-21", + "how_it_violates": "The const-prop lints (unconditional_panic, arithmetic_overflow) ran in mir_drops_elaborated_and_const_checked, which was only forced lazily, so their errors could fire after codegen had already written the CGU object files. rustc then aborted without deleting them, leaving `..rcgu.o` / `..-cgu.N.rcgu.o` in the output directory (target/debug/deps under cargo). PR #110107 ('Ensure mir_drops_elaborated_and_const_checked when requiring codegen') makes sure that query runs before codegen. Its description says: 'may emit errors while codegen has started, and the compiler would exit leaving object code files around. Found by @cuviper in #109731' (cuviper's comment there: 'each time I try one of these failing tests, it's leaving temporary *.rcgu.o files around'). Issue #107001 is the standalone report with this exact repro. It is still marked open, but its repro stops leaking at this PR. I bisected the nightlies myself and the boundary is exactly this merge: nightly-2023-04-21 (8bdcc62cb) leaks and nightly-2023-04-22 (fec9adcdb) is clean. The commit range between them contains #110107.", + "before_toolchain": "nightly-2023-04-21", + "after_toolchain": "nightly-2023-04-22", + "reproducible_cheaply": true, + "files": [ + { + "path": "code.rs", + "content": "fn main() {\n let a = [1, 2, 3, 4, 5];\n let _x = a[9];\n}\n" + } + ], + "commands": "mkdir out\nrustc +TOOLCHAIN code.rs --out-dir out; echo \"exit=$?\"\nls -A out", + "observe": "Both toolchains fail the same way: 'error: this operation will panic at runtime ... index out of bounds: the length is 5 but the index is 9' with `#[deny(unconditional_panic)]`, exit=1. On nightly-2023-04-21 (also stable 1.70.0), `ls -A out` lists 6 leftover object files, e.g. `code.1tgaf0fuackrygys.rcgu.o code.code.56d798bc-cgu.0.rcgu.o ...`. On nightly-2023-04-22 (also stable 1.71.0 and everything since, through nightly-2026-10-06), `out` is empty. Verified locally." + }, + { + "issue": 139899, + "title": "rustdoc --test leaves rustdoctest* temporary directories behind whenever a doctest fails (compile error or panic)", + "fix_pr": 140706, + "merged": "2025-05-08", + "how_it_violates": "When any doctest failed, rustdoc (or libtest) called process::exit, so the TempDir destructor never ran and the `rustdoctestXXXXXX` directory was never removed. Issue #139899 reports 6197 of them piling up in /tmp. PR #140706 ('[rustdoc] Ensure that temporary doctest folder is correctly removed even if doctests failed') adds a libtest hook that runs after all tests and cleans the folder up. It also adds the regression test tests/run-make/rustdoc/doctest/tempdir-removal, whose two input files are used verbatim below. The leftovers go to TMPDIR, not to target/, so this hits P7 only when TMPDIR points inside the build tree. It is still a real, fixed 'temp dir left behind on the failure path' bug in the toolchain that `cargo test --doc` runs. Note: this is rustdoc, not rustc.", + "before_toolchain": "nightly-2025-05-07", + "after_toolchain": "nightly-2025-05-09", + "reproducible_cheaply": true, + "files": [ + { + "path": "compile-error.rs", + "content": "#![doc(test(attr(deny(warnings))))]\n\n//! ```\n//! let a = 12;\n//! ```\n" + }, + { + "path": "run-error.rs", + "content": "//! ```\n//! panic!();\n//! ```\n" + } + ], + "commands": "mkdir tmp\nfor f in compile-error.rs run-error.rs; do for ed in 2018 2024; do TMPDIR=$PWD/tmp rustdoc +TOOLCHAIN --test $f --edition $ed >/dev/null 2>&1; echo \"$f $ed exit=$?\"; done; done\nls -A tmp", + "observe": "Every run exits 101 (failed doctest) on both toolchains. On nightly-2025-05-07, `ls -A tmp` shows one leftover directory per run (4 in total), e.g. `rustdoctestLP6B1A rustdoctesteGUqc8 rustdoctest0kc32h rustdoctestqVdzVs`. On nightly-2025-05-09, `tmp` is empty. Verified locally for both editions (2018 per-test and 2024 merged doctests)." + } + ], + "notes": "Only two real bugs both fit P7 and have a fixing PR. I reproduced both locally on the before and after toolchains above. Fixed rustc bugs about leftover temporary files are rare: on success, rustc left nothing extra in the output directory on any stable from 1.20 to 1.96 in my matrix (crate types lib/bin/dylib/cdylib/staticlib/proc-macro; emit/LTO/incremental/split-debuginfo variants). All the leaks I found are on error paths or come from split-debuginfo=unpacked. None of the 11 bugs you already knew about (#122891, #130201, ...) is about P7; they are about metadata content and determinism.\n\nThe most useful P7 evidence for maintainers is three open bugs or still-live violations. All three reproduce today.\n\n(A) #161824 (OPEN), a regression caused by #139453 (merged 2025-04-11). With -Csplit-debuginfo=unpacked plus incremental, every rebuild writes a new set of *.rcgu.dwo files (Linux) or *.rcgu.o files (macOS) into target/debug/deps, named with a per-invocation random segment. The old set is never deleted, so files pile up without bound: the reporter had 3.29M files / 341 GB. Verified here on x86_64 Linux. In a fresh `cargo new t`, run `for i in 1 2 3; do touch src/main.rs; RUSTC_WRAPPER= CARGO_INCREMENTAL=1 CARGO_PROFILE_DEV_INCREMENTAL=true CARGO_PROFILE_DEV_SPLIT_DEBUGINFO=unpacked cargo +TOOLCHAIN build -q; find target/debug/deps -name '*.dwo' | wc -l; done`. Results: nightly-2025-04-11 prints 6 6 6; nightly-2025-04-12 prints 6 12 18; 1.88.0 and 1.98.0 print 6 12 18 24. On this VM, ~/.cargo/config.toml sets incremental=false and an sccache wrapper, which hides the bug; hence the env overrides.\n\n(B) #78074 (closed 'completed' by tmiasko on 2021-06-17 with no linked PR) asks the general question: should .rcgu.o files be left when compilation fails after codegen started? bjorn3 says keeping them is intended only for linker errors. The class is still live. An LLVM inline-asm error leaves the objects today: `mod evil { pub fn e() { unsafe { std::arch::asm!(\"missing_instruction\"); } } } fn main(){ evil::e(); }` then `rustc +TOOLCHAIN b.rs --out-dir out`. Leftover .rcgu.o files appear in out/ on 1.70.0, 1.96.1 and nightly-2026-10-06. With -Cincremental the names include #139453's random segment, so failed rebuilds pile up too. A second trigger, 'values of the type [u64; 1<<60] are too big' from `fn g() -> Option<[T; 1 << 60]> { None }` called as g::(), left 1 .rcgu.o file on 1.78 to 1.94. It is clean from nightly-2026-01-22 (last leaking nightly-2026-01-21, commits 5c49c4f7c..eda76d9d1). That looks incidental, probably #149209 'Move LTO to OngoingCodegen::join', but I did not bisect inside the rollup, so treat that attribution as unconfirmed. The original #78074 repro stopped erroring at all by 1.53, so it no longer demonstrates anything.\n\n(C) #143278 (OPEN): rustc does not delete partial outputs or .rcgu.o files when killed by SIGTERM. It needs a timing race, so it is not cheap to reproduce.\n\nAlso related but not leftovers: #148257 (OPEN, Windows only, flaky): 'failed to remove temporary directory ... .temp-archive' (os error 32), which can fail the build and leave the archive temp dir behind. Needs Windows plus a racing scanner or indexer, so not cheap. #111157 (closed, won't fix): `-o /dev/null` makes rustc try to create /dev/rmetaXXXXXX, which shows that the rmeta temp dir lives next to the output. #139963 (closed by PR #150311, 2025-12-24): the link temp dir moved from $TMPDIR into the output directory, so it is now in P7's scope. In my linker-failure test (-Clinker=false) on nightly-2026-10-06, rustc left only the intentional .rcgu.o files, not a rustc* directory.\n\nNot leftovers, for the record: incremental `s-*.lock` files next to finalized session dirs are by design, and so is an `s-*-working` dir left after a failed session; the next session garbage-collects it. mirth should whitelist both. GitHub search was rate-limited at times, so the coverage of closed issues is good but not exhaustive." + }, + { + "key": "tracked", + "bugs": [ + { + "issue": 89598, + "title": "VTable-related miscompilation with incremental compilation: untracked tcx vtable cache gives a stale vtable after an upstream crate reorders trait methods", + "fix_pr": 89619, + "merged": "2021-10-08", + "how_it_violates": "PR #86475 put a global vtable-allocation cache in the tcx that bypassed dependency tracking (issue #89598: 'circumvents incremental compilation's dependency tracking... object files might not get updated when methods within a trait get shuffled around... trait object method calls will invoke the wrong method'). The vtable layout comes from the trait's method order, and when the trait lives in another crate that order is read from that crate's metadata, but the result never recorded a dependency on it. When the upstream crate reorders the methods, the downstream incremental build reuses the stale CGU and calls the wrong method. Fixed by turning vtable_allocation() into a tracked query (#89619). The issue's own test is single-crate. The repro here is a cross-crate variant (trait in crate a, call site in bin), which I verified on both toolchains.", + "before_toolchain": "nightly-2021-10-07", + "after_toolchain": "nightly-2021-10-09", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.sh", + "content": "# $1 = 1 or 2: order of trait methods in upstream crate a\nif [ \"$1\" = 1 ]; then M=\" fn method1(&self) -> u32;\n fn method2(&self) -> u32;\"; else M=\" fn method2(&self) -> u32;\n fn method1(&self) -> u32;\"; fi\ncat > a.rs < u32 { 17 }\n fn method2(&self) -> u32 { 42 }\n}\nEOT\n" + }, + { + "path": "main.rs", + "content": "extern crate a;\nfn main() {\n let x: &dyn a::Foo = &0u32;\n println!(\"method1() returned {}\", mod1::foo(x));\n}\nmod mod1 {\n pub fn foo(x: &dyn a::Foo) -> u32 { x.method1() }\n}\n" + } + ], + "commands": "rm -rf inc-a inc-main *.rlib main\nfor v in 1 2; do sh gen.sh $v; rustc +TOOLCHAIN --crate-type rlib -C incremental=inc-a a.rs && rustc +TOOLCHAIN -C incremental=inc-main -L . main.rs && echo \"step$v: $(./main)\"; done", + "observe": "Before the fix (nightly-2021-10-07): 'step1: method1() returned 17' then 'step2: method1() returned 42'. After a reorder in the upstream crate, the incrementally rebuilt binary calls method2 through the stale vtable (a silent miscompile). After the fix (nightly-2021-10-09), both steps print 'method1() returned 17'." + }, + { + "issue": 82920, + "title": "Miscompilation with incr. comp.: predicates sorted by DefId, which is unstable for upstream items, so cached results disagree across sessions", + "fix_pr": 83074, + "merged": "2021-03-15", + "how_it_violates": "The query results (super predicates / bounds) were sorted by DefId. For items in another crate, the DefId index comes from that crate's metadata and changes when the upstream crate inserts or reorders items, while the DefPathHash, which is what incremental tracks, stays the same. So the green or cached result encodes an order that depends on untracked upstream metadata. The original report (rust-analyzer workspace, cross-crate) produced segfaults and out-of-bounds panics until cargo clean. #83074 sorts by DefPathHash instead ('Even if an item does not change between compilation sessions, it may end up with a different DefId... the query result should be unchanged if the query inputs are unchanged'). The repro is a cross-crate variant of tests/incremental/issue-82920-predicate-order-miscompile.rs, verified on both toolchains. On the old nightly it shows up as an unstable-fingerprint ICE and not as the silent miscompile.", + "before_toolchain": "nightly-2021-03-14", + "after_toolchain": "nightly-2021-03-17", + "reproducible_cheaply": true, + "files": [ + { + "path": "gen.sh", + "content": "# $1 = 1 or 2: swap the definition order of One/Two in upstream crate a\nif [ \"$1\" = 1 ]; then X='pub trait One { fn method_one(&self) -> usize; }\npub trait Two { fn method_two(&self) -> usize; }'; else X='pub trait Two { fn method_two(&self) -> usize; }\npub trait One { fn method_one(&self) -> usize; }'; fi\ncat > a.rs < One for T { fn method_one(&self) -> usize { 1 } }\nimpl Two for T { fn method_two(&self) -> usize { 2 } }\nEOT\n" + }, + { + "path": "main.rs", + "content": "use a::{One, Two};\ntrait MyTrait: One + Two {}\nimpl MyTrait for T {}\nfn main() {\n let a: &dyn MyTrait = &true;\n println!(\"method_one={} method_two={}\", a.method_one(), a.method_two());\n}\n" + } + ], + "commands": "rm -rf inc-a inc-main *.rlib main\nfor v in 1 2; do sh gen.sh $v; rustc +TOOLCHAIN --edition 2018 --crate-type rlib -C incremental=inc-a a.rs && rustc +TOOLCHAIN --edition 2018 -C incremental=inc-main --extern a=liba.rlib main.rs && echo \"step$v: $(./main)\"; done", + "observe": "Before the fix (nightly-2021-03-14): step1 prints 'method_one=1 method_two=2', then the step2 rebuild of main panics with \"found unstable fingerprints for super_predicates_that_define_assoc_type(...)\" (query stack: super_predicates_of `MyTrait`). The recomputed result differs from the cached one only because the upstream DefIds moved. After the fix (nightly-2021-03-17), step2 builds and prints 'method_one=1 method_two=2'." + }, + { + "issue": 84252, + "title": "ICE: found unstable fingerprints for has_global_allocator(): the query read untracked CStore state", + "fix_pr": 84260, + "merged": "2021-04-17", + "how_it_violates": "has_global_allocator was answered from the CStore (the crate loader's per-crate metadata state) without recording any dependency. PR #84260: 'This query reads from untracked global state in `CStore`'. It was fixed by making it eval_always. The repro is the regression test tests/incremental/issue-84252-global-alloc.rs: removing #[global_allocator] between sessions leaves the cached result stale, and recomputing it gives a different value. The sibling fix #83153 (extern_mod_stmt_cnum made eval_always for the same reason, issue #83126) is a related example. #83126 is still open on GitHub, so I left it out as a separate entry.", + "before_toolchain": "nightly-2021-04-16", + "after_toolchain": "nightly-2021-04-19", + "reproducible_cheaply": true, + "files": [ + { + "path": "lib.rs", + "content": "#![crate_type=\"lib\"]\n#![crate_type=\"cdylib\"]\n\n#[allow(unused_imports)]\nuse std::alloc::System;\n\n#[cfg(bpass1)]\n#[global_allocator]\nstatic ALLOC: System = System;\n" + } + ], + "commands": "rm -rf inc out; mkdir out\nrustc +TOOLCHAIN -C incremental=inc --out-dir out --cfg bpass1 lib.rs; echo rc1=$?\nrustc +TOOLCHAIN -C incremental=inc --out-dir out lib.rs; echo rc2=$?", + "observe": "Before the fix (nightly-2021-04-16): the second build panics with \"found unstable fingerprints for has_global_allocator(lib[8787]): false\" (assertion left/right Fingerprint mismatch in rustc_query_system plumbing.rs) and exits non-zero. After the fix (nightly-2021-04-19): rc1=0 and rc2=0." + }, + { + "issue": 111295, + "title": "debugger_visualizer: edits to the visualizer script are not picked up under incremental (stale content that is exported into metadata)", + "fix_pr": 111641, + "merged": "2023-05-19", + "how_it_violates": "The debugger_visualizers query result (the script contents, which are also encoded into crate metadata so downstream binaries embed upstream visualizers) did not depend on the script file. After the .py file was edited, an incremental rebuild reused the stale contents. #111641 ('Fix dependency tracking for debugger visualizers') made the query eval_always, and the fix also hashes the visualizer contents into crate_hash. The code comment says 'that content is exported into crate metadata, so any changes to it need to be reflected in the crate hash', so downstream crates that read it from metadata also see the change. The same PR fixes #111227 (an ICE from changing a natvis file with --crate-type=rlib) and #111226 (dep-info). The repro is single-crate, as in the issue. No gdb is needed: read the .debug_gdb_scripts section directly.", + "before_toolchain": "nightly-2023-05-18", + "after_toolchain": "nightly-2023-05-21", + "reproducible_cheaply": true, + "files": [ + { + "path": "foo.rs", + "content": "#![debugger_visualizer(gdb_script_file = \"foo.py\")]\n\nfn main() {\n let x = 1;\n println!(\"breakpoint {x}\");\n}\n" + } + ], + "commands": "rm -rf incremental foo\necho \"print('hello!')\" > foo.py\nrustc +TOOLCHAIN -Cincremental=incremental -g foo.rs\necho \"print('hello world')\" > foo.py\nrustc +TOOLCHAIN -Cincremental=incremental -g foo.rs\nobjcopy -O binary --only-section=.debug_gdb_scripts foo /dev/stdout | tr -c '[:print:]' ' ' | grep -o \"print('[^']*')\"", + "observe": "Before the fix (nightly-2023-05-18): the rebuilt binary's .debug_gdb_scripts still contains print('hello!'), the stale script. After the fix (nightly-2023-05-21): it contains print('hello world'). Linux/ELF only; needs binutils objcopy." + }, + { + "issue": 114669, + "title": "Metadata was never reused across incremental sessions: always re-encoded (fixed by #114669 'Make metadata a workproduct and reuse it', with prerequisite #143247 'Avoid depending on forever-red DepNode when encoding metadata')", + "fix_pr": 114669, + "merged": "2025-07-04", + "how_it_violates": "A performance violation of 'metadata is reused (not re-encoded) when nothing it depends on changed'. Before July 2025, metadata encoding depended on the forever-red DepNode (iter_local_def_id / def_path_table read DepNodeIndex::FOREVER_RED_NODE) and was not a dep-graph task, so every incremental session re-encoded the .rmeta from scratch. #143247 (merged 2025-07-04, split out of #114669 'for perf') removed the forever-red read by depending on `analysis` instead. #114669 (merged 2025-07-04) wraps encoding in a Metadata dep-node task, saves the rmeta as a work product ('metadata') in the incremental dir, and when the node is green it hardlinks or copies the saved file instead of encoding ('can yield substantial gains (~10%)... if all the changes are in upstream crates and have no effect on it'). Observable without logs: the reused rmeta is a hardlink of the work product, so its inode stays the same across no-op sessions. Verified on both toolchains. This is a numbered PR, not an issue: no separate GitHub issue exists, so the issue field holds the PR number.", + "before_toolchain": "nightly-2025-07-03", + "after_toolchain": "nightly-2025-07-06", + "reproducible_cheaply": true, + "files": [ + { + "path": "a.rs", + "content": "pub fn public() -> u32 { private() }\nfn private() -> u32 { 1 }\n" + } + ], + "commands": "rm -rf out inc; mkdir out\nfor i in 1 2; do rustc +TOOLCHAIN --edition 2021 --crate-type lib --emit=metadata,link -C incremental=inc --out-dir out a.rs; echo \"run$i inode=$(stat -c %i out/liba.rmeta) links=$(stat -c %h out/liba.rmeta)\"; done", + "observe": "Before (nightly-2025-07-03): run1 and run2 report different inodes and links=1, so the rmeta is freshly encoded on the unchanged rebuild. After (nightly-2025-07-06): links=2 (the file is hardlinked to the 'metadata' work product in the incremental dir) and run2 has the same inode as run1, so the metadata was reused rather than re-encoded. Note: after the fix, editing even a private fn body still changed the inode in my test (the Metadata node went red), so the no-op rebuild is the clean demonstration. Needs a filesystem with hardlinks; on one without them link_or_copy falls back to copying, and the inode check does not apply." + } + ], + "notes": "I ran every reproduction on both toolchains and saw the stated before/after behaviour (rustup nightlies, x86_64 Linux). Scratch dirs: /tmp/mr89598, /tmp/mr82920, /tmp/mr84252, /tmp/mr111295, /tmp/mr114669.\n\nThe #89598 and #82920 repros are my cross-crate variants of the issues' tests: the trait lives in an upstream rlib, and both crates are built with -C incremental. To get the stale results I edited source between sessions instead of changing --cfg, because some command-line changes reset the whole incremental cache. In #82920 the cross-crate variant ICEs (unstable fingerprint) on the old nightly; it does not silently miscompile. The original report, a rust-analyzer workspace, did miscompile.\n\n#83126 (extern_mod_stmt_cnum unstable fingerprints, fixed in spirit by #83153 'Mark extern_mod_stmt_cnum as eval_always', merged 2021-03-16) is a sibling of #84252: an untracked CStore read. I left it out as a separate entry because the issue is still open on GitHub.\n\n#158955 ('Incremental ThinLTO doesn't reuse unchanged CGUs originating from upstream crates') is still open and unfixed. It fits the 'work products never reused' half of the property, but there is no fixed toolchain to compare against. Its symptom, from the issue: with -Clto=thin -Cincremental, upstream CGUs are re-optimized every session.\n\nRelated but not included: #154724 (merged 2026-09-26) computes crate_hash from encoded metadata instead of HIR, related to open issue #94878. Its aim is to make the SVH, which downstream incremental sessions use to detect upstream changes, cover everything in metadata. It is a hardening change, not a fix for a reported stale-build bug.\n\nThe GitHub search API was rate-limited for most of this session, so the survey came from label listings (A-incr-comp, closed, non-ICE), the tests/incremental directory, and blame of crate_hash. There may be more cross-crate stale-result bugs that I did not find." + } +] \ No newline at end of file diff --git a/docs/motivating/run.py b/docs/motivating/run.py new file mode 100755 index 0000000..b6a124d --- /dev/null +++ b/docs/motivating/run.py @@ -0,0 +1,54 @@ +#!/usr/bin/env python3 +"""Run the reproductions in bugs.json on the toolchain from before each fix and the +one after it, and save both outputs. + + docs/motivating/run.py [issue…] + +bugs.json is a list of {issue, before_toolchain, after_toolchain, files: [{path, +content}], commands, observe, ...}; `commands` uses the word TOOLCHAIN where the +toolchain goes. Outputs go to out//{before,after}.txt. +""" + +import json +import os +import subprocess +import sys +import tempfile +from pathlib import Path + +here = Path(__file__).resolve().parent +bugs = json.loads((here / "bugs.json").read_text()) +only = {int(a) for a in sys.argv[1:]} + + +def install(toolchain): + r = subprocess.run(["rustup", "toolchain", "install", toolchain, "--profile", "minimal"], + capture_output=True, text=True) + return r.returncode == 0 + + +for bug in bugs: + if only and bug["issue"] not in only: + continue + out = here / "out" / str(bug["issue"]) + out.mkdir(parents=True, exist_ok=True) + for side in ("before", "after"): + toolchain = bug[f"{side}_toolchain"] + if not install(toolchain): + (out / f"{side}.txt").write_text(f"could not install {toolchain}\n") + continue + with tempfile.TemporaryDirectory() as d: + for f in bug["files"]: + path = Path(d) / f["path"] + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(f["content"]) + env = dict(os.environ, RUSTC_WRAPPER="", CARGO_TERM_COLOR="never") + script = bug["commands"].replace("TOOLCHAIN", toolchain) + try: + r = subprocess.run(["bash", "-c", script], cwd=d, env=env, capture_output=True, + text=True, timeout=900) + text = f"$ toolchain {toolchain}\n{r.stdout}\n--- stderr ---\n{r.stderr}\nexit {r.returncode}\n" + except subprocess.TimeoutExpired: + text = f"$ toolchain {toolchain}\ntimed out\n" + (out / f"{side}.txt").write_text(text) + print(f"#{bug['issue']}: done", flush=True) diff --git a/docs/motivating/toolchains.txt b/docs/motivating/toolchains.txt new file mode 100644 index 0000000..bc1996c --- /dev/null +++ b/docs/motivating/toolchains.txt @@ -0,0 +1,47 @@ +1.45.0 +1.46.0 +1.84.0 +1.85.0 +nightly-2016-08-27 +nightly-2016-08-30 +nightly-2017-11-17 +nightly-2017-11-20 +nightly-2019-10-05 +nightly-2019-10-08 +nightly-2020-01-22 +nightly-2020-01-24 +nightly-2021-03-13 +nightly-2021-03-17 +nightly-2021-04-16 +nightly-2021-04-19 +nightly-2021-04-29 +nightly-2021-05-01 +nightly-2021-10-07 +nightly-2021-10-10 +nightly-2023-04-21 +nightly-2023-04-22 +nightly-2023-05-18 +nightly-2023-05-21 +nightly-2023-11-26 +nightly-2023-11-28 +nightly-2024-01-01 +nightly-2024-01-05 +nightly-2024-03-24 +nightly-2024-03-26 +nightly-2024-09-17 +nightly-2024-09-19 +nightly-2025-03-28 +nightly-2025-03-30 +nightly-2025-04-10 +nightly-2025-04-13 +nightly-2025-05-07 +nightly-2025-05-09 +nightly-2025-07-03 +nightly-2025-07-06 +nightly-2025-07-24 +nightly-2025-07-26 +nightly-2025-12-10 +nightly-2026-07-20 +nightly-2026-09-25 +nightly-2026-10-01 +nightly-2026-10-03 diff --git a/docs/plan.md b/docs/plan.md index 0d78f8a..fa4594c 100644 --- a/docs/plan.md +++ b/docs/plan.md @@ -77,6 +77,9 @@ frame lookup_deprecation_entry(dep) decode tables.lookup_deprecation x1 | P6 | an incremental rebuild after an edit to a fixture gives the same `.rmeta` bytes as a clean build | incremental bugs | | P7 | nothing is left in the target directory beyond the expected outputs | dangling files | +[`motivating.md`](motivating.md) lists the real rustc bugs that violated each of these, +each reproduced on a toolchain from before its fix. + ## Edits that break it Each is a small patch to rustc's metadata code, of the kind a contributor diff --git a/docs/properties.md b/docs/properties.md new file mode 100644 index 0000000..27288a7 --- /dev/null +++ b/docs/properties.md @@ -0,0 +1,623 @@ +# Properties from 1,000 rustc bugs + +The 1,000 most recently closed rust-lang/rust issues labelled `C-bug` that a merged +PR closed (2025-12-10 to 2026-10-05; 480 of them ICEs) were read by ten agents, +100 each, looking for invariants that a past bug violated and that a mechanical check +could test on any crate, without knowing the bug. 94 candidates were merged into 30 +properties. A second pass checked every cited issue against its text and kept only the +issues that really violate the property: 107 of 156 citations survived. The crash +baseline (no ICE, hang or stack overflow; 125 issues) was not checked this way, since any +harness detects a crash. + +The agents were small models and the checking was done by small models too. Treat the +issue lists as leads, not proof: the checking rejected most of the original citations +for some properties, and it also accepted at least one citation that does not fit +(#159677 under "incremental equals clean": it is about the library search path). + +| # | property | kind | issues | cost | mirth | +|---|---|---|---|---|---| +| 1 | [Interned type-system values are well-formed](#1) | invariant | 12 | cheap | Good | +| 2 | [Old and next trait solvers agree](#2) | differential | 11 | moderate | Good | +| 3 | [extern "C" ABI matches the platform C compiler](#3) | differential | 9 | moderate | Low | +| 4 | [MIR passes validation after every pass at every mir-opt-level](#4) | invariant | 8 | cheap | Moderate | +| 5 | [Library operations stay sound under injected panics and allocation failures](#5) | semantic | 7 | expensive | Poor | +| 6 | [rustdoc accepts every crate rustc accepts](#6) | differential | 6 | cheap | Low to moderate | +| 7 | [Every span is valid](#7) | invariant | 6 | cheap | Good for the metadata half (hook span encoding) | +| 8 | [Semantically neutral edits do not change results](#8) | differential | 5 | moderate | Moderate | +| 9 | [Layout views agree](#9) | invariant | 4 | cheap | Good | +| 10 | [rustdoc output is the same however an item is re-exported](#10) | differential | 4 | cheap | Poor | +| 11 | [P6+ incremental equals clean (outputs and diagnostics)](#11) | differential | 3 | moderate | Excellent | +| 12 | [Parallel front end is deterministic](#12) | differential | 3 | cheap | Good | +| 13 | [Results agree across opt-level, mir-opt-level, codegen-units and LTO](#13) | semantic | 3 | expensive | Low | +| 14 | [A successful compile contains no error types](#14) | invariant | 3 | cheap | Good | +| 15 | [Each encoded record is written once and hashed after it is final](#15) | invariant | 3 | cheap | Excellent | +| 16 | [#[expect] is fulfilled exactly when the lint fires](#16) | differential | 3 | moderate | Low | +| 17 | [Metadata reads hit entries the writer wrote](#17) | invariant | 2 | moderate | Excellent | +| 18 | [Query results contain no inference variables](#18) | invariant | 2 | cheap | Excellent | +| 19 | [Machine-applicable suggestions apply cleanly](#19) | differential | 2 | moderate | Poor | +| 20 | [Polonius accepts at least what NLL accepts, and no more than is sound](#20) | differential | 2 | moderate | Low | +| 21 | [Unstable syntax is gated before expansion](#21) | invariant | 2 | cheap | Poor | +| 22 | [P5+ metadata determinism under irrelevant perturbation](#22) | differential | 1 | cheap | Excellent | +| 23 | [P4+ no query reads untracked state](#23) | invariant | 1 | cheap | Excellent | +| 24 | [Symbol names are injective and stable](#24) | invariant | 1 | cheap | Good | +| 25 | [dyn types have a dyn-compatible principal](#25) | invariant | 1 | cheap | Good | +| 26 | [Built-in attributes reject malformed arguments](#26) | invariant | 1 | cheap | Poor | +| 27 | [Advertised target features exist in LLVM](#27) | invariant | 1 | cheap | Poor | +| 28 | [Eq and Hash agree for std types](#28) | semantic | 1 | cheap | None | +| 29 | [Each diagnostic is emitted once](#29) | invariant | 0 | cheap | Poor | + + +## 1. Interned type-system values are well-formed + +**Statement.** Every interned GenericArgs-carrying value (TraitRef, AliasTy, FnDef, etc.) must have arguments whose count and kind match the corresponding generics_of definition. Every interned const value must have a valtree whose structure matches its type. No ParamEnv must contain two predicates with the same parameter (e.g., ConstArgHasType) that specify different constraints. + +**Check.** In an instrumented release-mode rustc, run the existing debug-only checks at every intern site: debug_assert_args_compatible, a valtree-vs-type conformance check, and a ParamEnv duplicate/conflict scan when param_env is built. Run on the corpus plus the UI suite. + +**Cost:** cheap per crate (checks at intern sites). **Kind:** invariant. **Subsystem:** type-system / const generics. + +**Tested today?** Some of these checks exist as debug_assert in debug compilers only. Valtree/type conformance and ParamEnv conflicts are not checked at all. The issues were found as ICEs later in the pipeline. + +**mirth.** Good. Inject the checks at the mk_* or intern call sites and record the value, the caller's query and the DefId on failure. This turns a late ICE into a report at the point of creation. + +**Bugs that violated it:** + +- [#150506](https://github.com/rust-lang/rust/issues/150506) ICE: valtree: `expected leaf, got Value` +- [#150712](https://github.com/rust-lang/rust/issues/150712) ICE: `expected branch, got Leaf` +- [#150734](https://github.com/rust-lang/rust/issues/150734) ICE `expected leaf, got Value` +- [#151126](https://github.com/rust-lang/rust/issues/151126) ICE with --emit=mir :` expected ConstKind::Value, got X/#0` +- [#158675](https://github.com/rust-lang/rust/issues/158675) [ICE]: did not expect duplicate `ConstParamHasTy` for `N/#1` in param-env: ParamEnv { +- [#157189](https://github.com/rust-lang/rust/issues/157189) [ICE]: `args not compatible with generics for Borrow` +- [#137084](https://github.com/rust-lang/rust/issues/137084) mgca: index out of bounds +- [#150673](https://github.com/rust-lang/rust/issues/150673) ICE: delegation: index out of bounds +- [#150714](https://github.com/rust-lang/rust/issues/150714) ICE: abi: index out of bounds (`offsets[FieldIdx::new(i)]`) +- [#150841](https://github.com/rust-lang/rust/issues/150841) ICE `const tuple must have a tuple type` +- [#151024](https://github.com/rust-lang/rust/issues/151024) ICE `const array must have an array type` +- [#151186](https://github.com/rust-lang/rust/issues/151186) ICE:index out of bounds: the len is 0 but the index is 0 + + +## 2. Old and next trait solvers agree + +**Statement.** The next-solver and old solver must behave identically on any crate: accepting and rejecting the same items, inferring the same types for each body (up to region erasure), and selecting the same impls. The cited issues are cases where this invariant was violated, with the next-solver either panicking (ICE) or disagreeing with the old solver on acceptance/rejection of code. + +**Check.** Build the corpus with both solvers. Compare exit status and diagnostics, and per body DefPath compare typeck_results (node types, method resolutions) and codegen Instances. Every disagreement is a bug in one of the two solvers. + +**Cost:** moderate (two full typechecks). **Kind:** differential. **Subsystem:** trait-system. + +**Tested today?** The next-solver CI job runs the UI suite with the next solver, and crater runs happen occasionally. Per-body inferred types are never compared. + +**mirth.** Good. mirth can record typeck_results and selected impls per DefPath in both configurations, which gives a precise diff instead of only accept/reject. + +**Bugs that violated it:** + +- [#102580](https://github.com/rust-lang/rust/issues/102580) Overflow when deriving Clone on a struct with a recursive GAT +- [#90950](https://github.com/rust-lang/rust/issues/90950) HRTB bounds not resolving correctly (take 3, lifetimes on the RHS) +- [#152789](https://github.com/rust-lang/rust/issues/152789) `-Znext-solver`: trait object candidate ICE "could not replace AliasTerm" +- [#151329](https://github.com/rust-lang/rust/issues/151329) [ICE]: `could not replace AliasTerm` (unsatisifed bounds) +- [#151957](https://github.com/rust-lang/rust/issues/151957) ICE: `entered unreachable code: PointeeSized is removed during lowering` with `-Z next-solver=globally` and recursive associated type bound +- [#151323](https://github.com/rust-lang/rust/issues/151323) [ICE]: !tcx.next_trait_solver_globally() +- [#151322](https://github.com/rust-lang/rust/issues/151322) [ICE]: !self.tcx.next_trait_solver_globally() +- [#151318](https://github.com/rust-lang/rust/issues/151318) [ICE]: error performing operation: query type op +- [#138274](https://github.com/rust-lang/rust/issues/138274) [bug] When I Use tauri-plugin-http and reqwest either, I got a panic +- [#137916](https://github.com/rust-lang/rust/issues/137916) ICE Unsize coercion, but `Box<{async block@file.rs}>` isn't coercible to `Box` +- [#151462](https://github.com/rust-lang/rust/issues/151462) [ICE]: `Inconsistent rustc_transmute::is_transmutable(...) result, got Yes` + + +## 3. extern "C" ABI matches the platform C compiler + +**Statement.** For every target, a function signature with C-compatible types (unions, structs with floats, bool, homogeneous vector aggregates, and other aggregates) used in extern \"C\" functions must be lowered to argument passing and return value handling (register vs stack, indirect vs direct, sign/zero extension) that matches the behavior of clang/gcc for equivalent C declarations. + +**Check.** Generate random C-compatible type signatures, emit matching Rust and C, and cross-call them in both directions (abi-cafe style). Run natively or under qemu for each tier-1/2 target. Optionally diff rustc's FnAbi against clang's LLVM IR attributes per parameter. + +**Cost:** moderate (cross targets need qemu). **Kind:** differential. **Subsystem:** codegen / abi. + +**Tested today?** tests/ui/abi/compatibility.rs and per-target codegen tests check fixed cases. abi-cafe is not in CI. Most of these bugs were reported by users on less common targets (SPARC, PowerPC, LoongArch). + +**mirth.** Low. This is an output-level differential with C compilers. mirth could dump the FnAbi of every extern fn to drive the comparison, but the core harness is abi-cafe plus qemu. + +**Bugs that violated it:** + +- [#121408](https://github.com/rust-lang/rust/issues/121408) Clang vs wasm32-{emscripten,wasi} rustc C ABI mismatch w.r.t. "singleton" unions +- [#162011](https://github.com/rust-lang/rust/issues/162011) ABI mismatch on powerpc64 (elfv1) for union containing floats +- [#163074](https://github.com/rust-lang/rust/issues/163074) `check_abi.rs` fails due to ArgAttributes mismatch with the ABI on LoongArch64 +- [#122620](https://github.com/rust-lang/rust/issues/122620) sparc64 has incorrect ABI for struct containing f64 and f32 +- [#115399](https://github.com/rust-lang/rust/issues/115399) ICE in sparc64 `fn_abi_of_instance` +- [#147883](https://github.com/rust-lang/rust/issues/147883) "Size::sub: 0 - 8 would result in negative size" ICE on sparc +- [#159244](https://github.com/rust-lang/rust/issues/159244) Miscompilation with FFI `bool` return type on AArch64 +- [#43894](https://github.com/rust-lang/rust/issues/43894) struct pass-by-value failing on SPARC +- [#161382](https://github.com/rust-lang/rust/issues/161382) Mishandling of AAPCS64 Homogeneous Vector Aggregates (HVAs) on aarch64-unknown-linux-gnu + + +## 4. MIR passes validation after every pass at every mir-opt-level + +**Statement.** MIR optimization passes must preserve well-formedness invariants: const arguments must have the types their parameters declare, locals must only be accessed between their StorageLive and StorageDead markers, and place types as well as const values must be fully normalized. This holds at all -Zmir-opt-level settings (0, 2, 4) and with -Zinline-mir, and can be verified mechanically by running -Zvalidate-mir -Zlint-mir on real crates. + +**Check.** Build every crate in the corpus with -Zvalidate-mir -Zlint-mir -Zmir-opt-level={0,2,4} -Zinline-mir and -Cdebug-assertions on the compiler. Any validation failure counts. This uses existing flags; you only need to run them on real crates. + +**Cost:** cheap (existing flags; one extra build per opt-level). **Kind:** invariant. **Subsystem:** mir-build / mir-opt / const-eval. + +**Tested today?** mir-opt tests use the validator, but UI tests and real crates do not run with -Zvalidate-mir by default, and mir-opt-level=4 is rarely exercised on real code. + +**mirth.** Moderate. No MIR rewriting is needed because the flags exist. mirth helps by recording which pass first broke validation, and by giving every report one shared corpus and harness. + +**Bugs that violated it:** + +- [#156409](https://github.com/rust-lang/rust/issues/156409) [ICE]: `CTFE tried to evaluate type-const` +- [#154750](https://github.com/rust-lang/rust/issues/154750) [ICE]: `attempting to project to field at offset 0 with size 8 into immediate with layout TyAndLayout` +- [#154748](https://github.com/rust-lang/rust/issues/154748) [ICE]: invalid field access on immediate +- [#152962](https://github.com/rust-lang/rust/issues/152962) [ICE]: mgca: broken mir `Failed subtyping u8 and usize` +- [#158231](https://github.com/rust-lang/rust/issues/158231) SimplifyComparisonIntegral introduces access to a dead local variable +- [#151647](https://github.com/rust-lang/rust/issues/151647) ICE: mGCA+GCI: Broken MIR: equate_normalized_input_or_output: NoSolution +- [#151579](https://github.com/rust-lang/rust/issues/151579) ICE with `-Znext-solver` when accessing hir place +- [#120811](https://github.com/rust-lang/rust/issues/120811) ICE: Broken MIR: NoSolution + + +## 5. Library operations stay sound under injected panics and allocation failures + +**Statement.** Library operations (particularly BTreeMap, Arc, Box, and array functions) fail to maintain soundness invariants when user callbacks (comparators, closures) panic or when allocations fail. Violations include: improper drops of unprocessed values, state corruption leading to double-frees, use-after-free from incorrect strong count management, and aliasing violations when custom allocators or mutable reference derivation are involved. Additionally, allocators that return pointers with provenance exceeding the requested size can be unsoundly treated as having restricted provenance, violating the deallocation contract. + +**Check.** Run library tests and property tests under Miri with fault injection: closures and comparators that panic on the Nth call (for every N), a failing allocator, and a counting Drop type. Check for no Miri UB, drops equal to constructions, and a structure that is still usable or droppable afterwards. + +**Cost:** expensive (Miri, N-sweep). **Kind:** semantic. **Subsystem:** library (alloc/core/std). + +**Tested today?** Library tests run under Miri in CI, but there is no systematic sweep that panics at every callback position or fails allocation at every point. + +**mirth.** Poor. This tests the library, not the compiler. Use Miri plus a fault-injection harness. mirth is not needed. + +**Bugs that violated it:** + +- [#162720](https://github.com/rust-lang/rust/issues/162720) `Arc::new_cyclic_in` uses pointer derived from mutable reference unsoundly +- [#162719](https://github.com/rust-lang/rust/issues/162719) `Box::into_unique(self) -> (Unique, A)` goes through `&mut *ptr` +- [#158165](https://github.com/rust-lang/rust/issues/158165) BTreeMap::split_off is not panic-safe leading to a potential double-free +- [#157203](https://github.com/rust-lang/rust/issues/157203) Unsoundness in `Arc::make_mut` if `handle_alloc_error` unwinds +- [#155746](https://github.com/rust-lang/rust/issues/155746) UAF in allocator-backed `Arc::make_mut` after caught unwind +- [#152211](https://github.com/rust-lang/rust/issues/152211) `array::map` and `array::try_map` do not drop ZSTs properly +- [#160815](https://github.com/rust-lang/rust/issues/160815) `Condvar` and `Mutex` in `std::sys::pal::unix::sync` violate aliasing rules + + +## 6. rustdoc accepts every crate rustc accepts + +**Statement.** rustdoc must not ICE on any Rust code that passes `cargo check`, particularly code involving type-relative paths and anon consts with `--generate-link-to-definition`, and must not allow proc-macro-generated spans to overwrite source spans in documentation links + +**Check.** For each crate in the corpus, run cargo check, then cargo rustdoc with each mode. Any rustdoc failure on a crate that checks is a bug. Validate the JSON output against rustdoc-types. + +**Cost:** cheap. **Kind:** differential. **Subsystem:** rustdoc. + +**Tested today?** docs.rs and crater cover the default HTML mode. --generate-link-to-definition and JSON on real crates are barely exercised. + +**mirth.** Low to moderate. This is mostly a harness differential. mirth could record which TypeckResults body was consulted for each path in order to check that it is the owning body. + +**Bugs that violated it:** + +- [#156418](https://github.com/rust-lang/rust/issues/156418) rustdoc ICEs on free calls, method calls & type-relative paths in anon consts under `--generate-link-to-definition` +- [#149089](https://github.com/rust-lang/rust/issues/149089) ICE when documenting embedded-io with nightly: `node HirId(...) cannot be placed in TypeckResults` +- [#150153](https://github.com/rust-lang/rust/issues/150153) ICE: rustdoc: `node HirId cannot be placed in TypeckResults with hir_owner DefId` +- [#147882](https://github.com/rust-lang/rust/issues/147882) Internal compiler error when building docs for `serde_with@3.14.1` on nightly +- [#147057](https://github.com/rust-lang/rust/issues/147057) rustdoc: ICE: [trying to look up a HirId in the wrong context] +- [#158050](https://github.com/rust-lang/rust/issues/158050) [rustdoc] link to definition doesn't generate link to item's doc when clicking on its name + + +## 7. Every span is valid + +**Statement.** Every span in diagnostics (including suggestion parts) and in encoded metadata has lo <= hi (non-empty), lies inside an existing SourceFile with matching context, falls on UTF-8 character boundaries, and each suggestion's substitutions do not overlap. + +**Check.** Two checks. (1) Post-process --error-format=json for every corpus crate, plus a corpus with multibyte identifiers and strings: check byte ranges against the file contents and check that suggestion parts are disjoint. (2) With mirth, validate every span encoded into rmeta against the SourceMap when it is written. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** span / diagnostics / metadata. + +**Tested today?** There is an internal assertion when a bad span is sliced. Nothing validates spans systematically, and multibyte source is rare in tests. + +**mirth.** Good for the metadata half (hook span encoding). The diagnostic half needs only a JSON post-processor. + +**Bugs that violated it:** + +- [#151610](https://github.com/rust-lang/rust/issues/151610) [ICE]: `Span must not be empty and have no suggestion` +- [#151607](https://github.com/rust-lang/rust/issues/151607) [ICE]: ` all spans must be disjoint` +- [#147339](https://github.com/rust-lang/rust/issues/147339) ICE: `span context mismatch` +- [#131292](https://github.com/rust-lang/rust/issues/131292) ICE: `bpos.to_u32() >= mbc.pos.to_u32() + mbc.bytes as u32` +- [#156316](https://github.com/rust-lang/rust/issues/156316) [ICE]: `bpos.to_u32() >= mbc.pos.to_u32() + mbc.bytes as u32` +- [#155037](https://github.com/rust-lang/rust/issues/155037) [ICE]: `bpos.to_u32() >= mbc.pos.to_u32() + mbc.bytes as u32` + + +## 8. Semantically neutral edits do not change results + +**Statement.** The Rust compiler may incorrectly identify certain edits as semantically neutral (particularly parentheses in patterns, braces affecting resource lifetime scope, and glob re-exports) when they are actually required for compilation or correct behavior, and may fail to recognize trait implementations on type aliases and projections that normalize to an underlying type as equivalent to direct implementations on that type. + +**Check.** Apply automatic rewrites from a fixed catalogue to corpus crates. Compare exit status, the multiset of diagnostics keyed by (code, lint, item DefPath), and rmeta after span normalization. + +**Cost:** moderate. **Kind:** differential. **Subsystem:** resolve / lints / trait-system / macros. + +**Tested today?** UI tests check isolated forms. No transformation-based differential testing exists. + +**mirth.** Moderate. mirth can compare per-query results keyed by DefPath between the original and the rewritten build, which is finer-grained than diagnostics. + +**Bugs that violated it:** + +- [#86959](https://github.com/rust-lang/rust/issues/86959) Unnecessary parentheses warning for (A | B) as :pat in 2018 edition +- [#160741](https://github.com/rust-lang/rust/issues/160741) "unnecessary braces around `for` iterator expression" have effect on program behavior +- [#157758](https://github.com/rust-lang/rust/issues/157758) False positive of lint `missing_debug_implementations` if `Debug` is implemented for an alias type (e.g., projection) that can get normalized to the relevant type +- [#157757](https://github.com/rust-lang/rust/issues/157757) False positive of lint `missing_debug_implementations` if `Debug` is implemented for a free alias type that expands to the relevant type +- [#152004](https://github.com/rust-lang/rust/issues/152004) Regression on nightly: unused pub(crate) use::*; is not actually unused + + +## 9. Layout views agree + +**Statement.** Unsafe binder layout checks and discriminant helpers must use the inner type's layout view. Transmute's SizeSkeleton check must account for repr(align/packed) differences that affect a type's actual size. repr(transparent) wrappers have the same ABI as their non-ZST field. + +**Check.** With mirth, record layout_of results per type. At each transmute check and each transparent or wrapper type, compare against layout_of of the inner type and flag disagreements. Also feed a generated corpus of repr(align/packed/C/transparent) types. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** layout / type-system. + +**Tested today?** There are debug assertions in layout code and transmute UI tests. Nothing cross-checks SizeSkeleton against layout_of. + +**mirth.** Good. Hook layout_of and SizeSkeleton::compute and compare their results. + +**Bugs that violated it:** + +- [#154426](https://github.com/rust-lang/rust/issues/154426) [ICE]: `Scalar` layout for non-primitive non-enum type unsafe +- [#154424](https://github.com/rust-lang/rust/issues/154424) [ICE]: discriminant_for_variant() is None +- [#155412](https://github.com/rust-lang/rust/issues/155412) transmute size check is wrong for overaligned newtype +- [#88290](https://github.com/rust-lang/rust/issues/88290) Transmute special-case doesn't take into consideration alignment or enum repr. + + +## 10. rustdoc output is the same however an item is re-exported + +**Statement.** Re-exported items may fail to preserve the documentation and cfg badges of their original definitions, particularly: multi-level re-exports may lose documentation from intermediate levels, glob re-exports may fail to include items that appear in named re-exports, and cfg badges may not propagate correctly to glob re-exports or re-exported type aliases. + +**Check.** Use rustdoc JSON. For every re-exported item in the corpus, compare the inlined item's docs, cfg and signature with the original item's (outer docs and cfg added at the re-export are allowed). + +**Cost:** cheap. **Kind:** differential. **Subsystem:** rustdoc. + +**Tested today?** rustdoc tests cover specific re-export shapes only. + +**mirth.** Poor. It is a JSON post-processor. + +**Bugs that violated it:** + +- [#81893](https://github.com/rust-lang/rust/issues/81893) Documentation of a re-export doesn't appear on level-two re-export +- [#53724](https://github.com/rust-lang/rust/issues/53724) `pub use serde::*` doesn't show traits in `cargo doc` +- [#96166](https://github.com/rust-lang/rust/issues/96166) doc(cfg) doesn't work on glob reexports +- [#154921](https://github.com/rust-lang/rust/issues/154921) `doc(auto_cfg)` and `doc(cfg)` don't add cfgs to re-exported type aliases + + +## 11. P6+ incremental equals clean (outputs and diagnostics) + +**Statement.** Incremental builds can produce different .rmeta bytes (due to iteration-order-dependent content like DocLinkResMap when encountering unrelated files in the library search path) and duplicate diagnostics compared to clean builds (due to span-based deduplication issues). + +**Check.** Use a script per crate: clean build; then incremental builds after (a) a no-op touch, (b) a body edit, (c) a signature edit, (d) reverting to the original. Byte-compare each result with a clean build of the same source. Also compare the --error-format=json diagnostic multisets. A diagnostic that is missing or duplicated only in the incremental build is a failure. + +**Cost:** moderate (several builds per crate). **Kind:** differential. **Subsystem:** incremental. + +**Tested today?** tests/incremental checks that specific revisions compile or fail and that specific nodes are clean or dirty. It does not compare bytes against a clean build, and it does not compare diagnostics. mirth's P6 does byte comparison for single edits only. + +**mirth.** Excellent. P6 exists. Extend it with revert/no-op sequences and a diagnostic diff. mirth's record of which queries were re-executed versus loaded from cache attributes a divergence to a cached query. + +**Bugs that violated it:** + +- [#162901](https://github.com/rust-lang/rust/issues/162901) Diagnostic deduplication breaks with incr comp +- [#159677](https://github.com/rust-lang/rust/issues/159677) `.rmeta` contents depend on unrelated files in the library search path +- [#106571](https://github.com/rust-lang/rust/issues/106571) Regression: duplicate messages appear in --error-format=json + + +## 12. Parallel front end is deterministic + +**Statement.** The parallel front end with -Zthreads > 1 produces non-deterministic behavior in: (1) encoding of syntax contexts affecting derived code generation, (2) LLVM inline asm location cookies in bitcode/LTO output, and (3) query cycle handling in the deadlock resolver. These cause byte-identical outputs and repeated runs to diverge from both -Zthreads=1 and each other. + +**Check.** Build the corpus with -Zthreads=1 once and -Zthreads=8 three times. Byte-compare outputs and diff the JSON diagnostics. When they differ, compare mirth's per-table write logs. + +**Cost:** cheap. **Kind:** differential. **Subsystem:** parallel front end. + +**Tested today?** A parallel-rustc CI job runs some UI tests. Outputs are not compared against single-threaded builds. + +**mirth.** Good. It reuses the P5 harness with a different flag, and mirth's write logs localize the divergence. + +**Bugs that violated it:** + +- [#129094](https://github.com/rust-lang/rust/issues/129094) derives: parallel compiler makes builds irreproducible +- [#150451](https://github.com/rust-lang/rust/issues/150451) parallel compiler: `threads::spawn`ning loop not reproducible +- [#153391](https://github.com/rust-lang/rust/issues/153391) [ICE]: parallel: None in compiler/rustc_type_ir/src/ty_kind.rs + + +## 13. Results agree across opt-level, mir-opt-level, codegen-units and LTO + +**Statement.** Compiler optimization passes and LTO can silently change program semantics: a program's observable behavior (exit status, panic status, and linkability) may differ across -Copt-level 0/3, -Zmir-opt-level 0/4, codegen-units 1/16, and lto off/thin/fat settings, when it should remain identical. + +**Check.** Run cargo test for crates in the corpus under a configuration matrix and diff per-test outcomes. Flag link failures that appear in only one configuration. Run Miri on a subset as the reference. + +**Cost:** expensive (matrix of full builds plus test runs). **Kind:** semantic. **Subsystem:** mir-opt / codegen / LTO. + +**Tested today?** Individual mir-opt tests exist. There is no systematic matrix over real crates' test suites, and LTO/codegen-unit link failures are found by users. + +**mirth.** Low. It is an output-level differential. mirth could record which MIR pass changed a function that later misbehaves, which helps bisect the cause. + +**Bugs that violated it:** + +- [#163779](https://github.com/rust-lang/rust/issues/163779) SsaRangePropagation propagates range information from optional asserts +- [#162348](https://github.com/rust-lang/rust/issues/162348) 1.99 beta crater regression: SIGSEGV in LLVM in release mode +- [#153645](https://github.com/rust-lang/rust/issues/153645) Externally Implementable Items: `error: undefined symbol` when opt-level >= 1 + + +## 14. A successful compile contains no error types + +**Statement.** If a compilation exits 0, no ty::Error, ConstKind::Error, or ErrorGuaranteed-carrying value should appear in compiler internals (valtree constants, transmute layout checking, MIR, query results, or metadata). Error types are created only after errors are emitted; violations occur when compiler subsystems construct these types during failed operations (const evaluation, layout normalization, or delegation) without emitting corresponding errors or errors without reporting. + +**Check.** With mirth, hook the construction of Ty::new_error, Const::new_error and region errors and record the creation site. When the session ends with zero errors, any recorded creation is a bug. Also scan encoded metadata for error types. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** type-system / diagnostics. + +**Tested today?** ErrorGuaranteed makes forging an error hard. Delayed bugs ICE at session end only if they were delayed. Silent paths that drop an error are not checked. + +**mirth.** Good. It needs one hook at the error constructors and a check at session end. + +**Bugs that violated it:** + +- [#150969](https://github.com/rust-lang/rust/issues/150969) ICE: valtrees: 'called `Result::unwrap()` on an `Err` value: ReferencesError(ErrorGuaranteed(()))' +- [#149588](https://github.com/rust-lang/rust/issues/149588) layout errors in transmute checking don't get emitted +- [#154780](https://github.com/rust-lang/rust/issues/154780) [ICE]: delegation: `TyKind::Error constructed but no error reported` + + +## 15. Each encoded record is written once and hashed after it is final + +**Statement.** The metadata encoder writes each dep node at most once under concurrent execution, and crate_hash is not computed before metadata encoding finishes + +**Check.** With mirth, log (table, key) and (dep node) writes and fail on any duplicate. Log when crate_hash and the svh are computed relative to the encoder finishing, and fail if hashing happens first. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** metadata / incremental. + +**Tested today?** There are a few assertions in the dep-graph encoder. The ordering of hashing relative to encoding is not checked. + +**mirth.** Excellent. It uses the same write hooks as the reader-writer property. + +**Bugs that violated it:** + +- [#150018](https://github.com/rust-lang/rust/issues/150018) assertion failed: trying to encode a dep node twice +- [#142778](https://github.com/rust-lang/rust/issues/142778) ICE: rustc_query_system: dep_graph: assertion failed (dep node index out of range) +- [#163426](https://github.com/rust-lang/rust/issues/163426) [ICE]: rustdoc ICE with -Zmetrics-dir: crate_hash(LOCAL_CRATE) called before metadata encoding + + +## 16. #[expect] is fulfilled exactly when the lint fires + +**Statement.** #[expect(L)] fulfillment status is not always correctly aligned with whether #[warn(L)] would actually emit L. Cases involving match guards, derives, and cfg_attr show mismatches where expectations are reported as unfulfilled despite the lint firing, or vice versa. + +**Check.** For each #[allow] or #[expect] in the corpus, generate the #[warn] and #[expect] variants. Build both and check the pairing: lint emitted exactly when the expectation is fulfilled. + +**Cost:** moderate (one build per attribute; batchable). **Kind:** differential. **Subsystem:** lints. + +**Tested today?** There are UI tests for specific cases. No automatic pairing check exists. + +**mirth.** Low. A source-rewrite harness is enough. + +**Bugs that violated it:** + +- [#152004](https://github.com/rust-lang/rust/issues/152004) Regression on nightly: unused pub(crate) use::*; is not actually unused +- [#151983](https://github.com/rust-lang/rust/issues/151983) Missing "unused variable" warning when using a match guard +- [#152401](https://github.com/rust-lang/rust/issues/152401) unfulfilled-lint-expectations for missing_docs + + +## 17. Metadata reads hit entries the writer wrote + +**Statement.** Reads of metadata table entries fail when the upstream encoder never wrote those entries, causing ICEs when dependent crates attempt to access missing (table, DefIndex) pairs or when compiler passes try to read unwritten entries from internal lookup tables + +**Check.** With mirth, log every table write (table, index) in the writer process and every lookup in reader processes across the whole cargo build. Join the logs. A read of an unwritten entry is a bug, except for tables documented as default-on-absent, which need an allowlist. + +**Cost:** moderate (whole-build logs). **Kind:** invariant. **Subsystem:** metadata. + +**Tested today?** Not tested. Missing entries usually decode as defaults with no error. + +**mirth.** Excellent. It needs both writer and reader instrumentation across one cargo build, which no other tool can do. + +**Bugs that violated it:** + +- [#163426](https://github.com/rust-lang/rust/issues/163426) [ICE]: rustdoc ICE with -Zmetrics-dir: crate_hash(LOCAL_CRATE) called before metadata encoding +- [#159233](https://github.com/rust-lang/rust/issues/159233) [ICE]: resolve: `no entry found for key` + + +## 18. Query results contain no inference variables + +**Statement.** Inference variables leak into the const literal lowering logic (lit_to_const), appearing in cached query results where they cause hashing panics. + +**Check.** With mirth, at every query return and every HashStable call, check has_infer() and has_placeholders() on the value and record the query and key on failure. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** query system / type inference. + +**Tested today?** There are a few debug assertions at specific sites. No global check exists. + +**mirth.** Excellent. One generic hook at the query provider return. + +**Bugs that violated it:** + +- [#153525](https://github.com/rust-lang/rust/issues/153525) [ICE]: `type variables should not be hashed` +- [#153524](https://github.com/rust-lang/rust/issues/153524) [ICE]: `const variables should not be hashed` + + +## 19. Machine-applicable suggestions apply cleanly + +**Statement.** Applying every MachineApplicable suggestion (rustfix) gives code that compiles and the diagnostic is gone, provided that suggestions are generated with correctly-calculated non-empty, non-overlapping spans. Violations occur when the suggestion system generates suggestions with malformed spans that cause compiler panics. + +**Check.** Run cargo fix --broken-code on corpus crates with all warn-by-default lints enabled, then rebuild. Check that it compiles and that the original lint no longer fires. + +**Cost:** moderate. **Kind:** differential. **Subsystem:** diagnostics / lints. + +**Tested today?** run-rustfix UI tests cover curated cases. + +**mirth.** Poor. It is a cargo-level differential. + +**Bugs that violated it:** + +- [#161213](https://github.com/rust-lang/rust/issues/161213) invalid range in macro triggers compiler panic at assertion `left == right` failed: suggestion must not have overlapping parts +- [#161472](https://github.com/rust-lang/rust/issues/161472) [ICE]: Macro captures a list meta item with $m:meta and passes it directly to #[derive] causes `must not be empty and have no suggestion` + + +## 20. Polonius accepts at least what NLL accepts, and no more than is sound + +**Statement.** -Zpolonius=next has soundness bugs in opaque type region handling that allow acceptance of programs with undefined behavior, violating the guarantee that Polonius-only-accepted programs should be sound under Miri. + +**Check.** Borrow-check the corpus with both. Any NLL-accept/Polonius-reject is a bug. For Polonius-only accepts, which mostly come from the UI suite, run under Miri. + +**Cost:** moderate. **Kind:** differential. **Subsystem:** borrowck. + +**Tested today?** There is a polonius compare-mode on some UI tests. + +**mirth.** Low. It is an accept/reject differential. + +**Bugs that violated it:** + +- [#153215](https://github.com/rust-lang/rust/issues/153215) free region visitor for liveness marking regions dead and polonius alpha soundness +- [#160669](https://github.com/rust-lang/rust/issues/160669) Zpolonius=next soundness bug: defining use of an opaque type discards the first one's region + + +## 21. Unstable syntax is gated before expansion + +**Statement.** Every unstable syntactic form must be rejected at pre-expansion on stable without its feature gate, even inside #[cfg(FALSE)] blocks or unused macro_rules arms, to prevent code breakage when feature gates are removed. + +**Check.** Take each feature-gate UI test that exercises syntax, wrap the gated syntax in #[cfg(FALSE)], and compile without the feature. The compile must fail. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** parser / feature gates. + +**Tested today?** Some features have cfg(FALSE) tests. It is not enforced for each one. + +**mirth.** Poor. It is a source-rewrite harness. + +**Bugs that violated it:** + +- [#152501](https://github.com/rust-lang/rust/issues/152501) `try bikeshed $ty { … }` is not pre-expansion gated (affects beta+nightly) +- [#152499](https://github.com/rust-lang/rust/issues/152499) Inline const patterns are no longer pre-expansion gated + + +## 22. P5+ metadata determinism under irrelevant perturbation + +**Statement.** .rmeta contents must not depend on unrelated rlib files present in the library search path (-L), i.e., two builds with identical source, flags, and target produce byte-identical .rmeta even when the set of available (but unused) libraries in the search path differs. + +**Check.** For every crate in a cargo build, run the build twice from clean. Between the runs, change one perturbation: add a decoy rlib to the search path, use a different target dir, set junk env vars, or shift the clock with faketime. Byte-compare every output. When they differ, use mirth's table-write recording to find the first table or row that diverges, then the query that produced it. + +**Cost:** cheap (2 builds; perturbations are free). **Kind:** differential. **Subsystem:** metadata / const-eval / whole-compiler. + +**Tested today?** rustc has tests/run-make reproducibility tests on a few small fixtures and one perturbation each. mirth already checks P5 on whole cargo builds without perturbation. Search-path and build-dir perturbations across real crates are not tested. + +**mirth.** Excellent. P5 already exists. Adding perturbations is a harness change. mirth's per-table write log turns a byte diff into the table, the row and the query that wrote it. + +**Bugs that violated it:** + +- [#159677](https://github.com/rust-lang/rust/issues/159677) `.rmeta` contents depend on unrelated files in the library search path + + +## 23. P4+ no query reads untracked state + +**Statement.** No query reads HashMap/HashSet iteration order (including symbol interner order) in ways that affect incremental compilation artifacts (like .rmeta) unless the dependency on which symbols get interned/which HashMap entries are iterated is recorded in the dep graph. + +**Check.** Use mirth instrumentation of env::var*, SystemTime/Instant, fs::metadata, and iteration of std/Fx HashMaps. Attribute every read to the query on top of the stack and flag reads that are not covered by a tracking query. Confirm a finding by perturbing that input and checking whether the query's fingerprint changes. + +**Cost:** cheap once instrumented. **Kind:** invariant. **Subsystem:** incremental / query system. + +**Tested today?** Upstream has a lint against unordered iteration (rustc::potential_query_instability) with many allow exceptions. mirth's P4 covers metadata encoding only. + +**mirth.** Excellent. This is mirth's core use case: extend the P4 hooks from encode_metadata to every query frame. + +**Bugs that violated it:** + +- [#159677](https://github.com/rust-lang/rust/issues/159677) `.rmeta` contents depend on unrelated files in the library search path + + +## 24. Symbol names are injective and stable + +**Statement.** The v0 and legacy symbol mangling schemes must encode all type attributes that the type system considers semantically distinct, including splat on function types, to ensure monomorphized instances never share symbol names. V0 symbols must round-trip through rustc-demangle to the instance's path." + +**Check.** With mirth, record (Instance, symbol_name) pairs in every codegen process of a cargo build. Check that the mapping is injective across crates and round-trip each v0 symbol through rustc-demangle against the printed instance. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** symbol mangling. + +**Tested today?** Mangling UI tests cover specific types. There is no whole-build injectivity check. + +**mirth.** Good. Hook symbol_name, which is easy, and join the results across crates. + +**Bugs that violated it:** + +- [#158644](https://github.com/rust-lang/rust/issues/158644) Splat is ignored in symbol mangling, leading to symbol clashes + + +## 25. dyn types have a dyn-compatible principal + +**Statement.** In successful compilations, every TyKind::Dynamic that reaches typeck, MIR, or metadata must have a principal trait that is explicitly marked dyn-compatible; allowing Dynamic types with non-dyn-compatible principals (like DerefPure) enables unsound behavior. + +**Check.** With mirth, when a Dynamic type is interned, record its principal. At the end of a successful session, assert is_dyn_compatible for each one. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** trait-system. + +**Tested today?** Dyn-compatibility is checked at direct dyn-syntax sites. No global check exists. + +**mirth.** Good. One intern hook plus a check at session end. + +**Bugs that violated it:** + +- [#154619](https://github.com/rust-lang/rust/issues/154619) `deref_patterns` is unsound due to `dyn` of subtrait of `DerefPure` + + +## 26. Built-in attributes reject malformed arguments + +**Statement.** Every built-in attribute rejects argument forms its template does not allow by producing a diagnostic error, rather than silently ignoring excess or incorrectly-formed arguments. + +**Check.** From the BUILTIN_ATTRIBUTES templates, generate malformed forms for each attribute (extra arguments, list instead of word, name-value instead of word) and assert that a diagnostic is produced. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** attribute parsing. + +**Tested today?** There are per-attribute tests, but not generated from the templates. + +**mirth.** Poor. + +**Bugs that violated it:** + +- [#154977](https://github.com/rust-lang/rust/issues/154977) Invalid value accepted for `#[macro_export(local_inner_macros)]` + + +## 27. Advertised target features exist in LLVM + +**Statement.** Every target feature that rustc lists for a target, and every feature that an asm register class or intrinsic requires, must be recognized by LLVM and must be consistent with the target's baseline—that is, register class and intrinsic feature requirements cannot exceed what is available for that target's baseline configuration. + +**Check.** For every target, enumerate rustc's feature table and query LLVM's feature list. Compile an empty function with each feature enabled and confirm there is no 'unknown feature' warning or spurious requirement. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** codegen / targets. + +**Tested today?** There is a partial tidy-style check. + +**mirth.** Poor. + +**Caveat.** The check judged this not general: The single cited issue provides clear evidence of a violation. Register classes were requiring features in a manner inconsistent with the target's baseline capabilities. The fix involved both adjusting feature requirements and adding missing implied features, confirming the underlying inconsistency that the property prohibits. + +**Bugs that violated it:** + +- [#159976](https://github.com/rust-lang/rust/issues/159976) `armv8r-none-eabihf` target cannot use FPU instructions in inline asm + + +## 28. Eq and Hash agree for std types + +**Statement.** For every std type implementing both Eq and Hash, a == b implies hash(a) == hash(b), for every Hasher. (The property statement is precise and matches the evidence; the violation in Path on Windows with verbatim paths is a concrete instance of this general invariant being broken.) + +**Check.** Run property tests over generated values of std types, including borrowed forms through Borrow, with several hashers. + +**Cost:** cheap. **Kind:** semantic. **Subsystem:** library. + +**Tested today?** There are ad hoc tests only. + +**mirth.** None. It is a library property test. + +**Bugs that violated it:** + +- [#161651](https://github.com/rust-lang/rust/issues/161651) Eq returns true but hashes are different for some verbatim paths on Windows + + +## 29. Each diagnostic is emitted once + +**Statement.** When the same diagnostic (identified by identical level, error code, message text, and primary source span) is emitted more than once in a single compilation session, it appears multiple times in --error-format=json output. This can be detected by extracting diagnostics from JSON and checking for exact duplicates by the tuple (level, code, message, span). + +**Check.** Post-process JSON diagnostics from corpus builds and from the UI suite run with --error-format=json, and look for duplicate keys. + +**Cost:** cheap. **Kind:** invariant. **Subsystem:** diagnostics. + +**Tested today?** The emitter deduplicates some cases. UI .stderr snapshots would show duplicates, but nothing asserts they are absent. + +**mirth.** Poor. It is an output check. + +**Bugs that violated it:** + +- none of the cited issues held up + + diff --git a/docs/scale.md b/docs/scale.md new file mode 100644 index 0000000..4b958e9 --- /dev/null +++ b/docs/scale.md @@ -0,0 +1,162 @@ +# Scaling up: history replay, an edit fuzzer, and a survey of properties + +Three ways to find more bugs than one fixture and ten hand-written edits can: + +1. **Replay real history.** `rustc/replay.py` walks a crate's git history, oldest + first, building each commit incrementally on top of the last and again from scratch, + and compares them (P6). Ten crates, up to 2,000 commits each. +2. **Fuzz edits on a large fixture.** `fixtures/sink` is a five-crate workspace with as + many stable language features as fit. `rustc/fuzz.py` makes random mechanical edits to it + and checks every incremental rebuild against a clean build. +3. **Survey past bugs for properties.** Agents read 1,000 fixed rustc bugs and extracted + the invariants they violated; a second pass checked every citation. The result is + [`properties.md`](properties.md). + +All of it runs against a compiler with the candidate fixes for the bugs already known +([`hunt.md`](hunt.md)), so those do not drown new ones; each new finding is then confirmed +on the official nightly. + +## What was found + +| finding | how | status | +|---|---|---| +| Incremental rebuilds republish stale metadata when an edit moves no span | fuzzer and replay, independently | **new**; root-caused, regression from #114669 (1.90); fixed and tested ([report](hunt/issue-stale-metadata-reuse.md)) | +| The candidate fix for the literal bug over-deduplicated | replay (serde, one commit) | my own mistake; the fix was narrowed ([report](hunt/issue-literal-dedup.md)) | +| Three untracked options that change reused output | closed-bug queries and an option audit ([`ur-queries.md`](ur-queries.md)) | new instances of #66955's class; report drafted ([draft](hunt/issue-untracked-options.md)) | + +**Stale metadata.** Since #114669, an incremental session reuses the saved `.rmeta` when +its dep-node is green. But the metadata's source map records every source file's length, +line table and content hash, read straight from the session's source map, which nothing +tracks. An edit that changes no query result (a comment after the last item, a typo fixed +inside a comment without moving any span, a change inside `#[cfg]`'d-out code) leaves the +node green, and the previous session's metadata is republished. A dependent then cannot match +the dependency's source: its diagnostics lose the snippet. Real histories hit it on ordinary +commits: memchr (5), serde, smallvec (2), hashbrown (1), anyhow (3), regex (several), all on +comment or docstring edits. The fix makes the metadata task read an `eval_always` query +fingerprinting the local source files, so reuse still happens when the files are truly +unchanged (the blessed touch-only rebuild still shows it). + +**A flaw in my own fix.** The first fix for the literal bug deduplicated every immutable +allocation on decode. The serde replay showed P6 failing on commit `2f58a20` with it applied: +`-Zmeta-stats` put the whole 59-byte difference in `interpret-alloc-index`, because +allocations a clean session keeps apart were merged. The fix now records, at encode time, +whether an allocation was created through deduplication, and repeats only that. The same +replay window is clean with it. + +With all three fixes, each regression test passes and each fails without its own fix, and +rustc's incremental, metadata UI and run-make, consts/statics/const-generics UI and +codegen-llvm tests pass. + +## History replay + +`rustc/replay.py --rustc --repo --work [--commits N] [--from i --to j]` + +For each first-parent commit, oldest first: check it out, `cargo build --lib` +incrementally on the previous commit's target directory, then build the same source from +scratch at the same path. Registry dependencies are not compiled incrementally by Cargo, so +the clean build starts from a copy of the target directory with every path package cleaned +(`cargo clean -p`) and the incremental cache removed. It compares every `.rmeta` Cargo +reports for a path package, and flags ICEs, builds where only one side fails, and clean +builds that did not actually recompile ("stale", which would compare a file with itself). +Every commit is built with `--cap-lints=warn`, so old commits that deny warnings still build. + +Crates: regex, itertools, smallvec, hashbrown, indexmap, memchr, bitflags, serde, anyhow, +thiserror. On the compiler with all three fixes: + +| crate | commits | both built | P6 | ICE | one side failed | +|---|---:|---:|---:|---:|---:| +| anyhow | 669 | 524 | 0 | 0 | 0 | +| bitflags | 301 | 300 | 0 | 0 | 0 | +| hashbrown | 477 | 456 | 0 | 0 | 0 | +| indexmap | 465 | 415 | 0 | 0 | 0 | +| itertools | 1,358 | 942 | 0 | 0 | 0 | +| memchr | 267 | 267 | 0 | 0 | 0 | +| regex | 700 of 1,368 | 590 | 0 | 0 | 0 | +| smallvec | 391 | 390 | 0 | 0 | 0 | +| serde | 806 of 2,000 | 425 | 0 | 0 | 0 | +| thiserror | 556 | 553 | 0 | 0 | 0 | + +4,862 commits built on both sides and none differed. Commits that failed on both sides are +mostly old code the current compiler rejects. regex and serde stopped part-way when the +disk filled. This run compared `.rmeta` only; the rlib and diagnostic comparisons were +added after it started. + +Before the stale-metadata fix, the same replay found that bug on ordinary commits in +memchr, smallvec, hashbrown, anyhow and regex: comment and docstring edits that moved no +span. One P6 failure, on serde, came from my own first fix for the literal bug. + +A positive control: on the unpatched nightly, every incremental rebuild of itertools' +last 12 commits differs from a clean build (the `param_def_id_to_index` bug); on the patched +compiler none do. + +Harness problems found and fixed on the way, each of which produced false findings or +skipped commits silently: artifacts of earlier commits counted as differences; path +dependencies that are not workspace members left uncleaned; `cargo clean` refusing a +directory without `CACHEDIR.TAG`. + +## The fixture and the fuzzer + +`fixtures/sink`, about 1,200 lines in five crates: + +- `sink-macros`, a proc-macro crate with no dependencies: a derive with a helper + attribute, an attribute macro, a function-like macro; +- `sink-core`, with a build script generating code and a cfg, features, traits with + associated types and consts, generic associated types, blanket impls, trait objects and + upcasting, operators, `impl Trait` and `async fn` in traits, async closures, const + generics, const fn and const blocks, statics, thread locals, unions, `repr`, `Drop`, + raw pointers, `MaybeUninit`, FFI, `#[track_caller]`, inline attributes, deprecation, + error types, iterators, closures, patterns, labelled blocks, `let else`, `if let` chains, + collections, `Rc`/`RefCell`/`Arc`/`Mutex`, scoped threads, exported and internal + `macro_rules!`; +- `sink-mid`, using all three kinds of proc macro and the exported macros, with nested + modules and restricted visibility, re-exports, and generic code dependents instantiate; +- `sink-dy`, built as a dylib too; +- `sink`, a binary that exercises everything and checks 60 results, so a miscompile fails. + +`rustc/fuzz.py --rustc --fixture fixtures/sink --work [--workers N]` + +Each worker keeps one evolving copy of the fixture. It makes a random edit, chosen from 16 +kinds (a comment, a blank line, an indented line, swapped or moved or deleted items, a +duplicated function, a new item of one of 12 shapes, a changed number or string, an +inline attribute, a doc comment, reordered derives, narrowed visibility, a renamed local), +and builds incrementally with `-Zincremental-verify-ich`. If the edit does not compile, it +is reverted; the next build then also exercises recovery from a failed session. Otherwise +the same source is built from scratch at the same path, and every `.rmeta` and every proc +macro's embedded metadata are compared (P6), the two binaries are run and their output +compared, and ICEs, hangs and one-sided failures are reported. Every 40 kept edits the worker +starts again from the pristine fixture. A finding keeps every edit since the last reset, and +`rustc/fuzz-replay.py` replays it exactly. + +Throughput on this 16-core machine: about 2 edits a second with six workers, about +170,000 a day, while other work shared the machine. + +| compiler | edits | built and compared | findings | +|---|---:|---:|---| +| with the first two fixes | 4,009 | 3,336 | 60 stale metadata reuses (the third bug) | +| with all three fixes | 10,717 | 8,936 | none | + +The last 1,455 of those comparisons also checked object code, binaries and diagnostics. + +Millions of edits means about a week here, or several machines. + +## The survey + +[`properties.md`](properties.md) has 29 properties plus the crash baseline, ranked by how +many of the 1,000 bugs violated each, after every citation was checked against the issue +text (107 of 156 citations held). The top: intern-time well-formedness of type-system values +(12 bugs), the old and new trait solvers agreeing (11), `extern "C"` lowering matching C +(9), MIR validating at every opt level on real crates (8), library soundness under injected +panics (7). The incremental and metadata properties mirth already checks rank lower by bug +count (3 and 1), yet they are where all four bugs found in this work came from, which says +more about which bugs get reported and fixed than about where bugs are. + +Small models did both the reading and the checking. The issue lists are leads, not proof. + +## Limits + +- Everything ran on Linux x86_64, single-threaded unless stated. +- `-Zthreads` is not fuzzed: the fixture has `impl Trait` in traits, and #162202 makes + every threaded build differ. +- The fuzzer edits text, not syntax trees; about 8% of edits do not compile and are + reverted, and some kinds of edit (deleting items, narrowing visibility) almost never do. +- The replay compares `.rmeta` only; the binaries' code and debug info are not compared. diff --git a/docs/shadow-mode.md b/docs/shadow-mode.md new file mode 100644 index 0000000..74544bf --- /dev/null +++ b/docs/shadow-mode.md @@ -0,0 +1,53 @@ +# Checking reuse inside rustc + +**The idea.** When an incremental session reuses something from the cache, rustc could +also compute it afresh and compare the two, behind a `-Z` flag, or always in debug builds of +the compiler. Every CI job and every user who turns the flag on would then run P6 on their +own code, at the moment of reuse, instead of only in mirth's builds. + +## Is anyone doing it? + +Not that we could find (October 2026): + +- **`-Zincremental-verify-ich`** is the nearest thing. When a query result is loaded from + the incremental cache, it hashes the loaded value again and compares it with the hash the + previous session stored. Without the flag it checks a rotating 1-in-32 subset + (`should_verify_loaded_value`, `rustc_query_impl/src/incremental.rs`). It never recomputes + a result, so it cannot see a value that was wrong when stored. It ignores fields marked + `#[stable_hash(ignore)]`: `Generics::param_def_id_to_index`, the field behind the first + bug in [`hunt/`](hunt), is one. And it never looks at work products: reused metadata and + object files are not query results. The fuzzer ran every build with it, and it reported + none of the three bugs. +- **`#[rustc_clean]`** and `-Zquery-dep-graph` assert in `tests/incremental` which nodes + are reused or recomputed, for hand-written cases. +- **The 2026 project goal ["Incremental Systems Rethought"](https://goals.rust-lang.org/2026/incremental-system-rethought.html)** + plans *more* reuse: `cargo build` reusing `cargo check`'s work, and data dependencies and + diffing. Its plan does not mention verifying reused results. More reuse is more places for + stale reuse, so a check like this would back it up. +- Searches of rust-lang/rust issues and PRs for verifying reused work products, recomputing + green queries, or shadow verification found nothing relevant. The one open issue nearby + is #162601, an ICE when an LTO work product is missing. + +## What it would check + +Each kind of reuse has a natural comparison: + +| reused | how it is reused today | the shadow check | +|---|---|---| +| a query result marked green | loaded from the cache, or kept without loading | recompute it (force the query as if red) and compare with the loaded value, by value, not by stable hash | +| metadata | the saved `.rmeta` hard-linked or copied when its node is green (`encode_metadata`) | encode it again and compare bytes; this would have caught the stale-metadata bug the first time it happened | +| an object file (codegen unit) | the saved `.o` reused when its CGU is green | codegen it again and compare, or compare the LLVM IR | +| diagnostics | replayed from the cache | compare with the diagnostics the recomputation emits | + +Recomputing everything doubles a build's cost, so the flag would usually sample, as +`-Zincremental-verify-ich` does: a deterministic subset per session, rotating so that every +reused item is checked over many sessions, with an option to check everything. + +## Where it would go first + +Metadata is the cheapest to start with and has the most recent bug: the reuse decision is +one function (`encode_metadata` in `rustc_metadata/src/rmeta/encoder.rs`), and encoding +again into a temporary file and comparing bytes is a small change. A failure would name +the first differing byte, which `-Zmeta-stats`'s sections place in a table. + +Not started. This page is the record of what was looked at. diff --git a/docs/survey/candidates.json b/docs/survey/candidates.json new file mode 100644 index 0000000..70e8a11 --- /dev/null +++ b/docs/survey/candidates.json @@ -0,0 +1,2072 @@ +{ + "merged": { + "properties": [ + { + "name": "P5+ metadata determinism under irrelevant perturbation", + "statement": "Two clean builds with the same source, dependencies, flags and target produce byte-identical .rmeta (and .rlib/.o), even when inputs that should not matter change: the HashMap seed, unrelated crates on the -L search path, the build directory (with --remap-path-prefix), environment variables that are not tracked, and the wall clock.", + "subsystem": "metadata / const-eval / whole-compiler", + "kind": "differential", + "how_to_check": "For every crate in a cargo build, run the build twice from clean. Between the runs, change one perturbation: add a decoy rlib to the search path, use a different target dir, set junk env vars, or shift the clock with faketime. Byte-compare every output. When they differ, use mirth's table-write recording to find the first table or row that diverges, then the query that produced it.", + "cost": "cheap (2 builds; perturbations are free)", + "issues": [ + 159677, + 157743, + 157747, + 158602, + 89911, + 150409, + 150419, + 106571, + 151537, + 142152, + 141540, + 141313, + 138910, + 138089, + 133966, + 151625, + 150983 + ], + "currently_tested": "rustc has tests/run-make reproducibility tests on a few small fixtures and one perturbation each. mirth already checks P5 on whole cargo builds without perturbation. Search-path and build-dir perturbations across real crates are not tested.", + "mirth_fit": "Excellent. P5 already exists. Adding perturbations is a harness change. mirth's per-table write log turns a byte diff into the table, the row and the query that wrote it.", + "rank_reason": "Most bugs per check, cheapest of all (two builds), mirth already has the infrastructure, and the search-path and HashMap-order cases (e.g. DocLinkResMap, #159677) are untested upstream. Some of the const-eval issue IDs folded in here were not checked against the issue text." + }, + { + "name": "P6+ incremental equals clean (outputs and diagnostics)", + "statement": "After any sequence of edits, including a no-op edit and an edit followed by its revert, an incremental build produces the same .rmeta/.rlib bytes and the same set of diagnostics as a clean build of the final source.", + "subsystem": "incremental", + "kind": "differential", + "how_to_check": "Use a script per crate: clean build; then incremental builds after (a) a no-op touch, (b) a body edit, (c) a signature edit, (d) reverting to the original. Byte-compare each result with a clean build of the same source. Also compare the --error-format=json diagnostic multisets. A diagnostic that is missing or duplicated only in the incremental build is a failure.", + "cost": "moderate (several builds per crate)", + "issues": [ + 162407, + 162901, + 162585, + 125564, + 135062, + 159677, + 150464, + 106571 + ], + "currently_tested": "tests/incremental checks that specific revisions compile or fail and that specific nodes are clean or dirty. It does not compare bytes against a clean build, and it does not compare diagnostics. mirth's P6 does byte comparison for single edits only.", + "mirth_fit": "Excellent. P6 exists. Extend it with revert/no-op sequences and a diagnostic diff. mirth's record of which queries were re-executed versus loaded from cache attributes a divergence to a cached query.", + "rank_reason": "Many silent wrong-output bugs found this way. Moderate cost. Upstream tests never compare incremental output with a clean build, so the gap is large." + }, + { + "name": "MIR passes validation after every pass at every mir-opt-level", + "statement": "MIR is well-formed after every pass. Operand types match their places, every local is used only between its StorageLive and StorageDead, no projections remain unnormalized after borrowck, and const arguments have the types their parameters declare. This holds at -Zmir-opt-level=0..4 and with -Zinline-mir.", + "subsystem": "mir-build / mir-opt / const-eval", + "kind": "invariant", + "how_to_check": "Build every crate in the corpus with -Zvalidate-mir -Zlint-mir -Zmir-opt-level={0,2,4} -Zinline-mir and -Cdebug-assertions on the compiler. Any validation failure counts. This uses existing flags; you only need to run them on real crates.", + "cost": "cheap (existing flags; one extra build per opt-level)", + "issues": [ + 158037, + 156409, + 154750, + 154748, + 152962, + 158231, + 151647, + 151579, + 152278, + 120811 + ], + "currently_tested": "mir-opt tests use the validator, but UI tests and real crates do not run with -Zvalidate-mir by default, and mir-opt-level=4 is rarely exercised on real code.", + "mirth_fit": "Moderate. No MIR rewriting is needed because the flags exist. mirth helps by recording which pass first broke validation, and by giving every report one shared corpus and harness.", + "rank_reason": "Many bugs, very cheap, and it catches type-mismatch and storage-liveness bugs before they become miscompiles. Upstream runs it only on curated tests." + }, + { + "name": "extern \"C\" ABI matches the platform C compiler", + "statement": "For every target, a function signature with C-compatible types (structs, unions, floats, bool, small aggregates, varargs) is lowered to the same argument and return passing (registers vs stack, sign/zero extension, splitting) as clang/gcc for the equivalent C declaration.", + "subsystem": "codegen / abi", + "kind": "differential", + "how_to_check": "Generate random C-compatible type signatures, emit matching Rust and C, and cross-call them in both directions (abi-cafe style). Run natively or under qemu for each tier-1/2 target. Optionally diff rustc's FnAbi against clang's LLVM IR attributes per parameter.", + "cost": "moderate (cross targets need qemu)", + "issues": [ + 121408, + 162011, + 163173, + 163074, + 163075, + 160827, + 151791, + 122620, + 115399, + 147883, + 158897, + 159244, + 159116, + 43894, + 161382 + ], + "currently_tested": "tests/ui/abi/compatibility.rs and per-target codegen tests check fixed cases. abi-cafe is not in CI. Most of these bugs were reported by users on less common targets (SPARC, PowerPC, LoongArch).", + "mirth_fit": "Low. This is an output-level differential with C compilers. mirth could dump the FnAbi of every extern fn to drive the comparison, but the core harness is abi-cafe plus qemu.", + "rank_reason": "High bug count, all silent miscompiles. Moderate cost. Weak upstream coverage. Ranked below the first three because mirth adds little. The FFI-bool and volatile issue IDs in this list were not individually verified." + }, + { + "name": "Interned type-system values are well-formed", + "statement": "Every interned TraitRef, AliasTy, FnDef or other GenericArgs-carrying value has arguments that match generics_of in count and kind. Every const value's valtree shape matches its type. No ParamEnv contains two conflicting predicates for one parameter (e.g. two ConstArgHasType predicates with different types).", + "subsystem": "type-system / const generics", + "kind": "invariant", + "how_to_check": "In an instrumented release-mode rustc, run the existing debug-only checks at every intern site: debug_assert_args_compatible, a valtree-vs-type conformance check, and a ParamEnv duplicate/conflict scan when param_env is built. Run on the corpus plus the UI suite.", + "cost": "cheap per crate (checks at intern sites)", + "issues": [ + 150506, + 150712, + 150734, + 151126, + 158675, + 158362, + 157189, + 137084, + 150673, + 150714, + 150841, + 151024, + 151186 + ], + "currently_tested": "Some of these checks exist as debug_assert in debug compilers only. Valtree/type conformance and ParamEnv conflicts are not checked at all. The issues were found as ICEs later in the pipeline.", + "mirth_fit": "Good. Inject the checks at the mk_* or intern call sites and record the value, the caller's query and the DefId on failure. This turns a late ICE into a report at the point of creation.", + "rank_reason": "Many bugs and a cheap check. However, most of these bugs come from MGCA and other unstable features that real crates do not use, so on a crate corpus it would catch fewer of them than the count suggests. Running it on the UI suite helps." + }, + { + "name": "Old and next trait solvers agree", + "statement": "On any crate, -Znext-solver=globally and the old solver accept and reject the same items, infer the same types for each body (up to region erasure), and select the same impls.", + "subsystem": "trait-system", + "kind": "differential", + "how_to_check": "Build the corpus with both solvers. Compare exit status and diagnostics, and per body DefPath compare typeck_results (node types, method resolutions) and codegen Instances. Every disagreement is a bug in one of the two solvers.", + "cost": "moderate (two full typechecks)", + "issues": [ + 102580, + 90950, + 152789, + 151329, + 151957, + 151323, + 151322, + 151318, + 138274, + 137916, + 151462 + ], + "currently_tested": "The next-solver CI job runs the UI suite with the next solver, and crater runs happen occasionally. Per-body inferred types are never compared.", + "mirth_fit": "Good. mirth can record typeck_results and selected impls per DefPath in both configurations, which gives a precise diff instead of only accept/reject.", + "rank_reason": "Many bugs, and it is the main verification gap for the solver migration. Moderate cost." + }, + { + "name": "Library operations stay sound under injected panics and allocation failures", + "statement": "std/alloc operations, including ones given user comparators, closures or allocators, never cause UB, double drops or leaks of already-moved values when a user callback panics or allocation fails. Allocators return pointers with provenance over the full requested size.", + "subsystem": "library (alloc/core/std)", + "kind": "semantic", + "how_to_check": "Run library tests and property tests under Miri with fault injection: closures and comparators that panic on the Nth call (for every N), a failing allocator, and a counting Drop type. Check for no Miri UB, drops equal to constructions, and a structure that is still usable or droppable afterwards.", + "cost": "expensive (Miri, N-sweep)", + "issues": [ + 162720, + 162719, + 158165, + 157203, + 155746, + 152211, + 114581, + 160815, + 161018, + 43894, + 161382 + ], + "currently_tested": "Library tests run under Miri in CI, but there is no systematic sweep that panics at every callback position or fails allocation at every point.", + "mirth_fit": "Poor. This tests the library, not the compiler. Use Miri plus a fault-injection harness. mirth is not needed.", + "rank_reason": "A fair number of real soundness bugs, but it is expensive and outside mirth's scope." + }, + { + "name": "P4+ no query reads untracked state", + "statement": "No query, not only metadata encoding, reads environment variables, the clock, filesystem metadata or HashMap/HashSet iteration order unless that read is recorded in the dep graph (env_var_os query, tracked file, sorted/stable-hash containers).", + "subsystem": "incremental / query system", + "kind": "invariant", + "how_to_check": "Use mirth instrumentation of env::var*, SystemTime/Instant, fs::metadata, and iteration of std/Fx HashMaps. Attribute every read to the query on top of the stack and flag reads that are not covered by a tracking query. Confirm a finding by perturbing that input and checking whether the query's fingerprint changes.", + "cost": "cheap once instrumented", + "issues": [ + 162901, + 162407, + 159677, + 160255 + ], + "currently_tested": "Upstream has a lint against unordered iteration (rustc::potential_query_instability) with many allow exceptions. mirth's P4 covers metadata encoding only.", + "mirth_fit": "Excellent. This is mirth's core use case: extend the P4 hooks from encode_metadata to every query frame.", + "rank_reason": "Fewer issues are attributed here, but it is the root cause behind most P5 and P6 failures. It is cheap and has mirth's best fit, so it is ranked high despite the count." + }, + { + "name": "Metadata reads hit entries the writer wrote", + "statement": "Every (table, DefIndex) lookup and lazy decode that a dependent crate performs reads an entry the upstream crate's encoder actually wrote. It never silently gets a default or empty value.", + "subsystem": "metadata", + "kind": "invariant", + "how_to_check": "With mirth, log every table write (table, index) in the writer process and every lookup in reader processes across the whole cargo build. Join the logs. A read of an unwritten entry is a bug, except for tables documented as default-on-absent, which need an allowlist.", + "cost": "moderate (whole-build logs)", + "issues": [ + 163426, + 159233, + 159703 + ], + "currently_tested": "Not tested. Missing entries usually decode as defaults with no error.", + "mirth_fit": "Excellent. It needs both writer and reader instrumentation across one cargo build, which no other tool can do.", + "rank_reason": "It catches silent cross-crate bugs that nothing else sees. Few attributed issues, but very poor upstream coverage." + }, + { + "name": "Parallel front end is deterministic", + "statement": "-Zthreads=1 and -Zthreads=N give byte-identical outputs and identical diagnostics, and repeated -Zthreads=N runs agree with each other.", + "subsystem": "parallel front end", + "kind": "differential", + "how_to_check": "Build the corpus with -Zthreads=1 once and -Zthreads=8 three times. Byte-compare outputs and diff the JSON diagnostics. When they differ, compare mirth's per-table write logs.", + "cost": "cheap", + "issues": [ + 129094, + 150451, + 153391 + ], + "currently_tested": "A parallel-rustc CI job runs some UI tests. Outputs are not compared against single-threaded builds.", + "mirth_fit": "Good. It reuses the P5 harness with a different flag, and mirth's write logs localize the divergence.", + "rank_reason": "Cheap, a big upstream gap, and it will matter more as the parallel front end matures. Low current issue count." + }, + { + "name": "rustdoc accepts every crate rustc accepts", + "statement": "If rustc check succeeds on a crate, rustdoc in every output mode (HTML, --generate-link-to-definition, --output-format json, --document-private-items) also succeeds. Span-map and type-relative path resolution use the body that owns the path, and proc-macro spans do not overwrite source spans.", + "subsystem": "rustdoc", + "kind": "differential", + "how_to_check": "For each crate in the corpus, run cargo check, then cargo rustdoc with each mode. Any rustdoc failure on a crate that checks is a bug. Validate the JSON output against rustdoc-types.", + "cost": "cheap", + "issues": [ + 156418, + 156327, + 149089, + 150153, + 147882, + 147057, + 158050, + 157756 + ], + "currently_tested": "docs.rs and crater cover the default HTML mode. --generate-link-to-definition and JSON on real crates are barely exercised.", + "mirth_fit": "Low to moderate. This is mostly a harness differential. mirth could record which TypeckResults body was consulted for each path in order to check that it is the owning body.", + "rank_reason": "Several bugs, cheap, and modes that are clearly undertested. It finds crashes, not silent errors." + }, + { + "name": "Results agree across opt-level, mir-opt-level, codegen-units and LTO", + "statement": "A program's observable behaviour (test results, stdout, exit and panic status) and its linkability are the same at -Copt-level 0/3, -Zmir-opt-level 0/4, codegen-units 1/16, lto off/thin/fat, and incremental on/off. Where Miri can run the program, it agrees too.", + "subsystem": "mir-opt / codegen / LTO", + "kind": "semantic", + "how_to_check": "Run cargo test for crates in the corpus under a configuration matrix and diff per-test outcomes. Flag link failures that appear in only one configuration. Run Miri on a subset as the reference.", + "cost": "expensive (matrix of full builds plus test runs)", + "issues": [ + 163779, + 163220, + 162348, + 153645, + 153451 + ], + "currently_tested": "Individual mir-opt tests exist. There is no systematic matrix over real crates' test suites, and LTO/codegen-unit link failures are found by users.", + "mirth_fit": "Low. It is an output-level differential. mirth could record which MIR pass changed a function that later misbehaves, which helps bisect the cause.", + "rank_reason": "Miscompiles are severe, but the check is expensive and few issues are attributed." + }, + { + "name": "Every span is valid", + "statement": "Every span in diagnostics (including suggestion parts) and in encoded metadata has lo <= hi, lies inside an existing SourceFile, and falls on UTF-8 character boundaries. Each suggestion's substitutions do not overlap.", + "subsystem": "span / diagnostics / metadata", + "kind": "invariant", + "how_to_check": "Two checks. (1) Post-process --error-format=json for every corpus crate, plus a corpus with multibyte identifiers and strings: check byte ranges against the file contents and check that suggestion parts are disjoint. (2) With mirth, validate every span encoded into rmeta against the SourceMap when it is written.", + "cost": "cheap", + "issues": [ + 151610, + 151607, + 147339, + 131292, + 156316, + 155037 + ], + "currently_tested": "There is an internal assertion when a bad span is sliced. Nothing validates spans systematically, and multibyte source is rare in tests.", + "mirth_fit": "Good for the metadata half (hook span encoding). The diagnostic half needs only a JSON post-processor.", + "rank_reason": "Cheap and broad. Moderate count. Coverage of multibyte source is weak." + }, + { + "name": "Semantically neutral edits do not change results", + "statement": "Edits that cannot change meaning do not change accept/reject, the lint set or the metadata (apart from spans). Examples: adding an unused `use` of a non-trait item, wrapping an item in an identity macro_rules, replacing a type with a projection or alias that normalizes to it, enabling an unrelated feature gate, or reordering items.", + "subsystem": "resolve / lints / trait-system / macros", + "kind": "differential", + "how_to_check": "Apply automatic rewrites from a fixed catalogue to corpus crates. Compare exit status, the multiset of diagnostics keyed by (code, lint, item DefPath), and rmeta after span normalization.", + "cost": "moderate", + "issues": [ + 156004, + 86959, + 160741, + 157758, + 157757, + 152004, + 151983 + ], + "currently_tested": "UI tests check isolated forms. No transformation-based differential testing exists.", + "mirth_fit": "Moderate. mirth can compare per-query results keyed by DefPath between the original and the rewritten build, which is finer-grained than diagnostics.", + "rank_reason": "Catches lint false positives and inference instability that usually need a human to spot. Moderate cost. The rewrite catalogue must be kept truly semantics-preserving." + }, + { + "name": "A successful compile contains no error types", + "statement": "If a compilation exits 0, no ty::Error, ConstKind::Error or ErrorGuaranteed-carrying value appears in any query result, MIR body or encoded metadata. Equivalently, every error type is created only after an error has been emitted.", + "subsystem": "type-system / diagnostics", + "kind": "invariant", + "how_to_check": "With mirth, hook the construction of Ty::new_error, Const::new_error and region errors and record the creation site. When the session ends with zero errors, any recorded creation is a bug. Also scan encoded metadata for error types.", + "cost": "cheap", + "issues": [ + 150969, + 149588, + 151299, + 154780 + ], + "currently_tested": "ErrorGuaranteed makes forging an error hard. Delayed bugs ICE at session end only if they were delayed. Silent paths that drop an error are not checked.", + "mirth_fit": "Good. It needs one hook at the error constructors and a check at session end.", + "rank_reason": "Cheap and catches silent acceptance of bad code. Few attributed issues." + }, + { + "name": "Layout views agree", + "statement": "Every component's view of a type's size and layout agrees with layout_of: the transmute SizeSkeleton check, const-eval, and wrapper types such as unsafe binders and repr(transparent). A repr(transparent) wrapper has the same ABI as its non-ZST field. A transmute that SizeSkeleton rejects as size-mismatched really does have different concrete sizes.", + "subsystem": "layout / type-system", + "kind": "invariant", + "how_to_check": "With mirth, record layout_of results per type. At each transmute check and each transparent or wrapper type, compare against layout_of of the inner type and flag disagreements. Also feed a generated corpus of repr(align/packed/C/transparent) types.", + "cost": "cheap", + "issues": [ + 154426, + 154424, + 155412, + 88290 + ], + "currently_tested": "There are debug assertions in layout code and transmute UI tests. Nothing cross-checks SizeSkeleton against layout_of.", + "mirth_fit": "Good. Hook layout_of and SizeSkeleton::compute and compare their results.", + "rank_reason": "Cheap and catches both false rejects and possible unsoundness. Low count." + }, + { + "name": "Each encoded record is written once and hashed after it is final", + "statement": "The metadata encoder and the on-disk query cache write each table entry and each dep node at most once. No fingerprint or crate_hash is computed over an output before its encoder has finished.", + "subsystem": "metadata / incremental", + "kind": "invariant", + "how_to_check": "With mirth, log (table, key) and (dep node) writes and fail on any duplicate. Log when crate_hash and the svh are computed relative to the encoder finishing, and fail if hashing happens first.", + "cost": "cheap", + "issues": [ + 150018, + 142778, + 163426 + ], + "currently_tested": "There are a few assertions in the dep-graph encoder. The ordering of hashing relative to encoding is not checked.", + "mirth_fit": "Excellent. It uses the same write hooks as the reader-writer property.", + "rank_reason": "Cheap and reuses existing hooks. Low count." + }, + { + "name": "rustdoc output is the same however an item is re-exported", + "statement": "An item documented through a direct definition, a named re-export, a glob re-export or a multi-level re-export shows the same docs, cfg badges and signature, apart from additions made at the re-export site.", + "subsystem": "rustdoc", + "kind": "differential", + "how_to_check": "Use rustdoc JSON. For every re-exported item in the corpus, compare the inlined item's docs, cfg and signature with the original item's (outer docs and cfg added at the re-export are allowed).", + "cost": "cheap", + "issues": [ + 119965, + 81893, + 53724, + 96166, + 154921 + ], + "currently_tested": "rustdoc tests cover specific re-export shapes only.", + "mirth_fit": "Poor. It is a JSON post-processor.", + "rank_reason": "Cheap with a moderate count, but documentation-only impact." + }, + { + "name": "#[expect] is fulfilled exactly when the lint fires", + "statement": "#[expect(L)] is reported unfulfilled if and only if replacing it with #[warn(L)] would emit no L diagnostic in that scope. This still holds after derive, cfg_attr and macro expansion.", + "subsystem": "lints", + "kind": "differential", + "how_to_check": "For each #[allow] or #[expect] in the corpus, generate the #[warn] and #[expect] variants. Build both and check the pairing: lint emitted exactly when the expectation is fulfilled.", + "cost": "moderate (one build per attribute; batchable)", + "issues": [ + 152004, + 151983, + 152401, + 152289 + ], + "currently_tested": "There are UI tests for specific cases. No automatic pairing check exists.", + "mirth_fit": "Low. A source-rewrite harness is enough.", + "rank_reason": "Mechanical, but narrow. The issue attribution is shared with the neutral-edits property." + }, + { + "name": "Query results contain no inference variables", + "statement": "No value that is returned from a query, hashed into the dep graph or stored in the on-disk cache contains inference variables (type, const or region) or placeholders from an inference context.", + "subsystem": "query system / type inference", + "kind": "invariant", + "how_to_check": "With mirth, at every query return and every HashStable call, check has_infer() and has_placeholders() on the value and record the query and key on failure.", + "cost": "cheap", + "issues": [ + 153525, + 153524 + ], + "currently_tested": "There are a few debug assertions at specific sites. No global check exists.", + "mirth_fit": "Excellent. One generic hook at the query provider return.", + "rank_reason": "Very cheap and fully general, but few attributed bugs." + }, + { + "name": "Each diagnostic is emitted once", + "statement": "No diagnostic with the same level, code, message and primary span is emitted twice in one session, in either clean or incremental mode.", + "subsystem": "diagnostics", + "kind": "invariant", + "how_to_check": "Post-process JSON diagnostics from corpus builds and from the UI suite run with --error-format=json, and look for duplicate keys.", + "cost": "cheap", + "issues": [ + 115376, + 106571 + ], + "currently_tested": "The emitter deduplicates some cases. UI .stderr snapshots would show duplicates, but nothing asserts they are absent.", + "mirth_fit": "Poor. It is an output check.", + "rank_reason": "Cheap, low impact, low count." + }, + { + "name": "Symbol names are injective and stable", + "statement": "Distinct monomorphized instances never get the same symbol name, under both v0 and legacy mangling, including for feature-gated type constructors. A v0 symbol demangles to the instance's path.", + "subsystem": "symbol mangling", + "kind": "invariant", + "how_to_check": "With mirth, record (Instance, symbol_name) pairs in every codegen process of a cargo build. Check that the mapping is injective across crates and round-trip each v0 symbol through rustc-demangle against the printed instance.", + "cost": "cheap", + "issues": [ + 158644 + ], + "currently_tested": "Mangling UI tests cover specific types. There is no whole-build injectivity check.", + "mirth_fit": "Good. Hook symbol_name, which is easy, and join the results across crates.", + "rank_reason": "Low count, but symbol clashes cause silent link-time misbinding." + }, + { + "name": "dyn types have a dyn-compatible principal", + "statement": "In a compilation that succeeds, every TyKind::Dynamic that reaches typeck results, MIR or metadata has a dyn-compatible principal trait, including subtraits and associated-type bindings.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "With mirth, when a Dynamic type is interned, record its principal. At the end of a successful session, assert is_dyn_compatible for each one.", + "cost": "cheap", + "issues": [ + 154619 + ], + "currently_tested": "Dyn-compatibility is checked at direct dyn-syntax sites. No global check exists.", + "mirth_fit": "Good. One intern hook plus a check at session end.", + "rank_reason": "A soundness property and cheap, but few bugs are attributed." + }, + { + "name": "Machine-applicable suggestions apply cleanly", + "statement": "Applying every MachineApplicable suggestion (rustfix) gives code that compiles, and the diagnostic that produced the suggestion is gone.", + "subsystem": "diagnostics / lints", + "kind": "differential", + "how_to_check": "Run cargo fix --broken-code on corpus crates with all warn-by-default lints enabled, then rebuild. Check that it compiles and that the original lint no longer fires.", + "cost": "moderate", + "issues": [ + 161213, + 161472 + ], + "currently_tested": "run-rustfix UI tests cover curated cases.", + "mirth_fit": "Poor. It is a cargo-level differential.", + "rank_reason": "Low count, user-facing breakage only." + }, + { + "name": "Polonius accepts at least what NLL accepts, and no more than is sound", + "statement": "-Zpolonius=next accepts every program that NLL accepts. Programs that only Polonius accepts run without UB under Miri.", + "subsystem": "borrowck", + "kind": "differential", + "how_to_check": "Borrow-check the corpus with both. Any NLL-accept/Polonius-reject is a bug. For Polonius-only accepts, which mostly come from the UI suite, run under Miri.", + "cost": "moderate", + "issues": [ + 153215, + 160669 + ], + "currently_tested": "There is a polonius compare-mode on some UI tests.", + "mirth_fit": "Low. It is an accept/reject differential.", + "rank_reason": "Low count. Polonius is experimental." + }, + { + "name": "Unstable syntax is gated before expansion", + "statement": "Every unstable syntactic form is rejected on stable without its feature even inside #[cfg(FALSE)] or unused macro_rules arms, so that removing the gate later cannot break code.", + "subsystem": "parser / feature gates", + "kind": "invariant", + "how_to_check": "Take each feature-gate UI test that exercises syntax, wrap the gated syntax in #[cfg(FALSE)], and compile without the feature. The compile must fail.", + "cost": "cheap", + "issues": [ + 152501, + 152499 + ], + "currently_tested": "Some features have cfg(FALSE) tests. It is not enforced for each one.", + "mirth_fit": "Poor. It is a source-rewrite harness.", + "rank_reason": "Cheap but narrow." + }, + { + "name": "Built-in attributes reject malformed arguments", + "statement": "Every built-in attribute rejects argument forms its template does not allow. It errors or lints instead of silently ignoring them.", + "subsystem": "attribute parsing", + "kind": "invariant", + "how_to_check": "From the BUILTIN_ATTRIBUTES templates, generate malformed forms for each attribute (extra arguments, list instead of word, name-value instead of word) and assert that a diagnostic is produced.", + "cost": "cheap", + "issues": [ + 154977 + ], + "currently_tested": "There are per-attribute tests, but not generated from the templates.", + "mirth_fit": "Poor.", + "rank_reason": "Cheap, generated mechanically, but low impact." + }, + { + "name": "Advertised target features exist in LLVM", + "statement": "Every target feature that rustc lists for a target, and every feature that an asm register class or intrinsic requires, is recognized by LLVM for that target and is consistent with the target's baseline.", + "subsystem": "codegen / targets", + "kind": "invariant", + "how_to_check": "For every target, enumerate rustc's feature table and query LLVM's feature list. Compile an empty function with each feature enabled and confirm there is no 'unknown feature' warning or spurious requirement.", + "cost": "cheap", + "issues": [ + 159976 + ], + "currently_tested": "There is a partial tidy-style check.", + "mirth_fit": "Poor.", + "rank_reason": "Cheap but narrow." + }, + { + "name": "Eq and Hash agree for std types", + "statement": "For every std type implementing both Eq and Hash, a == b implies hash(a) == hash(b), for every Hasher.", + "subsystem": "library", + "kind": "semantic", + "how_to_check": "Run property tests over generated values of std types, including borrowed forms through Borrow, with several hashers.", + "cost": "cheap", + "issues": [ + 161651 + ], + "currently_tested": "There are ad hoc tests only.", + "mirth_fit": "None. It is a library property test.", + "rank_reason": "One attributed bug and library-only." + }, + { + "name": "Crash-freedom baseline (no ICE, hang or stack overflow)", + "statement": "rustc never panics, hangs or overflows its stack on any input, valid or invalid. For bad input it emits diagnostics and returns a normal error exit.", + "subsystem": "whole compiler", + "kind": "invariant", + "how_to_check": "Run fuzzers (icemaker, fuzz-rustc, tree-splicer) over the UI suite and the corpus, with a debug-assertions compiler, a timeout, and a reduced RUST_MIN_STACK. Exit code 101, a timeout, or SIGSEGV counts as a failure.", + "cost": "expensive (fuzzing campaigns)", + "issues": [ + 162147, + 162146, + 161837, + 161770, + 161101, + 161062, + 160994, + 160798, + 159939, + 159890, + 159889, + 159559, + 155802, + 155405, + 151304, + 150969, + 149703, + 148630, + 146834, + 141804, + 140107, + 160171, + 159989, + 160255, + 159252, + 159261, + 159299, + 159323, + 159591, + 159685, + 159815, + 159867, + 160008, + 160024, + 159914, + 159958, + 160060, + 160539, + 160314, + 160390, + 153601, + 153599, + 153539, + 153502, + 153499, + 153433, + 153420, + 153391, + 153390, + 153389, + 153388, + 153354, + 153351, + 153236, + 153199, + 153198, + 152895, + 152797, + 152744, + 152684, + 152683, + 152682, + 152663, + 152654, + 152653, + 152633, + 152606, + 152601, + 152595, + 152545, + 152518, + 151631, + 151477, + 150457, + 149695, + 149643, + 146754, + 143498, + 128801, + 127971, + 127423, + 125805, + 113870, + 157516, + 153205, + 157937, + 152716, + 157950, + 157853, + 157888, + 149821, + 150927, + 150928, + 150976, + 151008, + 151037, + 151213, + 152340, + 151878, + 151708, + 151331, + 151273, + 151027, + 150960, + 148953, + 152030, + 151591, + 151814, + 137582, + 146984, + 152244, + 152158, + 152309, + 156482, + 158439, + 158429, + 137190, + 135470, + 149920, + 154820, + 154539, + 123629, + 153744, + 153743, + 152936 + ], + "currently_tested": "Fuzzers run continuously by volunteers and most of these issues came from them. Crater catches ICEs on published crates.", + "mirth_fit": "Poor. An ICE is already its own oracle (exit code 101), so mirth adds no detection. At most it gives earlier localization when an invariant hook above fires first.", + "rank_reason": "By raw count this is the largest group. It is ranked last because it is the existing fuzzing baseline, not a property in the requested sense. Most of these bugs need fuzzer-generated invalid code, so a corpus-wide check would not have caught them." + } + ], + "notes": "On your question (is mirth good at catching ICEs?): No, and you are right. An ICE is already its own oracle (exit code 101 plus a panic message), and nearly all of the ~130 ICE issues here came from fuzzers feeding invalid or unstable code. mirth cannot see them unless its inputs trigger them, and a corpus of real crates rarely does. mirth's strength is bugs that never crash: nondeterminism, incremental divergence, writer/reader mismatch, untracked reads, silent error-dropping, and layout or ABI disagreement. It can also turn debug-only assertions into always-on checks in a release-built instrumented compiler, so a violation is reported where it is created instead of as a later ICE. That only helps when the corpus exercises the code. For the MGCA/next-solver invariants, run the UI suite through the instrumented compiler as well as your chosen crates.\n\nConsolidation: 94 candidates became 29 properties.\n- Merged: five determinism candidates into P5+; four incremental ones into P6+; three untracked-state ones into P4+.\n- MIR validator, storage markers, type-const normalization and borrowck normalization became one property: run -Zvalidate-mir and -Zlint-mir on real crates.\n- Valtree, generic-argument count, ParamEnv and const-generic bounds candidates became one intern-time well-formedness property.\n- Lint false-positive and equivalent-form candidates became \"neutral edits do not change results\".\n- Error-not-emitted and TyKind::Error became \"a successful compile contains no error types\".\n\nDropped as one-off fixes or not mechanically checkable: performance consistency, Rc/Arc pointer-compare, become/rust-call, enum-variant generic args, pin_v2 gating, volatile store.\n\nFolded into the crash-freedom baseline because their invariant is already enforced by a panic and mirth adds no detection: HIR-for-queried-DefIds, expect_local on non-local DefIds, DefKind handling, parser recovery (3), glob imports, trait aliases, THIR categorization, typing mode, vtable supertraits, transmute-before-const-eval, diagnostic-arg uniqueness.\n\nCaveats:\n1. Agents assigned issue IDs without checking the issue text against the property. Some lists look wrong:\n - the const-eval list folded into P5+ (138089, 133966, etc. look like ICEs);\n - the volatile/FFI IDs folded into ABI;\n - 43894 and 161382 appear under both ABI and library soundness;\n - 106571 appears under both determinism and duplicate diagnostics.\n Rankings by count should be read with this in mind. Verify before citing any number.\n2. Two differentials from the task's own list had no candidate, so no issues support them: check-vs-build diagnostics, and a crate compiled alone vs as a dependency. Both fit mirth's whole-cargo-build model and are worth adding later.\n3. As you requested, rerun the existing P5 and P6 baseline on the patched compiler before extending them. The perturbations and edit sequences in P5+ and P6+ are harness changes on top of what exists.\n\nBest mirth return per unit of effort, given the existing hooks:\n- P5+ and P6+ (extend harnesses);\n- P4+ (move hooks from encode_metadata to every query frame);\n- the read-hits-written and write-once/hash-after-final properties (reuse the table-write logs);\n- no inference variables in query results, and no error types in a successful compile (one hook each).\n\nABI, library fault-injection and rustdoc are real gaps, but they belong to other tools (abi-cafe, Miri, JSON post-processors)." + }, + "raw": [ + { + "issues_read": 100, + "candidates": [ + { + "name": "Codegen output determinism", + "statement": "The compiler must produce byte-identical output files when run multiple times on the same source, regardless of parallel execution level.", + "subsystem": "codegen", + "kind": "differential", + "how_to_check": "Build a crate with `-Zthreads=1` and `-Zthreads=N` (where N>1), then byte-compare the resulting .rmeta and .o/.rlib files. Repeat the build 2-3 times with same thread count and compare outputs.", + "cost": "cheap", + "issues": [ + 129094, + 150451 + ], + "currently_tested": "Partially - rustc has some determinism tests but not comprehensive across all subsystems with parallel execution" + }, + { + "name": "Incremental build equivalence", + "statement": "An incremental rebuild after a no-op (unchanged source) must produce output files byte-identical to a clean build of the same source.", + "subsystem": "incremental", + "kind": "differential", + "how_to_check": "Build a crate normally, then rebuild incrementally with `-Cincremental=` and no source changes. Byte-compare the .rmeta files. Also compare with fresh clean build without incremental enabled.", + "cost": "moderate", + "issues": [ + 162407, + 162901, + 162585 + ], + "currently_tested": "Partially - incremental tests exist but focused on correctness; deterministic output comparison less common" + }, + { + "name": "ABI compatibility with platform C standard", + "statement": "The compiler must emit LLVM IR for function ABIs that matches the platform's standard C calling convention as implemented by major C compilers (GCC/Clang).", + "subsystem": "codegen", + "kind": "differential", + "how_to_check": "For a test function with various struct/union/float parameters, compile with rustc to LLVM IR and with Clang/GCC to LLVM IR. Compare parameter passing conventions (register vs stack, field layout interpretation) for equivalent C and Rust code.", + "cost": "moderate", + "issues": [ + 121408, + 162011, + 163173, + 163074, + 163075 + ], + "currently_tested": "Yes - ABI tests exist (check_abi.rs) but may not catch all edge cases like unions with floats" + }, + { + "name": "MIR optimization semantic preservation", + "statement": "All MIR-level optimizations must preserve the observable semantics of the program, including panic behavior, overflow behavior, and computed values.", + "subsystem": "mir-opt", + "kind": "semantic", + "how_to_check": "For optimized code, run the compiled binary against a reference (unoptimized build or Miri interpretation) and compare output and exit codes. Test with various opt-levels.", + "cost": "expensive", + "issues": [ + 163779, + 163220, + 162348 + ], + "currently_tested": "Partially - MIR-opt tests check output but don't systematically run programs; no regression test suite comparing opt-level outputs" + }, + { + "name": "Custom allocator interface soundness", + "statement": "When using the allocator_api, the pointer returned from allocate() must have provenance over the entire allocated size, and deallocate() must be called with a pointer having compatible provenance.", + "subsystem": "allocators", + "kind": "invariant", + "how_to_check": "Use Miri with stacked borrows to verify that all allocator calls maintain pointer provenance correctly. Build stdlib and crates that use custom allocators with Miri.", + "cost": "expensive", + "issues": [ + 162720, + 162719 + ], + "currently_tested": "No - allocator soundness is not systematically tested; would require instrumented Miri runs" + }, + { + "name": "Metadata encoding completeness", + "statement": "All compiler-internal queries that read another crate's metadata must have their dependencies on that metadata recorded such that the metadata reader can find all necessary data.", + "subsystem": "metadata", + "kind": "invariant", + "how_to_check": "Instrument the metadata encoder to verify: (1) before reading crate_hash, metadata encoding is complete; (2) every metadata read operation has corresponding encoder writes; (3) incremental invalidation correctly invalidates encoded data.", + "cost": "moderate", + "issues": [ + 163426 + ], + "currently_tested": "No - metadata invariants are not systematically verified; would require compiler instrumentation" + }, + { + "name": "Concurrency-free incremental compilation state", + "statement": "The compiler must not read environment variables, filesystem metadata, or clock values during incremental query computation in a way that creates dependencies between the query result and that untracked state.", + "subsystem": "incremental", + "kind": "invariant", + "how_to_check": "Mock/intercept filesystem and clock reads during incremental compilation, and verify that query results are not influenced by values read. Compare incremental build hashes when these are changed but source is identical.", + "cost": "expensive", + "issues": [ + 162901, + 162407 + ], + "currently_tested": "No - untracked state reads are not systematically verified; would require compiler instrumentation" + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "Deterministic metadata encoding", + "statement": "Two consecutive clean builds of the same crate should produce byte-identical .rmeta files. Metadata encoding should not depend on untracked state like environment variables, random HashMap iteration order, or clock values.", + "subsystem": "metadata", + "kind": "differential", + "how_to_check": "Build the same crate twice cleanly with identical inputs, compare the resulting .rmeta files byte-by-byte. Use mirth to instrument what state is read during metadata encoding.", + "cost": "moderate", + "issues": [ + 106571 + ], + "currently_tested": "Partially - rustc has some determinism tests but mirth would provide more comprehensive instrumentation of what state is accessed during encoding" + }, + { + "name": "No spurious diagnostics from overlapping suggestions", + "statement": "When emitting a diagnostic suggestion, the compiler should not panic or emit incorrect output if suggestion spans overlap. Error messages should remain well-formed even when suggestions cannot be properly generated.", + "subsystem": "diagnostics", + "kind": "invariant", + "how_to_check": "Compile code that triggers overlapping suggestion spans, verify no ICE occurs and error message is sensible. Instrument rustc_errors suggestion generation code.", + "cost": "cheap", + "issues": [ + 161213, + 161472 + ], + "currently_tested": "Minimally - mostly caught by individual test cases, not systematic checking of suggestion validity" + }, + { + "name": "Lint conditions are accurate and not overly broad", + "statement": "Lint rules should fire consistently - if a pattern matches the lint's condition in one context, it should match in all equivalent contexts. False positives should not occur based on syntactic form.", + "subsystem": "resolve", + "kind": "differential", + "how_to_check": "Write equivalent code in different forms (macros vs inline, 2018 vs 2024 edition, different placement), verify lint fires/doesn't fire consistently. Compile with rustc and check lint output across variants.", + "cost": "moderate", + "issues": [ + 86959, + 160741 + ], + "currently_tested": "Via UI test suite, but the suite may not cover all equivalent forms systematically" + }, + { + "name": "Soundness: No unsound behavior in safe code", + "statement": "Code marked #![forbid(unsafe_code)] or code using only safe APIs should never cause undefined behavior, memory corruption, or use-after-free. This includes panic-safety of data structure operations.", + "subsystem": "codegen", + "kind": "semantic", + "how_to_check": "Run Miri on code marked #![forbid(unsafe_code)], check for UB errors. Also run under Valgrind or ASan on compiled binaries to detect memory safety violations.", + "cost": "expensive", + "issues": [ + 43894, + 114581, + 158165, + 160815, + 161018, + 161382 + ], + "currently_tested": "Partially via Miri for some crates, but not systematically for all safe code paths" + }, + { + "name": "Panic-safe state consistency", + "statement": "If a panicking comparator or allocator is used during a destructive operation, the data structure should either complete the operation or abort entirely - it should never leave the structure in a partially-modified state that causes double-free or data loss.", + "subsystem": "codegen", + "kind": "invariant", + "how_to_check": "Instrument alloc code to trigger panics at various points during split_off, insert, etc. Verify no double-free occurs and data is either fully updated or fully reverted.", + "cost": "expensive", + "issues": [ + 158165 + ], + "currently_tested": "Via Miri for some code, but panic-safety during comparisons is not systematically tested" + }, + { + "name": "Eq-Hash consistency", + "statement": "If two values of a type compare as equal via Eq, they must have the same hash value. If hash(a) != hash(b), then a != b must be true.", + "subsystem": "codegen", + "kind": "invariant", + "how_to_check": "Compare Eq and Hash implementations for all std types. Build simple programs that hash and compare values, verify consistency. Check with different hasher implementations.", + "cost": "cheap", + "issues": [ + 161651 + ], + "currently_tested": "Not systematically - caught only when regressions manifest in real code" + }, + { + "name": "ABI consistency within and across platforms", + "statement": "The same type should use the same calling convention and layout across all targets that support it. platform-specific ABIs should be consistent with reference implementations (Clang, GCC).", + "subsystem": "codegen", + "kind": "differential", + "how_to_check": "Use abi-cafe or similar cross-platform ABI test suite. Compile FFI test cases targeting different platforms, verify struct layout and calling conventions match reference implementations.", + "cost": "expensive", + "issues": [ + 43894, + 161382, + 160827 + ], + "currently_tested": "Not systematically - mostly caught by abi-cafe crater runs and user reports" + }, + { + "name": "No diagnostic duplicates across incremental rebuilds", + "statement": "When running in incremental mode with --error-format=json, each unique diagnostic should appear exactly once, not duplicated based on Span::parent relationships.", + "subsystem": "incremental", + "kind": "differential", + "how_to_check": "Compile a crate with --error-format=json in incremental mode, parse JSON output and verify no duplicate error codes/messages. Compare against clean build output.", + "cost": "cheap", + "issues": [ + 106571 + ], + "currently_tested": "Minimally - test suite doesn't systematically check for diagnostic duplicates in JSON output" + }, + { + "name": "Trait solver produces correct results", + "statement": "The trait solver should accept or reject the same code consistently. GATs, HRTB bounds, and associated type resolution should work correctly and not require workarounds.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "Compile trait-heavy codebases (yoke, futures) with old and new solver. Verify same code is accepted/rejected consistently. Check that derive(Clone) works on recursive GATs.", + "cost": "expensive", + "issues": [ + 102580, + 90950, + 102580 + ], + "currently_tested": "Partially - next-solver fixes are tested, but old solver regressions not always caught" + }, + { + "name": "Type inference produces deterministic results", + "statement": "Type inference should produce the same inferred types regardless of what other types are in scope. Importing a type should not change inference for unrelated type variables.", + "subsystem": "trait-system", + "kind": "differential", + "how_to_check": "Build code with and without importing unrelated types (e.g., serde_json::Value), verify same types are inferred in both cases. Check error messages are helpful.", + "cost": "moderate", + "issues": [ + 156004 + ], + "currently_tested": "Via UI tests for specific issues, but not systematically for all inference contexts" + }, + { + "name": "No spurious lifetime diagnostics", + "statement": "The borrow checker should emit exactly the diagnostic errors needed to explain the problem, not spurious duplicates or unrelated lifetime requirements.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "Compile async functions with lifetime issues, verify exactly one error per constraint violation. Instrument borrowck diagnostics to catch duplicates.", + "cost": "cheap", + "issues": [ + 115376 + ], + "currently_tested": "Via UI tests, but the test suite doesn't systematically verify absence of spurious diagnostics" + }, + { + "name": "Rustdoc generation is consistent across equivalent code", + "statement": "Documentation should be generated consistently whether a type is re-exported once, twice, or defined directly. Inner and outer doc comments should not interact incorrectly.", + "subsystem": "resolve", + "kind": "differential", + "how_to_check": "Generate rustdoc for equivalent code with different re-export chains (direct def, one level, two levels of re-export). Compare HTML output and verify consistency.", + "cost": "moderate", + "issues": [ + 119965, + 81893, + 53724 + ], + "currently_tested": "Via rustdoc UI tests, but not systematically for multi-level re-exports" + }, + { + "name": "Performance is consistent for equivalent code", + "statement": "Two semantically equivalent pieces of code should have similar performance characteristics. Refactorings that should have no semantic impact should not regress performance significantly.", + "subsystem": "codegen", + "kind": "differential", + "how_to_check": "Compile and benchmark code before and after refactorings (e.g., before/after BorrowedBuf changes). Measure wall-clock time and memory usage, verify no regression >10%.", + "cost": "expensive", + "issues": [ + 158008 + ], + "currently_tested": "Not systematically - performance regressions found via manual benchmarking and crater" + }, + { + "name": "ICEs do not occur on valid code", + "statement": "The compiler should never panic on valid code. All user code that type-checks should compile without internal compiler errors.", + "subsystem": "codegen", + "kind": "invariant", + "how_to_check": "Compile a diverse set of valid Rust code (benchmark suite, tokio, serde, etc), verify no ICEs occur. Use RUSTFLAGS=-Ztreat-err-as-bug to catch any panics.", + "cost": "expensive", + "issues": [ + 162147, + 162146, + 161837, + 161770, + 161213, + 161101, + 161062, + 160994, + 160798, + 159939, + 159890, + 159889, + 159559, + 155802, + 155405 + ], + "currently_tested": "Partially via test suite and crater, but new code patterns can still trigger ICEs" + }, + { + "name": "Feature gates enforce declared feature availability", + "statement": "A target should only advertise features in its feature list if those features are actually available. Intrinsics and inline asm register classes should only require features that the target actually has.", + "subsystem": "codegen", + "kind": "invariant", + "how_to_check": "Cross-compile inline asm code to all targets, verify no spurious feature requirement errors. Check declared target features match what LLVM actually provides.", + "cost": "moderate", + "issues": [ + 159976 + ], + "currently_tested": "Minimally - mostly caught when users report issues on specific platforms" + }, + { + "name": "Macro expansions preserve semantic meaning", + "statement": "When a macro expands and is used differently (by position, edition, context), the semantic meaning should remain constant. A pattern that requires parens in one edition should not be differently required in another.", + "subsystem": "macros", + "kind": "differential", + "how_to_check": "Expand macros into different contexts (different editions, different scopes), verify the resulting code has the same meaning and parsing rules apply consistently.", + "cost": "moderate", + "issues": [ + 86959 + ], + "currently_tested": "Via UI tests for specific macro bugs, not systematically for macro semantics" + }, + { + "name": "Compilation with -Zpolonius produces sound results", + "statement": "The polonius borrow checker variant should never unsoundly accept code that violates lifetime requirements, even when multiple defining uses of the same opaque type occur.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "Compile code with -Zpolonius=next, run under Miri, verify no UB occurs. Check that dangling references are rejected.", + "cost": "expensive", + "issues": [ + 160669 + ], + "currently_tested": "Partially via Miri tests, but not systematically for all trait/opaque interactions" + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "Panic instead of error", + "statement": "When presented with invalid code, the compiler should emit a proper diagnostic error rather than panicking (ICEing). If the compiler encounters a situation it cannot handle, it must either recover gracefully or report the error, never panic.", + "subsystem": "general/diagnostics", + "kind": "invariant", + "how_to_check": "Run the compiler on a collection of invalid-but-minimal code samples (like those in the issues), capture both exit code and stderr. No panic messages should appear; instead only error diagnostics with E-codes. Can be tested with: for each issue ICE, extract the minimal reproducer, compile, verify clean error message (no 'panicked at' or 'thread ... panicked').", + "cost": "expensive", + "issues": [ + 151304, + 150969, + 149703, + 148630, + 146834, + 141804, + 140107, + 160171, + 159989, + 160255, + 159252, + 159261, + 159299, + 159323, + 159591, + 159685, + 159815, + 159867, + 160008, + 160024, + 159914, + 159958, + 160060, + 160539, + 160314, + 160390 + ], + "currently_tested": "Partially - the rustc test suite has some ICE tests but these are specific regression tests. The issue is that many latent panicking paths aren't exercised by the existing test suite until novel code shapes trigger them. A property-based approach testing structural invariants would be more comprehensive." + }, + { + "name": "Error created but not emitted", + "statement": "Whenever a {type error} or error constant is created during compilation (e.g., during failed type normalization), a corresponding diagnostic error must be emitted to the user. Silent error suppression hides actual problems from the programmer.", + "subsystem": "diagnostics/type-checking", + "kind": "invariant", + "how_to_check": "Instrument the compiler to track when error types are created (in type normalization, const evaluation, trait solving). Cross-check against emitted diagnostics. Code that creates {type error} should have a matching E-diagnostic emission nearby. Test case: compile code with unsolvable trait bounds and check that diagnostic is emitted, not {type error} left silent.", + "cost": "expensive", + "issues": [ + 150969, + 149588, + 151299 + ], + "currently_tested": "No - error silencing is typically only caught when output differs from expected diagnostic text. A focused test would instrument error sites and verify every creation has a corresponding emission." + }, + { + "name": "Metadata determinism", + "statement": "Across multiple clean builds of the same code, .rmeta files must be byte-identical. The contents should not depend on: unrelated files in the library search path, HashMap iteration order, environment variables, or random state.", + "subsystem": "metadata/incremental", + "kind": "differential", + "how_to_check": "Build the same crate multiple times with: (1) different library search paths containing unrelated crates, (2) different -Zrandomize-layout/HashMap orderings. Compare the resulting .rmeta file hashes with `sha256sum`. All must match. Code: for a test crate, run rustc 10 times with randomized HashMap seeds, hash the rmeta output, verify all hashes are identical.", + "cost": "cheap", + "issues": [ + 159677 + ], + "currently_tested": "Yes - there are determinism tests (e.g., P5 mentioned in the task description: 'two clean builds produce byte-identical .rmeta'). However, the specific case of search-path-independence and HashMap-order-independence for DocLinkResMap is now fixed but wasn't being tested before. P6 (incremental vs clean agreement) is related." + }, + { + "name": "Incremental vs clean agreement", + "statement": "An incremental rebuild after editing code should produce identical metadata and compiled output as a clean rebuild of the edited source. Incremental caches must not introduce divergence.", + "subsystem": "incremental", + "kind": "differential", + "how_to_check": "For a test crate: (1) perform a clean build, (2) make a semantically meaningful edit, (3) perform an incremental rebuild, (4) compare resulting .rmeta and .rlib hashes. Automate by: create a test crate, build it, touch/modify a file, rebuild with `cargo build` (incremental), then compare against `cargo clean && cargo build` (clean). Hashes must match.", + "cost": "moderate", + "issues": [ + 125564, + 135062, + 159677 + ], + "currently_tested": "Yes - property P6 from the task background: 'an incremental rebuild after an edit produces the same metadata as a clean build of the edited source.' This is tested but only as a regression test for specific fixed bugs, not as a systematic property check across all crates." + }, + { + "name": "Codegen correctness for volatile and FFI", + "statement": "Volatile memory operations and FFI function calls must generate the correct instruction sequences. Volatile stores must use volatile flags, bool types in FFI must use the platform-correct ABI calling convention, and other memory/calling-convention semantics must match the source intent.", + "subsystem": "codegen/LLVM", + "kind": "semantic", + "how_to_check": "Generate assembly code for programs using volatile_store and FFI bool returns, inspect the LLVM IR or final assembly. For volatile: check that LLVM IR includes volatile flag (not plain store), for FFI bool: check calling convention matches platform ABI (e.g., signext/zeroext on some platforms). Test: compile a fn with `volatile_store::()` and inspect codegen; compile FFI bool return and verify ABI extension.", + "cost": "moderate", + "issues": [ + 159815, + 159867, + 158897, + 159244, + 159116 + ], + "currently_tested": "Partially - LLVM IR tests exist for some operations (codegen-llvm tests), but they are regression tests for fixed bugs. Systematic checking of volatile and FFI bool codegen across platforms would require architecture-specific test matrices." + }, + { + "name": "No untracked state in metadata encoding", + "statement": "Metadata encoding must not depend on untracked compiler state: environment variables (except those explicitly tracked), clock values, randomly-seeded data structures (like HashMap iteration order), or values from unrelated library files. Only explicitly tracked query results should affect metadata.", + "subsystem": "metadata/incremental", + "kind": "invariant", + "how_to_check": "Instrument the compiler to track what data is read during metadata encoding. Cross-check against the incremental dependency graph. Any reads of environment, clock, or HashMap iteration should be logged. Run a metadata encoding pass with different environment states and verify output is identical. Code: use mirth or similar instrumentation to record 'what did the metadata writer read'; verify all reads are from the tracked dependency graph.", + "cost": "expensive", + "issues": [ + 159677, + 160255 + ], + "currently_tested": "Yes - property P4 from the task background: 'no environment variable, clock, or randomly seeded HashMap is read while encoding metadata.' This was tested via instrumentation in earlier work; issue #159677 fixed a case where DocLinkResMap HashMap iteration was observable." + }, + { + "name": "Reader sees written data from writer", + "statement": "Every DefId that a reader queries for must have been written by the writer. If a reader asks for metadata about a definition, that definition must exist in the metadata produced by the writer. No missing or dangling references.", + "subsystem": "metadata/incremental", + "kind": "invariant", + "how_to_check": "Instrument metadata reader and writer. Record all DefIds written during the writer phase. During reader phase, log all DefIds requested. Verify: every DefId in the reader's requests was written by the writer, no mismatches. Can be tested with a fuzzer that generates arbitrary DefId references and checks if they resolve correctly.", + "cost": "expensive", + "issues": [ + 159233, + 159703 + ], + "currently_tested": "Yes - property 'tracked': 'every query that reads a dependency's metadata records a dependency on that crate.' This is enforced through the dependency tracking system, but edge cases emerge (issues #159233, #159703) when new language features or query paths aren't properly integrated." + }, + { + "name": "Parser error recovery does not cross statement boundaries", + "statement": "When the parser encounters a syntax error in one statement and enters recovery mode, that recovery state must be cleared before parsing the next statement. Stale recovery state from a previous statement must not affect parsing of subsequent statements.", + "subsystem": "parser", + "kind": "invariant", + "how_to_check": "Feed the parser a sequence of malformed statements, each on separate lines. Verify that errors are reported only for the statement containing the syntax error, not for subsequent correct statements. Test case: parse code with `fn main() { for<> || {}; for<'a> || |_| {} }` - the error in the first should not affect the second. Automated: run parser on programs with multi-statement error sequences and verify error positions are correct.", + "cost": "cheap", + "issues": [ + 160171, + 159989 + ], + "currently_tested": "No - this is a parser internals property that would require parser-specific test infrastructure to verify that recovery state transitions are correct across statement boundaries." + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "MIR storage marker invariant", + "statement": "MIR code must not reference a local variable outside its StorageLive..StorageDead region. Every local use must occur within a region where StorageLive has been executed but StorageDead has not.", + "subsystem": "MIR opts", + "kind": "invariant", + "how_to_check": "Instrument MIR passes to track StorageLive/StorageDead for each local and assert that all Local references occur within valid storage regions. Run on all optimized MIR.", + "cost": "cheap", + "issues": [ + 158231 + ], + "currently_tested": "Partially: tests may catch crashes, but there is no systematic pass to validate storage invariants across all optimization passes." + }, + { + "name": "Symbol mangling feature completeness", + "statement": "All type constructors that are syntactically distinct (including feature-gated variations like splat) must produce distinct v0 mangled symbols. Symbol clash indicates a missing mangling component.", + "subsystem": "metadata", + "kind": "invariant", + "how_to_check": "Build two crates using different feature combinations (e.g., with/without splat, splatted vs non-splatted function signatures). Extract exported symbols and verify that semantically different types have different mangled names.", + "cost": "moderate", + "issues": [ + 158644 + ], + "currently_tested": "Partially: there are mangling tests but they do not systematically cover all feature combinations or validate against symbol clashes across feature variants." + }, + { + "name": "Compiler termination on degenerate obligations", + "statement": "The trait solver and type system must terminate on all input, including degenerate trait obligations, recursive type constraints, and self-referential type cycles. The compiler must not hang due to infinite loops in obligation processing or layout computation.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "Run rustc with timeout on test cases involving GAT projections, recursive trait bounds, layout cycles, and self-referential types. Check that all cases either error or complete within a reasonable time bound.", + "cost": "moderate", + "issues": [ + 157516, + 153205, + 157937 + ], + "currently_tested": "Partially: specific hangs may be caught if tests timeout, but there is no systematic property checking for termination." + }, + { + "name": "Const evaluation rejects generics in discriminant contexts", + "statement": "When evaluating const expressions in discriminants, array lengths, and const pattern matching, the compiler must not attempt to const-evaluate expressions containing unresolved type parameters. It should reject or error rather than panic.", + "subsystem": "const-eval", + "kind": "invariant", + "how_to_check": "Attempt const pattern matching on const values with generic types (e.g., StructuralPartialEq impls with PhantomData where T is a type parameter). Verify no ICE occurs and appropriate error is emitted.", + "cost": "cheap", + "issues": [ + 150296, + 148891 + ], + "currently_tested": "Partially: some const evaluation errors are caught, but there is no systematic check that generics in const paths are rejected before const evaluation attempts them." + }, + { + "name": "Parameter environment consistency", + "statement": "A parameter environment must not contain duplicate or conflicting predicates. Specifically, there must not be multiple ConstArgHasType predicates for the same const parameter with different types.", + "subsystem": "type-system", + "kind": "invariant", + "how_to_check": "After constructing parameter environments, scan for duplicates and conflicts. Build parameter environments from function delegations, trait impls, and generic contexts, and assert no duplicates exist.", + "cost": "cheap", + "issues": [ + 158675, + 158362 + ], + "currently_tested": "Partially: the ICE was caught in this batch, but there is no assertion preventing duplicate predicates during param-env construction." + }, + { + "name": "MIR typing mode correctness in body context", + "statement": "When building MIR for a function body, the typing environment must use body-level typing mode (reveal opaque types in this crate), not non_body_analysis mode. Using the wrong mode causes assertion failures and incorrect type normalization.", + "subsystem": "MIR opts", + "kind": "invariant", + "how_to_check": "Before MIR construction, assert that the typing environment is set to the correct mode for the current context. Test with next-solver enabled and type_alias_impl_trait features.", + "cost": "cheap", + "issues": [ + 158439, + 158429 + ], + "currently_tested": "Partially: the specific context (MIR building with next-solver) may not be systematically tested in the test suite." + }, + { + "name": "Opaque type region liveness captures outlived regions", + "statement": "Region liveness analysis for opaque types must include not just the regions that the opaque outlives, but also any regions that outlive the opaque type and are constrained by use<> bounds or captured in the opaque's definition.", + "subsystem": "borrow-checker", + "kind": "invariant", + "how_to_check": "Run borrowck with polonius next on code with opaque types capturing outer lifetimes. Verify that loan invalidation and region liveness correctly account for captured regions beyond just outlives bounds.", + "cost": "moderate", + "issues": [ + 153215 + ], + "currently_tested": "Partially: opaque type handling is tested but not systematically for region liveness with polonius mode enabled." + }, + { + "name": "Arc/Rc make_mut panic safety under alloc failure", + "statement": "Arc::make_mut and Rc::make_mut must not leave the Arc/Rc in an inconsistent state if allocation fails or handle_alloc_error unwinds. The strong/weak count must remain valid for further operations.", + "subsystem": "allocator", + "kind": "invariant", + "how_to_check": "Build with custom allocator that fails on demand. Call Arc::make_mut and Rc::make_mut with weak references held, trigger allocation failure, catch panic, and verify the Arc/Rc can still be safely dropped and counted.", + "cost": "moderate", + "issues": [ + 157203, + 155746 + ], + "currently_tested": "Partially: these issues were caught but require nightly and custom allocator hooks to reproduce; general test suite does not systematically test panic-safety of Arc::make_mut." + }, + { + "name": "Metadata determinism across clean rebuilds", + "statement": "Two clean builds of the same crate with identical sources and compiler state must produce byte-identical .rmeta files. Non-determinism in metadata indicates untracked state reads (HashMap iteration, environment variables, clocks, or random sources).", + "subsystem": "metadata", + "kind": "differential", + "how_to_check": "Build the same crate twice cleanly and compare .rmeta byte-for-byte. Use instrumented rustc to verify no untracked state is read during metadata encoding.", + "cost": "cheap", + "issues": [ + 157743, + 157747, + 158602 + ], + "currently_tested": "Partially: determinism is tested in existing tests but not as a systematic property check on all builds." + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "MIR type validity before codegen", + "statement": "Every value in MIR before codegen must have a concrete type that was assigned during type inference or explicit monomorphization. Type errors must be reported before MIR validation, not discovered during codegen.", + "subsystem": "codegen, mir-opt, const-eval", + "kind": "invariant", + "how_to_check": "Instrument rustc to track when types are assigned to MIR values vs. when MIR is validated. Check that any type mismatch is detected during type-checking or normalization, before reaching MIR validation. Build affected crates (e.g., regex_syntax with various flags) and verify no span_bug fires for type mismatches.", + "cost": "moderate", + "issues": [ + 158037, + 156409, + 154750, + 154748 + ], + "currently_tested": "No; MIR validation ICEs on bad types happen at validation time, not earlier. There's no systematic pre-validation type-checking." + }, + { + "name": "Parser recovery should not panic", + "statement": "Parser error recovery paths must emit diagnostics and stop gracefully, never reaching assertion failures or panic points. Any recovery path that encounters tokens it cannot handle should close the recovery scope cleanly.", + "subsystem": "parser", + "kind": "invariant", + "how_to_check": "Fuzz the parser with incomplete/malformed input targeting recovery paths (e.g., `&raw` without `const`/`mut`, trailing commas in recovery, nested delimiters). Verify all result in user errors, not ICE. Test specifically: `&raw x,)`, `&raw 2` in arrays, `&raw` in various nesting contexts.", + "cost": "cheap", + "issues": [ + 157950, + 157853, + 157888 + ], + "currently_tested": "Partially; parser has isolated tests but recovery is not systematically fuzzed end-to-end." + }, + { + "name": "Lint checks must consult trait solver when correctness matters", + "statement": "Lints that claim to check a semantic property (e.g., trait implementation) must query the trait solver, not use fast-path lookups that can miss alias types, projections, or blanket implementations.", + "subsystem": "lints, trait-system", + "kind": "invariant", + "how_to_check": "Test lints against code where a trait is implemented via an alias type, projection, or blanket impl. For `missing_debug_implementations`, check impl on `::Output` where `Identity` projects to the type. Verify lint does not fire falsely.", + "cost": "moderate", + "issues": [ + 157758, + 157757 + ], + "currently_tested": "No; lints skip trait solving in some cases as an optimization, leading to false positives." + }, + { + "name": "Generic argument counts must match trait generics in suggestions", + "statement": "When building a trait ref for diagnostics (e.g., method suggestions), the argument list must be compatible with the trait's generic parameters. Fresh inference variables must be provided for missing generics.", + "subsystem": "trait-system, diagnostics", + "kind": "invariant", + "how_to_check": "Call a method that doesn't exist on a type, where the matching trait (e.g., `Borrow`) has extra generics beyond `Self`. Verify suggestion builds trait ref correctly without ICE. Test: `.borrow()` on dyn closure type.", + "cost": "cheap", + "issues": [ + 157189 + ], + "currently_tested": "No; only caught by debug assertions in trait ref construction." + }, + { + "name": "Type const normalization must preserve type invariants", + "statement": "Normalizing a type const (in const-arg context) must not change the semantic type of the resulting value. If normalization changes the type, a `ConstArgHasType` mismatch error must be reported during normalization, before MIR validation.", + "subsystem": "const-eval, trait-system", + "kind": "invariant", + "how_to_check": "Compile code with type consts that normalize differently when used in const arguments vs. their definition site. Example: `foo::<{projection_const}>` where projection normalizes to a different type. Verify error is reported at normalization time, not at MIR validation.", + "cost": "moderate", + "issues": [ + 152962, + 154750, + 154748 + ], + "currently_tested": "Partially; `ConstArgHasType` exists but normalization ordering can skip it." + }, + { + "name": "Spans from synthesized code must not override source spans", + "statement": "When a proc macro generates code (e.g., impl blocks), the spans it assigns must not propagate backward to override the spans of the original item being defined. Intra-doc links and other source-based lookups must use the original source span, not a synthesized span.", + "subsystem": "macros, rustdoc", + "kind": "invariant", + "how_to_check": "Compile code with a `derive` proc macro that generates impl blocks referencing the item. Check that intra-doc links still resolve to the correct item's documentation page, not the generated impl. Test with `--generate-link-to-definition`.", + "cost": "moderate", + "issues": [ + 158050, + 157756 + ], + "currently_tested": "No; rustdoc span-map management is not systematically validated against proc-macro-generated code." + }, + { + "name": "Expression categorization must be consistent in THIR and MIR", + "statement": "If an expression is categorized as a `Place` in THIR, its lowering to MIR must follow the place-lowering path. If it's `Rvalue`, it must follow the rvalue path. No expression should be categorized differently, leading to unreachable code in place-lowering or rvalue-lowering.", + "subsystem": "mir-build, thir", + "kind": "invariant", + "how_to_check": "Build MIR for expressions with uncommon categorization (e.g., reborrow, as_place calls on certain types). Instrument expr_as_place to log all paths taken. Verify categorization matches actual lowering for reborrow and other special expressions. Test generic reborrow expressions.", + "cost": "moderate", + "issues": [ + 156482 + ], + "currently_tested": "No; THIR and MIR lowering are separate, and consistency is not monitored." + }, + { + "name": "Recovery diagnostics must not leave parsing state ambiguous", + "statement": "After a recovery path closes, the parser's token position and state must be well-defined. The next token must either be a close delimiter, EOF, or something the parser can interpret without ambiguity in the current context.", + "subsystem": "parser", + "kind": "invariant", + "how_to_check": "Test recovery scenarios where the recovery path stops before reaching a delimiter, then check that parsing resumes correctly. Example: `&raw x,)` recovery should consume `x` and the comma, leaving `)` as the next token, not causing parse_token_tree to be called on it.", + "cost": "cheap", + "issues": [ + 157950, + 157888 + ], + "currently_tested": "No; parser state after recovery is not systematically validated." + }, + { + "name": "Unsafe semantics must be enforced through feature gates", + "statement": "Features that require unsafe invariants (e.g., `#[pin_v2]` for structural pinning) must enforce those invariants wherever the feature is used. Implicit paths (e.g., match-ergonomics) and explicit paths (e.g., `&pin` patterns) must both check the same gate.", + "subsystem": "trait-system, mir-build", + "kind": "invariant", + "how_to_check": "Compile code using pin-ergonomics patterns on a type that does not opt into `#[pin_v2]`, in both explicit (`&pin mut`) and implicit reborrow-through-match forms. Verify both reject the code if `#[pin_v2]` is missing.", + "cost": "cheap", + "issues": [ + 157634, + 157011 + ], + "currently_tested": "Partially; `#[pin_v2]` checking exists but didn't cover explicit patterns initially." + }, + { + "name": "FFI function parameters with unusual ABIs must reject unsound patterns", + "statement": "Tail calls (`become`) via `extern \"rust-call\"` or other ABIs that don't support explicit stack manipulation must be rejected by the compiler, not allowed to compile and miscompile at runtime.", + "subsystem": "codegen, abi", + "kind": "invariant", + "how_to_check": "Compile code using `become` with a `rust-call` extern function and check that it's rejected. Test also with default ABI. Verify at -O0 and -O3, no miscompilation occurs (wrong values computed).", + "cost": "cheap", + "issues": [ + 158017 + ], + "currently_tested": "No; `become` support is incomplete and tail calls are not validated against ABIs before codegen." + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "rustdoc type-relative path body owner correctness", + "statement": "When rustdoc resolves type-dependent paths (method calls, type-relative paths) in non-body items nested in bodies, it must use the correct body owner, not an enclosing body's TypeckResults.", + "subsystem": "rustdoc", + "kind": "invariant", + "how_to_check": "Instrument rustdoc's SpanMapVisitor to track the active body context and verify that type-relative path resolution always uses the correct body. Assert that TypeckResults queries are never made for paths in a different body than they belong to.", + "cost": "moderate", + "issues": [ + 156418, + 156327, + 149089, + 150153, + 147882, + 147057 + ], + "currently_tested": "Covered by rustdoc tests but not as a systematic invariant check. Regression tests exist for individual cases but no property-level checking." + }, + { + "name": "attribute argument validation completeness", + "statement": "All attributes where arguments are not expected must validate and reject any provided arguments, emitting clear errors rather than silently ignoring them.", + "subsystem": "attribute parsing", + "kind": "invariant", + "how_to_check": "After attribute parsing, assert via a debug assertion that every meta-item in an attribute that claims to accept no arguments has had its arguments checked. Build a debug harness that parses attributes and verifies this invariant on all builtin attributes.", + "cost": "cheap", + "issues": [ + 154977 + ], + "currently_tested": "Individual attributes have their own tests but no systematic property-level checking. Some attributes silently ignored invalid args before the fix." + }, + { + "name": "transmute size check accounts for repr flags", + "statement": "Transmute size checks must account for all repr flags (align, pack, etc.) that affect type size, not just for bare struct/enum cases.", + "subsystem": "type system", + "kind": "invariant", + "how_to_check": "Build a test harness that generates transmute expressions between types with various repr flags and verifies that rejected transmutes are those with provably different sizes. Compare output against a reference implementation that correctly accounts for all repr flags.", + "cost": "moderate", + "issues": [ + 155412, + 88290 + ], + "currently_tested": "Some transmute tests exist but they do not comprehensively cover repr flag interactions. No systematic property-level checking." + }, + { + "name": "UTF-8 span boundaries never mid-sequence", + "statement": "All span calculations must result in spans whose boundaries never point into the middle of a UTF-8 character sequence.", + "subsystem": "span", + "kind": "invariant", + "how_to_check": "Instrument span construction to validate that byte positions always align with UTF-8 character boundaries. Test with source files containing multibyte Unicode characters at span boundaries and verify no out-of-bounds assertions occur.", + "cost": "cheap", + "issues": [ + 156316, + 155037 + ], + "currently_tested": "Partially covered by parser error recovery tests but not as a systematic invariant. No comprehensive check for all span construction paths." + }, + { + "name": "rustdoc cfg propagation to all reexports", + "statement": "rustdoc must propagate cfg attributes consistently to all reexported items, including those from glob reexports, with the same visibility and formatting as explicitly named reexports.", + "subsystem": "rustdoc", + "kind": "invariant", + "how_to_check": "Generate test cases with various reexport patterns (glob, explicit, inline) and verify that cfg badges appear consistently. Compare rendered HTML for equivalent item references via different reexport paths.", + "cost": "moderate", + "issues": [ + 96166, + 154921 + ], + "currently_tested": "Covered by rustdoc rendering tests but not as a systematic property. Regression tests added after fixes but no property-level checking." + }, + { + "name": "vtable construction handles unsatisfied supertraits", + "statement": "VTable iteration and construction must not panic when a supertrait of a trait object is not implemented; it should emit a proper error instead.", + "subsystem": "trait system", + "kind": "invariant", + "how_to_check": "Construct trait objects with impossible trait bounds (supertrait constraints that cannot be satisfied) and verify that compilation fails with a user-visible error, not an ICE.", + "cost": "moderate", + "issues": [ + 137190, + 135470 + ], + "currently_tested": "Regression tests exist but not as a property-level invariant. The fix checked for this specific case but no broader invariant checking." + }, + { + "name": "transmute validation before const evaluation", + "statement": "Transmute expressions with mismatched sizes must be validated and rejected during type-checking, before const evaluation runs, to prevent ICEs in the interpreter.", + "subsystem": "const eval", + "kind": "invariant", + "how_to_check": "Compile transmute expressions with size mismatches under -Z mir-enable-passes and other const eval triggers. Verify that a proper diagnostic is emitted, not an interpreter ICE.", + "cost": "moderate", + "issues": [ + 149920 + ], + "currently_tested": "No systematic check. Regression tests exist but the property is not checked during normal compilation." + }, + { + "name": "array map/try_map drops all elements", + "statement": "Array operations like map and try_map must drop all unmapped elements when the closure panics or early-returns, including ZSTs.", + "subsystem": "library", + "kind": "semantic", + "how_to_check": "Build arrays of types with side-effectful Drop implementations, call map/try_map with closures that panic or return early, and verify drop is called the correct number of times using a drop counter.", + "cost": "cheap", + "issues": [ + 152211 + ], + "currently_tested": "Covered by library tests but not as a property-level invariant across all array methods. The fix uses a strategy from slice::IterMut that should apply broadly." + }, + { + "name": "Rc/Arc pointer comparison for all sized types", + "statement": "Rc and Arc equality should use pointer comparison as an optimization for all types, including unsized types (DSTs), not just sized types.", + "subsystem": "library", + "kind": "invariant", + "how_to_check": "Create Rc and Arc instances pointing to the same allocation (via Rc::clone and Arc::clone) with unsized inner types (str, dyn Trait, etc.), and verify pointer comparison is used rather than deep equality.", + "cost": "cheap", + "issues": [ + 154998 + ], + "currently_tested": "Library tests exist but do not systematically cover DST cases. No property-level checking for pointer comparison optimization." + }, + { + "name": "enum variant path generic args only on enum segment", + "statement": "Generic arguments in enum variant paths (e.g., `Enum::::Variant`) should only be accepted on the enum segment; they must not be accepted on module or other non-enum path segments even if they reexport the variant.", + "subsystem": "HIR type lowering", + "kind": "invariant", + "how_to_check": "Construct variant paths through module reexports and attempt to apply generic arguments to the module segment. Verify that a clear error is emitted, not silently accepted.", + "cost": "cheap", + "issues": [ + 154962 + ], + "currently_tested": "Regression test added but not as a systematic property. The fix validates the penultimate segment's def for enum-ness when args are present." + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "HIR-for-queried-items", + "statement": "Every DefId that a query accesses must have fully lowered HIR available, even if code containing that DefId had errors during lowering.", + "subsystem": "metadata", + "kind": "invariant", + "how_to_check": "Before any query that accesses a DefId (layout, discriminant, generics_of, etc.), verify the item's HIR was lowered. If a query on DefId D succeeds, then HIR lookup for D must not panic. Instrument with mirth to record all DefId queries and HIR lookups; cross-check that no query accesses unmapped DefIds.", + "cost": "moderate", + "issues": [ + 154820, + 154539, + 123629, + 153744, + 153743 + ], + "currently_tested": "The compiler has assertions on some query boundaries but not comprehensive coverage. rustc_hir::map lookups panic on missing entries, but these panics aren't caught before they occur. No continuous check across all queries." + }, + { + "name": "TyKind-Error-error-emitted", + "statement": "Whenever TyKind::Error is constructed, an error diagnostic must have been emitted before that point.", + "subsystem": "type-system", + "kind": "invariant", + "how_to_check": "Instrument ty::Error construction to require a preceding error emission. Use mirth to track all TyKind::Error creations and cross-reference with the diagnostic buffer; for each Error, verify at least one error was emitted in an earlier pass or same pass before the type was created.", + "cost": "cheap", + "issues": [ + 154780 + ], + "currently_tested": "Partial: rustc has a span_bug! that fires when TyKind::Error appears without an error, but only in debug builds and only when that specific condition is checked. No systematic instrumentation of all Error creations." + }, + { + "name": "Layout-wrapper-consistency", + "statement": "Layout computation on wrapper types (like UnsafeBinder) must compute the layout of the inner type, not of the wrapper itself.", + "subsystem": "codegen", + "kind": "invariant", + "how_to_check": "For types with transparent wrappers, compare the reported layout of the wrapper against the layout of the inner type: they must match. Build with mirth to instrument layout queries and verify layout(Wrapper) == layout(T) where Wrapper is defined as transparent.", + "cost": "cheap", + "issues": [ + 154426, + 154424 + ], + "currently_tested": "Debug-only assertions exist in layout.rs, but these don't fire until the layout is actually used. No ahead-of-time check that layout rules are consistent." + }, + { + "name": "Diagnostic-arg-uniqueness", + "statement": "Diagnostic arguments must not be added with duplicate names in a single diagnostic.", + "subsystem": "diagnostics", + "kind": "invariant", + "how_to_check": "In the diagnostic builder, before emitting, scan all added arguments for duplicate names. On duplicate, use mirth to record the call stack and issue that produced it. Simple check: before emit(), call set of arg names must equal total count.", + "cost": "cheap", + "issues": [ + 152936 + ], + "currently_tested": "Yes, partially. rustc_errors has a debug assertion that fires on duplicate arg names, but only in debug builds and only on emit(). Should be checked earlier in the build phase." + }, + { + "name": "Dyn-incompatible-no-trait-objects", + "statement": "Traits marked with #[rustc_dyn_incompatible_trait] cannot be used as trait object bounds.", + "subsystem": "trait-system", + "kind": "semantic", + "how_to_check": "When a trait object is created (via dyn T or type Assoc = dyn Trait), check that the trait is dyn-compatible. Use mirth to instrument dyn-compatibility checks and verify that dyn_incompatible traits never appear in object bounds in final metadata or code.", + "cost": "cheap", + "issues": [ + 154619 + ], + "currently_tested": "Partially. Dyn-compatibility is checked in rustc_hir_analysis, but only for direct dyn usage, not for subtraits. The unsoundness came from creating trait objects of subtraits that weren't checked." + }, + { + "name": "Trait-object-bounds-consistency", + "statement": "When solving trait predicates involving trait objects, the solver must use the trait object's declared bounds (from the dyn syntax), not bounds from the proof goal context.", + "subsystem": "trait-system", + "kind": "differential", + "how_to_check": "Compile code with both old and next solvers; trait object predicate solutions must agree. For any dyn T that implements goal G, verify the solution comes from bounds declared on T (or T's supertrait bounds), not from unrelated goals. Build a test suite of trait object code and run with two different solver backends.", + "cost": "moderate", + "issues": [ + 152789, + 151329 + ], + "currently_tested": "No. The trait system has no automated cross-solver verification. The bug was caught only because code panicked when the solver couldn't prove something it thought it should." + }, + { + "name": "Type-wrapper-query-forwarding", + "statement": "Query results for wrapper types (e.g., unsafe binders, type aliases) must forward to inner/aliased types for predicates like discriminant_for_variant() and layout operations.", + "subsystem": "type-system", + "kind": "invariant", + "how_to_check": "For each wrapper/alias type, call discriminant/layout queries on both wrapper and inner. Results must agree (for transparent wrappers) or the wrapper must have its own explicit implementation. Use mirth to instrument query calls and verify that wrapper queries don't panic when the inner type query would succeed.", + "cost": "moderate", + "issues": [ + 154424 + ], + "currently_tested": "Partially via debug assertions in specific query implementations, but no comprehensive check across all query types." + }, + { + "name": "No-untracked-state-in-encoded-metadata", + "statement": "Metadata encoding must not depend on untracked state: environment variables, system clock, HashMap iteration order, or randomness seeded outside the compiler's control.", + "subsystem": "metadata", + "kind": "invariant", + "how_to_check": "Instrument metadata encoding with mirth to log all external state reads (env vars, clock, RNG). Compare two metadata encodes in different processes with different clocks and env; .rmeta bytes must match if inputs are identical. Use a determinism test harness.", + "cost": "expensive", + "issues": [], + "currently_tested": "Partially. The compiler has some checks (P4 mentioned in context), but these are not comprehensive. The batch doesn't include specific P4 violations, but it's a known property from earlier testing." + }, + { + "name": "Incremental-matches-clean-rebuild", + "statement": "An incremental rebuild after a source edit, plus rebuild of the original, must produce byte-identical .rmeta and intermediate metadata files.", + "subsystem": "incremental", + "kind": "differential", + "how_to_check": "Given a crate, perform three builds: (1) clean, (2) clean again. Then (3) edit a non-semantic line, (4) incremental rebuild, (5) revert edit, (6) incremental rebuild. Compare byte-wise: rmeta from step 6 must equal rmeta from step 1. Use mirth to capture dependency graphs and metadata writes to verify they match.", + "cost": "expensive", + "issues": [], + "currently_tested": "Partially. The P6 property mentioned in context exists as a property idea but not as a continuous check. The batch doesn't contain specific P6 violations detected here." + }, + { + "name": "Deterministic-metadata-across-builds", + "statement": "Two independent clean builds of the same crate with identical inputs (source, deps, target) must produce byte-identical .rmeta files.", + "subsystem": "metadata", + "kind": "invariant", + "how_to_check": "Build the same crate twice in separate processes with identical inputs; compare .rmeta files byte-wise. Use mirth to instrument metadata writing and verify the sequence of writes is identical. Test on crates of various sizes and complexity.", + "cost": "moderate", + "issues": [], + "currently_tested": "Partially. This is mentioned as P5 in context and is tested in some cases, but not as a continuous property check. No regression seen in this batch." + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "No unwrap/expect panics on reachable code paths", + "statement": "Critical compiler algorithms should not panic via unwrap/expect/assert on inputs that are syntactically and semantically valid according to Rust's rules.", + "subsystem": "MIR optimization, trait resolution, metadata, codegen, incremental", + "kind": "invariant", + "how_to_check": "Instrument rustc with mirth to record all unwrap/expect/assert calls. Compile a large test suite (e.g., cargo/clippy/miri) and flag any panics that occurred on code that passed the parser. Verify the panic happened on a path that should not be reachable for valid code.", + "cost": "expensive", + "issues": [ + 153601, + 153599, + 153539, + 153525, + 153524, + 153502, + 153499, + 153433, + 153420, + 153391, + 153390, + 153389, + 153388, + 153354, + 153351, + 153236, + 153199, + 153198, + 152895, + 152797, + 152744, + 152684, + 152683, + 152682, + 152663, + 152654, + 152653, + 152633, + 152606, + 152601, + 152595, + 152545, + 152518, + 151631, + 151477, + 150457, + 149695, + 149643, + 146754, + 143498, + 128801, + 127971, + 127423, + 125805, + 113870 + ], + "currently_tested": "No. The rustc test suite catches some panics but many ICEs only appear on specific crate combinations or with specific compiler flags (-Zthreads=N, -Znext-solver, etc.). A comprehensive instrumentation would catch many more." + }, + { + "name": "DefId classification before local-only operations", + "statement": "DefIds must be correctly classified as local vs non-local before operations like expect_local() are called. A non-local DefId (from external crates) should never reach a code path that asserts it is local.", + "subsystem": "resolve, stability, delegation, trait-system", + "kind": "invariant", + "how_to_check": "Instrument the compiler to track DefId origins (local vs non-local). Before any call to expect_local or similar local-only operations, verify the DefId is actually local. Flag any violations where non-local DefIds reach these operations. Test on code using features that reference external crates (fn_delegation with std items, staged_api with dependency crates).", + "cost": "moderate", + "issues": [ + 153599, + 153502, + 153433, + 153499, + 143498, + 152633 + ], + "currently_tested": "Partially. The stability and delegation systems test with standard library items, but not systematically across all code paths." + }, + { + "name": "Inference variables resolved before hashing/caching", + "statement": "Inference variables (type variables ?0t, const variables ?0c) must never be included in values that are hashed or entered into the query cache. All inference must be resolved before the value is cached.", + "subsystem": "type inference, const generics, query system", + "kind": "invariant", + "how_to_check": "Instrument const/type lowering and query caching to detect when inference variables appear in values being hashed. Compile test cases with const generics featuring min_generic_const_args and generic_const_parameter_types. Any hash/cache operation on a value containing ?Nt or ?Nc should be flagged.", + "cost": "cheap", + "issues": [ + 153525, + 153524 + ], + "currently_tested": "No. The incremental compilation testing doesn't specifically check for inference variable leakage into the cache." + }, + { + "name": "Symbol availability after optimization", + "statement": "All symbols that the linker expects to find (because they were in the object files or referenced in metadata) must remain present after optimization passes, even with aggressive LTO.", + "subsystem": "codegen, LTO, symbol export", + "kind": "differential", + "how_to_check": "Compile crates with different optimization levels and LTO settings (opt-level 0/1/3, lto off/thin/fat, codegen-units 1/16). Verify all symbols referenced in metadata are present in final artifacts. Use nm/llvm-nm to check symbol tables. Flag symbols that disappear with higher optimization or LTO.", + "cost": "moderate", + "issues": [ + 153645, + 153451 + ], + "currently_tested": "Partially. The testsuite checks that EII code works, but doesn't systematically compare symbol presence across LTO settings." + }, + { + "name": "Recursion depth limits on deeply nested types", + "statement": "Recursive compiler algorithms that traverse type structures must have explicit recursion depth limits to prevent stack overflow on valid code with arbitrarily deep nesting.", + "subsystem": "trait resolution, type normalization", + "kind": "invariant", + "how_to_check": "Construct deeply nested type expressions (e.g., nested type aliases, complex trait bounds with projections). Try bounds like `T: for<'a> Proj<'a, Assoc = for<'b> fn(>::Assoc)>`. Monitor stack depth during compilation. Flag any algorithm that recurses proportionally to type nesting without a explicit depth check.", + "cost": "cheap", + "issues": [ + 152716 + ], + "currently_tested": "No. There are only a few existing tests for stack overflow on nested types." + }, + { + "name": "Deterministic compilation with parallel threads", + "statement": "Parallel and single-threaded compilation of the same code must produce identical outputs. No race conditions or non-deterministic state reads.", + "subsystem": "parallel front end, incremental", + "kind": "differential", + "how_to_check": "Compile a test suite with -Zthreads=1 and -Zthreads=8. Compare the resulting .rmetadata, .o files, and binaries byte-for-byte. Flag any differences. Also instrument for HashMap iteration order and timestamp reads during compilation.", + "cost": "expensive", + "issues": [ + 153391 + ], + "currently_tested": "Partially. The incremental test suite runs with threads but doesn't always compare against single-threaded output." + }, + { + "name": "Pre-expansion feature gating for new syntax", + "statement": "New syntactic forms must be gated with pre-expansion gating (checked before macro expansion) rather than post-expansion, to prevent forward compatibility issues with #[cfg(false)] code.", + "subsystem": "parser, feature gates", + "kind": "invariant", + "how_to_check": "For each new syntax feature, verify it is gated with check_crate_and_resolve_attrs or equivalent pre-expansion checks. Test that syntax is rejected in #[cfg(false)] blocks and in macros without the feature. Parse the AST before macro expansion to ensure gating happens early.", + "cost": "cheap", + "issues": [ + 152501, + 152499 + ], + "currently_tested": "Partially. The feature gate tests check syntax is gated, but don't systematically verify pre-expansion vs post-expansion timing for all features." + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "Deterministic const evaluation", + "statement": "Const evaluation should produce byte-identical results across multiple builds of the same crate, with no dependence on build order, hashmap iteration order, or other untracked state.", + "subsystem": "const-eval", + "kind": "differential", + "how_to_check": "Run `rustc` twice on the same crate with the same inputs and compare the encoded const values in rmeta. Compute hashes of const evaluation results during monomorphization. No difference should exist between runs.", + "cost": "cheap", + "issues": [ + 151537, + 142152, + 141540, + 141313, + 138910, + 138089, + 133966, + 151625, + 150983 + ], + "currently_tested": "const eval has test suite but not differential builds; incremental testing exists but doesn't specifically check const determinism across clean builds" + }, + { + "name": "Complete trait solver on all goals", + "statement": "The next trait solver should not panic or ICE when evaluating valid trait goals, including goals involving stalled coroutine obligations and type alias impl trait.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "Compile crates with -Znext-solver=globally enabled and instrument the solver to track all goals attempted. Check that no panic/ICE occurs; goals should either resolve or emit proper diagnostic errors, never internal crashes.", + "cost": "moderate", + "issues": [ + 151957, + 151323, + 151322, + 151318, + 138274, + 137916, + 151462 + ], + "currently_tested": "next-solver has limited test coverage; stalled obligation tests exist but not comprehensive across all trait patterns" + }, + { + "name": "Type checking handles all DefKinds", + "statement": "Every DefKind that appears in a crate should be correctly handled in type checking, const eval, and code generation without calling invalid query functions or panicking.", + "subsystem": "type-system", + "kind": "invariant", + "how_to_check": "Instrument type checking and const eval phases to log which DefKinds are encountered. Verify that no internal API contracts are violated (e.g., codegen_fn_attrs called only on items that have codegen attrs, generics_of called on valid parents). Compile diverse crates and check that all DefKinds are handled.", + "cost": "moderate", + "issues": [ + 152340, + 151878, + 151708, + 151331, + 151273, + 151027, + 150960, + 148953 + ], + "currently_tested": "type checking tests cover common cases but not comprehensive DefKind coverage; const items are newer so less tested" + }, + { + "name": "Graceful layout computation error handling", + "statement": "Layout computation should never panic on invalid layouts (overflow, unsized, etc); it should return error types or emit diagnostic errors instead.", + "subsystem": "layout", + "kind": "invariant", + "how_to_check": "Compile crates with intentionally invalid layouts (oversized arrays, unsized in repr(C), etc) and verify that the compiler either reports an error diagnostic or gracefully handles the layout without panicking. Instrument layout computation to track panics.", + "cost": "cheap", + "issues": [ + 152030, + 151591, + 151814, + 137582, + 146984 + ], + "currently_tested": "layout tests exist but don't comprehensively cover edge cases that cause panics; no CI test for panic-free error handling" + }, + { + "name": "Consistent trait alias resolution", + "statement": "Trait aliases should be expanded identically in all compilation phases (resolve_bound_vars, HIR lowering, trait checking). An unexpanded trait alias in one phase should not cause an ICE in another.", + "subsystem": "trait-system", + "kind": "invariant", + "how_to_check": "Compile crates using trait aliases (especially with return type notation and complex bounds) and instrument trait resolution to verify that trait aliases are expanded at all sites where they are encountered. Check that phase-specific resolution doesn't skip alias expansion.", + "cost": "moderate", + "issues": [ + 152244, + 152158, + 152309 + ], + "currently_tested": "trait alias tests exist but limited coverage for cross-phase consistency; no test for RTN with trait aliases" + }, + { + "name": "Normalization completeness in borrowck", + "statement": "When normalizing projections in borrowck, both type AND const projections must be fully normalized to completion, not partially left as unevaluated consts.", + "subsystem": "trait-system", + "kind": "differential", + "how_to_check": "Compile crates that use generic const items with associated const references in borrowck context. Verify that when borrowck normalizes, const projections are fully evaluated. Compare incremental vs clean builds to ensure both produce identical normalized forms.", + "cost": "moderate", + "issues": [ + 151647, + 151579, + 152278, + 120811 + ], + "currently_tested": "normalization tests exist but don't cover const projection normalization; incremental vs clean comparison not done for this property" + }, + { + "name": "Diagnostic span validity", + "statement": "All diagnostic spans must be non-empty (unless explicitly allowed) and non-overlapping. Spans must originate from files that exist in the source tree.", + "subsystem": "diagnostics", + "kind": "invariant", + "how_to_check": "Instrument all diagnostic creation to validate that spans satisfy invariants: non-empty, disjoint, valid file references. Compile large crates and collect all emitted diagnostics, verifying span properties. Check that no panic occurs from span construction.", + "cost": "cheap", + "issues": [ + 151610, + 151607, + 147339, + 131292 + ], + "currently_tested": "diagnostic tests exist but don't validate span invariants comprehensively; no instrumentation to catch invalid spans" + }, + { + "name": "Deterministic ABI and layout computation", + "statement": "ABI and layout computation should produce identical results across builds, with no dependence on floating-point precision details, hashmap iteration order, or platform-specific undefined behavior.", + "subsystem": "codegen", + "kind": "differential", + "how_to_check": "Compile the same crate twice and compare the computed ABIs and layouts for all types. Use structured instrumentation to track ABI computation for all function signatures and struct layouts. No differences should exist between runs.", + "cost": "moderate", + "issues": [ + 151791, + 122620, + 115399, + 147883 + ], + "currently_tested": "ABI tests exist but don't compare across full builds; determinism not checked; SPARC ABI issues show it was never deterministic" + }, + { + "name": "Metadata encoder state consistency", + "statement": "Each dependency node should be encoded exactly once into the metadata. Attempting to encode the same node twice indicates a compiler bug.", + "subsystem": "incremental", + "kind": "invariant", + "how_to_check": "Instrument the metadata encoder to track which dep nodes have been written. Before writing each node, assert that it hasn't already been encoded. Run builds and verify no assertions fire. Compare with clean builds to ensure no spurious duplicate-write assertions.", + "cost": "cheap", + "issues": [ + 150018, + 142778 + ], + "currently_tested": "incremental tests exist but don't instrument encoder for duplicate-write detection; query system tests don't validate this invariant" + }, + { + "name": "Lint analysis consistency across features", + "statement": "Lint detection should be consistent: if a lint triggers on code without experimental features, it should trigger identically on the same code with those features enabled (or explicitly exclude them). Lint expectations should not be lost during macro expansion.", + "subsystem": "lints", + "kind": "differential", + "how_to_check": "Compile the same crate with and without various feature flags (e.g., min_generic_const_args). Collect lint violations and verify they match. Check that #[expect] attributes survive derive and cfg_attr expansion without creating spurious unfulfilled-lint-expectations.", + "cost": "moderate", + "issues": [ + 152004, + 151983, + 152401, + 152289 + ], + "currently_tested": "lint tests exist but don't compare across feature flags; attribute survival through macros not tested systematically" + } + ] + }, + { + "issues_read": 100, + "candidates": [ + { + "name": "Glob import processing does not panic", + "statement": "The glob import resolution and overwriting process must not panic or fail assertions on any valid Rust code, including code with macro-expanded extern crates and complex glob hierarchies.", + "subsystem": "resolve", + "kind": "invariant", + "how_to_check": "Compile a set of crates with complex glob import patterns (macro-expanded items, nested glob imports, glob overwriting) using a rustc instrumented to catch panics and assertions; verify zero panics. Use existing test crates like polars-plan, chumsky, diesel that triggered original issues.", + "cost": "moderate", + "issues": [ + 149821, + 150927, + 150928, + 150976, + 151008, + 151037, + 151213 + ], + "currently_tested": "Partially through UI tests but not systematically as a panic-checking property" + }, + { + "name": "Const generic bounds checking", + "statement": "Generic argument processing must perform bounds checks before array indexing to prevent panics when processing const generics with function pointers, arrays, or tuples.", + "subsystem": "const-eval", + "kind": "invariant", + "how_to_check": "Compile Rust code using const generic features (min_generic_const_args) with function pointers, arrays, and tuples as arguments; instrument generic argument processing with bounds assertions; verify no index-out-of-bounds panics occur.", + "cost": "cheap", + "issues": [ + 137084, + 150673, + 150714, + 150841, + 151024, + 151186 + ], + "currently_tested": "Some through const-generics UI tests but not as systematic invariant" + }, + { + "name": "Valtree type consistency", + "statement": "Const value tree (valtree) construction and pretty-printing must validate that the value's type matches the expected type before encoding or printing, preventing mismatches like 'expected leaf got branch'.", + "subsystem": "const-eval", + "kind": "invariant", + "how_to_check": "Compile Rust code with MGCA (min_generic_const_args) features and complex const generics; instrument valtree pretty-printing code to add type assertions before format operations; verify assertions pass on all code.", + "cost": "cheap", + "issues": [ + 150506, + 150712, + 150734, + 151126 + ], + "currently_tested": "Limited; mostly found through ICE regression tests" + }, + { + "name": "Compiler operations do not panic on input errors", + "statement": "The compiler should emit proper errors instead of panics when encountering type mismatches in const evaluation, invalid const generics, or malformed patterns. Type errors should be converted to errors immediately rather than delayed bugs.", + "subsystem": "hir-analysis", + "kind": "semantic", + "how_to_check": "Compile invalid Rust code that triggers type mismatches in const contexts; verify compiler emits E-series errors (not panics or delayed bugs) that can be caught by CI systems and reported to users.", + "cost": "cheap", + "issues": [ + 151024, + 150841, + 150506, + 150657, + 149746 + ], + "currently_tested": "Yes, through error codegen, but not systematically checked that all ICE-prone paths error instead" + }, + { + "name": "Incremental builds match clean builds", + "statement": "Incremental builds after edits must produce the same compiled metadata as clean builds of the edited source, particularly for const evaluation and const block promotion.", + "subsystem": "incremental", + "kind": "differential", + "how_to_check": "For const-heavy crates, build with -Cincr=full and with clean, comparing the produced query results and metadata checksums; flag any differences.", + "cost": "moderate", + "issues": [ + 150464 + ], + "currently_tested": "Yes through incremental test suite but not focused on const blocks" + }, + { + "name": "Two clean builds produce identical output", + "statement": "Two consecutive clean builds of the same Rust code must produce byte-identical .rmeta files, indicating deterministic compilation without HashMap iteration order or other non-deterministic behavior.", + "subsystem": "metadata", + "kind": "differential", + "how_to_check": "Build each crate twice cleanly and compare .rmeta file contents; additionally instrument to detect HashMap iteration, environment variable reads, or clock reads during metadata encoding.", + "cost": "moderate", + "issues": [ + 89911, + 150409, + 150419 + ], + "currently_tested": "P5 property mentioned in memory; some tooling exists" + } + ] + } + ] +} \ No newline at end of file diff --git a/docs/survey/fetch.py b/docs/survey/fetch.py new file mode 100644 index 0000000..93a1dec --- /dev/null +++ b/docs/survey/fetch.py @@ -0,0 +1,37 @@ +import json, subprocess, sys +Q = '''query($q:String!, $after:String) { search(query:$q, type:ISSUE, first:50, after:$after) { + issueCount pageInfo { hasNextPage endCursor } + nodes { ... on Issue { number title closedAt body + labels(first:15) { nodes { name } } + timelineItems(itemTypes:[CLOSED_EVENT], last:1) { nodes { ... on ClosedEvent { + closer { ... on PullRequest { number title body merged files(first:40) { nodes { path } } } + ... on Commit { associatedPullRequests(first:1) { nodes { number title body merged files(first:40) { nodes { path } } } } } } } } } } } } }''' +import datetime +out = [] +end = datetime.date(2026, 10, 7) +while len(out) < 1000 and end.year > 2020: + start = end - datetime.timedelta(days=30) + q = f"repo:rust-lang/rust is:issue is:closed reason:completed label:C-bug closed:{start}..{end}" + after = None + while True: + args = ['gh', 'api', 'graphql', '-f', f'query={Q}', '-f', f'q={q}'] + if after: args += ['-f', f'after={after}'] + r = json.loads(subprocess.run(args, capture_output=True, text=True, env={**__import__('os').environ, 'GH_HOST': 'github.com'}).stdout) + s = r['data']['search'] + for n in s['nodes']: + closers = [t['closer'] for t in n['timelineItems']['nodes'] if t.get('closer')] + pr = closers[0] if closers else None + if pr and 'associatedPullRequests' in pr: + prs = pr['associatedPullRequests']['nodes'] + pr = prs[0] if prs else None + if not pr or not pr.get('merged'): continue + out.append({'issue': n['number'], 'title': n['title'], 'closed': n['closedAt'][:10], + 'labels': [l['name'] for l in n['labels']['nodes']], + 'body': (n['body'] or '')[:2500], + 'pr': pr['number'], 'pr_title': pr['title'], 'pr_body': (pr['body'] or '')[:2000], + 'pr_files': [f['path'] for f in pr['files']['nodes']]}) + print(len(out), s['issueCount'], file=sys.stderr) + if not s['pageInfo']['hasNextPage']: break + after = s['pageInfo']['endCursor'] + end = start - datetime.timedelta(days=1) +json.dump(out[:1000], open('issues.json', 'w')) diff --git a/docs/survey/unroll.py b/docs/survey/unroll.py new file mode 100644 index 0000000..11efa84 --- /dev/null +++ b/docs/survey/unroll.py @@ -0,0 +1,27 @@ +import json, subprocess, os, re, time +env = {**os.environ, 'GH_HOST': 'github.com'} +d = json.load(open('issues.json')) +Q = '''query($n:Int!) { repository(owner:"rust-lang", name:"rust") { pullRequest(number:$n) { body } } }''' +Q2 = '''query($n:Int!) { repository(owner:"rust-lang", name:"rust") { pullRequest(number:$n) { number title body merged files(first:40) { nodes { path } } } } }''' +def gql(q, n): + r = subprocess.run(['gh', 'api', 'graphql', '-f', f'query={q}', '-F', f'n={n}'], capture_output=True, text=True, env=env) + return json.loads(r.stdout)['data']['repository']['pullRequest'] +cache = {} +fixed = 0 +for x in d: + if not x['pr_title'].startswith('Rollup'): continue + if x['pr'] not in cache: + body = gql(Q, x['pr'])['body'] or '' + cache[x['pr']] = sorted(set(int(m) for m in re.findall(r'#(\d{5,6})', body))) + for n in cache[x['pr']]: + key = ('sub', n) + if key not in cache: + cache[key] = gql(Q2, n) + p = cache[key] + if p and re.search(rf'#{x["issue"]}\b|issues/{x["issue"]}\b', p['body'] or ''): + x.update(pr=p['number'], pr_title=p['title'], pr_body=(p['body'] or '')[:2000], + pr_files=[f['path'] for f in p['files']['nodes']]) + fixed += 1 + break +json.dump(d, open('issues.json', 'w')) +print('unrolled', fixed, 'still rollup', sum(x['pr_title'].startswith('Rollup') for x in d)) diff --git a/docs/survey/verified.json b/docs/survey/verified.json new file mode 100644 index 0000000..9ea2b92 --- /dev/null +++ b/docs/survey/verified.json @@ -0,0 +1,1043 @@ +[ + { + "i": 0, + "name": "P5+ metadata determinism under irrelevant perturbation", + "citations": [ + { + "n": 159677, + "supports": true, + "reason": "Issue explicitly demonstrates .rmeta non-determinism: the same crate produces different .rmeta when unrelated `libfoo_extra` is added to the library search path. This directly violates the invariant that .rmeta should be byte-identical regardless of unrelated crates in -L." + }, + { + "n": 157743, + "supports": false, + "reason": "Performance regression in generated code for MaybeUninit initialization\u2014unrelated to .rmeta determinism; about optimization behavior, not output consistency." + }, + { + "n": 157747, + "supports": false, + "reason": "Flaky rustdoc GUI test due to DOM timing; test infrastructure issue unrelated to .rmeta or build output determinism." + }, + { + "n": 158602, + "supports": false, + "reason": "Soundness issue in test harness environment clearing; about test harness safety, not about .rmeta output determinism." + }, + { + "n": 89911, + "supports": false, + "reason": "Reproducible build test failure with debuginfo=2 on Linux. While titled 'reproducible build,' the issue text does not explicitly demonstrate that .rmeta outputs differed between builds; the failure mechanics are not documented as a metadata determinism violation." + }, + { + "n": 150409, + "supports": false, + "reason": "ICE (internal compiler error) in const hashing; a crash during compilation, not a demonstration of .rmeta output non-determinism." + }, + { + "n": 150419, + "supports": false, + "reason": "LLVM instruction selection error for AArch64 SVE debuginfo spilling; compilation error unrelated to .rmeta determinism." + }, + { + "n": 106571, + "supports": false, + "reason": "Duplicate messages in diagnostic JSON output; about diagnostic formatting and deduplication, not .rmeta or build output determinism." + }, + { + "n": 151537, + "supports": false, + "reason": "ICE in const evaluation of SIMD types; a compiler crash unrelated to .rmeta output determinism." + }, + { + "n": 142152, + "supports": false, + "reason": "ICE in query system dep graph under concurrent access; a crash with race condition but the issue text does not demonstrate that .rmeta outputs were inconsistent between builds." + }, + { + "n": 141540, + "supports": false, + "reason": "Duplicate dep graph nodes under concurrent coloring; while this race could affect output, the issue text reports a compilation error/panic, not a demonstrated difference in .rmeta outputs." + }, + { + "n": 141313, + "supports": false, + "reason": "GVN optimization creating overlapping assignments leading to UB; about MIR optimization correctness, not .rmeta output determinism." + }, + { + "n": 138910, + "supports": false, + "reason": "Linker-plugin LTO documentation issue; about documentation completeness, not a violation of .rmeta determinism." + }, + { + "n": 138089, + "supports": false, + "reason": "ICE in inherent associated type normalization; a compiler crash unrelated to .rmeta determinism." + }, + { + "n": 133966, + "supports": false, + "reason": "ICE in const evaluation of wide pointer handling; a compiler crash unrelated to .rmeta output determinism." + }, + { + "n": 151625, + "supports": false, + "reason": "ICE from invalid const literal validation; a compiler crash unrelated to .rmeta determinism." + }, + { + "n": 150983, + "supports": false, + "reason": "ICE from generic const item type validation; a compiler crash unrelated to .rmeta determinism." + } + ], + "general": true, + "sharper_statement": ".rmeta contents must not depend on unrelated rlib files present in the library search path (-L), i.e., two builds with identical source, flags, and target produce byte-identical .rmeta even when the set of available (but unused) libraries in the search path differs.", + "comment": "Only citation #159677 provides concrete evidence of the property being violated. It demonstrates non-determinism in .rmeta production caused by the mere presence of an unrelated library in the search path. The original property statement claims coverage of five sources of irrelevant perturbation (HashMap seed, search path, build directory, environment variables, wall clock), but only the search path violation is actually supported by the cited issues. Most other citations are unrelated bugs: ICEs, soundness issues, test flakiness, optimization regressions, and diagnostic issues. This property is mechanically checkable by comparing .rmeta hashes across builds that differ only in search path contents." + }, + { + "i": 1, + "name": "P6+ incremental equals clean (outputs and diagnostics)", + "citations": [ + { + "n": 162407, + "supports": false, + "reason": "This is a build system invalidation issue where bootstrap fails to re-check, not about incremental vs clean outputs being different" + }, + { + "n": 162901, + "supports": true, + "reason": "Explicitly shows incremental produces 4 duplicated diagnostics while clean produces 1 - direct violation of property" + }, + { + "n": 162585, + "supports": false, + "reason": "This is a compiler crash/ICE issue, not about different outputs from incremental vs clean builds" + }, + { + "n": 125564, + "supports": false, + "reason": "This is a compiler ICE issue, not about incremental vs clean outputs being different" + }, + { + "n": 135062, + "supports": false, + "reason": "This is about incorrect type checking behavior; the issue does not explicitly compare incremental vs clean diagnostics or outputs" + }, + { + "n": 159677, + "supports": true, + "reason": ".rmeta contents differ based on presence of unrelated files in library search path - direct violation of property that .rmeta bytes should match" + }, + { + "n": 150464, + "supports": false, + "reason": "This is a compiler error/cycle detection issue, not about different outputs from incremental vs clean builds" + }, + { + "n": 106571, + "supports": true, + "reason": "Duplicate diagnostic messages appear in JSON format - direct violation of property that incremental and clean should have the same diagnostics" + } + ], + "general": true, + "sharper_statement": "Incremental builds can produce different .rmeta bytes (due to iteration-order-dependent content like DocLinkResMap when encountering unrelated files in the library search path) and duplicate diagnostics compared to clean builds (due to span-based deduplication issues).", + "comment": "Of the 8 cited issues, 3 clearly support the property: #162901 and #106571 show diagnostic deduplication failures producing duplicates in incremental builds, and #159677 shows .rmeta content varying based on unrelated files. The other 5 issues are compiler crashes, build system issues, or type checking problems that do not directly demonstrate violations of the incremental-equals-clean property. The property itself is general and mechanically checkable as described in the how_to_check section." + }, + { + "i": 2, + "name": "MIR passes validation after every pass at every mir-opt-level", + "citations": [ + { + "n": 158037, + "supports": false, + "reason": "Text is truncated and does not clearly demonstrate a MIR validation failure of the stated property; fix addresses MIR phase assignment but property violation is not evidenced." + }, + { + "n": 156409, + "supports": true, + "reason": "Shows type const with mismatched type (usize assigned i32), violating the requirement that const arguments have correct types; fix normalizes type const patterns to prevent this." + }, + { + "n": 154750, + "supports": true, + "reason": "Fix explicitly states that normalizing a type const can change the value's type, causing MIR to become ill-formed; directly demonstrates const argument type violations break well-formedness." + }, + { + "n": 154748, + "supports": true, + "reason": "Same as 154750: shows type const with wrong type (bool assigned i32), demonstrating that const argument type mismatches cause ill-formed MIR." + }, + { + "n": 152962, + "supports": true, + "reason": "MIR validation catches broken MIR with failed subtyping (u8 vs usize) caused by type const with wrong type; demonstrates validation catching const argument violations." + }, + { + "n": 158231, + "supports": true, + "reason": "Shows a local variable accessed after StorageDead and before next StorageLive (undefined behavior in MIR), directly violating the storage lifetime invariant; demonstrates pass-introduced violations of the property." + }, + { + "n": 151647, + "supports": true, + "reason": "MIR validation reports broken MIR due to unevaluated const not being normalized, causing type equating to fail; demonstrates const normalization violations break well-formedness." + }, + { + "n": 151579, + "supports": true, + "reason": "Shows broken MIR with NoSolution error related to unnormalized place types in captures; fix normalizes place types to restore well-formedness." + }, + { + "n": 152278, + "supports": false, + "reason": "not in dataset" + }, + { + "n": 120811, + "supports": true, + "reason": "Same pattern as 151579: broken MIR due to unnormalized place types; same fix (normalize capture place types) demonstrates the property violation." + } + ], + "general": true, + "sharper_statement": "MIR optimization passes must preserve well-formedness invariants: const arguments must have the types their parameters declare, locals must only be accessed between their StorageLive and StorageDead markers, and place types as well as const values must be fully normalized. This holds at all -Zmir-opt-level settings (0, 2, 4) and with -Zinline-mir, and can be verified mechanically by running -Zvalidate-mir -Zlint-mir on real crates.", + "comment": "Eight of ten citations confirm the property. Issue 158037's truncated text prevents clear verification. Issue 152278 is not in dataset. The supporting issues collectively demonstrate violations in three areas: const argument type mismatches, storage lifetime violations, and unnormalized place/const types. The property is general and mechanically checkable using existing compiler validation flags." + }, + { + "i": 3, + "name": "extern \"C\" ABI matches the platform C compiler", + "citations": [ + { + "n": 121408, + "supports": true, + "reason": "Union passed as scalar instead of indirectly by pointer, violating wasm C ABI specification." + }, + { + "n": 162011, + "supports": true, + "reason": "Union with floats passed in FPRs instead of as aggregate, violating powerpc64 ELFv1 ABI." + }, + { + "n": 163173, + "supports": false, + "reason": "Inline asm register naming issue, not about function parameter or return passing convention." + }, + { + "n": 163074, + "supports": true, + "reason": "ArgAttributes mismatch detected in ABI test, indicating incorrect argument attribute handling." + }, + { + "n": 163075, + "supports": false, + "reason": "not in dataset" + }, + { + "n": 160827, + "supports": false, + "reason": "Missing wcslen builtin function, not a function signature ABI matching issue." + }, + { + "n": 151791, + "supports": false, + "reason": "Internal Rust type layout calculation for Archive trait, not extern \"C\" function ABI." + }, + { + "n": 122620, + "supports": true, + "reason": "Struct with f64 and f32 has incorrect ABI on sparc64, showing register/layout passing violation." + }, + { + "n": 115399, + "supports": true, + "reason": "ICE in ABI calculation (fn_abi_of_instance) when computing struct passing on sparc64." + }, + { + "n": 147883, + "supports": true, + "reason": "ICE in ABI calculation for struct with f64 and f32 on sparc, indicating broken ABI logic." + }, + { + "n": 158897, + "supports": false, + "reason": "Intrinsic behavior issue with unaligned_volatile_store, not about extern \"C\" function ABI." + }, + { + "n": 159244, + "supports": true, + "reason": "Bool return type emitted as i1 zeroext which violates aarch64 ABI guarantees, causing miscompilation." + }, + { + "n": 159116, + "supports": false, + "reason": "Const evaluation and internal memory layout corruption, not extern \"C\" function signature handling." + }, + { + "n": 43894, + "supports": true, + "reason": "Struct pass-by-value handling on SPARC violates C ABI, mismatching with clang output." + }, + { + "n": 161382, + "supports": true, + "reason": "Homogeneous Vector Aggregates passed indirectly instead of directly in registers, violating AAPCS64." + } + ], + "general": true, + "sharper_statement": "For every target, a function signature with C-compatible types (unions, structs with floats, bool, homogeneous vector aggregates, and other aggregates) used in extern \\\"C\\\" functions must be lowered to argument passing and return value handling (register vs stack, indirect vs direct, sign/zero extension) that matches the behavior of clang/gcc for equivalent C declarations.", + "comment": "The property is general and mechanically checkable as described: cross-compile C and Rust with matching signatures and verify calling conventions match. The supporting citations show concrete violations across multiple targets (wasm, powerpc64, sparc64, aarch64) where rustc diverges from platform C compilers in how it passes/returns types." + }, + { + "i": 4, + "name": "Interned type-system values are well-formed", + "citations": [ + { + "n": 150506, + "supports": true, + "reason": "ICE from const value's valtree shape not matching type: expected scalar leaf but got struct value in Assume field" + }, + { + "n": 150712, + "supports": true, + "reason": "ICE from valtree shape mismatch: Container generic arg received malformed valtree structure" + }, + { + "n": 150734, + "supports": true, + "reason": "ICE from valtree shape mismatch: Outer struct const arg had invalid valtree structure" + }, + { + "n": 151126, + "supports": true, + "reason": "ICE from const valtree shape mismatch: unevaluated const treated as if evaluated in array context" + }, + { + "n": 158675, + "supports": true, + "reason": "Directly violates the third part of the property: ParamEnv contained two conflicting ConstArgHasType predicates for the same const parameter" + }, + { + "n": 158362, + "supports": false, + "reason": "not in dataset" + }, + { + "n": 157189, + "supports": true, + "reason": "GenericArgs-carrying value (Borrow TraitRef) had argument count not matching generics_of: passed only receiver when trait had additional parameters" + }, + { + "n": 137084, + "supports": true, + "reason": "Index out of bounds from args-generics mismatch: empty args when generics expected, shows violation of count-matching invariant" + }, + { + "n": 150673, + "supports": true, + "reason": "Index out of bounds from incorrect generics index mapping in delegation, caused by failure to maintain args-generics correspondence" + }, + { + "n": 150714, + "supports": true, + "reason": "ICE from const arg valtree mismatch: accessing struct field offsets failed because const arg structure didn't match its type" + }, + { + "n": 150841, + "supports": true, + "reason": "Const value's valtree shape (empty tuple literal) does not match declared type (non-tuple), violating the property" + }, + { + "n": 151024, + "supports": true, + "reason": "Const value's valtree shape (empty array literal) does not match declared type (non-array), violating the property" + }, + { + "n": 151186, + "supports": true, + "reason": "Index out of bounds from args-generics count mismatch: empty args when function pointer type generics exist" + } + ], + "general": true, + "sharper_statement": "Every interned GenericArgs-carrying value (TraitRef, AliasTy, FnDef, etc.) must have arguments whose count and kind match the corresponding generics_of definition. Every interned const value must have a valtree whose structure matches its type. No ParamEnv must contain two predicates with the same parameter (e.g., ConstArgHasType) that specify different constraints.", + "comment": "All 12 cited issues (excluding the not-in-dataset one) demonstrate clear violations of this invariant, supporting all three facets of the property. The issues fall into two categories: valtree shape mismatches with const types (the majority), and GenericArgs count/kind mismatches with their generics definitions. One issue directly shows ParamEnv predicate conflicts. The property is general across all interned type-system values and mechanically checkable via the debug assertions described in how_to_check." + }, + { + "i": 5, + "name": "Old and next trait solvers agree", + "citations": [ + { + "n": 102580, + "supports": true, + "reason": "Shows derive(Clone) overflow in old solver but not next-solver, documenting solver disagreement on acceptance" + }, + { + "n": 90950, + "supports": true, + "reason": "HRTB bounds issue marked as fixed-by-next-solver, showing solvers disagreed on what's accepted" + }, + { + "n": 152789, + "supports": true, + "reason": "Explicitly states code compiles with old solver but ICEs with -Znext-solver" + }, + { + "n": 151329, + "supports": true, + "reason": "ICE with next-solver (could not replace AliasTerm), showing solver disagreement" + }, + { + "n": 151957, + "supports": true, + "reason": "Explicitly states compiles with old solver but ICEs with -Znext-solver" + }, + { + "n": 151323, + "supports": true, + "reason": "ICE with next-solver, showing solver disagreement on handling stalled coroutine obligations" + }, + { + "n": 151322, + "supports": true, + "reason": "ICE with next-solver, showing solver disagreement on opaque type handling" + }, + { + "n": 151318, + "supports": true, + "reason": "ICE with -Znext-solver on region-dependent goals, showing solver disagreement" + }, + { + "n": 138274, + "supports": true, + "reason": "Fixed by solver fix PR #152327, indicating solver-related disagreement in async block handling" + }, + { + "n": 137916, + "supports": true, + "reason": "ICE with next-solver on unsize coercion, showing solver disagreement" + }, + { + "n": 151462, + "supports": true, + "reason": "ICE with -Znext-solver in transmutability error reporting, showing solver disagreement" + } + ], + "general": true, + "sharper_statement": "The next-solver and old solver must behave identically on any crate: accepting and rejecting the same items, inferring the same types for each body (up to region erasure), and selecting the same impls. The cited issues are cases where this invariant was violated, with the next-solver either panicking (ICE) or disagreeing with the old solver on acceptance/rejection of code.", + "comment": "All 11 citations demonstrate clear violations of the property by showing cases where the two solvers disagreed. Most show next-solver panicking on code the old solver accepts; one shows the reverse (old solver overflowing where next-solver succeeds). This is a mechanically checkable property: compile with both solvers and compare exit status, diagnostics, typeck_results, and codegen instances." + }, + { + "i": 6, + "name": "Library operations stay sound under injected panics and allocation failures", + "citations": [ + { + "n": 162720, + "supports": true, + "reason": "Issue text demonstrates allocator returning pointer with full provenance (10 extra bytes) being incorrectly restricted when Arc derives through mutable reference, causing UB on deallocation." + }, + { + "n": 162719, + "supports": true, + "reason": "Issue text shows Box::into_unique restricting pointer provenance through mutable reference, violating the contract that deallocate receives pointer with same provenance as allocate." + }, + { + "n": 158165, + "supports": true, + "reason": "Issue text and reproducer show key comparator panic during split_off unwinds mid-restructuring, leaving tree inconsistent with stale length, causing double-free on safe iteration." + }, + { + "n": 157203, + "supports": true, + "reason": "Issue text shows allocation failure causing handle_alloc_error to unwind, leaving Arc with strong count 0, creating use-after-free." + }, + { + "n": 155746, + "supports": true, + "reason": "Issue text demonstrates allocator-backed Arc::make_mut panic-unsafety where allocator panic during clone leaves data freed but reachable, causing UAF on later access." + }, + { + "n": 152211, + "supports": true, + "reason": "Issue text and examples show closure panic in array::map/try_map leaves unprocessed elements undroppd, violating drop invariant." + }, + { + "n": 114581, + "supports": false, + "reason": "This is a generic Stacked Borrows violation in SGX-specific unsafe code, not about library operations receiving user callbacks, allocators panicking, or allocation failures." + }, + { + "n": 160815, + "supports": true, + "reason": "Issue text describes aliasing violations when using Condvar/Mutex with custom allocators and Arc::get_mut_unchecked, causing mutable reference creation that violates aliasing rules." + }, + { + "n": 161018, + "supports": false, + "reason": "This is a threading/cleanup issue where process::exit from non-main thread incorrectly unmaps main thread's stack. Not about user callbacks, allocators, or allocation failures." + }, + { + "n": 43894, + "supports": false, + "reason": "SPARC struct ABI code generation issue completely unrelated to panic safety or allocation failures." + }, + { + "n": 161382, + "supports": false, + "reason": "AAPCS64 HVA ABI handling issue completely unrelated to panic safety or allocation failures." + } + ], + "general": true, + "sharper_statement": "Library operations (particularly BTreeMap, Arc, Box, and array functions) fail to maintain soundness invariants when user callbacks (comparators, closures) panic or when allocations fail. Violations include: improper drops of unprocessed values, state corruption leading to double-frees, use-after-free from incorrect strong count management, and aliasing violations when custom allocators or mutable reference derivation are involved. Additionally, allocators that return pointers with provenance exceeding the requested size can be unsoundly treated as having restricted provenance, violating the deallocation contract.", + "comment": "7 of 11 citations support the property. The property is general (applies across multiple library functions and failure modes) and mechanically checkable via Miri with fault injection. The unsupported issues (#114581, #161018, #43894, #161382) are unrelated soundness problems in different subsystems: generic unsafe code, threading cleanup, and ABI code generation respectively." + }, + { + "i": 7, + "name": "P4+ no query reads untracked state", + "citations": [ + { + "n": 162901, + "supports": false, + "reason": "Diagnostic deduplication logic produces different hashes due to how Span::parent is considered; not about untracked state reads (env vars, clock, filesystem, HashMap iteration)" + }, + { + "n": 162407, + "supports": false, + "reason": "Build system (bootstrap) dependency tracking issue, not a rustc query reading untracked state; the property is specifically about queries in incremental compilation" + }, + { + "n": 159677, + "supports": true, + "reason": "Directly demonstrates HashMap iteration order (symbol interning) affecting .rmeta output when unrelated files in library search path are scanned, without that filesystem dependency being tracked in dep graph" + }, + { + "n": 160255, + "supports": false, + "reason": "Type resolution ICE in borrowck when handling const with fn pointer types; not about env vars, clock, filesystem metadata, or HashMap iteration order" + } + ], + "general": true, + "sharper_statement": "No query reads HashMap/HashSet iteration order (including symbol interner order) in ways that affect incremental compilation artifacts (like .rmeta) unless the dependency on which symbols get interned/which HashMap entries are iterated is recorded in the dep graph.", + "comment": "Only one of four citations supports the property. Issue #159677 is strong evidence: rustc's proc-macro dependency search scans unrelated crates in the filesystem, interns their symbols as a side effect, and since Symbol hashes the interner index, .rmeta hashing depends on which files were present without that dependency being tracked. The other three citations address different classes of bugs (deduplication logic, build system invalidation, type resolution ICE) unrelated to untracked state reads." + }, + { + "i": 8, + "name": "Metadata reads hit entries the writer wrote", + "citations": [ + { + "n": 163426, + "supports": true, + "reason": "The ICE directly shows reading an SVH from the metadata table that was never written because metadata generation was skipped for rustdoc runs" + }, + { + "n": 159233, + "supports": true, + "reason": "The panic shows reading invocation_parents[invoc_id] for an eager invocation that never went through collection, so the entry was never written" + }, + { + "n": 159703, + "supports": false, + "reason": "Issue text is truncated and does not explain what metadata entry was read that was never written" + } + ], + "general": true, + "sharper_statement": "Reads of metadata table entries fail when the upstream encoder never wrote those entries, causing ICEs when dependent crates attempt to access missing (table, DefIndex) pairs or when compiler passes try to read unwritten entries from internal lookup tables", + "comment": "Two of three cited issues strongly support the property with clear evidence of reads hitting unwritten table entries. The third issue has insufficient text to verify. The property is mechanically checkable via mirth's table write/read logging approach." + }, + { + "i": 9, + "name": "Parallel front end is deterministic", + "citations": [ + { + "n": 129094, + "supports": true, + "reason": "PR #161450 fixes non-deterministic encoding of syntax contexts in parallel compilation, confirming that parallel compiler produced non-deterministic outputs for derived trait code" + }, + { + "n": 150451, + "supports": true, + "reason": "Issue directly shows repeated compilations with -Zthreads=3 producing different output checksums (ddeb820..., 0604d96..., 3c39057...), proving repeated runs do not agree" + }, + { + "n": 153391, + "supports": true, + "reason": "PR #153694 fixes arbitrary rotation of cycle[0] by parallel deadlock handler, demonstrating non-deterministic behavior in parallel compiler's query system affecting output consistency" + } + ], + "general": true, + "sharper_statement": "The parallel front end with -Zthreads > 1 produces non-deterministic behavior in: (1) encoding of syntax contexts affecting derived code generation, (2) LLVM inline asm location cookies in bitcode/LTO output, and (3) query cycle handling in the deadlock resolver. These cause byte-identical outputs and repeated runs to diverge from both -Zthreads=1 and each other.", + "comment": "All three cited issues demonstrate genuine violations of the determinism property. The property is general and mechanically checkable through byte-comparing outputs and comparing diagnostics across runs with different thread settings. The issues show the parallel front end fails this property across multiple subsystems: syntax context encoding, LLVM cookies, and query cycle handling." + }, + { + "i": 10, + "name": "rustdoc accepts every crate rustc accepts", + "citations": [ + { + "n": 156418, + "supports": true, + "reason": "Text shows rustdoc ICEs on valid anon const code with type-relative paths when rustc check would succeed, directly violating the core property" + }, + { + "n": 156327, + "supports": false, + "reason": "not in dataset" + }, + { + "n": 149089, + "supports": true, + "reason": "Text shows rustdoc ICEs when documenting a real crate (embedded-io) that rustc accepts, directly violating the property" + }, + { + "n": 150153, + "supports": true, + "reason": "Text shows rustdoc ICEs on build-pass code involving type-relative paths in nested items, violating the property" + }, + { + "n": 147882, + "supports": true, + "reason": "Text shows rustdoc ICEs building docs for serde_with dependency when rustc would accept it, violating the property" + }, + { + "n": 147057, + "supports": true, + "reason": "Text shows rustdoc ICEs on workspace code with type-relative path lookups, violating the property" + }, + { + "n": 158050, + "supports": true, + "reason": "Text and PR show proc-macro-generated impl blocks overwriting source spans in intra-doc links, directly violating the proc-macro span requirement" + }, + { + "n": 157756, + "supports": false, + "reason": "This is about rustc's orphan checking being too strict, not about rustdoc failing when rustc succeeds; a different bug in the same subsystem" + } + ], + "general": true, + "sharper_statement": "rustdoc must not ICE on any Rust code that passes `cargo check`, particularly code involving type-relative paths and anon consts with `--generate-link-to-definition`, and must not allow proc-macro-generated spans to overwrite source spans in documentation links", + "comment": "Seven of eight citations support the property with clear violations. The supporting issues show two categories of failure: (1) rustdoc ICEs on code that rustc accepts (issues #156418, #149089, #150153, #147882, #147057), primarily related to type-relative path resolution in nested items, and (2) proc-macro spans overwriting source spans in doc links (#158050). Issue #157756 about orphan checking is unrelated\u2014it's about rustc being too strict, not about the rustdoc-accepts-everything-rustc-accepts relationship. The property is mechanically testable: run cargo check, then all rustdoc modes on a corpus, and verify no failures." + }, + { + "i": 11, + "name": "Results agree across opt-level, mir-opt-level, codegen-units and LTO", + "citations": [ + { + "n": 163779, + "supports": true, + "reason": "Shows observable behavior (return value and potential panic) differs when optimizations are enabled vs disabled, with the assert being incorrectly eliminated at higher opt-levels" + }, + { + "n": 163220, + "supports": false, + "reason": "This is a performance regression issue (throughput measurements), but the property explicitly specifies 'observable behaviour (test results, stdout, exit and panic status)' which are semantic concerns, not performance metrics" + }, + { + "n": 162348, + "supports": true, + "reason": "Shows exit status differs based on opt-level: crashes with SIGILL in release mode (optimizations enabled) but presumably succeeds without optimization" + }, + { + "n": 153645, + "supports": true, + "reason": "Shows linkability fails when opt-level >= 1 (due to thin-local LTO eliminating necessary symbols), but succeeds at opt-level = 0, codegen-units = 1, or lto = off, directly violating the linkability invariant" + }, + { + "n": 153451, + "supports": false, + "reason": "While a linking issue, it's platform-specific (FreeBSD environ symbol) rather than clearly dependent on the property's configuration parameters (opt-level, mir-opt-level, codegen-units, lto)" + } + ], + "general": true, + "sharper_statement": "Compiler optimization passes and LTO can silently change program semantics: a program's observable behavior (exit status, panic status, and linkability) may differ across -Copt-level 0/3, -Zmir-opt-level 0/4, codegen-units 1/16, and lto off/thin/fat settings, when it should remain identical.", + "comment": "Three issues confirm the property: a range propagation pass incorrectly eliminates overflow checks (#163779), optimization passes can trigger crashes (#162348), and LTO can eliminate necessary function symbols (#153645). The property is indeed general and mechanically checkable via regression testing across configuration matrices." + }, + { + "i": 12, + "name": "Every span is valid", + "citations": [ + { + "n": 151610, + "supports": true, + "reason": "The fix reveals spans becoming empty (lo >= hi) when stmt.span == tail_expr.span, violating the lo <= hi requirement" + }, + { + "n": 151607, + "supports": true, + "reason": "SubstitutionParts have overlapping byte ranges (5:13-5:16 overlaps with 5:15-5:16), violating the non-overlap requirement" + }, + { + "n": 147339, + "supports": true, + "reason": "Span context mismatch between body_span (#5) and covspan.span (#6) indicates spans not properly within the same SourceFile context" + }, + { + "n": 131292, + "supports": true, + "reason": "Assertion failure on byte position arithmetic indicates span boundaries are invalid, violating UTF-8 character boundary constraints" + }, + { + "n": 156316, + "supports": true, + "reason": "Span created by subtracting BytePos(1) from multibyte character tokens points into the middle of UTF-8 sequences, violating character boundary requirement" + }, + { + "n": 155037, + "supports": true, + "reason": "span_extend_prev_while using hardcoded byte offset 1 instead of c.len_utf8() caused spans to point into middle of UTF-8 multibyte characters" + } + ], + "general": true, + "sharper_statement": "Every span in diagnostics (including suggestion parts) and in encoded metadata has lo <= hi (non-empty), lies inside an existing SourceFile with matching context, falls on UTF-8 character boundaries, and each suggestion's substitutions do not overlap.", + "comment": "All six cited issues represent concrete violations of the span invariant: invalid empty spans, overlapping suggestion parts, context mismatches across files, and byte positions pointing into the middle of multibyte UTF-8 characters. The property is general and mechanically checkable as described in the how_to_check field." + }, + { + "i": 13, + "name": "Semantically neutral edits do not change results", + "citations": [ + { + "n": 156004, + "supports": false, + "reason": "About improving compiler diagnostics for type inference, not about source code edit neutrality" + }, + { + "n": 86959, + "supports": true, + "reason": "Shows parentheses marked as unnecessary are actually required in 2018 edition for pattern alternation, violating the property's guarantee that neutral edits don't change compilation" + }, + { + "n": 160741, + "supports": true, + "reason": "Shows braces marked as unnecessary actually affect lock release timing and program behavior, demonstrating the linter incorrectly classifies a non-neutral edit as neutral" + }, + { + "n": 157758, + "supports": true, + "reason": "Shows Debug implemented on a normalized projection type is not recognized as Debug being implemented for the underlying type, violating the property's expectation that normalized aliases are equivalent" + }, + { + "n": 157757, + "supports": true, + "reason": "Shows Debug implemented on a free type alias is not recognized as equivalent to Debug on the normalized target type, contradicting the property's treatment of normalizing aliases" + }, + { + "n": 152004, + "supports": true, + "reason": "Shows pub(crate) glob re-export marked as unused is actually required for compilation, demonstrating the linter incorrectly identifies a non-neutral edit as removable" + }, + { + "n": 151983, + "supports": false, + "reason": "Shows missing lint detection for unused variables in match guards, not about false claims that edits are neutral or incorrect equivalence handling" + } + ], + "general": true, + "sharper_statement": "The Rust compiler may incorrectly identify certain edits as semantically neutral (particularly parentheses in patterns, braces affecting resource lifetime scope, and glob re-exports) when they are actually required for compilation or correct behavior, and may fail to recognize trait implementations on type aliases and projections that normalize to an underlying type as equivalent to direct implementations on that type.", + "comment": "5 out of 7 citations support the property by demonstrating violations where the compiler incorrectly claims edits are neutral or fails to recognize equivalent types. The property is general (applies to any neutral edits) and mechanically checkable (compare outputs after applying a catalogue of neutral rewrites), though it requires precise definition of what constitutes semantic neutrality. The violations occur in the linting system's classification of necessity and in trait resolution's handling of normalized aliases." + }, + { + "i": 14, + "name": "A successful compile contains no error types", + "citations": [ + { + "n": 150969, + "supports": true, + "reason": "ErrorGuaranteed values were being processed as regular const values instead of being detected as errors, demonstrating that error types appeared in internal structures (ConstValue) without proper error checking" + }, + { + "n": 149588, + "supports": true, + "reason": "Type error placeholders ({type error}) were created in transmute layout checking when normalization failed, but these errors weren't emitted properly, violating the invariant that error types should only appear after errors are reported" + }, + { + "n": 151299, + "supports": false, + "reason": "Text is truncated and does not clearly show violation of the invariant; fix is merely a regression test without explanation of the root cause" + }, + { + "n": 154780, + "supports": true, + "reason": "ICE message explicitly states 'TyKind::Error constructed but no error reported', directly demonstrating the property violation in delegation's handling of generic parameters" + } + ], + "general": true, + "sharper_statement": "If a compilation exits 0, no ty::Error, ConstKind::Error, or ErrorGuaranteed-carrying value should appear in compiler internals (valtree constants, transmute layout checking, MIR, query results, or metadata). Error types are created only after errors are emitted; violations occur when compiler subsystems construct these types during failed operations (const evaluation, layout normalization, or delegation) without emitting corresponding errors or errors without reporting.", + "comment": "Three of four issues support the property with clear violations where error types were constructed or improperly handled without corresponding error emission. Issue #151299 cannot be verified due to truncated text. The property is mechanically checkable via mirth as described in how_to_check." + }, + { + "i": 15, + "name": "Layout views agree", + "citations": [ + { + "n": 154426, + "supports": true, + "reason": "Issue shows unsafe binder layout check was examining outer type when it should use inner type layout, directly violating the agreement property" + }, + { + "n": 154424, + "supports": true, + "reason": "Issue shows discriminant helpers were not using unsafe binder's erased inner type view, violating layout agreement across components" + }, + { + "n": 155412, + "supports": true, + "reason": "Issue demonstrates transmute size check allowed invalid transmutes by ignoring repr(align) differences, violating agreement between SizeSkeleton and actual layout" + }, + { + "n": 88290, + "supports": true, + "reason": "Issue shows transmute check ignoring alignment and enum repr, causing SizeSkeleton to disagree with actual type layouts" + } + ], + "general": true, + "sharper_statement": "Unsafe binder layout checks and discriminant helpers must use the inner type's layout view. Transmute's SizeSkeleton check must account for repr(align/packed) differences that affect a type's actual size. repr(transparent) wrappers have the same ABI as their non-ZST field.", + "comment": "All four citations support the property. The unsafe binder citations (#154426, #154424) show failures where layout checks examined the wrapper instead of delegating to the inner type. The transmute citations (#155412, #88290) show the SizeSkeleton check was incomplete, failing to reject transmutes when repr(align/packed) changed the size. The property is mechanically checkable by recording layout_of per type and comparing against what each component (transmute checks, wrapper handling, discriminant helpers) assumes about layout." + }, + { + "i": 16, + "name": "Each encoded record is written once and hashed after it is final", + "citations": [ + { + "n": 150018, + "supports": true, + "reason": "Issue text explicitly shows `trying to encode a dep node twice` panic message, directly violating the write-once invariant" + }, + { + "n": 142778, + "supports": true, + "reason": "Fixed by same PR #151509 addressing duplicate dep node writes; the bounds-check assertion failure is a symptom of the invariant violation" + }, + { + "n": 163426, + "supports": true, + "reason": "Issue text shows `crate_hash(LOCAL_CRATE) called before metadata encoding` panic, directly violating the constraint that crate_hash not be computed before encoder finishes" + } + ], + "general": true, + "sharper_statement": "The metadata encoder writes each dep node at most once under concurrent execution, and crate_hash is not computed before metadata encoding finishes", + "comment": "All three citations directly or indirectly support the property. The property is general and mechanically checkable via mirth logging and instrumentation as specified in how_to_check. The concrete violations are: (1) concurrent race conditions causing duplicate dep node writes to metadata, and (2) premature crate_hash computation before metadata encoding." + }, + { + "i": 17, + "name": "rustdoc output is the same however an item is re-exported", + "citations": [ + { + "n": 119965, + "supports": false, + "reason": "Issue is about intra-doc link resolution with mixed inner/outer doc comments, not about re-exported items showing different documentation" + }, + { + "n": 81893, + "supports": true, + "reason": "Directly demonstrates multi-level re-exports losing documentation from intermediate levels" + }, + { + "n": 53724, + "supports": true, + "reason": "Shows glob re-exports failing to include items that appear in named re-exports" + }, + { + "n": 96166, + "supports": true, + "reason": "Demonstrates cfg badges failing to propagate to glob re-exported items while appearing on individual re-exports" + }, + { + "n": 154921, + "supports": true, + "reason": "Shows cfg attributes being lost when type aliases are re-exported, violating the property that re-exports show the same cfg badges" + } + ], + "general": true, + "sharper_statement": "Re-exported items may fail to preserve the documentation and cfg badges of their original definitions, particularly: multi-level re-exports may lose documentation from intermediate levels, glob re-exports may fail to include items that appear in named re-exports, and cfg badges may not propagate correctly to glob re-exports or re-exported type aliases.", + "comment": "Four of five citations directly demonstrate violations. The property is mechanically checkable via rustdoc JSON comparison across re-export forms." + }, + { + "i": 18, + "name": "#[expect] is fulfilled exactly when the lint fires", + "citations": [ + { + "n": 152004, + "supports": true, + "reason": "False positive in unused imports lint demonstrates lint firing incorrectly, violating the property's assumption that lint emission reliably indicates expectation fulfillment." + }, + { + "n": 151983, + "supports": true, + "reason": "Missing unused_variables lint in match guards shows the differential check failing: #[expect] would be unfulfilled but #[warn] would emit the lint." + }, + { + "n": 152401, + "supports": true, + "reason": "Direct violation: #[expect(missing_docs)] reported unfulfilled even though removing it shows the lint does fire, proving expectation status does not match lint emission." + }, + { + "n": 152289, + "supports": false, + "reason": "not in dataset" + } + ], + "general": true, + "sharper_statement": "#[expect(L)] fulfillment status is not always correctly aligned with whether #[warn(L)] would actually emit L. Cases involving match guards, derives, and cfg_attr show mismatches where expectations are reported as unfulfilled despite the lint firing, or vice versa.", + "comment": "All three available citations demonstrate violations of the property's core claim about the correspondence between #[expect] fulfillment and lint emission. The property is mechanically testable through differential checking but the cited issues prove the implementation does not reliably implement this invariant, particularly in complex scoping contexts." + }, + { + "i": 19, + "name": "Query results contain no inference variables", + "citations": [ + { + "n": 153525, + "supports": true, + "reason": "Issue shows ICE when type inference variable (?0t) appears in cached query result (lit_to_const) and query system attempts to hash it, directly violating the property" + }, + { + "n": 153524, + "supports": true, + "reason": "Issue shows ICE when const inference variable (?0c) appears in cached query result and cannot be hashed, demonstrating the exact invariant violation described in the property" + } + ], + "general": true, + "sharper_statement": "Inference variables leak into the const literal lowering logic (lit_to_const), appearing in cached query results where they cause hashing panics.", + "comment": "Both cited issues demonstrate inference variables escaping into the const literal lowering phase and then being encountered during query result caching/hashing, causing ICE panics. The property is mechanically checkable as described in the how_to_check field using tools like mirth to inspect query returns and HashStable calls." + }, + { + "i": 20, + "name": "Each diagnostic is emitted once", + "citations": [ + { + "n": 115376, + "supports": false, + "reason": "Issue shows errors with different codes (E0310 vs E0311) at the same span. Property requires identical level, code, message, AND primary span for a duplicate violation; different codes means these are distinct diagnostics by the property's definition" + }, + { + "n": 106571, + "supports": false, + "reason": "Issue text in dataset is incomplete/truncated; does not show the actual duplicate diagnostics despite title suggesting duplicates appeared in JSON output" + } + ], + "general": true, + "sharper_statement": "When the same diagnostic (identified by identical level, error code, message text, and primary source span) is emitted more than once in a single compilation session, it appears multiple times in --error-format=json output. This can be detected by extracting diagnostics from JSON and checking for exact duplicates by the tuple (level, code, message, span).", + "comment": "The property is mechanically checkable by post-processing JSON diagnostics and building a key from (level, code, message, primary_span), then detecting keys with count > 1. However, neither cited issue actually demonstrates this specific violation: #115376 shows different error codes despite same span, and #106571's issue text is truncated, preventing verification of the actual duplicates despite the suggestive title about \"duplicate messages\"." + }, + { + "i": 21, + "name": "Symbol names are injective and stable", + "citations": [ + { + "n": 158644, + "supports": true, + "reason": "Issue shows splat attributes missing from v0 mangling, causing distinct function types (splatted vs non-splatted) to generate the same symbol name, directly violating the injectivity property" + } + ], + "general": true, + "sharper_statement": "The v0 and legacy symbol mangling schemes must encode all type attributes that the type system considers semantically distinct, including splat on function types, to ensure monomorphized instances never share symbol names. V0 symbols must round-trip through rustc-demangle to the instance's path.\"", + "comment": "The cited issue (158644) is an exemplary violation: distinct types in the Rust type system (splat vs no splat on function pointer types) were compressed to the same symbol due to a mangling bug. The property as stated is mechanically verifiable via mirth's instance tracking and symbol injection/round-trip checks. The evidence is concrete: a reproducible test case, explicit design decision reference (#153697), and a fixing PR (#158890)." + }, + { + "i": 22, + "name": "dyn types have a dyn-compatible principal", + "citations": [ + { + "n": 154619, + "supports": true, + "reason": "Issue demonstrates a concrete violation: dyn Trait where Trait: DerefPure (a non-dyn-compatible principal) was allowed to compile and caused unsoundness; the PR fix explicitly prevents dyn-incompatible traits from being principals of trait objects" + } + ], + "general": true, + "sharper_statement": "In successful compilations, every TyKind::Dynamic that reaches typeck, MIR, or metadata must have a principal trait that is explicitly marked dyn-compatible; allowing Dynamic types with non-dyn-compatible principals (like DerefPure) enables unsound behavior.", + "comment": "The property is well-supported by the single citation. Issue #154619 is a direct counterexample to the property that was caught in practice and fixed. The property is mechanically checkable by asserting dyn-compatibility markers on all principals of Dynamic types in a successful compilation session." + }, + { + "i": 23, + "name": "Machine-applicable suggestions apply cleanly", + "citations": [ + { + "n": 161213, + "supports": true, + "reason": "Shows a violation where the const-extraction suggestion was generated with overlapping component spans, causing a compiler panic that prevents clean application" + }, + { + "n": 161472, + "supports": true, + "reason": "Shows a violation where a suggestion was generated with an incorrect empty span, causing a compiler panic that prevents processing and application" + } + ], + "general": true, + "sharper_statement": "Applying every MachineApplicable suggestion (rustfix) gives code that compiles and the diagnostic is gone, provided that suggestions are generated with correctly-calculated non-empty, non-overlapping spans. Violations occur when the suggestion system generates suggestions with malformed spans that cause compiler panics.", + "comment": "Both cited issues demonstrate concrete violations of the property by showing cases where suggestions were generated with defective spans (overlapping or empty), causing compiler panics rather than clean application. The property is general and mechanically checkable via cargo fix, but the cited issues reveal specific preconditions: suggestion span correctness is critical to the property holding." + }, + { + "i": 24, + "name": "Polonius accepts at least what NLL accepts, and no more than is sound", + "citations": [ + { + "n": 153215, + "supports": true, + "reason": "Shows Polonius misses errors in opaque type region liveness tracking that NLL correctly handles, allowing unsound code to be accepted." + }, + { + "n": 160669, + "supports": true, + "reason": "Demonstrates Polonius accepting code with undefined behavior (segfault confirmed by Miri) that should be rejected, violating soundness." + } + ], + "general": true, + "sharper_statement": "-Zpolonius=next has soundness bugs in opaque type region handling that allow acceptance of programs with undefined behavior, violating the guarantee that Polonius-only-accepted programs should be sound under Miri.", + "comment": "Both issues are violations of the second part of the property: that programs accepted only by Polonius run without UB under Miri. The property is mechanically checkable by differential borrow-checking with both NLL and Polonius, followed by Miri testing of Polonius-only-accepting programs." + }, + { + "i": 25, + "name": "Unstable syntax is gated before expansion", + "citations": [ + { + "n": 152501, + "supports": true, + "reason": "The issue demonstrates `try bikeshed () {}` inside `#[cfg(false)]` compiling without a pre-expansion gate warning, directly violating the property that unstable syntax must be rejected pre-expansion even in dead code." + }, + { + "n": 152499, + "supports": true, + "reason": "The issue demonstrates `const { 0 }` patterns (unstable syntax) being accepted pre-expansion in macro patterns and let bindings when they should be rejected, violating the property that all unstable syntax must be pre-expansion gated." + } + ], + "general": true, + "sharper_statement": "Every unstable syntactic form must be rejected at pre-expansion on stable without its feature gate, even inside #[cfg(FALSE)] blocks or unused macro_rules arms, to prevent code breakage when feature gates are removed.", + "comment": "Both cited issues show violations where unstable syntax forms were accepted pre-expansion when they should have been rejected. The property is general and mechanically checkable via the described UI test procedure." + }, + { + "i": 26, + "name": "Built-in attributes reject malformed arguments", + "citations": [ + { + "n": 154977, + "supports": true, + "reason": "The issue demonstrates that `#[macro_export(local_inner_macros)]` silently ignored malformed arguments like `(false)` and `= false` instead of erroring, violating the property that built-in attributes must reject argument forms their template does not allow. The fix (PR #155193) explicitly adds checking to reject such malformed arguments." + } + ], + "general": true, + "sharper_statement": "Every built-in attribute rejects argument forms its template does not allow by producing a diagnostic error, rather than silently ignoring excess or incorrectly-formed arguments.", + "comment": "The property is supported by a clear violation: a built-in attribute accepted and silently ignored arguments it should reject. The property is general and mechanically checkable\u2014the how_to_check guidance correctly describes a systematic approach: generate malformed argument forms for each attribute (extra arguments, wrong syntactic form, etc.) and verify that diagnostics are produced for all of them. This requires generating test cases for each BUILTIN_ATTRIBUTE template and ensuring proper argument validation." + }, + { + "i": 27, + "name": "Advertised target features exist in LLVM", + "citations": [ + { + "n": 159976, + "supports": true, + "reason": "Issue demonstrates that register classes required target features (vfp2) inconsistently with the target baseline. The fix adjusted register class requirements and added missing implied features, directly correcting a violation of the property's consistency requirement." + } + ], + "comment": "The single cited issue provides clear evidence of a violation. Register classes were requiring features in a manner inconsistent with the target's baseline capabilities. The fix involved both adjusting feature requirements and adding missing implied features, confirming the underlying inconsistency that the property prohibits.", + "general": false, + "sharper_statement": "Every target feature that rustc lists for a target, and every feature that an asm register class or intrinsic requires, must be recognized by LLVM and must be consistent with the target's baseline\u2014that is, register class and intrinsic feature requirements cannot exceed what is available for that target's baseline configuration." + }, + { + "i": 28, + "name": "Eq and Hash agree for std types", + "citations": [ + { + "n": 161651, + "supports": true, + "reason": "Issue shows Path::new(r\"\\\\?\\C:/foo\") == Path::new(r\"\\\\?\\C:\\foo\") is true, but their hashes differ, directly violating the Eq/Hash agreement invariant for the std Path type." + } + ], + "general": true, + "sharper_statement": "For every std type implementing both Eq and Hash, a == b implies hash(a) == hash(b), for every Hasher. (The property statement is precise and matches the evidence; the violation in Path on Windows with verbatim paths is a concrete instance of this general invariant being broken.)", + "comment": "The property is a fundamental requirement of the Rust hash contract. The cited issue directly demonstrates a violation where Path types on Windows violated this invariant due to inconsistent handling of slashes in verbatim path prefixes. The fix addresses the root cause by making path normalization consistent. This is a general, mechanically checkable property that applies to all std types implementing both Eq and Hash." + } +] \ No newline at end of file diff --git a/docs/ur-queries.md b/docs/ur-queries.md new file mode 100644 index 0000000..765289e --- /dev/null +++ b/docs/ur-queries.md @@ -0,0 +1,167 @@ +# Bug patterns as Ur queries + +Each bug mirth found came from a pattern in rustc's source, not a one-off mistake. Written +as a query, a pattern can be run over the whole compiler to find the other places it +occurs. [`ur/rustc/RoundTrip.rsc`](../ur/rustc/RoundTrip.rsc) holds one query per bug, +written as classifiers for [Ur](https://github.com/PowderworksCode/codebase/tree/main/projects/ur)'s +`ur rewrite --classify`. Each is made as general as it can be while still finding its bug. + +```sh +ur rewrite --classify --rules ur/rustc/RoundTrip.rsc --report report.json path/to/rust/compiler +``` + +On rustc `ea137335b`'s `compiler/` (2,215 files) the queries run in 0.9 seconds with +Ur `ba84ca7`. + +## The queries and what they found + +| query | the pattern | its bug | sites | after triage | +|---|---|---|---:|---| +| `hashOrderEncoded` | a field whose type is a hash-ordered collection (`FxHashMap`, `FxHashSet`, `HashMap`, `HashSet`, `UnordMap`, `UnordSet`) in a type that derives an encoder | `Generics::param_def_id_to_index` ([report](hunt/issue-generics-order.md)) | 12 | the bug; the rest never reach `.rmeta` | +| `hashOrderAlias` | a type alias for such a collection, which anything encoding a value of it writes in iteration order | #159677 (`DocLinkResMap`), from before its fix | 14 | none reach an encoder today | +| `decodedFresh` | a decoder calling a function that reserves a fresh identity (`reserve*`, `fresh*`, `*next_id*`) where neither its name nor its definition deduplicates | the string literal decoded twice ([report](hunt/issue-literal-dedup.md)) | 1 | the bug | +| `untrackedWhileEncoding` | inside an encoder (an `Encode*` impl or a function named `encode*`), a read of the session, the environment or the clock | stale metadata from the untracked source map ([report](hunt/issue-stale-metadata-reuse.md)) | 23 | the bug, and one more read of the same data; the rest benign | + +**Each query finds the bug it came from.** `hashOrderAlias` also finds #159677, a bug fixed +in July 2026: with `DocLinkResMap` put back to the `UnordMap` it was before #159718, the +query flags it. It was written for bug 1 and not tuned to that one. + +**Triage.** No new bug turned up. Every other site was checked by hand: + +- `hashOrderEncoded`: `UnordMap.inner` and `UnordSet.inner` show that every `UnordMap` and + `UnordSet` is encoded in iteration order, so any reaching metadata would be this bug + again. The hash-ordered query results that metadata reads are encoded sorted + (`stability_implications` uses `to_sorted_stable_ord`), and `doc_link_resolutions` is an + `FxIndexMap` since #159718. The other fields are in the incremental cache's footer, + `TypeckResults`, `WorkProduct` and codegen's `CrateInfo`, which never reach `.rmeta`; the + cache does not have to be reproducible. `FormatArguments.names` is in the AST, which + metadata does not encode. +- `hashOrderAlias`: fed back into the field query, no alias is a field of an encoded type; + `UnhashMap`, the one used in `rustc_metadata`, is a decoder's lookup table. +- `decodedFresh`: before it checked the callee's definition it also flagged + `reserve_and_set_fn_alloc`, `_vtable_alloc`, `_static_alloc` and `_type_id_alloc`, which + deduplicate inside. It now flags only the bug. +- `untrackedWhileEncoding`: 16 reads of `sess.opts`. Each option that decides what is + encoded is `[TRACKED]` (`metadata_crate_hash`, `force_unstable_if_unmarked`, + `embed_metadata`, `optimize`, `output_types`, `target_triple`). The `[UNTRACKED]` ones only + print statistics (`meta_stats`), keep temporary files (`save_temps`), prefetch (`jobs`), or + differ between incremental and non-incremental builds and not within either + (`incremental`). The second `source_map()` read (encoding spans) reads the same untracked + file data as the bug, which its fix's fingerprint covers. + `gather_enabled_denied_partial_mitigations()` looked like the same bug: the + `-Zallow-partial-mitigations` and `-Zdeny-partial-mitigations` options are `[UNTRACKED]` + and kept outside the dependency-tracking hash. But the list it encodes holds every + mitigation kind with its level, which come from tracked options (`stack_protector`, + `control_flow_guard`); the untracked ones decide only whether rustc reports an error, + which a build with only `-Zdeny-partial-mitigations` changed between sessions confirms. + `proc_macro_quoted_spans()` is session state that expansion rebuilds in every session. + `sess.target` is tracked, `dcx()` and `prof` only report, and the hit in + `rustc_sanitizers` is a type-id encoder, not metadata: the query's notion of an encoder + is a name. + +## What Ur would need to do this better + +The queries are syntactic. Ur's [rewriting design](https://github.com/PowderworksCode/codebase/blob/main/projects/ur-docs/ur/rewrite.md) +plans most of what they lacked: + +- **Types** (its section 4). "A value of a hash-ordered type reaches an encoder" is the + real property; without types, the queries approximate it by fields of types that derive + encoders and by aliases, and the triage above did the rest by hand. +- **Names across files** (section 3). `decodedFresh` decides whether a callee deduplicates + by looking for its definition in the same file. +- **Properties as queries** (section 9). These are that section's checks, run as + classifiers. A check that fails when its count changes would let CI run them on every + nightly. +- **Matching inside names.** A method name cannot be a pattern hole, and `contains` on + unparsed text matched `LazyValue` for the alias `Value`. + +# Closed bugs as queries + +The same approach, applied to bugs rustc has already fixed: the bugs behind mirth's +properties ([`motivating.md`](motivating.md)). [`ur/rustc/ClosedBugs.rsc`](../ur/rustc/ClosedBugs.rsc) +has one query per bug's pattern, each generalized as far as it still finds the code its +fix changed. [`ur/verify-closed.py`](../ur/verify-closed.py) checks that: for each bug it +fetches the files the fixing PR changed, as they were before the fix, runs the query, and +looks for the site. + +| bug | query | the pattern | finds the fixed site | +|---|---|---|---| +| #82920 | `sortByDefId` | sorting or deduplicating by `DefId`, which is not stable across sessions | `dedup_by_key` in `conv_object_ty_poly_trait_ref` | +| #89598 | `contextCache` | a cache in interior-mutable state on a context, outside the dependency graph | `GlobalCtxt.vtables_cache` | +| #84252 | `untrackedCrateStore` | reading the crate store directly | the `has_global_allocator` provider | +| #40364 | `envRead` | reading an environment variable | `env::var` in `expand_env` | +| #111227, #111295 | `fileRead` | reading a file | `std::fs::read` in `check_for_debugger_visualizer` | +| #45841 | `writeInPlace` | creating an output file directly instead of renaming a finished one into place | `fs::File::create(out_filename)` in `emit_metadata` | +| #117254 | `encoderNotFinished` | a function that creates a `FileEncoder` and never finishes it | `encode_metadata_impl` | +| #119456 | `writeErrorNotFatal` | a failed write reported with `emit_err`, after which compilation goes on | `encode_metadata` | +| #34902 | `hashIterated` | iterating a hash-ordered collection declared in the same file | `xrefs.into_iter()` in `encode_xrefs` | +| #65036 | `hashOrderAlias` (in `RoundTrip.rsc`) | an alias for a hash-ordered collection | `Resolutions = FxHashMap` | +| #66955 | `optionRead` | reading an option, joined afterwards with the options marked `[UNTRACKED]` | `remap_path_prefix` | + +These cover 13 of the 27 bugs, and #159677 is the first section's `hashOrderAlias`. The +others were not written as queries: + +- **Writer and reader out of step** (#122859, #130201, #144004): a table one crate reads + and another never writes. The check compares the tables declared in `rmeta/mod.rs`, the + `record!` calls in the encoder and the providers in `cstore_impl.rs`, which are all + inside macro invocations, which Ur does not parse into trees yet. +- **Not a pattern in the source**: the parallel front end's nondeterminism (#129094, + #140413, #150451), the new trait solver accepting overlapping impls incrementally + (#135514), inline assembly from a failed session leaking into a later one (#139407), + diagnostics deduplicated wrongly with incremental compilation (#162901), and temporary + files left after an error (#107001, #139899). +- **Outside rustc's source**: #138678's randomly seeded map was inside `pulldown-cmark`; + #68149, a dependent opening a dependency's in-flight `.rlib` while searching for crates, + depends on timing between processes. +- #114669 is the feature metadata reuse came from, not a bug. +Getting all eleven to find their site took four changes to the queries and one to the +inputs: arguments and match patterns are lists in Ur's trees, so the queries read their +text; a provider is labelled by the `Providers` field it is given as; `hashIterated` also +reads parameters; `optionRead` also reads `sopts` and `self` inside `impl Options`; and +the old unstable `crate` visibility, which Ur's grammars lack, is rewritten to `pub(crate)`. + +## What they found in today's compiler + +On rustc `ea137335b`'s `compiler/` they run in 2.3 seconds and report 1,400 sites. Not +every site was triaged; what was: + +**Three untracked options that change reused output.** `optionRead`, joined with the +options marked `[UNTRACKED]`, gave 74 such options read outside the session and the +driver. Most only affect linking, which runs every session, or debugging output. To test +the rest without judging each by hand, [`rustc/audit-options.py`](../rustc/audit-options.py) +builds a crate incrementally without an option, then with it, and compares with a clean +build that has it, as a comment on rust-lang/rust#84232 ("Audit all UNTRACKED options", +open since 2021) suggests. Of 50 boolean options, three change what an incremental session +reuses, on the official nightly as on the patched compiler: + +| option | a clean build | the incremental rebuild | +|---|---|---| +| `-Zemit-stack-sizes` | 42 of 44 objects get a `.stack_sizes` section | the previous objects, without it | +| `-Zcodegen-source-order` | one object's code in source order | the previous object | +| `-Zbuild-sdylib-interface` | an interface without function bodies: different metadata, 44 different objects, 27 warnings | the previous full build, 5 warnings | + +`-Csave-temps` also writes no temporary files for reused codegen units, which is expected +for a debugging option. Values of `-Ccodegen-units`, `-Zthreads` and three other +value-taking options made no difference. Report draft: +[`hunt/issue-untracked-options.md`](hunt/issue-untracked-options.md). + +**Checked and fine:** + +- `untrackedCrateStore`: every query provider that reads the crate store is `eval_always`. +- `encoderNotFinished`: the three functions hand their encoder to code that finishes it. +- `writeErrorNotFatal` in `rustc_metadata/src/fs.rs`: reports a failure to copy metadata + to standard output, where there is no published file. +- The offload manifest that the mono-item collector reads (`fileRead`) is recorded in + dep-info, and its query, `collect_and_partition_mono_items`, is `eval_always`. +- `sortByDefId`: four of five sites order diagnostics, debug output or suggestions. + +**Leads, not confirmed:** + +- `dependency_formats` is not `eval_always` but reads `CStore::injected_panic_runtime()` + directly, #84252's pattern. Triggering it would need the injected panic runtime to change + while every tracked input stays the same, which normally changes the crate list too. +- `specialization_graph_provider` sorts trait impls by `CrateNum` and `DefIndex`, #82920's + pattern; its comment says the order is deliberate. + +**Not triaged:** most of `envRead` (118), `fileRead` (49), `writeInPlace` (29), +`contextCache` (26), `hashIterated` (14), and the other `writeErrorNotFatal` sites (23). diff --git a/fixtures/audit/lib.rs b/fixtures/audit/lib.rs new file mode 100644 index 0000000..0934e34 --- /dev/null +++ b/fixtures/audit/lib.rs @@ -0,0 +1,57 @@ +//! A crate for rustc/audit-options.py: generic code, code of its own for codegen +//! options to change, a closure, a static, and warnings for diagnostic options. + +use std::collections::BTreeMap; +use std::fmt::Display; + +pub trait Shape { + fn area(&self) -> f64; + fn name(&self) -> String { + format!("shape of area {:.2}", self.area()) + } +} + +#[derive(Debug, Clone, Copy, PartialEq)] +pub struct Square(pub f64); + +impl Shape for Square { + fn area(&self) -> f64 { + self.0 * self.0 + } +} + +pub fn largest(shapes: &[T]) -> Option<&T> { + shapes.iter().max_by(|a, b| a.area().total_cmp(&b.area())) +} + +pub fn describe(items: &[T]) -> String { + items.iter().map(|i| i.to_string()).collect::>().join(", ") +} + +/// Code of its own, with a stack frame worth measuring. +pub fn checksum(values: &[u64]) -> u64 { + let mut buffer = [0u64; 64]; + for (i, v) in values.iter().enumerate() { + buffer[i % 64] = buffer[i % 64].wrapping_mul(31).wrapping_add(*v); + } + buffer.iter().fold(0, |acc, x| acc ^ x.rotate_left(7)) +} + +pub fn histogram(words: &str) -> BTreeMap { + let mut counts = BTreeMap::new(); + for w in words.split_whitespace() { + *counts.entry(w.len()).or_insert(0) += 1; + } + counts +} + +pub fn multiplier(n: u32) -> impl Fn(u32) -> u32 { + move |x| x.wrapping_mul(n) +} + +pub static TABLE: [u32; 4] = [1, 2, 3, 4]; + +fn unused_helper() -> u32 { + let unused_value = 3; + 4 +} diff --git a/fixtures/sink/Cargo.lock b/fixtures/sink/Cargo.lock new file mode 100644 index 0000000..c544752 --- /dev/null +++ b/fixtures/sink/Cargo.lock @@ -0,0 +1,36 @@ +# This file is automatically @generated by Cargo. +# It is not intended for manual editing. +version = 4 + +[[package]] +name = "sink" +version = "0.0.0" +dependencies = [ + "sink-core", + "sink-dy", + "sink-macros", + "sink-mid", +] + +[[package]] +name = "sink-core" +version = "0.0.0" + +[[package]] +name = "sink-dy" +version = "0.0.0" +dependencies = [ + "sink-core", +] + +[[package]] +name = "sink-macros" +version = "0.0.0" + +[[package]] +name = "sink-mid" +version = "0.0.0" +dependencies = [ + "sink-core", + "sink-macros", +] diff --git a/fixtures/sink/Cargo.toml b/fixtures/sink/Cargo.toml new file mode 100644 index 0000000..973b528 --- /dev/null +++ b/fixtures/sink/Cargo.toml @@ -0,0 +1,14 @@ +[workspace] +members = ["macros", "core", "mid", "dy"] + +[package] +name = "sink" +version = "0.0.0" +edition = "2024" +publish = false + +[dependencies] +sink-core = { path = "core" } +sink-mid = { path = "mid" } +sink-macros = { path = "macros" } +sink-dy = { path = "dy" } diff --git a/fixtures/sink/core/Cargo.toml b/fixtures/sink/core/Cargo.toml new file mode 100644 index 0000000..c6204b4 --- /dev/null +++ b/fixtures/sink/core/Cargo.toml @@ -0,0 +1,10 @@ +[package] +name = "sink-core" +version = "0.0.0" +edition = "2024" +publish = false +build = "build.rs" + +[features] +default = ["extra"] +extra = [] diff --git a/fixtures/sink/core/build.rs b/fixtures/sink/core/build.rs new file mode 100644 index 0000000..3964e6e --- /dev/null +++ b/fixtures/sink/core/build.rs @@ -0,0 +1,21 @@ +// Generates a lookup table and a cfg for lib.rs. +use std::fmt::Write; + +fn main() { + let mut out = String::from("/// Generated by build.rs.\npub const PRIMES: [u32; 12] = ["); + let mut found = 0; + let mut n = 2u32; + while found < 12 { + if (2..n).all(|d| n % d != 0) { + write!(out, "{n}, ").unwrap(); + found += 1; + } + n += 1; + } + out.push_str("];\npub const GENERATED_NAME: &str = \"sink\";\n"); + let dir = std::env::var("OUT_DIR").unwrap(); + std::fs::write(format!("{dir}/generated.rs"), out).unwrap(); + println!("cargo::rustc-check-cfg=cfg(sink_generated)"); + println!("cargo::rustc-cfg=sink_generated"); + println!("cargo::rerun-if-changed=build.rs"); +} diff --git a/fixtures/sink/core/src/algo.rs b/fixtures/sink/core/src/algo.rs new file mode 100644 index 0000000..e884d34 --- /dev/null +++ b/fixtures/sink/core/src/algo.rs @@ -0,0 +1,209 @@ +//! Iterators, closures, patterns, control flow, collections, smart pointers. + +use std::borrow::Cow; +use std::cell::RefCell; +use std::collections::{BTreeMap, BTreeSet, HashMap, VecDeque}; +use std::rc::Rc; +use std::sync::{Arc, Mutex}; + +/// A boxed closure pipeline. +pub struct Pipeline { + stages: Vec T + Send + Sync>>, +} + +impl core::fmt::Debug for Pipeline { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + write!(f, "Pipeline({} stages)", self.stages.len()) + } +} + +impl Pipeline { + pub fn new() -> Self { + Pipeline { stages: Vec::new() } + } + + pub fn then(mut self, f: impl Fn(T) -> T + Send + Sync + 'static) -> Self { + self.stages.push(Box::new(f)); + self + } + + pub fn run(&self, input: T) -> T { + self.stages.iter().fold(input, |acc, f| f(acc)) + } +} + +impl Default for Pipeline { + fn default() -> Self { + Self::new() + } +} + +/// Patterns: slices, ranges, bindings, guards, nested enums. +pub fn classify(values: &[i32]) -> &'static str { + match values { + [] => "empty", + [x] if *x < 0 => "one negative", + [_] => "one", + [first, .., last] if first == last => "same ends", + [0..=9, rest @ ..] if rest.len() > 2 => "small head, long", + [_, _] => "pair", + _ => "other", + } +} + +#[derive(Debug, Clone, PartialEq, Eq, Hash, PartialOrd, Ord)] +pub enum Token { + Num(i64), + Word(String), + Op(char), + Group(Vec), +} + +pub fn eval(tokens: &[Token]) -> Option { + let mut stack: Vec = Vec::new(); + for token in tokens { + match token { + Token::Num(n) => stack.push(*n), + Token::Op(op @ ('+' | '-' | '*')) => { + let (Some(b), Some(a)) = (stack.pop(), stack.pop()) else { + return None; + }; + stack.push(match op { + '+' => a + b, + '-' => a - b, + _ => a * b, + }); + } + Token::Group(inner) => stack.push(eval(inner)?), + Token::Word(w) if w == "dup" => { + let top = *stack.last()?; + stack.push(top); + } + Token::Word(_) | Token::Op(_) => return None, + } + } + stack.pop() +} + +/// Labelled breaks, loops with values, `if let` chains. +pub fn find_pair(grid: &[Vec], target: i32) -> Option<(usize, usize)> { + let found = 'outer: { + for (i, row) in grid.iter().enumerate() { + for (j, &v) in row.iter().enumerate() { + if v == target { + break 'outer Some((i, j)); + } + } + } + None + }; + if let Some((i, j)) = found + && i <= j + { + return Some((i, j)); + } + found +} + +pub fn collatz(mut n: u64) -> u32 { + let mut steps = 0; + loop { + if n == 1 { + break steps; + } + n = if n % 2 == 0 { n / 2 } else { 3 * n + 1 }; + steps += 1; + } +} + +/// Iterator adaptors and a custom iterator. +#[derive(Debug)] +pub struct Fib { + a: u64, + b: u64, +} + +impl Iterator for Fib { + type Item = u64; + fn next(&mut self) -> Option { + let r = self.a; + self.a = self.b; + self.b += r; + Some(r) + } +} + +pub fn fib() -> Fib { + Fib { a: 0, b: 1 } +} + +pub fn stats(words: &str) -> BTreeMap> { + let mut by_len: BTreeMap> = BTreeMap::new(); + for w in words.split_whitespace().filter(|w| !w.is_empty()) { + by_len.entry(w.len()).or_default().push(w); + } + by_len.values_mut().for_each(|v| v.sort_unstable()); + by_len +} + +pub fn histogram(values: impl IntoIterator) -> HashMap { + let mut h = HashMap::new(); + for v in values { + *h.entry(v).or_insert(0) += 1; + } + h +} + +pub fn unique_sorted(items: &[T]) -> Vec { + items.iter().cloned().collect::>().into_iter().collect() +} + +pub fn rotate(mut q: VecDeque, n: usize) -> VecDeque { + q.rotate_left(n % q.len().max(1)); + q +} + +/// Cow, Rc/RefCell, Arc/Mutex and threads. +pub fn normalize(s: &str) -> Cow<'_, str> { + if s.chars().any(char::is_uppercase) { Cow::Owned(s.to_lowercase()) } else { Cow::Borrowed(s) } +} + +#[derive(Debug, Default)] +pub struct Node { + pub value: i32, + pub children: Vec>>, +} + +pub fn tree_sum(node: &Rc>) -> i32 { + let n = node.borrow(); + n.value + n.children.iter().map(tree_sum).sum::() +} + +pub fn parallel_sum(chunks: Vec>) -> u64 { + let total = Arc::new(Mutex::new(0)); + std::thread::scope(|s| { + for chunk in &chunks { + let total = Arc::clone(&total); + s.spawn(move || *total.lock().unwrap() += chunk.iter().sum::()); + } + }); + *total.lock().unwrap() +} + +/// Closures capturing by move, by ref and by mut ref; returning closures. +pub fn counter() -> impl FnMut() -> u32 { + let mut n = 0; + move || { + n += 1; + n + } +} + +pub fn compose(f: impl Fn(A) -> B, g: impl Fn(B) -> C) -> impl Fn(A) -> C { + move |x| g(f(x)) +} + +/// Sorting with closures and keys. +pub fn sort_people(people: &mut [(String, u32)]) { + people.sort_by(|a, b| b.1.cmp(&a.1).then_with(|| a.0.cmp(&b.0))); +} diff --git a/fixtures/sink/core/src/asyncs.rs b/fixtures/sink/core/src/asyncs.rs new file mode 100644 index 0000000..fc21e34 --- /dev/null +++ b/fixtures/sink/core/src/asyncs.rs @@ -0,0 +1,69 @@ +//! async fn, async blocks, async closures, async fn in traits, a tiny executor. + +use core::future::Future; +use core::pin::{Pin, pin}; +use core::task::{Context, Poll, Waker}; + +pub fn block_on(future: F) -> F::Output { + let mut future = pin!(future); + let mut cx = Context::from_waker(Waker::noop()); + loop { + if let Poll::Ready(v) = future.as_mut().poll(&mut cx) { + return v; + } + } +} + +/// Yields once before finishing. +#[derive(Debug, Default)] +pub struct YieldOnce(bool); + +impl Future for YieldOnce { + type Output = (); + fn poll(mut self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<()> { + if self.0 { + Poll::Ready(()) + } else { + self.0 = true; + Poll::Pending + } + } +} + +pub async fn add_later(a: u32, b: u32) -> u32 { + YieldOnce::default().await; + a + b +} + +#[allow(async_fn_in_trait)] +pub trait Fetch { + async fn fetch(&self, key: &str) -> Option; +} + +#[derive(Debug)] +pub struct Static(pub &'static [(&'static str, &'static str)]); + +impl Fetch for Static { + async fn fetch(&self, key: &str) -> Option { + YieldOnce::default().await; + self.0.iter().find(|(k, _)| *k == key).map(|(_, v)| v.to_string()) + } +} + +/// An async closure that borrows what it captures. +pub fn greeter(name: String) -> impl AsyncFn(&str) -> String { + async move |greeting: &str| format!("{greeting}, {name}") +} + +pub async fn call_twice(f: impl AsyncFn(&str) -> String) -> String { + let a = f("hello").await; + let b = f("bye").await; + format!("{a}; {b}") +} + +pub fn boxed_future(n: u32) -> Pin + Send>> { + Box::pin(async move { + YieldOnce::default().await; + n * 2 + }) +} diff --git a/fixtures/sink/core/src/consts.rs b/fixtures/sink/core/src/consts.rs new file mode 100644 index 0000000..7a9a6c1 --- /dev/null +++ b/fixtures/sink/core/src/consts.rs @@ -0,0 +1,89 @@ +//! Const generics, const fn, const evaluation, inline const, associated consts. + +#[derive(Debug, Clone, Copy, PartialEq)] +pub struct Matrix { + pub cells: [[i64; C]; R], +} + +impl Matrix { + pub const ZERO: Self = Matrix { cells: [[0; C]; R] }; + pub const SIZE: usize = R * C; + + pub const fn identity_like() -> Self { + let mut m = Self::ZERO; + let mut i = 0; + while i < R && i < C { + m.cells[i][i] = 1; + i += 1; + } + m + } + + pub fn transpose(&self) -> Matrix { + let mut t = Matrix::::ZERO; + for i in 0..R { + for j in 0..C { + t.cells[j][i] = self.cells[i][j]; + } + } + t + } + + pub fn mul(&self, other: &Matrix) -> Matrix { + let mut out = Matrix::::ZERO; + for i in 0..R { + for k in 0..K { + out.cells[i][k] = (0..C).map(|j| self.cells[i][j] * other.cells[j][k]).sum(); + } + } + out + } +} + +pub const fn factorial(n: u64) -> u64 { + if n == 0 { 1 } else { n * factorial(n - 1) } +} + +pub const FACT_10: u64 = factorial(10); +pub static TABLE: [u64; 6] = { + let mut t = [0; 6]; + let mut i = 0; + while i < 6 { + t[i] = factorial(i as u64); + i += 1; + } + t +}; + +pub fn first_n(v: &[u8]) -> Option<[u8; N]> { + v.get(..N)?.try_into().ok() +} + +pub fn inline_const() -> [Vec; 3] { + [const { Vec::new() }; 3] +} + +pub trait HasId { + const ID: u32; + fn id(&self) -> u32 { + Self::ID + } +} + +#[derive(Debug)] +pub struct A; +#[derive(Debug)] +pub struct B; +impl HasId for A { + const ID: u32 = 1; +} +impl HasId for B { + const ID: u32 = A::ID + 1; +} + +pub const STR: &str = "a constant string"; +pub const BYTES: &[u8] = b"bytes\x00\xff"; +pub const C_STR: &core::ffi::CStr = c"c string"; +pub const WIDE: i128 = i128::MAX / 3; +pub const FLOAT: f64 = 1.0 / 3.0; +pub const CHAR: char = '\u{1F980}'; diff --git a/fixtures/sink/core/src/errors.rs b/fixtures/sink/core/src/errors.rs new file mode 100644 index 0000000..f49d947 --- /dev/null +++ b/fixtures/sink/core/src/errors.rs @@ -0,0 +1,57 @@ +//! Error types, `?`, `From` conversions, `dyn Error`, `Result` chains. + +use core::fmt; +use std::error::Error; +use std::num::ParseIntError; + +#[derive(Debug)] +pub enum SinkError { + Parse(ParseIntError), + Range { value: i64, max: i64 }, + Other(Box), +} + +impl fmt::Display for SinkError { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + match self { + SinkError::Parse(e) => write!(f, "parse: {e}"), + SinkError::Range { value, max } => write!(f, "{value} is over {max}"), + SinkError::Other(e) => write!(f, "other: {e}"), + } + } +} + +impl Error for SinkError { + fn source(&self) -> Option<&(dyn Error + 'static)> { + match self { + SinkError::Parse(e) => Some(e), + _ => None, + } + } +} + +impl From for SinkError { + fn from(e: ParseIntError) -> Self { + SinkError::Parse(e) + } +} + +pub fn parse_bounded(s: &str, max: i64) -> Result { + let value: i64 = s.trim().parse()?; + if value > max { + return Err(SinkError::Range { value, max }); + } + Ok(value) +} + +pub fn sum_all(items: &[&str]) -> Result> { + let mut total = 0; + for item in items { + total += parse_bounded(item, 1000)?; + } + Ok(total) +} + +pub fn first_char_upper(s: &str) -> Option { + Some(s.chars().next()?.to_ascii_uppercase()) +} diff --git a/fixtures/sink/core/src/lib.rs b/fixtures/sink/core/src/lib.rs new file mode 100644 index 0000000..e0cf785 --- /dev/null +++ b/fixtures/sink/core/src/lib.rs @@ -0,0 +1,51 @@ +//! The bottom of the sink: as many stable language features as fit, for +//! dependents to read through metadata. +#![warn(missing_debug_implementations)] +#![allow(dead_code)] + +extern crate alloc; + +pub mod algo; +pub mod asyncs; +pub mod consts; +pub mod errors; +pub mod memory; +pub mod shapes; +#[macro_use] +mod macros; + +pub use shapes::{Area, Circle, Shape, Square}; + +include!(concat!(env!("OUT_DIR"), "/generated.rs")); + +use core::sync::atomic::AtomicUsize; + +/// Counted by `#[sink_macros::traced]`. +pub static TRACE: AtomicUsize = AtomicUsize::new(0); + +/// Implemented by `#[derive(sink_macros::Describe)]`. +pub trait Describe { + const NAME: &'static str; + const FIELDS: usize; + + fn tag(&self) -> &'static str; + + fn describe(&self) -> String { + format!("{} with {} fields, tagged {}", Self::NAME, Self::FIELDS, self.tag()) + } +} + +#[cfg(sink_generated)] +pub fn generated() -> bool { + true +} + +#[cfg(feature = "extra")] +pub fn extra() -> &'static str { + "extra" +} + +/// A doc link to [`Shape`], [`algo::Pipeline`] and [`crate::consts::Matrix`]. +pub fn version() -> (u32, &'static str) { + (PRIMES[3], GENERATED_NAME) +} diff --git a/fixtures/sink/core/src/macros.rs b/fixtures/sink/core/src/macros.rs new file mode 100644 index 0000000..844070f --- /dev/null +++ b/fixtures/sink/core/src/macros.rs @@ -0,0 +1,48 @@ +//! Declarative macros: recursive, exported, with hygiene. + +/// Sums its arguments at compile time. +#[macro_export] +macro_rules! sum { + () => { 0 }; + ($head:expr $(, $tail:expr)* $(,)?) => { $head + $crate::sum!($($tail),*) }; +} + +/// Builds a `HashMap` from `key => value` pairs. +#[macro_export] +macro_rules! map { + ($($key:expr => $value:expr),* $(,)?) => {{ + let mut map = ::std::collections::HashMap::new(); + $(map.insert($key, $value);)* + map + }}; +} + +/// Defines a newtype with `Deref`, used within this crate only. +macro_rules! newtype { + ($(#[$attr:meta])* $vis:vis $name:ident($inner:ty)) => { + $(#[$attr])* + #[derive(Debug, Clone, PartialEq)] + $vis struct $name(pub $inner); + + impl ::core::ops::Deref for $name { + type Target = $inner; + fn deref(&self) -> &$inner { + &self.0 + } + } + }; +} + +newtype!( + /// Metres, as a newtype. + pub Metres(f64) +); + +/// Hygiene: the `x` inside does not see the caller's `x`. +#[macro_export] +macro_rules! shadow { + ($e:expr) => {{ + let x = 10; + $e + x + }}; +} diff --git a/fixtures/sink/core/src/memory.rs b/fixtures/sink/core/src/memory.rs new file mode 100644 index 0000000..b1e9b60 --- /dev/null +++ b/fixtures/sink/core/src/memory.rs @@ -0,0 +1,131 @@ +//! Unsafe code, raw pointers, unions, repr, Drop, statics, thread locals, FFI. + +use core::cell::Cell; +use core::mem::{ManuallyDrop, MaybeUninit}; +use core::sync::atomic::{AtomicU64, Ordering}; + +#[repr(C)] +#[derive(Clone, Copy)] +pub union IntOrFloat { + pub i: u64, + pub f: f64, +} + +impl core::fmt::Debug for IntOrFloat { + fn fmt(&self, f: &mut core::fmt::Formatter<'_>) -> core::fmt::Result { + write!(f, "IntOrFloat({:#x})", unsafe { self.i }) + } +} + +pub fn float_bits(x: f64) -> u64 { + unsafe { IntOrFloat { f: x }.i } +} + +#[repr(u8)] +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum Level { + Low = 1, + Mid = 5, + High = 9, +} + +#[non_exhaustive] +#[derive(Debug)] +pub struct Config { + pub level: Level, + pub name: &'static str, +} + +impl Config { + pub fn new() -> Self { + Config { level: Level::Mid, name: "default" } + } +} + +impl Default for Config { + fn default() -> Self { + Self::new() + } +} + +pub static DROPS: AtomicU64 = AtomicU64::new(0); + +#[derive(Debug)] +pub struct Loud(pub u32); + +impl Drop for Loud { + fn drop(&mut self) { + DROPS.fetch_add(u64::from(self.0), Ordering::SeqCst); + } +} + +thread_local! { + pub static CALLS: Cell = const { Cell::new(0) }; +} + +pub fn count_call() -> u32 { + CALLS.with(|c| { + c.set(c.get() + 1); + c.get() + }) +} + +/// Raw pointers and slices built by hand. +pub fn reverse_in_place(v: &mut [u32]) { + let len = v.len(); + let p = v.as_mut_ptr(); + for i in 0..len / 2 { + unsafe { core::ptr::swap(p.add(i), p.add(len - 1 - i)) }; + } +} + +pub fn init_array() -> [u16; 4] { + let mut a: [MaybeUninit; 4] = [MaybeUninit::uninit(); 4]; + for (i, slot) in a.iter_mut().enumerate() { + slot.write(i as u16 * 7); + } + unsafe { core::mem::transmute::<[MaybeUninit; 4], [u16; 4]>(a) } +} + +pub fn keep_alive(x: String) -> usize { + let m = ManuallyDrop::new(x); + let n = m.len(); + drop(ManuallyDrop::into_inner(m)); + n +} + +/// Exported with a C ABI. +#[unsafe(no_mangle)] +pub extern "C" fn sink_core_add(a: i32, b: i32) -> i32 { + a.wrapping_add(b) +} + +unsafe extern "C" { + safe fn abs(x: i32) -> i32; +} + +pub fn c_abs(x: i32) -> i32 { + abs(x) +} + +#[track_caller] +pub fn caller_line() -> u32 { + core::panic::Location::caller().line() +} + +#[must_use] +#[inline(always)] +pub fn always_inline(x: u32) -> u32 { + x.rotate_left(3) +} + +#[inline(never)] +#[cold] +pub fn never_inline(x: u32) -> u32 { + x.reverse_bits() +} + +#[deprecated(since = "0.0.0", note = "use `always_inline`")] +pub fn old(x: u32) -> u32 { + always_inline(x) +} diff --git a/fixtures/sink/core/src/shapes.rs b/fixtures/sink/core/src/shapes.rs new file mode 100644 index 0000000..a56b4fe --- /dev/null +++ b/fixtures/sink/core/src/shapes.rs @@ -0,0 +1,198 @@ +//! Traits with associated types and consts, default methods, supertraits, +//! generic and blanket impls, trait objects and upcasting, operators. + +use core::fmt; +use core::ops::{Add, Index, Mul}; + +pub trait Area { + fn area(&self) -> f64; +} + +pub trait Shape: Area + fmt::Debug { + type Unit: Copy + fmt::Display; + + fn name(&self) -> &'static str; + + fn perimeter(&self) -> f64 { + 0.0 + } + + fn scaled(&self, by: f64) -> Box> + where + Self: Sized + Clone + 'static, + { + let _ = by; + Box::new(self.clone()) + } +} + +#[derive(Debug, Clone, Copy, PartialEq, PartialOrd, Default)] +pub struct Square { + pub side: f64, +} + +#[derive(Debug, Clone, Copy, PartialEq, PartialOrd, Default)] +pub struct Circle { + pub radius: f64, +} + +impl Area for Square { + fn area(&self) -> f64 { + self.side * self.side + } +} + +impl Area for Circle { + fn area(&self) -> f64 { + core::f64::consts::PI * self.radius * self.radius + } +} + +impl Shape for Square { + type Unit = f64; + + fn name(&self) -> &'static str { + "square" + } + + fn perimeter(&self) -> f64 { + 4.0 * self.side + } +} + +impl Shape for Circle { + type Unit = f64; + + fn name(&self) -> &'static str { + "circle" + } +} + +/// An associated const, kept apart from `Shape` so that `dyn Shape` is allowed. +pub trait Sides { + const SIDES: u32; +} + +impl Sides for Square { + const SIDES: u32 = 4; +} + +/// A blanket impl over references. +impl Area for &T { + fn area(&self) -> f64 { + (**self).area() + } +} + +/// A blanket impl over boxes of trait objects. +impl Area for Box> { + fn area(&self) -> f64 { + (**self).area() + } +} + +/// Upcasting a trait object to its supertrait. +pub fn as_area(shape: &dyn Shape) -> &dyn Area { + shape +} + +pub fn total_area<'a, I>(shapes: I) -> f64 +where + I: IntoIterator>, +{ + shapes.into_iter().map(|s| s.area()).sum() +} + +/// A 2D vector with operators. +#[derive(Debug, Clone, Copy, PartialEq, Default)] +pub struct V2 { + pub x: f64, + pub y: f64, +} + +impl Add for V2 { + type Output = V2; + fn add(self, o: V2) -> V2 { + V2 { x: self.x + o.x, y: self.y + o.y } + } +} + +impl Mul for V2 { + type Output = V2; + fn mul(self, k: f64) -> V2 { + V2 { x: self.x * k, y: self.y * k } + } +} + +impl Index for V2 { + type Output = f64; + fn index(&self, i: usize) -> &f64 { + match i { + 0 => &self.x, + 1 => &self.y, + _ => panic!("V2 has two components"), + } + } +} + +impl fmt::Display for V2 { + fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result { + write!(f, "({:.1}, {:.1})", self.x, self.y) + } +} + +impl From<(f64, f64)> for V2 { + fn from((x, y): (f64, f64)) -> V2 { + V2 { x, y } + } +} + +/// A trait with a generic method and a generic associated type. +pub trait Container { + type Item<'a> + where + Self: 'a; + type Iter<'a>: Iterator> + where + Self: 'a; + + fn items(&self) -> Self::Iter<'_>; + + fn first_matching) -> bool>(&self, mut p: P) -> Option> { + self.items().find(|i| p(i)) + } +} + +impl Container for Vec { + type Item<'a> + = &'a T + where + T: 'a; + type Iter<'a> + = core::slice::Iter<'a, T> + where + T: 'a; + + fn items(&self) -> Self::Iter<'_> { + self.iter() + } +} + +/// Return-position `impl Trait` in a trait. +pub trait Labels { + fn labels(&self) -> impl Iterator + '_; +} + +impl Labels for [Square] { + fn labels(&self) -> impl Iterator + '_ { + self.iter().map(|s| format!("square {}", s.side)) + } +} + +/// Higher-ranked trait bounds. +pub fn apply_to_all(items: &[String], f: F) -> Vec +where + F: for<'a> Fn(&'a str) -> &'a str, +{ + items.iter().map(|s| f(s).len()).collect() +} diff --git a/fixtures/sink/dy/Cargo.toml b/fixtures/sink/dy/Cargo.toml new file mode 100644 index 0000000..310796d --- /dev/null +++ b/fixtures/sink/dy/Cargo.toml @@ -0,0 +1,11 @@ +[package] +name = "sink-dy" +version = "0.0.0" +edition = "2024" +publish = false + +[lib] +crate-type = ["rlib", "dylib"] + +[dependencies] +sink-core = { path = "../core" } diff --git a/fixtures/sink/dy/src/lib.rs b/fixtures/sink/dy/src/lib.rs new file mode 100644 index 0000000..1671d4f --- /dev/null +++ b/fixtures/sink/dy/src/lib.rs @@ -0,0 +1,11 @@ +//! Built as a dylib too, so its metadata is also embedded in a shared library. + +use sink_core::consts::Matrix; + +pub fn spin(m: &Matrix<2, 3>) -> Matrix<3, 2> { + m.transpose() +} + +pub fn twice(v: u32) -> u32 { + sink_core::memory::always_inline(v) * 2 +} diff --git a/fixtures/sink/macros/Cargo.toml b/fixtures/sink/macros/Cargo.toml new file mode 100644 index 0000000..626a8fb --- /dev/null +++ b/fixtures/sink/macros/Cargo.toml @@ -0,0 +1,8 @@ +[package] +name = "sink-macros" +version = "0.0.0" +edition = "2024" +publish = false + +[lib] +proc-macro = true diff --git a/fixtures/sink/macros/src/lib.rs b/fixtures/sink/macros/src/lib.rs new file mode 100644 index 0000000..4137339 --- /dev/null +++ b/fixtures/sink/macros/src/lib.rs @@ -0,0 +1,80 @@ +//! Procedural macros without dependencies: a derive with a helper attribute, +//! an attribute macro and a function-like macro. + +use proc_macro::{Delimiter, Group, Ident, Literal, Punct, Spacing, Span, TokenStream, TokenTree}; + +fn item_name(input: TokenStream) -> (String, usize) { + let mut tokens = input.into_iter().peekable(); + let mut fields = 0; + let mut name = None; + while let Some(token) = tokens.next() { + match &token { + TokenTree::Ident(ident) if name.is_none() => { + let word = ident.to_string(); + if word == "struct" || word == "enum" { + name = tokens.next().map(|t| t.to_string()); + } + } + TokenTree::Group(group) if name.is_some() && group.delimiter() == Delimiter::Brace => { + fields = group.stream().into_iter().filter(|t| matches!(t, TokenTree::Punct(p) if p.as_char() == ':')).count(); + } + _ => {} + } + } + (name.expect("a struct or an enum"), fields) +} + +/// `#[derive(Describe)]` implements `sink_core::Describe`, with the name, the +/// number of named fields, and an optional `#[describe(tag = "...")]`. +#[proc_macro_derive(Describe, attributes(describe))] +pub fn derive_describe(input: TokenStream) -> TokenStream { + let text = input.to_string(); + let tag = text + .split("tag = \"") + .nth(1) + .and_then(|rest| rest.split('"').next()) + .unwrap_or("untagged") + .to_string(); + let (name, fields) = item_name(input); + format!( + "impl ::sink_core::Describe for {name} {{ + const NAME: &'static str = \"{name}\"; + const FIELDS: usize = {fields}; + fn tag(&self) -> &'static str {{ \"{tag}\" }} + }}" + ) + .parse() + .unwrap() +} + +/// `#[traced]` makes a function count its calls in `sink_core::TRACE`. +#[proc_macro_attribute] +pub fn traced(_attr: TokenStream, item: TokenStream) -> TokenStream { + let mut out: Vec = item.into_iter().collect(); + if let Some(TokenTree::Group(body)) = out.pop() { + let prefix: TokenStream = + "::sink_core::TRACE.fetch_add(1, ::core::sync::atomic::Ordering::Relaxed);".parse().unwrap(); + let mut stream = prefix; + stream.extend(body.stream()); + let mut group = Group::new(Delimiter::Brace, stream); + group.set_span(body.span()); + out.push(TokenTree::Group(group)); + } + out.into_iter().collect() +} + +/// `squares!(a, b, c)` expands to an array of the squares, built token by token. +#[proc_macro] +pub fn squares(input: TokenStream) -> TokenStream { + let mut items = Vec::new(); + for token in input { + if let TokenTree::Literal(lit) = token { + let n: i64 = lit.to_string().trim_end_matches("i64").parse().unwrap(); + items.push(TokenTree::Literal(Literal::i64_suffixed(n * n))); + items.push(TokenTree::Punct(Punct::new(',', Spacing::Alone))); + } + } + let group = Group::new(Delimiter::Bracket, items.into_iter().collect()); + let _ = Ident::new("unused", Span::call_site()); + TokenStream::from(TokenTree::Group(group)) +} diff --git a/fixtures/sink/mid/Cargo.toml b/fixtures/sink/mid/Cargo.toml new file mode 100644 index 0000000..f4282a3 --- /dev/null +++ b/fixtures/sink/mid/Cargo.toml @@ -0,0 +1,9 @@ +[package] +name = "sink-mid" +version = "0.0.0" +edition = "2024" +publish = false + +[dependencies] +sink-core = { path = "../core" } +sink-macros = { path = "../macros" } diff --git a/fixtures/sink/mid/src/lib.rs b/fixtures/sink/mid/src/lib.rs new file mode 100644 index 0000000..beade37 --- /dev/null +++ b/fixtures/sink/mid/src/lib.rs @@ -0,0 +1,115 @@ +//! The middle: implements `sink-core`'s traits, uses all three kinds of +//! procedural macro and the exported declarative macros, re-exports, and +//! generic code that dependents instantiate. + +pub use sink_core::shapes::{Container, Labels, V2}; +pub use sink_core::{Area, Describe, Shape, Square, sum}; + +use sink_core::asyncs::Fetch; +use sink_macros::{Describe, squares, traced}; + +pub mod nested { + pub mod deeper { + pub(crate) fn hidden() -> u8 { + 7 + } + pub(in crate::nested) fn semi() -> u8 { + super::super::nested::deeper::hidden() + 1 + } + pub fn visible() -> u8 { + semi() * 2 + } + } + pub use self::deeper::visible as shown; +} + +/// A shape defined downstream, with a derive and a helper attribute. +#[derive(Debug, Clone, PartialEq, Describe)] +#[describe(tag = "polygon")] +pub struct Polygon { + pub sides: u32, + pub length: f64, + pub label: &'static str, +} + +#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, PartialOrd, Ord, Describe)] +pub enum Colour { + Red, + Green, + Blue, +} + +impl Area for Polygon { + fn area(&self) -> f64 { + let n = f64::from(self.sides); + n * self.length * self.length / (4.0 * (core::f64::consts::PI / n).tan()) + } +} + +impl Shape for Polygon { + type Unit = f64; + + fn name(&self) -> &'static str { + self.label + } + + fn perimeter(&self) -> f64 { + f64::from(self.sides) * self.length + } +} + +#[traced] +pub fn traced_square(x: i64) -> i64 { + x * x +} + +pub fn squared_table() -> [i64; 4] { + squares!(1, 2, 3, 4) +} + +pub fn macro_sum() -> i32 { + sum!(1, 2, 3, 4, 5) +} + +pub fn shadowing() -> i32 { + let x = 1; + sink_core::shadow!(x) +} + +/// Generic code that dependents instantiate, with `#[inline]` so its MIR is in the metadata. +#[inline] +pub fn largest(shapes: &[&S]) -> Option { + shapes.iter().map(|s| s.area()).reduce(f64::max) +} + +pub struct Wrapper { + pub items: [T; N], +} + +impl Wrapper { + pub fn filled(t: T) -> Self { + Wrapper { items: core::array::from_fn(|_| t.clone()) } + } + + pub fn describe(&self) -> String { + format!("{} items: {:?}", N, self.items) + } +} + +pub async fn lookup(store: &impl Fetch, keys: &[&str]) -> Vec { + let mut out = Vec::new(); + for key in keys { + if let Some(v) = store.fetch(key).await { + out.push(v); + } + } + out +} + +pub fn labels(squares: &[Square]) -> Vec { + squares.labels().collect() +} + +pub fn colours() -> Vec<(&'static str, &'static str)> { + [Colour::Red, Colour::Green, Colour::Blue].iter().map(|c| (Colour::NAME, c.tag())).collect() +} diff --git a/fixtures/sink/src/main.rs b/fixtures/sink/src/main.rs new file mode 100644 index 0000000..3da1c57 --- /dev/null +++ b/fixtures/sink/src/main.rs @@ -0,0 +1,111 @@ +//! Runs everything and checks the results, so a miscompile shows as a failure. + +use std::collections::VecDeque; + +use sink_core::algo::{self, Token}; +use sink_core::asyncs::{self, Static}; +use sink_core::consts::{self, HasId, Matrix}; +use sink_core::memory; +use sink_core::shapes::{self, Container, V2}; +use sink_core::{Circle, Describe, Shape, Square, map}; +use sink_mid::{Colour, Polygon}; + +fn main() { + let mut checks: Vec<(&str, bool)> = Vec::new(); + let mut check = |name, ok| checks.push((name, ok)); + + let sq = Square { side: 2.0 }; + let ci = Circle { radius: 1.0 }; + let poly = Polygon { sides: 6, length: 1.0, label: "hexagon" }; + let shapes: Vec<&dyn Shape> = vec![&sq, &ci, &poly]; + check("total area", (shapes::total_area(shapes.iter().copied()) - 9.739).abs() < 0.01); + check("upcast", shapes::as_area(&sq).area() == 4.0); + check("largest", sink_mid::largest(&[&sq, &sq]) == Some(4.0)); + check("perimeter", poly.perimeter() == 6.0 && sq.perimeter() == 8.0); + check("derive", poly.describe() == "Polygon with 3 fields, tagged polygon"); + check("derive enum", Colour::NAME == "Colour" && sink_mid::colours().len() == 3); + check("operators", (V2 { x: 1.0, y: 2.0 } + V2::from((1.0, 1.0))) * 2.0 == V2 { x: 4.0, y: 6.0 }); + check("index", V2 { x: 3.0, y: 4.0 }[1] == 4.0); + check("gat", vec![1, 2, 3].first_matching(|x| **x > 1) == Some(&2)); + check("rpitit", sink_mid::labels(&[sq]) == vec!["square 2".to_string()]); + check("hrtb", shapes::apply_to_all(&["ab".into(), "cde".into()], |s| s.trim()) == vec![2, 3]); + + check("pipeline", algo::Pipeline::new().then(|x: i32| x + 1).then(|x| x * 10).run(1) == 20); + check("classify", algo::classify(&[1, 5, 6, 7]) == "small head, long" && algo::classify(&[3, 9, 3]) == "same ends"); + let program = [Token::Num(2), Token::Group(vec![Token::Num(3), Token::Word("dup".into()), Token::Op('*')]), Token::Op('+')]; + check("eval", algo::eval(&program) == Some(11)); + check("labelled break", algo::find_pair(&[vec![1, 2], vec![3, 4]], 4) == Some((1, 1))); + check("loop value", algo::collatz(27) == 111); + check("fib", algo::fib().skip(10).next() == Some(55)); + check("btree", algo::stats("a bb cc d").get(&2).map(Vec::len) == Some(2)); + check("histogram", algo::histogram([1, 1, 2]).get(&1) == Some(&2)); + check("unique", algo::unique_sorted(&[3, 1, 3, 2]) == vec![1, 2, 3]); + check("deque", algo::rotate(VecDeque::from(vec![1, 2, 3]), 1) == VecDeque::from(vec![2, 3, 1])); + check("cow", algo::normalize("Hi") == "hi" && matches!(algo::normalize("hi"), std::borrow::Cow::Borrowed(_))); + let leaf = std::rc::Rc::new(std::cell::RefCell::new(algo::Node { value: 2, children: vec![] })); + let root = std::rc::Rc::new(std::cell::RefCell::new(algo::Node { value: 1, children: vec![leaf.clone(), leaf] })); + check("rc tree", algo::tree_sum(&root) == 5); + check("threads", algo::parallel_sum(vec![vec![1, 2], vec![3], vec![4, 5, 6]]) == 21); + let mut c = algo::counter(); + c(); + check("closure state", c() == 2); + check("compose", algo::compose(|x: u8| x as u32 + 1, |y| y * 3)(4) == 15); + let mut people = vec![("b".to_string(), 3), ("a".to_string(), 3), ("c".to_string(), 9)]; + algo::sort_people(&mut people); + check("sort", people[0].0 == "c" && people[1].0 == "a"); + + check("async", asyncs::block_on(asyncs::add_later(2, 3)) == 5); + let store = Static(&[("k", "v"), ("x", "y")]); + check("afit", asyncs::block_on(sink_mid::lookup(&store, &["x", "missing", "k"])) == vec!["y", "v"]); + check("async closure", asyncs::block_on(asyncs::call_twice(asyncs::greeter("sink".into()))) == "hello, sink; bye, sink"); + check("boxed future", asyncs::block_on(asyncs::boxed_future(21)) == 42); + + let m = Matrix::<2, 3> { cells: [[1, 2, 3], [4, 5, 6]] }; + check("const generics", m.mul(&m.transpose()).cells == [[14, 32], [32, 77]]); + check("dylib", sink_dy::spin(&m).cells[2] == [3, 6] && sink_dy::twice(1) == 16); + check("const fn", consts::FACT_10 == 3_628_800 && consts::TABLE[5] == 120); + check("identity", Matrix::<3, 3>::identity_like().cells[2][2] == 1 && Matrix::<2, 5>::SIZE == 10); + check("first_n", consts::first_n::<2>(&[9, 8, 7]) == Some([9, 8])); + check("inline const", consts::inline_const().iter().all(Vec::is_empty)); + check("assoc const", consts::B.id() == 2); + check("literals", consts::BYTES.len() == 7 && consts::C_STR.to_bytes().len() == 8 && consts::CHAR.len_utf8() == 4); + + check("union", memory::float_bits(1.0) == 0x3ff0_0000_0000_0000); + check("repr", memory::Level::High as u8 == 9); + { + let _a = memory::Loud(2); + let _b = memory::Loud(3); + } + check("drop", memory::DROPS.load(std::sync::atomic::Ordering::SeqCst) == 5); + memory::count_call(); + check("thread local", memory::count_call() == 2); + let mut v = vec![1, 2, 3, 4, 5]; + memory::reverse_in_place(&mut v); + check("raw pointers", v == [5, 4, 3, 2, 1]); + check("maybeuninit", memory::init_array() == [0, 7, 14, 21]); + check("manuallydrop", memory::keep_alive("four".into()) == 4); + check("ffi", memory::sink_core_add(2, 3) == 5 && memory::c_abs(-4) == 4); + check("track caller", memory::caller_line() == line!()); + check("inline attrs", memory::never_inline(memory::always_inline(1)) == 0x1000_0000); + + check("errors", sink_core::errors::sum_all(&["1", "2"]).ok() == Some(3)); + check("error kinds", sink_core::errors::parse_bounded("5000", 10).unwrap_err().to_string() == "5000 is over 10"); + check("try option", sink_core::errors::first_char_upper("abc") == Some('A')); + + check("proc macro attr", sink_mid::traced_square(3) == 9 && sink_core::TRACE.load(std::sync::atomic::Ordering::Relaxed) == 1); + check("proc macro fn", sink_mid::squared_table() == [1, 4, 9, 16]); + check("macro_rules", sink_mid::macro_sum() == 15 && sink_mid::shadowing() == 11); + check("map macro", map! { "a" => 1, "b" => 2 }.len() == 2); + check("build script", sink_core::generated() && sink_core::version() == (7, "sink")); + check("feature", sink_core::extra() == "extra"); + check("visibility", sink_mid::nested::shown() == 16); + check("wrapper", sink_mid::Wrapper::::filled(1).describe() == "2 items: [1, 1]"); + + let failed: Vec<_> = checks.iter().filter(|(_, ok)| !ok).map(|(name, _)| *name).collect(); + if failed.is_empty() { + println!("{} checks passed", checks.len()); + } else { + println!("failed: {failed:?}"); + std::process::exit(1); + } +} diff --git a/rustc/artifacts.py b/rustc/artifacts.py new file mode 100644 index 0000000..f8ebad7 --- /dev/null +++ b/rustc/artifacts.py @@ -0,0 +1,108 @@ +"""What a Cargo build produced, in a form two builds can be compared in. + +Shared by fuzz.py and replay.py. From the JSON messages of `cargo build +--message-format=json-render-diagnostics` it collects, for packages built +from a path: + + rmeta every .rmeta, by path + rlib every .rlib's members by name, with the incremental session suffix + of codegen-unit object names removed (`...rcgu.o`), + since it differs between sessions while the objects are identical + exe every executable + diag every diagnostic, rendered, counted per crate +""" + +import hashlib +import json +import re +from collections import Counter +from pathlib import Path + +SESSION = re.compile(rb"\.[0-9a-z]{7}\.rcgu\.o") + + +def ar_members(data): + """The members of a Unix ar archive (GNU format), as {name: bytes}.""" + if not data.startswith(b"!\n"): + return {"": data} + members, names, pos = {}, b"", 8 + while pos + 60 <= len(data): + header = data[pos:pos + 60] + name = header[:16].rstrip() + size = int(header[48:58].strip() or 0) + body = data[pos + 60:pos + 60 + size] + pos += 60 + size + (size & 1) + if name == b"//": + names = body + continue + if name in (b"/", b"/SYM64/"): + continue + if name.startswith(b"/") and name[1:].isdigit(): + offset = int(name[1:]) + name = names[offset:names.index(b"/\n", offset)] + name = name.rstrip(b"/") + members[SESSION.sub(b".rcgu.o", name).decode("utf-8", "replace")] = body + return members + + +def normalized_rlib(path): + out = {} + for name, body in ar_members(Path(path).read_bytes()).items(): + if name == "lib.rmeta-link": + body = SESSION.sub(b".rcgu.o", body) + out[name] = hashlib.sha256(body).hexdigest() + return out + + +def collect(stdout, target): + """Artifacts from cargo's JSON messages, keyed by path relative to `target`.""" + target = Path(target) + found = {"rmeta": {}, "rlib": {}, "exe": {}, "diag": Counter()} + for line in stdout.splitlines(): + try: + msg = json.loads(line) + except ValueError: + continue + if "path+file" not in msg.get("package_id", ""): + continue + if msg.get("reason") == "compiler-message": + m = msg.get("message", {}) + text = m.get("rendered") or m.get("message") or "" + found["diag"][(msg["target"]["name"], text)] += 1 + continue + if msg.get("reason") != "compiler-artifact": + continue + for f in msg.get("filenames", []): + rel = str(Path(f).relative_to(target)) if f.startswith(str(target)) else f + if f.endswith(".rmeta"): + found["rmeta"][rel] = hashlib.sha256(Path(f).read_bytes()).hexdigest() + elif f.endswith(".rlib"): + found["rlib"][rel] = normalized_rlib(f) + if msg.get("executable"): + f = msg["executable"] + rel = str(Path(f).relative_to(target)) if f.startswith(str(target)) else f + found["exe"][rel] = hashlib.sha256(Path(f).read_bytes()).hexdigest() + return found + + +def compare(a, b): + """{kind: [what differs]} for the kinds that differ between two collections.""" + out = {} + for kind in ("rmeta", "exe"): + diff = sorted(k for k in set(a[kind]) | set(b[kind]) if a[kind].get(k) != b[kind].get(k)) + if diff: + out[kind] = diff + diff = [] + for rel in sorted(set(a["rlib"]) | set(b["rlib"])): + ma, mb = a["rlib"].get(rel, {}), b["rlib"].get(rel, {}) + members = sorted(m for m in set(ma) | set(mb) if ma.get(m) != mb.get(m)) + if members: + diff.append(f"{rel}: {', '.join(members[:5])}{' …' if len(members) > 5 else ''}") + if diff: + out["rlib"] = diff + if a["diag"] != b["diag"]: + only_a = a["diag"] - b["diag"] + only_b = b["diag"] - a["diag"] + out["diag"] = [f"{c} only in the first: {t[:200]!r}" for (c, t), n in only_a.items()] + \ + [f"{c} only in the second: {t[:200]!r}" for (c, t), n in only_b.items()] + return out diff --git a/rustc/audit-options.py b/rustc/audit-options.py new file mode 100755 index 0000000..18d7dd0 --- /dev/null +++ b/rustc/audit-options.py @@ -0,0 +1,119 @@ +#!/usr/bin/env python3 +"""Audit rustc's [UNTRACKED] options for stale incremental reuse. + +An option rustc marks [UNTRACKED] is left out of the dependency-tracking hash, +so changing it between incremental sessions reuses the previous session's +results. That is only correct if the option cannot change them. For each +untracked option that takes no value or a boolean, this builds a crate +incrementally without it, then again with it, and compares the result with a +clean build that has it: the .rmeta, each .rlib member (object code, with the +incremental session suffix removed from names), the diagnostics, and the files +written. A difference means the option changes output that incremental +compilation reuses: it should be tracked, or the reuse checked. + + rustc/audit-options.py --rustc --source --crate [--edition 2018] [ARG…] + +With ARGs, audits those arguments instead of the options found in the +checkout's compiler/rustc_session/src/options.rs. The crate should have code of +its own for codegen options to change (non-generic functions) and a warning or +two for diagnostic options to change. +""" + +import argparse +import os +import re +import shutil +import subprocess +import sys +import tempfile +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import artifacts # noqa: E402 + +p = argparse.ArgumentParser() +p.add_argument("--rustc", required=True) +p.add_argument("--source", help="a rust checkout, to read the untracked options from") +p.add_argument("--crate", required=True, help="the crate root to build, as a library") +p.add_argument("--crate-name", default="audited") +p.add_argument("--edition", default="2021") +p.add_argument("args", nargs="*") +args = p.parse_args() + +# Options that stop compilation or change how arguments are read. +SKIP = {"-Chelp", "-Zhelp", "-Zno-analysis", "-Zparse-crate-root-only=yes", "-Zshell-argfiles=yes"} + + +def untracked_boolean_options(source): + text = (Path(source) / "compiler/rustc_session/src/options.rs").read_text() + codegen = re.search(r"options! \{\s*CodegenOptions,", text).start() + unstable = re.search(r"options! \{\s*UnstableOptions,", text).start() + found = [] + for m in re.finditer(r"^\s{4}(\w+):\s*([^=\n]+?)\s*=\s*\(([^,]*),\s*(parse_\w+),\s*\[UNTRACKED\]", text, re.M): + name, _ty, default, parser = m.groups() + if parser not in ("parse_bool", "parse_no_value", "parse_opt_bool"): + continue + if m.start() < codegen: + continue + group = "Z" if m.start() > unstable else "C" + flag = f"-{group}{name.replace('_', '-')}" + if parser != "parse_no_value": + flag += "=no" if default.strip() in ("true", "Some(true)") else "=yes" + if flag not in SKIP: + found.append(flag) + return found + + +crate = Path(args.crate).resolve() +name = args.crate_name + + +def build(work, incremental, out, extra): + """One build, always from the same working directory, which rustc records.""" + (work / out).mkdir(exist_ok=True) + r = subprocess.run([args.rustc, "--edition", args.edition, "--crate-type", "lib", "--crate-name", name, + "--emit=metadata,link", f"-Cincremental={work / incremental}", "--out-dir", str(work / out), + str(crate)] + extra, capture_output=True, text=True, cwd=work) + diagnostics = sorted(l for l in r.stderr.splitlines() if l.startswith(("warning", "error")) or "-->" in l) + return r.returncode, diagnostics + + +def audit(flag): + work = Path(tempfile.mkdtemp(prefix="audit-")) + try: + extra = [flag] if flag else [] + rc0, _ = build(work, "i", "o1", []) + rc1, diag_inc = build(work, "i", "o1", extra) + rc2, diag_clean = build(work, "j", "o2", extra) + if rc0 or rc2: + return [f"the crate does not build (exit {rc0} without the option, {rc2} with it)"] + problems = [] + if rc1 != rc2: + problems.append(f"the incremental rebuild exits {rc1}, a clean build {rc2}") + if (work / f"o1/lib{name}.rmeta").read_bytes() != (work / f"o2/lib{name}.rmeta").read_bytes(): + problems.append("metadata") + a = artifacts.normalized_rlib(work / f"o1/lib{name}.rlib") + b = artifacts.normalized_rlib(work / f"o2/lib{name}.rlib") + members = [m for m in set(a) | set(b) if a.get(m) != b.get(m)] + if members: + problems.append(f"object code ({len(members)} rlib members differ or exist on one side)") + if diag_inc != diag_clean: + problems.append(f"diagnostics ({len(diag_inc)} lines incrementally, {len(diag_clean)} clean)") + missing = set(os.listdir(work / "o2")) - set(os.listdir(work / "o1")) + if missing: + problems.append(f"{len(missing)} files only a clean build writes") + return problems + finally: + shutil.rmtree(work) + + +flags = args.args or untracked_boolean_options(args.source) +results = {} +control = audit("") +print(f"{'(control: no option)':40} {'; '.join(control) or 'same'}", flush=True) +if control: + sys.exit("the control differs: incremental and clean builds disagree without any option") +for flag in flags: + results[flag] = audit(flag) + print(f"{flag:40} {'; '.join(results[flag]) or 'same'}", flush=True) +sys.exit(1 if any(results.values()) else 0) diff --git a/rustc/fuzz-replay.py b/rustc/fuzz-replay.py new file mode 100755 index 0000000..e965b03 --- /dev/null +++ b/rustc/fuzz-replay.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Replay a fuzz finding: apply its edits to the pristine fixture in order, +with an incremental build after each, then compare with a clean build. + + rustc/fuzz-replay.py --rustc --fixture fixtures/sink --finding --work [--upto N] + +Prints which .rmeta files differ. --upto replays only the first N edits that +were kept, to find where the difference appears. +""" + +import argparse +import json +import os +import shutil +import subprocess +from pathlib import Path + +p = argparse.ArgumentParser() +p.add_argument("--rustc", required=True) +p.add_argument("--fixture", required=True) +p.add_argument("--finding", required=True) +p.add_argument("--work", required=True) +p.add_argument("--upto", type=int, default=None) +p.add_argument("--toolchain", default="nightly-2026-10-06") +p.add_argument("--rustflags", default="-Zincremental-verify-ich") +p.add_argument("--quiet", action="store_true") +args = p.parse_args() + +work = Path(args.work).resolve() +src, target, inc_target = work / "src", work / "target", work / "target-inc" +env = dict(os.environ, RUSTC=args.rustc, RUSTC_WRAPPER="", CARGO_INCREMENTAL="1", RUSTFLAGS=args.rustflags) + + +def build(t): + r = subprocess.run(["cargo", f"+{args.toolchain}", "build", "--workspace", "--offline", "-j", "4", + "--target-dir", str(t), "--message-format=json-render-diagnostics"], + cwd=src, env=env, capture_output=True, text=True) + rmetas = {} + for line in r.stdout.splitlines(): + try: + msg = json.loads(line) + except ValueError: + continue + if msg.get("reason") == "compiler-artifact": + for f in msg["filenames"]: + if f.endswith(".rmeta"): + rmetas[str(Path(f).relative_to(t))] = Path(f).read_bytes() + return r.returncode == 0, rmetas + + +if work.exists(): + shutil.rmtree(work) +work.mkdir(parents=True) +shutil.copytree(Path(args.fixture).resolve(), src, ignore=shutil.ignore_patterns("target", "edits", "edit")) +build(target) + +history = json.loads((Path(args.finding) / "history.json").read_text()) +kept = 0 +last = None +for step in history: + if step["edit"] == "revert": + (src / last["file"]).write_text(last["before"]) + else: + if args.upto is not None and step["kept"] and kept >= args.upto: + break + (src / step["file"]).write_text(step["after"]) + last = step + kept += step["kept"] + ok, _ = build(target) + if not args.quiet: + print(f"{step['edit']:18} {step['file']:24} {'built' if ok else 'failed'}") + +ok, inc = build(target) +target.rename(inc_target) +ok2, clean = build(target) +differ = sorted(r for r in set(inc) | set(clean) if inc.get(r) != clean.get(r)) +print(json.dumps({"kept_edits": kept, "inc_ok": ok, "clean_ok": ok2, "differ": [Path(d).name for d in differ]})) +for d in differ: + name = Path(d).name + (work / (name + ".inc")).write_bytes(inc.get(d, b"")) + (work / (name + ".clean")).write_bytes(clean.get(d, b"")) diff --git a/rustc/fuzz.py b/rustc/fuzz.py new file mode 100755 index 0000000..a2ec444 --- /dev/null +++ b/rustc/fuzz.py @@ -0,0 +1,470 @@ +#!/usr/bin/env python3 +"""Make random edits to a fixture and check each incremental rebuild. + +Each worker keeps one copy of the fixture and repeats: + + 1. apply a random mechanical edit (a comment, a moved item, a changed + literal, a new function, ...); + 2. rebuild incrementally; if the edit does not compile, revert it (the next + build then also exercises recovery from a failed session); + 3. build the same source from scratch, at the same path; + 4. compare: + P6 every .rmeta Cargo reports for a workspace member + rlib every rlib's members, object code included (artifacts.py) + exe the binary's bytes + diag the diagnostics each crate printed + run the binaries' output and exit status + ICE neither build crashed the compiler + split both builds succeed or both fail + +Every --reset edits the worker starts again from the pristine fixture. A +finding keeps the diffs since the last reset, which replay it, and both +builds' differing files and logs. + + rustc/fuzz.py --rustc --fixture fixtures/sink --work [--workers 8] [--edits N] + +Stop it early by creating /STOP. Progress is in /stats.json. +""" + +import argparse +import difflib +import json +import multiprocessing +import os +import random +import re +import shutil +import subprocess +import sys +import time +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import artifacts # noqa: E402 + +p = argparse.ArgumentParser() +p.add_argument("--rustc", required=True) +p.add_argument("--fixture", required=True) +p.add_argument("--work", required=True) +p.add_argument("--workers", type=int, default=8) +p.add_argument("--edits", type=int, default=10**9, help="per worker") +p.add_argument("--reset", type=int, default=40) +p.add_argument("--keep", type=int, default=5, help="findings kept per kind") +p.add_argument("--toolchain", default="nightly-2026-10-06") +p.add_argument("--rustflags", default="-Zincremental-verify-ich") +p.add_argument("--seed", type=int, default=0) +p.add_argument("--timeout", type=int, default=180, help="seconds before a build counts as hung") +args = p.parse_args() + +WORK = Path(args.work).resolve() +FIXTURE = Path(args.fixture).resolve() +BIN = FIXTURE.name + +# ---------------------------------------------------------------- edits + + +def blocks(text): + """Top-level items: blank-line separated, continuation blocks merged.""" + out = [] + for b in text.split("\n\n"): + if out and (b[:1].isspace() or b.startswith("}") or b.startswith("where")): + out[-1] += "\n\n" + b + else: + out.append(b) + return out + + +def edit_lines(text, rng, f): + lines = text.split("\n") + r = f(lines, rng) + return None if r is None else "\n".join(r) + + +def comment_line(text, rng, n): + def f(lines, rng): + i = rng.randrange(len(lines) + 1) + indent = re.match(r"\s*", lines[i] if i < len(lines) else "").group(0) + return lines[:i] + [f"{indent}// fuzz {n}"] + lines[i:] + return edit_lines(text, rng, f) + + +def blank_line(text, rng, n): + return edit_lines(text, rng, lambda lines, rng: (lambda i: lines[:i] + [""] + lines[i:])(rng.randrange(len(lines) + 1))) + + +def remove_comment(text, rng, n): + def f(lines, rng): + idx = [i for i, l in enumerate(lines) if l.strip().startswith("//") and not l.strip().startswith("//!")] + if not idx: + return None + i = rng.choice(idx) + return lines[:i] + lines[i + 1:] + return edit_lines(text, rng, f) + + +def indent_line(text, rng, n): + def f(lines, rng): + idx = [i for i, l in enumerate(lines) if l.strip()] + i = rng.choice(idx) + lines[i] = " " + lines[i] + return lines + return edit_lines(text, rng, f) + + +def swap_items(text, rng, n): + b = blocks(text) + if len(b) < 3: + return None + i = rng.randrange(1, len(b) - 1) + b[i], b[i + 1] = b[i + 1], b[i] + return "\n\n".join(b) + + +def move_item_to_end(text, rng, n): + b = blocks(text) + if len(b) < 3: + return None + i = rng.randrange(1, len(b)) + item = b.pop(i) + return "\n\n".join(b + [item.rstrip("\n")]) + "\n" + + +def delete_item(text, rng, n): + b = blocks(text) + if len(b) < 3: + return None + b.pop(rng.randrange(1, len(b))) + return "\n\n".join(b) + + +def duplicate_fn(text, rng, n): + b = blocks(text) + fns = [i for i, x in enumerate(b) if re.search(r"^(pub(\([^)]*\))? )?(const )?(async )?fn \w+", x, re.M)] + if not fns: + return None + i = rng.choice(fns) + copy = re.sub(r"\bfn (\w+)", lambda m: f"fn {m.group(1)}_fuzz{n}", b[i], count=1) + b.insert(i + 1, copy) + return "\n\n".join(b) + + +ADDITIONS = [ + "fn fuzz_private_{n}() -> u32 {{ {n} }}", + "pub fn fuzz_public_{n}(x: u32) -> u32 {{ x.wrapping_mul({n}) }}", + "#[inline]\npub fn fuzz_inline_{n}(x: &T) -> (T, u32) {{ (x.clone(), {n}) }}", + "pub const FUZZ_{n}: &str = \"fuzz {n}\";", + "pub static FUZZ_STATIC_{n}: [u8; 3] = [{n} as u8, 1, 2];", + "#[derive(Debug, Clone, PartialEq)]\npub struct Fuzz{n} {{ pub items: [T; N], pub tag: &'static str }}", + "pub enum FuzzEnum{n} {{ A(u32), B {{ x: i64 }}, C }}", + "pub trait FuzzTrait{n} {{ fn go(&self) -> impl Sized; const K: u32 = {n}; }}", + "pub async fn fuzz_async_{n}() -> u32 {{ {n} }}", + "pub type FuzzAlias{n} = Vec<(T, u32)>;", + "macro_rules! fuzz_macro_{n} {{ ($e:expr) => {{ $e + {n} }}; }}", + "pub mod fuzz_mod_{n} {{ pub fn inner() -> &'static str {{ \"{n}\" }} }}", +] + + +def add_item(text, rng, n): + b = blocks(text) + i = rng.randrange(1, len(b) + 1) + b.insert(i, rng.choice(ADDITIONS).format(n=n)) + return "\n\n".join(b) + + +def int_literal(text, rng, n): + ms = [m for m in re.finditer(r"(?= args.keep: + return + kept[key] = kept.get(key, 0) + 1 + d = findings / f"w{k}-{stats['edits']:07}-{kind}" + d.mkdir(parents=True, exist_ok=True) + (d / "history.json").write_text(json.dumps(history, indent=1)) + (d / "finding.json").write_text(json.dumps({"kind": kind, "detail": detail, "extra": extra}, indent=1)) + (d / "inc.log").write_text(inc["log"][-8000:]) + (d / "clean.log").write_text(clean["log"][-8000:] if clean else "") + for rel in detail: + if rel in inc["rmetas"]: + (d / (Path(rel).name + ".inc")).write_bytes(inc["rmetas"][rel]) + if clean and rel in clean["rmetas"]: + (d / (Path(rel).name + ".clean")).write_bytes(clean["rmetas"][rel]) + shutil.make_archive(str(d / "src"), "gztar", src) + + if not reset(): + return stats + started = time.time() + while stats["edits"] < args.edits and not (WORK / "STOP").exists(): + if len([h for h in history if h["kept"]]) >= args.reset and not reset(): + break + n = stats["edits"] + stats["edits"] += 1 + paths = sorted(p for p in src.rglob("*.rs") if "target" not in p.parts) + path = rng.choice(paths) + old = path.read_text() + fn = rng.choices([e for e, _ in EDITS], weights=[w for _, w in EDITS])[0] + if fn in (str_literal, int_literal) and path.name == "build.rs": + # Cargo keeps stale OUT_DIR files, so renaming a generated file splits the builds, + # and a changed number can make the build script loop forever. + continue + new = fn(old, rng, n) + if new is None or new == old: + continue + path.write_text(new) + rel = str(path.relative_to(src)) + diff = "".join(difflib.unified_diff(old.splitlines(True), new.splitlines(True), "a/" + rel, "b/" + rel)) + inc = build(src, target) + history.append({"edit": fn.__name__, "file": rel, "diff": diff, "before": old, "after": new, "kept": inc["ok"]}) + by = stats["by_edit"].setdefault(fn.__name__, [0, 0]) + by[0] += 1 + if inc["ice"]: + report("ICE", [], inc, None) + if inc["hang"]: + report("hang", [], inc, None) + if not inc["ok"]: + stats["failed"] += 1 + path.write_text(old) + history.append({"edit": "revert", "file": rel, "diff": "", "kept": False}) + continue + by[1] += 1 + stats["built"] += 1 + + # The clean build, at the same path. + if inc_target.exists(): + shutil.rmtree(inc_target) + target.rename(inc_target) + clean = build(src, target) + if clean["ice"]: + report("ICE", ["clean"], inc, clean) + if clean["hang"]: + report("hang", ["clean"], inc, clean) + if not clean["ok"]: + report("split", [], inc, clean) + else: + stats["compared"] += 1 + a, b = inc["rmetas"], clean["rmetas"] + differ = sorted(r for r in set(a) | set(b) if a.get(r) != b.get(r)) + # Known: metadata reused unchanged from the previous session although a source + # file changed (its hash and length in the source map are stale). + stale = [r for r in differ if r in previous and previous[r] == a.get(r)] + if differ and stale: + stats["findings"]["P6-stale-reuse"] = stats["findings"].get("P6-stale-reuse", 0) + 1 + elif differ: + report("P6", differ, inc, clean) + # More oracles: object code in the rlibs, the binary, and the diagnostics. + for kind, detail in artifacts.compare(inc["art"], clean["art"]).items(): + if kind != "rmeta": + report(kind, detail[:10], inc, clean) + ra = run_exe(str(inc["exe"]).replace(str(target), str(inc_target)) if inc["exe"] else None) + rb = run_exe(clean["exe"]) + if ra != rb: + report("run", [], inc, clean, {"inc": ra, "clean": rb}) + shutil.rmtree(target) + inc_target.rename(target) + previous.clear() + previous.update(inc["rmetas"]) + + if stats["edits"] % 20 == 0: + stats["secs"] = time.time() - started + (WORK / f"stats-w{k}.json").write_text(json.dumps(stats)) + stats["secs"] = time.time() - started + (WORK / f"stats-w{k}.json").write_text(json.dumps(stats)) + return stats + + +if __name__ == "__main__": + WORK.mkdir(parents=True, exist_ok=True) + with multiprocessing.Pool(args.workers) as pool: + results = pool.map(worker, range(args.workers)) + total = {"edits": sum(r["edits"] for r in results), "built": sum(r["built"] for r in results), + "compared": sum(r["compared"] for r in results)} + print(json.dumps(total)) diff --git a/rustc/replay.py b/rustc/replay.py new file mode 100755 index 0000000..2619f2a --- /dev/null +++ b/rustc/replay.py @@ -0,0 +1,201 @@ +#!/usr/bin/env python3 +"""Replay a crate's git history through incremental compilation. + +For each first-parent commit, oldest first, the workspace is checked out and +built incrementally on top of the previous commit's build, then built again +from scratch, and the two are compared: + + P6 every .rmeta Cargo reports for a workspace member is identical + rlib every rlib's members are identical, object code included + diag both builds printed the same diagnostics + ICE neither build crashed the compiler + split both builds succeed or both fail + +Registry dependencies are not compiled incrementally by Cargo, so the clean +build starts from a copy of the incremental target directory with the +workspace members and the incremental cache removed, and only the members are +built again. Both builds use the same target directory path, since Cargo +derives a crate's identity from paths. + + rustc/replay.py --rustc --repo --work [--commits 2000] + +Writes /results.jsonl, one line per commit, and keeps the logs of every +problem and both .rmeta files of the first few differences per crate in +/findings. +""" + +import argparse +import json +import os +import shutil +import subprocess +import sys +import time +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import artifacts # noqa: E402 + +p = argparse.ArgumentParser() +p.add_argument("--rustc", required=True) +p.add_argument("--repo", required=True) +p.add_argument("--work", required=True) +p.add_argument("--commits", type=int, default=2000) +p.add_argument("--toolchain", default="nightly-2026-10-06", help="for cargo") +p.add_argument("--jobs", default="4") +p.add_argument("--keep", type=int, default=3, help="differences kept per crate") +p.add_argument("--from", dest="start", type=int, default=0, help="first commit index to replay") +p.add_argument("--to", dest="end", type=int, default=None, help="last commit index to replay") +args = p.parse_args() + +work = Path(args.work).resolve() +src, target, inc_target = work / "src", work / "target", work / "target-inc" +findings = work / "findings" +work.mkdir(parents=True, exist_ok=True) +findings.mkdir(exist_ok=True) +if not src.exists(): + subprocess.run(["git", "clone", "-q", args.repo, str(src)], check=True) + +env = dict(os.environ) +env.update( + RUSTC=args.rustc, + RUSTC_WRAPPER="", + CARGO_INCREMENTAL="1", + # A commit that denies warnings would otherwise stop building with a newer compiler. + RUSTFLAGS="--cap-lints=warn", + CARGO_TERM_COLOR="never", +) +cargo = ["cargo", f"+{args.toolchain}"] + + +def run(cmd, **kw): + return subprocess.run(cmd, cwd=src, env=env, capture_output=True, text=True, **kw) + + +def members(): + """Every package built from a path: workspace members and path dependencies.""" + r = run(cargo + ["metadata", "--format-version", "1"]) + if r.returncode != 0: + r = run(cargo + ["metadata", "--no-deps", "--format-version", "1"]) + if r.returncode != 0: + return [] + return sorted({pkg["name"] for pkg in json.loads(r.stdout)["packages"] if pkg.get("source") is None}) + + +def build(target_dir): + """Build, and collect the .rmeta files Cargo reports for workspace members.""" + t = time.time() + r = run(cargo + ["build", "--lib", "-j", args.jobs, "--target-dir", str(target_dir), + "--message-format=json-render-diagnostics"]) + rmetas, fresh = {}, [] + for line in r.stdout.splitlines(): + try: + msg = json.loads(line) + except ValueError: + continue + if msg.get("reason") != "compiler-artifact" or "path+file" not in msg.get("package_id", ""): + continue + for f in msg.get("filenames", []): + if f.endswith(".rmeta"): + rmetas[str(Path(f).relative_to(target_dir))] = Path(f).read_bytes() + if msg.get("fresh"): + fresh.append(msg["target"]["name"]) + log = r.stderr + return { + "ok": r.returncode == 0, + "ice": "internal compiler error" in log or "the compiler unexpectedly panicked" in log, + "secs": round(time.time() - t, 1), + "log": log[-6000:], + "rmetas": rmetas, + "fresh": fresh, + "art": artifacts.collect(r.stdout, target_dir), + } + + +def keep_dir(i, commit): + d = findings / f"{i:05}-{commit[:10]}" + d.mkdir(exist_ok=True) + return d + + +commits = run(["git", "rev-list", "--first-parent", "--reverse", f"--max-count={args.commits}", "origin/HEAD"]).stdout.split() +if not commits: + commits = run(["git", "rev-list", "--first-parent", "--reverse", f"--max-count={args.commits}", "HEAD"]).stdout.split() +results = work / "results.jsonl" +done = {json.loads(line)["commit"] for line in results.open()} if results.exists() else set() +kept = {} + +for i, commit in enumerate(commits): + if commit in done or i < args.start or (args.end is not None and i > args.end): + continue + run(["git", "checkout", "-q", "--force", commit]) + run(["git", "clean", "-fdxq"]) + date = run(["git", "log", "-1", "--format=%cs", commit]).stdout.strip() + names = members() + inc = build(target) + + # The clean build: the same target directory path, starting from a copy with the + # workspace members and the incremental cache removed. + if inc_target.exists(): + shutil.rmtree(inc_target) + if target.exists(): + target.rename(inc_target) + shutil.copytree(inc_target, target, symlinks=True) + else: + inc_target.mkdir() + # Cargo refuses to clean a directory it did not create; this one may have been + # created here when the first build failed early. + target.mkdir(exist_ok=True) + tag = target / "CACHEDIR.TAG" + if not tag.exists(): + tag.write_text("Signature: 8a477f597d28d172789f06886806bc55\n") + for name in names: + run(cargo + ["clean", "-p", name, "--target-dir", str(target)]) + for d in target.glob("*/incremental"): + shutil.rmtree(d) + clean = build(target) + + problems, differ = [], [] + if inc["ice"] or clean["ice"]: + problems.append("ICE") + if inc["ok"] != clean["ok"]: + problems.append("split") + if clean["fresh"]: + # A member the clean build did not compile again would be compared with itself. + problems.append("stale") + if inc["ok"] and clean["ok"] and not clean["fresh"]: + a, b = inc["rmetas"], clean["rmetas"] + differ = sorted(rel for rel in set(a) | set(b) if a.get(rel) != b.get(rel)) + # More oracles: object code in the rlibs and the diagnostics. + for kind, detail in artifacts.compare(inc["art"], clean["art"]).items(): + if kind in ("rlib", "diag"): + problems.append(kind) + (keep_dir(i, commit) / f"{kind}.txt").write_text("\n".join(detail)) + if differ: + problems.append("P6") + for rel in differ: + crate = Path(rel).name.split("-")[0] + if kept.get(crate, 0) < args.keep: + kept[crate] = kept.get(crate, 0) + 1 + d = keep_dir(i, commit) + (d / f"{Path(rel).name}.inc").write_bytes(a.get(rel, b"")) + (d / f"{Path(rel).name}.clean").write_bytes(b.get(rel, b"")) + if problems: + d = keep_dir(i, commit) + (d / "inc.log").write_text(inc["log"]) + (d / "clean.log").write_text(clean["log"]) + + # Continue incrementally from the incremental build. + shutil.rmtree(target, ignore_errors=True) + inc_target.rename(target) + record = { + "i": i, "commit": commit, "date": date, "members": names, + "inc": {k: inc[k] for k in ("ok", "ice", "secs")}, + "clean": {k: clean[k] for k in ("ok", "ice", "secs")}, + "compared": len(inc["rmetas"]), "differ": differ, "problems": problems, + } + with results.open("a") as f: + f.write(json.dumps(record) + "\n") + status = "both ok" if inc["ok"] and clean["ok"] else f"inc {'ok' if inc['ok'] else 'failed'}, clean {'ok' if clean['ok'] else 'failed'}" + print(f"{i:5} {date} {commit[:10]} {status}, {len(inc['rmetas'])} compared, " + f"{inc['secs']}s/{clean['secs']}s {' '.join(problems)}", flush=True) diff --git a/ur/rustc/ClosedBugs.rsc b/ur/rustc/ClosedBugs.rsc new file mode 100644 index 0000000..bad3f44 --- /dev/null +++ b/ur/rustc/ClosedBugs.rsc @@ -0,0 +1,124 @@ +module rustc::ClosedBugs + +// Patterns behind closed rustc bugs (docs/motivating.md in mirth), each made as +// general as it can be while still finding the code its fix changed. + +data Language = language(str name); +language("rust"); + +data Classify = classify(str function, str description); +classify("sortByDefId", "#82920: sorting or deduplicating by DefId, whose order is not stable across sessions"); +classify("contextCache", "#89598: a cache held on a context, outside the dependency graph"); +classify("untrackedCrateStore", "#84252: reading the crate store directly, which only an eval_always query may do"); +classify("envRead", "#40364: reading an environment variable, which dep-info must record"); +classify("fileRead", "#111227, #111295: reading a file, which dep-info and the dependency graph must record"); +classify("writeInPlace", "#45841: creating or writing a file directly, rather than renaming a finished temporary file into place"); +classify("writeErrorNotFatal", "#119456: a failed write reported with emit_err, after which compilation goes on and may publish a partial output"); +classify("encoderNotFinished", "#117254: a function that creates a FileEncoder and never finishes it, so write errors are lost"); +classify("writeResultIgnored", "#117254: the result of finishing, flushing or syncing a write thrown away"); +classify("hashIterated", "#34902, #65036: iterating a hash-ordered collection declared in the same file"); +classify("optionRead", "#66955: an option read where it can affect output; joined afterwards with the options marked [UNTRACKED]"); + +data Rewrite = rewrite(str function, str description); +rewrite("never", "rewrites nothing"); +str never((Atom)`NEVER_MATCHES_ANYTHING`) = "never"; + +// The provider a closure is given as (`Providers { name: |tcx, key| ... }`), else the function. +str enclosing(node n) = ps[0] when let ps = [unparse(a.name) | a <- ancestors(n), a is Field, a has name, contains(unparse(a), "|")], size(ps) > 0; +str enclosing(node n) = fs[0] when let fs = [unparse(a.name) | a <- ancestors(n), a is Function], size(fs) > 0; +default str enclosing(node _) = ""; + +set[str] sorts = {"sort_by_key", "sort_unstable_by_key", "sort_by_cached_key", "sort_by", "sort_unstable_by", "dedup_by_key", "dedup_by", "binary_search_by_key"}; + +str sortByDefId(MethodCall c) = unparse(c.method) + " in " + enclosing(c) + when unparse(c.method) in sorts, + let args = unparse(c), + contains(args, "def_id") || contains(args, "DefId") || contains(args, ".krate") || contains(args, "def_index") || contains(args, "local_def_index"); + +set[str] interior = {"Lock<", "RefCell<", "Cell<", "OnceCell<", "OnceLock<", "Sharded<", "RwLock<", "Mutex<", "AppendOnlyVec<", "FreezeLock<"}; +set[str] containers = {"Map<", "Set<", "Cache", "Vec<", "Map>", "Table"}; + +str contextCache(FieldDeclaration f) = owner + "." + unparse(f.name) + when let items = [a | a <- ancestors(f), a is Item], + size(items) > 0, + let item = items[0], + item.item has name, + let owner = unparse(item.item.name), + endsWith(owner, "Ctxt") || endsWith(owner, "Context") || endsWith(owner, "Session") || owner == "CStore", + let ty = unparse(f.type), + any(i <- interior, contains(ty, i)), + any(k <- containers, contains(ty, k)); + +str untrackedCrateStore(node c) = enclosing(c) + when c is Call || c is MethodCall, + let text = unparse(c), + !contains(substring(text, 1, size(text)), "CStore::from_tcx("), + (startsWith(text, "CStore::from_tcx(") || (c is MethodCall && unparse(c.method) in {"cstore_untracked", "untracked"})); + +str envRead(Call c) = callee + " in " + enclosing(c) + when let callee = unparse(c.operand), + callee in {"env::var", "env::var_os", "std::env::var", "std::env::var_os", "env::vars", "std::env::vars"}; + +str fileRead(Call c) = callee + " in " + enclosing(c) + when let callee = unparse(c.operand), + callee in {"fs::read", "fs::read_to_string", "std::fs::read", "std::fs::read_to_string", "File::open", "fs::File::open", "std::fs::File::open"}; + +str writeInPlace(Call c) = unparse(c) + " in " + enclosing(c) + when let callee = unparse(c.operand), + callee in {"File::create", "fs::File::create", "std::fs::File::create", "fs::write", "std::fs::write"}; + +bool isWrite(str text) = any(w <- {"finish", "flush", "write", "create", "rename", "encode", "sync", "persist", "save", "emit"}, contains(text, w)); + +str writeErrorNotFatal(Arm a) = enclosing(a) + when let t = unparse(a), + startsWith(t, "Err("), + contains(t, "emit_err("), + let matches = [m | m <- ancestors(a), m is Match], + size(matches) > 0, + isWrite(unparse(matches[0].scrutinee)); + +str writeErrorNotFatal(If i) = enclosing(i) + when let t = unparse(i), + startsWith(t, "if let Err("), + contains(t, "emit_err("), + isWrite(t); + +str encoderNotFinished(Function f) = unparse(f.name) + when let t = unparse(f), + contains(t, "FileEncoder::new("), + !contains(t, "finish("); + +str writeResultIgnored(Let s) = unparse(s.value) + " in " + enclosing(s) + when unparse(s.pattern) == "_", + s has value, + let v = s.value, + v is MethodCall, + unparse(v.method) in {"finish", "flush", "sync_all", "sync_data", "write_all", "persist"}; + +set[str] hashOrdered = {"FxHashMap<", "FxHashSet<", "HashMap<", "HashSet<", "FnvHashMap<", "FnvHashSet<", "UnordMap<", "UnordSet<"}; + +// Names declared with a hash-ordered type in this file: fields, lets with a type, aliases. +set[str] hashNames() = {unparse(d.name) | d <- descendants(root()), d is FieldDeclaration, any(h <- hashOrdered, contains(unparse(d.type), h))} + + {unparse(d.pattern) | d <- descendants(root()), d is Parameter, d has type, any(h <- hashOrdered, contains(unparse(d.type), h))} + + {unparse(d.name) | d <- descendants(root()), d is TypeAlias, any(h <- hashOrdered, contains(unparse(d.type), h))}; + +str hashIterated(MethodCall c) = recv + "." + unparse(c.method) + "() in " + enclosing(c) + when unparse(c.method) in {"iter", "into_iter", "keys", "values", "drain", "iter_mut"}, + let recv = unparse(c.operand), + let last = lastName(recv), + last in hashNames(); + +str lastName(str text) = parts[size(parts) - 1] + when let parts = [p | p <- splitDots(text)], size(parts) > 0; + +list[str] splitDots(str text) = [text]; + +set[str] optionHolders = {"opts", "sopts", "options", "unstable_opts", "debugging_opts", "cg"}; + +str optionRead(Field f) = unparse(f.field) + when let o = f.operand, + (o is Field && unparse(o.field) in optionHolders) || unparse(o) in optionHolders; +// Inside `impl Options`, the options are `self`. +str optionRead(Field f) = unparse(f.field) + when unparse(f.operand) == "self", + any(a <- ancestors(f), a is Impl, unparse(a.type) == "Options"); diff --git a/ur/rustc/RoundTrip.rsc b/ur/rustc/RoundTrip.rsc new file mode 100644 index 0000000..95bb897 --- /dev/null +++ b/ur/rustc/RoundTrip.rsc @@ -0,0 +1,63 @@ +module rustc::RoundTrip + +// Patterns behind three incremental-compilation bugs, generalized so that each +// still finds the bug it came from (docs/hunt/issue-*.md in mirth). + +data Language = language(str name); +language("rust"); + +data Classify = classify(str function, str description); +classify("hashOrderEncoded", "a hash-ordered collection in a type that derives an encoder: written in iteration order, which a round trip can change"); +classify("hashOrderAlias", "a type alias for a hash-ordered collection: anything encoding a value of it writes iteration order"); +classify("decodedFresh", "a decoder that reserves a fresh identity, where creation may have deduplicated"); +classify("untrackedWhileEncoding", "an encoder reading untracked state: the session, its source map, the environment, the clock"); + +data Rewrite = rewrite(str function, str description); +rewrite("never", "rewrites nothing"); +str never((Atom)`NEVER_MATCHES_ANYTHING`) = "never"; + +set[str] hashOrdered = {"FxHashMap", "FxHashSet", "HashMap", "HashSet", "UnordMap", "UnordSet"}; + +bool derivesEncoder(node item) = any(a <- item.attributes, contains(unparse(a), "derive"), contains(unparse(a), "Encodable")); + +str hashOrderEncoded(FieldDeclaration f) = owner + "." + unparse(f.name) + ": " + kind + when let items = [a | a <- ancestors(f), a is Item], + size(items) > 0, + let item = items[0], + derivesEncoder(item), + let ty = unparse(f.type), + let found = [k | k <- hashOrdered, contains(ty, k + "<")], + size(found) > 0, + let kind = found[0], + let owner = (item.item has name) ? unparse(item.item.name) : "?"; + +str hashOrderAlias(TypeAlias t) = unparse(t.name) + " = " + kind + when let ty = unparse(t.type), + let found = [k | k <- hashOrdered, contains(ty, k + "<")], + size(found) > 0, + let kind = found[0]; + +bool inDecoder(node n) = any(a <- ancestors(n), a is Function, contains(toLowerCase(unparse(a.name)), "decode")); + +str decodedFresh(MethodCall c) = name + when let name = unparse(c.method), + startsWith(name, "reserve") || startsWith(name, "fresh") || contains(name, "next_id"), + !contains(name, "dedup"), + !deduplicatesInThisFile(name), + inDecoder(c); + +// A function defined in the same file whose body deduplicates is not fresh. +bool deduplicatesInThisFile(str name) = any(f <- descendants(root()), f is Function, unparse(f.name) == name, contains(unparse(f.body), "dedup")); + +bool inEncoder(node n) = any(a <- ancestors(n), (a is Function && startsWith(unparse(a.name), "encode")) || (a is Impl && contains(unparse(a.type), "Encode"))); + +bool isSession(node o) = o is Field && unparse(o.field) == "sess"; + +str untrackedWhileEncoding(Field f) = "sess." + unparse(f.field) + when isSession(f.operand), inEncoder(f); +str untrackedWhileEncoding(MethodCall c) = "sess." + unparse(c.method) + "()" + when isSession(c.operand), inEncoder(c); +str untrackedWhileEncoding(Call c) = callee + when let callee = unparse(c.operand), + endsWith(callee, "env::var") || endsWith(callee, "env::var_os") || callee == "SystemTime::now" || callee == "Instant::now", + inEncoder(c); diff --git a/ur/verify-closed.py b/ur/verify-closed.py new file mode 100755 index 0000000..cc1c2e1 --- /dev/null +++ b/ur/verify-closed.py @@ -0,0 +1,79 @@ +#!/usr/bin/env python3 +"""Check that each closed-bug query still finds the code its bug's fix changed. + +For each case, fetches the files the fixing PR changed, as they were at the PR's +base commit (GH_HOST=github.com gh), runs the query over them with Ur, and looks +for a site whose label and file match. + + ur/verify-closed.py [work dir] +""" + +import json +import os +import re +import subprocess +import sys +from pathlib import Path + +UR = sys.argv[1] +WORK = Path(sys.argv[2] if len(sys.argv) > 2 else "verify-closed").resolve() +HERE = Path(__file__).resolve().parent +ENV = dict(os.environ, GH_HOST="github.com") + +# (bug, fixing PR, module, query, label contains, file contains, profile) +CASES = [ + ("#82920", 83074, "ClosedBugs", "sortByDefId", "dedup_by_key", "astconv", None), + ("#89598", 89619, "ClosedBugs", "contextCache", "vtables_cache", "context.rs", None), + ("#84252", 84260, "ClosedBugs", "untrackedCrateStore", "has_global_allocator", "cstore_impl", None), + ("#40364", 71858, "ClosedBugs", "envRead", "env::var", "env.rs", "rust-2018"), + ("#45841", 45899, "ClosedBugs", "writeInPlace", "out_filename", "link.rs", "rust-2015"), + ("#117254", 117301, "ClosedBugs", "encoderNotFinished", "", "encoder.rs", None), + ("#119456", 119510, "ClosedBugs", "writeErrorNotFatal", "encode_metadata", "encoder.rs", None), + ("#34902", 35984, "ClosedBugs", "hashIterated", "xrefs", "encoder.rs", "rust-2015"), + ("#65036", 65043, "RoundTrip", "hashOrderAlias", "Resolutions", "lib.rs", "rust-2018"), + ("#66955", 84233, "ClosedBugs", "optionRead", "remap_path_prefix", "", None), + ("#111227", 111641, "ClosedBugs", "fileRead", "", "debugger_visualizer", None), +] +# Files the fix did not change but the bug's site is in. +EXTRA = {84260: ["compiler/rustc_metadata/src/rmeta/decoder/cstore_impl.rs"]} + + +def gh(*args): + return subprocess.run(["gh", *args], capture_output=True, text=True, env=ENV, check=True).stdout + + +def fetch(pr): + d = WORK / str(pr) + if d.exists(): + return d + info = json.loads(gh("pr", "view", str(pr), "-R", "rust-lang/rust", "--json", "baseRefOid,files")) + paths = [f["path"] for f in info["files"] if f["path"].endswith(".rs") and not f["path"].startswith(("tests/", "src/test/"))] + for path in paths + EXTRA.get(pr, []): + target = d / path + target.parent.mkdir(parents=True, exist_ok=True) + try: + text = gh("api", f"repos/rust-lang/rust/contents/{path}?ref={info['baseRefOid']}", "-H", "Accept: application/vnd.github.raw") + except subprocess.CalledProcessError: + continue # added by the fix: it did not exist before + + # Ur's grammars lack the old unstable `crate` visibility; it is not what the queries test. + text = re.sub(r"^(\s*)crate (fn|struct|enum|type|mod|use|trait|const|static|unsafe fn) ", r"\1pub(crate) \2 ", text, flags=re.M) + target.write_text(text) + return d + + +failed = 0 +for bug, pr, module, query, label, file, profile in CASES: + d = fetch(pr) + report = WORK / "report.json" + cmd = [UR, "rewrite", "--classify", "--rules", str(HERE / f"rustc/{module}.rsc"), "--report", str(report), str(d)] + if profile: + cmd += ["--profile", profile] + subprocess.run(cmd, capture_output=True, text=True) + sites = [(l["label"], s["file"], s["span"]["line"]) for l in json.loads(report.read_text())["labels"] + if l["classifier"] == query for s in l["sites"]] + hit = next((s for s in sites if label in s[0] and file in s[1]), None) + failed += hit is None + where = f"{hit[0][:60]} ({Path(hit[1]).name}:{hit[2]})" if hit else "" + print(f"{bug:8} {query:20} {'finds it' if hit else 'MISSES IT'} {where}") +sys.exit(1 if failed else 0)