Three ways to find more bugs than one fixture and ten hand-written edits can:
- Replay real history.
rustc/replay.pywalks a crate's git history, oldest first, building each commit incrementally on top of the last and again from scratch, and compares them (P6). Ten crates, up to 2,000 commits each. - Fuzz edits on a large fixture.
fixtures/sinkis a five-crate workspace with as many stable language features as fit.rustc/fuzz.pymakes random mechanical edits to it and checks every incremental rebuild against a clean build. - Survey past bugs for properties. Agents read 1,000 fixed rustc bugs and extracted
the invariants they violated; a second pass checked every citation. The result is
properties.md.
All of it runs against a compiler with the candidate fixes for the bugs already known
(hunt.md), so those do not drown new ones; each new finding is then confirmed
on the official nightly.
| finding | how | status |
|---|---|---|
| Incremental rebuilds republish stale metadata when an edit moves no span | fuzzer and replay, independently | new; root-caused, regression from #114669 (1.90); fixed and tested (report) |
| The candidate fix for the literal bug over-deduplicated | replay (serde, one commit) | my own mistake; the fix was narrowed (report) |
| Three untracked options that change reused output | closed-bug queries and an option audit (ur-queries.md) |
new instances of #66955's class; report drafted (draft) |
Stale metadata. Since #114669, an incremental session reuses the saved .rmeta when
its dep-node is green. But the metadata's source map records every source file's length,
line table and content hash, read straight from the session's source map, which nothing
tracks. An edit that changes no query result (a comment after the last item, a typo fixed
inside a comment without moving any span, a change inside #[cfg]'d-out code) leaves the
node green, and the previous session's metadata is republished. A dependent then cannot match
the dependency's source: its diagnostics lose the snippet. Real histories hit it on ordinary
commits: memchr (5), serde, smallvec (2), hashbrown (1), anyhow (3), regex (several), all on
comment or docstring edits. The fix makes the metadata task read an eval_always query
fingerprinting the local source files, so reuse still happens when the files are truly
unchanged (the blessed touch-only rebuild still shows it).
A flaw in my own fix. The first fix for the literal bug deduplicated every immutable
allocation on decode. The serde replay showed P6 failing on commit 2f58a20 with it applied:
-Zmeta-stats put the whole 59-byte difference in interpret-alloc-index, because
allocations a clean session keeps apart were merged. The fix now records, at encode time,
whether an allocation was created through deduplication, and repeats only that. The same
replay window is clean with it.
With all three fixes, each regression test passes and each fails without its own fix, and rustc's incremental, metadata UI and run-make, consts/statics/const-generics UI and codegen-llvm tests pass.
rustc/replay.py --rustc <rustc> --repo <checkout> --work <dir> [--commits N] [--from i --to j]
For each first-parent commit, oldest first: check it out, cargo build --lib
incrementally on the previous commit's target directory, then build the same source from
scratch at the same path. Registry dependencies are not compiled incrementally by Cargo, so
the clean build starts from a copy of the target directory with every path package cleaned
(cargo clean -p) and the incremental cache removed. It compares every .rmeta Cargo
reports for a path package, and flags ICEs, builds where only one side fails, and clean
builds that did not actually recompile ("stale", which would compare a file with itself).
Every commit is built with --cap-lints=warn, so old commits that deny warnings still build.
Crates: regex, itertools, smallvec, hashbrown, indexmap, memchr, bitflags, serde, anyhow, thiserror. On the compiler with all three fixes:
| crate | commits | both built | P6 | ICE | one side failed |
|---|---|---|---|---|---|
| anyhow | 669 | 524 | 0 | 0 | 0 |
| bitflags | 301 | 300 | 0 | 0 | 0 |
| hashbrown | 477 | 456 | 0 | 0 | 0 |
| indexmap | 465 | 415 | 0 | 0 | 0 |
| itertools | 1,358 | 942 | 0 | 0 | 0 |
| memchr | 267 | 267 | 0 | 0 | 0 |
| regex | 700 of 1,368 | 590 | 0 | 0 | 0 |
| smallvec | 391 | 390 | 0 | 0 | 0 |
| serde | 806 of 2,000 | 425 | 0 | 0 | 0 |
| thiserror | 556 | 553 | 0 | 0 | 0 |
4,862 commits built on both sides and none differed. Commits that failed on both sides are
mostly old code the current compiler rejects. regex and serde stopped part-way when the
disk filled. This run compared .rmeta only; the rlib and diagnostic comparisons were
added after it started.
Before the stale-metadata fix, the same replay found that bug on ordinary commits in memchr, smallvec, hashbrown, anyhow and regex: comment and docstring edits that moved no span. One P6 failure, on serde, came from my own first fix for the literal bug.
A positive control: on the unpatched nightly, every incremental rebuild of itertools'
last 12 commits differs from a clean build (the param_def_id_to_index bug); on the patched
compiler none do.
Harness problems found and fixed on the way, each of which produced false findings or
skipped commits silently: artifacts of earlier commits counted as differences; path
dependencies that are not workspace members left uncleaned; cargo clean refusing a
directory without CACHEDIR.TAG.
fixtures/sink, about 1,200 lines in five crates:
sink-macros, a proc-macro crate with no dependencies: a derive with a helper attribute, an attribute macro, a function-like macro;sink-core, with a build script generating code and a cfg, features, traits with associated types and consts, generic associated types, blanket impls, trait objects and upcasting, operators,impl Traitandasync fnin traits, async closures, const generics, const fn and const blocks, statics, thread locals, unions,repr,Drop, raw pointers,MaybeUninit, FFI,#[track_caller], inline attributes, deprecation, error types, iterators, closures, patterns, labelled blocks,let else,if letchains, collections,Rc/RefCell/Arc/Mutex, scoped threads, exported and internalmacro_rules!;sink-mid, using all three kinds of proc macro and the exported macros, with nested modules and restricted visibility, re-exports, and generic code dependents instantiate;sink-dy, built as a dylib too;sink, a binary that exercises everything and checks 60 results, so a miscompile fails.
rustc/fuzz.py --rustc <rustc> --fixture fixtures/sink --work <dir> [--workers N]
Each worker keeps one evolving copy of the fixture. It makes a random edit, chosen from 16
kinds (a comment, a blank line, an indented line, swapped or moved or deleted items, a
duplicated function, a new item of one of 12 shapes, a changed number or string, an
inline attribute, a doc comment, reordered derives, narrowed visibility, a renamed local),
and builds incrementally with -Zincremental-verify-ich. If the edit does not compile, it
is reverted; the next build then also exercises recovery from a failed session. Otherwise
the same source is built from scratch at the same path, and every .rmeta and every proc
macro's embedded metadata are compared (P6), the two binaries are run and their output
compared, and ICEs, hangs and one-sided failures are reported. Every 40 kept edits the worker
starts again from the pristine fixture. A finding keeps every edit since the last reset, and
rustc/fuzz-replay.py replays it exactly.
Since shadow-mode.md, the fuzzer and the replay also run every build
with the compiler's own check of what it reused (RUSTC_VERIFY_REUSE, on a compiler with
hunt/verify-reuse.patch), and report what it prints.
Throughput on this 16-core machine: about 2 edits a second with six workers, about 170,000 a day, while other work shared the machine.
| compiler | edits | built and compared | findings |
|---|---|---|---|
| with the first two fixes | 4,009 | 3,336 | 60 stale metadata reuses (the third bug) |
| with all three fixes | 10,717 | 8,936 | none |
The last 1,455 of those comparisons also checked object code, binaries and diagnostics.
Millions of edits means about a week here, or several machines.
Short fuzz runs (two workers, 40 edits each) with one flag set added to every build, on the
compiler with the three fixes, the reuse check and
hunt/debuginfo-checksum-stopgap.patch:
| flags | compared | differences from a clean build |
|---|---|---|
-Copt-level=2 (before the stopgap) |
102 | binaries and objects in a third of the rebuilds: finding 6 |
-Copt-level=2 |
152 | none |
-Cinstrument-coverage (official nightly) |
126 | none besides the known metadata bugs |
-Cdebuginfo=line-tables-only |
71 | none |
-Cpanic=abort |
68 | none |
-Zshare-generics=yes -Copt-level=1 |
71 | none |
-Copt-level=s |
67 | none |
-Ccodegen-units=1 -Copt-level=3 |
69 | none |
-Zdwarf-version=5 |
66 | none |
-Csplit-debuginfo=unpacked, =packed |
74, 67 | every rebuild, but two clean builds differ too: objects name their .dwo files with the session suffix, so this oracle does not apply |
The reuse check printed only the known allocation-sharing pattern in all of them.
Threads. With -Zthreads=8, two clean builds of fixtures/sink already differ
(#162202: the definitions made for impl Trait and async fn in traits get indices in a
nondeterministic order). fixtures/sink-threads.patch
turns the two such trait methods into boxed iterators and futures, and the free async fns
into functions returning impl Future (once the fuzzer duplicated an async fn, clean
threaded builds of the crate differed six ways in six builds, as in #162202's first
example); with it, clean threaded builds agree. The fuzzer now builds clean a second time
whenever anything differs, and reports P5 rather than P6 if the two clean builds disagree.
That still left -Zthreads fuzzing finding #162202 again whenever an edit made another
impl Trait or async fn. hunt/threads-def-order-stopgap.patch,
a testing aid like the other stopgap, makes the definitions that queries create (RPITIT
associated types, captured lifetimes of opaque types, coroutine by-move bodies) on one
thread, in definition order, just before rustc's commit_end_of_determinism, when the
front end is parallel. With it, eight -Zthreads=8 builds of #162202's reproduction give one
metadata file instead of four, and threaded builds of the original fixtures/sink agree,
so the threaded fuzzer runs on the unmodified fixture. With only one of the two changed, clean builds agreed but incremental rebuilds
differed from clean ones about half the time: the same out-of-order indices, made in one
session and kept by the next (the incremental tables for the made-up associated types ended
two indices earlier). That is #162202 reaching incremental sessions, not a new bug.
properties.md has 29 properties plus the crash baseline, ranked by how
many of the 1,000 bugs violated each, after every citation was checked against the issue
text (107 of 156 citations held). The top: intern-time well-formedness of type-system values
(12 bugs), the old and new trait solvers agreeing (11), extern "C" lowering matching C
(9), MIR validating at every opt level on real crates (8), library soundness under injected
panics (7). The incremental and metadata properties mirth already checks rank lower by bug
count (3 and 1), yet they are where all four bugs found in this work came from, which says
more about which bugs get reported and fixed than about where bugs are.
Small models did both the reading and the checking. The issue lists are leads, not proof.
- Everything ran on Linux x86_64, single-threaded unless stated.
-Zthreadsis not fuzzed: the fixture hasimpl Traitin traits, and #162202 makes every threaded build differ.- The fuzzer edits text, not syntax trees; about 8% of edits do not compile and are reverted, and some kinds of edit (deleting items, narrowing visibility) almost never do.
- The replay compares
.rmetaonly; the binaries' code and debug info are not compared.