Skip to content

A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is an LLVM ScalarEvolution miscompile rustc 1.99.0 exposes, its job takes 1.98.1, and the fork's compiler has the fault on both architectures; the power-off line's issue is closed - #778

Merged
Japabu merged 11 commits into
mainfrom
wt/toyos-nightlymac
Oct 9, 2026

Conversation

@Japabu

@Japabu Japabu commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Head 82ee7ab86: e0a61d070 (six commits and a merge of origin/main, 6776714db, #769, on 809c33c0c) and four commits that change only issues/ (git diff e0a61d070 82ee7ab86 --stat names three files, all under issues/; git diff aeb396f32 82ee7ab86 --stat, two).

Open, and the toolchain's: the fork's compiler, which builds every ToyOS kernel, loader and program, turns a correct loop of safe Rust into an endless one on both architectures. Measured for a u16 range to its maximum on the three AArch64 targets ToyOS builds for, and for a u128 range to its maximum on aarch64-unknown-toyos and x86_64-unknown-toyos. The fault is in LLVM's ScalarEvolution and is target-independent; upstream has the fault and has merged no fix. Its open report llvm/llvm-project#175729 has the same symptom on another loop: that ToyOS's loop is that report's fault is the issue file's reading, and that the report is the known fault is a maintainer's "probably". Nothing on this branch fixes it: issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md is filed for it, held by the toolchain, with a fix in the fork's LLVM being built on wt/toyos-scevfix. "How far the fault reaches" below has the measurements.

Every gate at e0a61d070 was run by the implementer in this branch's worktree, on an Apple-silicon development machine, macOS 27.0.1; each result is the command's own exit code. The measurements of rounds 1 to 5 were run by the orchestrator on the same machine, and the bisection (at e52e6275b) and the root-cause work by agents of his. The implementer of the last two rounds ran none of those: he read both reports and checked each number written here against their logs, traces and saved upstream responses.

1. The nightly's portability-macos hang: an LLVM miscompile that rustc 1.99.0's core exposes

What the nightly showed before this branch

Run 37778826093, head 8a883a142, job 113317209315: in cargo run -- --ci host's step the workspace's host members, toyos-userbound/tests/firmware.rs's a_port_answers_as_its_declaration_says never reported, libtest said so once after 60 s, and the driver waited on cargo with no bound for 3 h 36 min until the job's timeout-minutes: 350. The same test is ok in the same run's Linux host job (113317209042). It arrived with #749 and was never green on that runner.

What the instrumented nightly showed

Run 37834436612, a dispatch of nightly.yml on this branch at c952668c9, read against review round 2's list:

the list the job
concludes failure, not cancelled failure, 19:46 to 22:23Z; host, tcg / suite, portability-linux and toolchain / build success
the refusal's own line 22:11:02 [ci] the workspace's host members: cargo test --workspace …: said nothing for 900s and was ended with its process group by SIGQUIT; the last it said: test a_port_answers_as_its_declaration_says has been running for over 60 seconds, 15 min 0 s after libtest's line at 21:56:02; no outside that group
later steps run, the summary printed === [ci] the kernel's library at 22:11:02 and every step after it; [ci] Host: 1 of 77 step(s) red
no orphan firmware-* Cleaning up orphan processes is followed by no Terminate orphan process line
the artifact macos-crash-reports uploaded, 80,792 bytes, with a firmware-* report

The report, thread by thread. Two threads. main, the faulting one, is parked: semaphore_wait_trap under Thread::park, the CompletedTest channel's recv and test::console::run_tests_console. The thread named a_port_answers_as_its_declaration_says is in the test's body and nowhere else: its program counter is at offset 0 of the test's own function (the closure's FnOnce::call_once), its link register in libtest's __rust_begin_short_backtrace, and the report lists as inlined frames at that one address standing (tests/firmware.rs:495), firmware::port (src/firmware.rs:344, the match in its loop) and the test's line 513, the loop over the ports nothing declared.

The same binary made here, and what it is

Cargo compiles without incremental state where CI is set. No earlier measurement here set it, which is why 40 runs of 40 at the job's tree, compiler and package selection ended. cargo +<toolchain> test --locked -p toyos-userbound --test firmware --no-run, then the binary whole, ended after 60 s if it had not ended:

toolchain environment profile build the binary
1.99.0 CI=true the tree's EXIT=0 hung, killed at 60 s: 20 tests ok, then libtest's line for this one
1.99.0 CARGO_INCREMENTAL=0 the tree's EXIT=0 hung, the same
1.99.0 CARGO_INCREMENTAL=0 opt-level = 1 EXIT=0 EXIT=0
1.99.0 CARGO_INCREMENTAL=0 opt-level = 0 EXIT=0 EXIT=0
1.98.1 CARGO_INCREMENTAL=0 the tree's EXIT=0 EXIT=0

The hung binary's test function is at the offset the runner's report names (0x1564, 5476), and its disassembly is one instruction, a branch to itself:

…_8firmware38a_port_answers_as_its_declaration_says0…FnOnce…call_once…:
0000000100001564	b	…_8firmware38a_port_answers_as_its_declaration_says0…call_once…

The hanging build repeated with the test's cases edited, the file restored each time: without (0xFFFC, Width::DWord) EXIT=0; with it and without (0xFFFF, Width::Byte) hung; with the byte and without the dword EXIT=0. So at opt-level = 2 for aarch64-apple-darwin the compiler ends firmware::port's inclusive range over ports 0xFFFC..=0xFFFF, inlined into the test, wrongly, and drops the rest of the function. The source ends: the crate forbids unsafe, and every other build above runs it to its end. The cause is in "How far the fault reaches".

What this branch does about it

  • portability-macos installs 1.98.1 where it installed stable. Under 1.98.1 the hanging build ends (the table's last row). That is a recorded compromise, and a narrower one than it looked: 1.98.1 carries the same faulty LLVM and escapes this test only because its core does not yet give the loop the shape the fault needs. issues/rustc-1-99-0-makes-an-endless-loop-of-a-port-test-on-apple-silicon.md is the pin's record, kind: tooling, assigned to the orchestrator: the runner's reading, the stable compilers' rows, what the pin is and is not worth, and its exit, the job back on stable and green in a nightly on main. No stable that meets it exists.
  • firmware::port is not rewritten. The loop is right; a compiler that drops a loop's exit here is not made safe by rewriting the one loop where it was seen.
  • The test stays, green on the compiler the job now takes. To be dispatched before this lands: nightly.yml on wt/toyos-nightlymac, at a head that differs from the one that lands only under issues/; portability-macos must conclude success with test a_port_answers_as_its_declaration_says ... ok in the workspace's host members.

The bound, at the wait that had none

src/ci.rs's cargo called Command::status() and cargo_logged read a pipe to its end. Both now run through heard:

  • A command that says nothing for QUIET, 15 minutes, is ended and the step is red with the last it said, counted in bytes and cut at 200 characters: libtest's test <name> has been running for over 60 seconds, or on one thread the unfinished test <name> ... . The driver goes on to its other steps and its summary. The runner's reading above is this, end to end.
  • The command leads a process group and the group is signalled: what hangs is a test binary cargo started.
  • SIGQUIT first, then SIGKILL 10 s later for what still holds the output, to the group and to the child by pid, which is always reaped. The command starts with SIGQUIT at its default whatever the driver inherited: measured, a spinning test started that way ends in the millisecond of the signal and leaves a report with both its threads' stacks, and started from a shell's background job without the reset it ignores the signal for the 30 s it was given and leaves none.
  • An interrupt of the driver ends the group. SIGINT, SIGTERM and SIGHUP are handed to the group and then taken at their default, unless the driver was started ignoring them. The handler writes the signal it took and then reads the group; heard reads the pending signal after every write of the group (before a spawn, after it, and before the reap), all sequentially consistent, so one of the two sees the other on whatever thread the handler runs. That rests on the argument, which review round 3 walked pair by pair, and on no test.
  • Why silence and not a ceiling on the step. Between consecutive log lines inside a green driver step: 78 s at most in a macOS host job (job 113189702343), under 60 s in the Linux host job and 120 s at most in tcg / suite, both of run 37778826093. 15 minutes is 7.5 times the longest, and above the guest harness's own longest ceiling (GUEST_WEDGED, 141 s).

nightly.yml, portability-macos: timeout-minutes: 90 on the --ci host step, for a step that hangs while it talks; and ~/Library/Logs/DiagnosticReports uploaded as macos-crash-reports when the job fails, by actions/upload-artifact at the commit guest.yml pins. Both stay: the report they kept is what named the cause.

What the two tests see that reading cannot

Both run a step that is this test binary, which starts a second process, prints a line and hangs by writing into a pipe its judge's side holds; both judge by a pipe's end, in a process of their own (issues/a-pipe-made-on-macos-can-leak-into-a-sibling-spawn-and-red-the-tether-tests.md is why). a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line: the refusal ends in the step's last line, by SIGQUIT, with no holder of the output left. an_interrupt_of_the_driver_ends_the_group_its_step_runs_in: a driver sent SIGTERM dies of it and the pipe its step and the step's child write into ends within 10 s. QUIET is 10 s under cfg(test).

2. The artifact's other reports

  • cargo, 22:11:02Z: SIGQUIT, sent by toyos-build, its main thread in __wait4 under cargo_test::run_unit_tests. This branch's own signal to the silent step's group.
  • loom_sleep, 22:15:40Z: SIGABRT, abort() called, under core::panicking::panic_in_cleanup. The control doorbell-kick-relaxed reaching its verdict: the log has the model's panic, thread caused non-unwinding panic. aborting., and [ci] control \doorbell-kick-relaxed`: 1 verdict(s) reached`. Nothing owed.
  • rust-objcopy, 22 reports (not 25: the artifact has 25 files in all), each ended by dyld at launch inside --build-only, which went on and succeeded: 13 after the first Building bootstrap (Library not loaded: @rpath/libLLVM.dylib) and 9 after the second (Symbol not found, expected in a libLLVM.dylib). The development Mac's own reports hold 18 more, of 4, 7 and 8 October, every one Library not loaded. Filed as a defect, not fixed: issues/the-bootstraps-own-build-cannot-start-rust-objcopy-on-an-apple-host.md.

3. The power-off line's issue is closed

issues/the-power-offs-global-lock-line-is-logged-after-the-last-console-drain.md is deleted. Its exit's first clause, acpi_mediated_access green in the nightly's tcg / suite on a main carrying #767: run 37778826093, nightly on main at 8a883a142, of which 15b8f9d25 (#767) is an ancestor; job 113320660798, success: PASS acpi_mediated_access (4s), PASS acpi_lock_given_back_on_one_cpu (4s), test result: ok. 34 passed, 34 total (598.5s; workers: 427s building, 171s testing). The file itself records #767 meeting the other two. No citation of the slug is left; its rule is at kernel/src/power.rs's shutdown.

How far the fault reaches

A bisection on the same machine at e52e6275b, the tree clean before and after and built only into a target directory of the bisection's own, and after it a root-cause investigation with LLVM 22.1.8's opt and the fork's compiler, which touched no ToyOS file. The issue file carries every table, both sources and the IR; what they say:

Which compilers. Each step installed one nightly with rustup, built CARGO_INCREMENTAL=0 cargo test --locked -p toyos-userbound --test firmware --no-run, ran the binary whole with a bound of 120 s, and uninstalled it (UNINSTALL EXIT=0; rustup toolchain list byte-identical before and after).

nightly rustc LLVM the binary
2026-07-09 1.99.0-nightly (14cae6813 2026-07-08) 22.1.8 EXIT=0
2026-07-10 1.99.0-nightly (af3d95584 2026-07-09) 22.1.8 hung, ended by PID at 120 s
2026-09-25 1.100.0-nightly (f7575a9da 2026-09-24) 23.1.1 hung; the test's function is b .
2026-10-08 1.101.0-nightly (1d81eb4ad 2026-10-07) 23.1.3 hung; the test's function is b .

Seven steps, the five good ones before 07-09 in the issue. Five nightlies from 07-10 on were run (07-10, 07-13, 07-22, 09-25, 10-08) and each hangs; none between them was. nightly-2026-10-08 was the newest there was.

What put the test's loop in the shape. Between the two adjacent nightlies are 106 commits, none under src/llvm-project; among them is rust-lang/rust #155114, which rewrote RangeInclusive's next onto Step::forward_overflowing. That code is correct. The same next written by hand, with no core range, is miscompiled by stable 1.91.0, 1.95.0, 1.98.0 and 1.98.1 (LLVM 21.1.2 to 22.1.8). 1.88.0 (LLVM 20.1.5) compiles it right, which is an escape by shape and no bound on the fault: in its trace the overflowing_add stays a call of llvm.uadd.with.overflow.i16, so the chain below has no add to start from. 1.98.1 passes the tree's test the same way.

The cause, four steps, the first two sound:

  1. CorrelatedValuePropagation marks the range's increment nuw: its wrapped result is used by nothing. Measured in the pass traces; why it may, read from LLVM's source.
  2. LoopRotate, after inlining with the constant start, makes that add a header phi's increment. Measured.
  3. ScalarEvolution copies the add's nuw onto the phi's recurrence without a condition: the fault. Measured in its own print, which gives {-3,+,1}<nuw> the range [-3,0) beside an exit count of 3, the iteration at which it is 0; the lines that copy (ScalarEvolution.cpp:5776-5783 in the fork's source), read.
  4. indvars asks whether the exit's compare can hold and is told no. Measured: -C llvm-args=-opt-bisect-limit under nightly-2026-07-22, limit 690 prints true, 691 prints false, pass 691 is indvars, and it replaces the loop's exit at 0xFFFF with false. The path inside is read from the source and not measured.

With no front end: 14 lines of IR, in the issue as text, through opt -passes=indvars (LLVM 22.1.8, nightly-2026-07-22's) come out with the header br i1 false under no data layout, AArch64's and x86-64's. The same with a second exit and without the nuw, or with the add still in the header, is not folded. No opt built from the fork's LLVM was run.

Upstream. llvm/llvm-project#175729 is open, labelled miscompilation, with the same symptom and the same source path on another loop. Pull request #118959, proposed as its fix, is an unmerged draft opened in December 2024. Not established: a maintainer's sentence on the issue is "I've only glanced at it, but this is probably the known issue"; nobody upstream has seen ToyOS's loop; and nobody, here or there, has built an LLVM with #118959 to see what it fixes. No rust-lang/rust report was found. None is sent.

The fork's compiler (rust/ at 6d6ad8c71906, rustc 1.99.0-dev, LLVM 22.1.8; it contains #155114), --emit llvm-ir,asm -C opt-level=2, each rustc exit 0, run from the sysroot key the ToyOS targets take. On the u16 reproducer:

target caller, which returns true
aarch64-unknown-toyos no ret; b .LBB0_1
aarch64-unknown-none-softfloat no ret; b .LBB0_1
aarch64-unknown-uefi no ret; b .LBB0_1
aarch64-apple-darwin no ret; b LBB0_1
x86_64-unknown-toyos, x86_64-unknown-none, x86_64-unknown-uefi ret i1 true; movb $1, %al, retq

And by width, the range ending at the type's maximum, caller in the IR:

type aarch64-unknown-toyos x86_64-unknown-toyos
u8, u32, u64 ret i1 true ret i1 true
u16 no ret ret i1 true
u128 no ret no ret

i16 gave a caller that is not constant, on both; whether it is right was not checked.

And on the tree's test, on aarch64-apple-darwin with the store's sysroot as RUSTUP_TOOLCHAIN: build EXIT=0, the binary hung and was ended at 120 s, the test's function one branch to itself.

Reach. The kernel, the loader and userland are exposed on both architectures. Nothing wrong has been found in them, and nothing was looked for; no statistic or remark of today's compiler would count it. An inclusive range to its type's maximum is one instance of a class: an increment whose wrapped result is dead, a rotation that makes it a header phi's, a start of known range, and a client of ScalarEvolution reasoning about another value with the same expression. Of 40 sources made while reducing, 18 are miscompiled for each of the two AArch64 targets swept and 0 for x86_64-unknown-none or x86_64-unknown-toyos; the u128 row is the counterexample that makes that no immunity. Why the forty escaped: none names an integer wider than 64 bits, and on x86-64, where 16 is a legal width, indvars first rewrites the loop's exit test onto the incremented value, which gives the add a second use so step 1 never marks it. That is read from the x86-64 trace of the one reproducer and inferred for the rest. With the passes after indvars withheld the outcome is a wrong value and not a hang, so a silent wrong answer is possible.

The x86-64 kernel's own instance of firmware::port has the loop's exit. CI=true cargo run -- --build-only at e0a61d070, EXIT=0, the x86_64-unknown-none kernel built from nothing and disassembled with the Xcode tools' objdump: with incremental state, a function of its own of 0x11a bytes; with CARGO_INCREMENTAL=0 beside CI=true, EXIT=0, inlined into acpi_mode::port, 0x13a bytes. In both the loop counts the width's bytes down in a 16-bit register and leaves at zero, every refusal jumps out, and no backward branch is unconditional. The AArch64 kernel has no instance: the one call is under arch/x86_64, by the source.

Stable 1.99.0 on x86-64 compiles this test right: --target x86_64-apple-darwin, build EXIT=0, the test's function six instructions ending in ret, the binary EXIT=0 under Rosetta with 21 passed.

One issue became two

  • issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md, new, kind: defect, status: assigned: the toolchain holds it, and a fix in the fork's LLVM is being built on wt/toyos-scevfix. Round 5's file was named for AArch64 and for an inclusive range; the u128 row refuted both, and it is renamed with its one citation moved. It is the graver of the two and outlives the pin: no job installs the fork's compiler by a version. Its exit is held to ScalarEvolution and to no client of it: a change to indvars that made the reproducers return would leave the fault. It is measured by four things together: the IR above through the fixed LLVM's opt -passes=indvars with no br i1 false left; caller as ret i1 true for aarch64-unknown-toyos and x86_64-unknown-toyos in the u16 and the u128 source from the sysroot key the ToyOS targets take, and for aarch64-unknown-none-softfloat and aarch64-unknown-uefi in the u16 source from the key the kernel and the loader take, which is not the key the table's rows for those two were made from; the tree's test binary exiting 0 on Apple silicon; and the tree's own functions read, by the fix behind an LLVM option, the tree built twice and compared function by function. The last is the measurement owed, and the only one that reads what ToyOS ships.
  • issues/rustc-1-99-0-makes-an-endless-loop-of-a-port-test-on-apple-silicon.md keeps its slug, which nightly.yml cites and which is still true, and becomes the pin's record, kind: tooling. Its row for nightly-2026-09-25 and its sentences that the nightly was good, that 1.98.1's LLVM is right and that a single file does not reproduce the fault are deleted. Its exit says a stable can meet it by core moving the loop's shape again, which fixes nothing.

The reproducers are text in the issue, and no test arrives here

A host check that compiles the reproducer with the tree's own compiler for aarch64-unknown-toyos and reads whether caller returns is the cheapest tier that reaches the defect: it needs no machine and no QEMU, and no type or reading sees what a compiler emits. It is not added now. It would be red on main from the commit that added it, and a red test is a defect that is fixed or deleted; the tree has no known-red test and this is no reason to invent one. So both sources, the IR, their commands and what each must show are in the defect's file, where the exit reads them, and the check arrives with the fix, green. The tree's existing test already goes red on it, but only where the host suite is built by the fork's compiler without incremental state on Apple silicon, which no job does.

Review round 3

what closes it state
run 37834436612 read against the list, the report thread by thread, the issue given that reading above; the issue is renamed for what it now says and rewritten
this pull request contains what the reading selects the report names the cause: the job's compiler is pinned, with the issue recording it. The measurement behind it, a dispatch whose portability-macos is green with the test in the tree, is still to be run. Round 5 found the pin does not cover the fork's compiler, which has the fault too
host and guest / suite at the head that lands not run: the pull request is a draft
the two tests and both mutations at that head below, at this head, after the merge of #769

Growth

git diff --shortstat origin/main...82ee7ab86: 6 files, 1047 insertions, 116 deletions. src/ci.rs production +224 −34, tests +118; nightly.yml +17 −2; issues/ +688 −80.

Gates at e0a61d070

gate result
cargo run -- --build-only EXIT=0, Build finished.
cargo test --lib -- --exact the two tests EXIT=0, 2 passed, finished in 10.07s
mutation signal-the-child-alone: applied EXIT=0, cargo test --lib --no-run EXIT=0 builds
the quiet step's test under it EXIT=101: said nothing for 10s and was ended with its process group by SIGKILL; a process outside that group still held its output after it
restored EXIT=0, git status --porcelain empty
mutation the-interrupt-is-not-handed-on: applied EXIT=0, builds EXIT=0 builds
the interrupt's test under it EXIT=101: the step's group outlived its driver by 10s
restored EXIT=0, git status --porcelain empty
cargo run -- --ci host EXIT=0. In its log: === [ci] the build system with test ci::tests::an_interrupt_of_the_driver_ends_the_group_its_step_runs_in ... ok and test ci::tests::a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line ... ok; [ci] Host: 78 step(s), all green

The local --ci host runs under this machine's rustc 1.98.1 with incremental compilation, so it says nothing of the hang; the table in section 1 does.

At aeb396f32, which differs from e0a61d070 only under issues/ (git diff f532a13a6 aeb396f32 --stat: three files, all under issues/): cargo test --lib sourcegate EXIT=0, 10 passed, no_tracked_file_identifies_a_machine_or_its_network ... ok among them; cargo test --lib -- userlandhost every_program_an_architecture_leaves_out EXIT=0, 8 passed, the two other gates that read issues/. git grep of the old slug's bare name at that head finds nothing. No other gate was rerun at it.

At 82ee7ab86, which differs from aeb396f32 in the two issue files only (+18 −9, round 6's three NOTEs): cargo test --lib sourcegate EXIT=0, 10 passed, no_tracked_file_identifies_a_machine_or_its_network ... ok among them. No other gate was rerun at it.

Unmeasured

  • portability-macos green under the pin. The dispatch named above.
  • --ci guest through heard, at this head. The suite's cargo and the harness are a process group under the silence bound; each guest's QEMU is not in it (the harness starts it through tether::spawn, in a session of its own, holding the step's output through its inherited stderr) and is ended by its terminal's hangup when the harness dies. Run 37834436612's tcg / suite succeeded at c952668c9, which is a suite that ends, through heard, one round short of this head; a silent or interrupted suite has not been run.
  • Whether any function of the ToyOS kernel, loader or userland is miscompiled, on either architecture. One function of one x86-64 kernel was read. Nothing in the tree would name a loop compiled this way that returns a wrong value instead of hanging. The defect's exit names the measurement that would.
  • Any LLVM with a fix, upstream's proposed one or another, and so whether it fixes this.
  • The u128 form for the kernel's and the loader's targets, and opt built from the fork's own LLVM on the IR.
  • Whether 1.98.1, 1.99.0 or the runners' own rustc compiles any other host code wrongly. guest.yml and portability-linux install stable and log its version (the tcg / suite of run 37778826093: rustc 1.99.0 (b940084d7 2026-09-28)); the two host jobs take their image's.
  • The reproducer under an affected compiler for aarch64-unknown-none, aarch64-unknown-linux-gnu and x86_64-unknown-linux-gnu: the x86-64 host build measured is x86_64-apple-darwin's.
  • The gates at 82ee7ab86 other than sourcegate: the two other gates that read issues/ are aeb396f32's, and the rest are e0a61d070's.
  • An interrupt of a real driver at a terminal.

Unsure of

  • The bisection's and the root-cause work's evidence is in a scratch directory on the development machine and nowhere else: the step logs, the 40 sources, the pass traces, the IR controls and the saved upstream responses. The issue file carries both reproducers, the IR, the hand-written iterator, the commands and every table, which is what a later fix is held to; the logs behind the rows are not posted.
  • Which sysroot the rows were made from rests on a report. The scripts took the compiler's path as an argument and no log records it; the issue file says so and gives the key.
  • wt/toyos-scevfix was not on the remote when this was written. The defect names it as where the fix is being built, as the orchestrator briefed; nothing here says what that fix is or that it works.
  • The pin. It keeps one test of one job green on a compiler whose LLVM has the same fault, and nothing else: the Linux jobs and every developer's machine take whatever their stable is. Whether the host's rustc should be declared once, as .github/qemu-version declares QEMU, is the owner's to decide and is not done here; its home is issues/the-host-job-runs-the-toolchain-the-runner-ships.md.
  • A hung step under a local --ci host leaves a report in that machine's ~/Library/Logs/DiagnosticReports, and on a Linux host that keeps cores, a core.
  • A step's cargo no longer writes to the driver's terminal, so a local --ci host loses cargo's colour and progress bar.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A

…wer-off issue is closed

The nightly's `portability-macos` sat in `cargo run -- --ci host` until its
`timeout-minutes: 350` cancelled it (run 37778826093, job 113317209315, head
8a883a1): `toyos-userbound`'s `a_port_answers_as_its_declaration_says` never
ended, libtest said so once after 60 s, and the driver waited on cargo with no
bound for the remaining 3 h 36 min. The test's cause is not in the log and is
filed with the measurement that settles it. The unbounded wait is the build
system's, and is fixed here.

`src/ci.rs`: `cargo` and `cargo_logged` both run through `heard`, which reads
the command's output through one pipe and ends a command that says nothing
for `QUIET`, 15 minutes: the command leads a process group, the group is
killed, and the step is red with the last line said. The longest silences
measured in green steps: 78 s in a macOS host job (run 37740449787, a
compile), and in run 37778826093 under 60 s in the Linux `host` and 120 s in
`tcg / suite` (a kernel build). `cargo` used to
hand its child the driver's own stdout and stderr; it now passes the lines
through as `cargo_logged` always did.

`nightly.yml`: the macOS `--ci host` step gets `timeout-minutes: 90`, the
limit the Linux `host` job has, for a step that hangs while it talks.

`issues/the-power-offs-global-lock-line-is-logged-after-the-last-console-drain.md`
is deleted, its exit met: the first nightly `tcg / suite` on a `main` carrying
#767 is run 37778826093 at 8a883a1 (job 113320660798), `PASS
acpi_mediated_access (4s)`, `PASS acpi_lock_given_back_on_one_cpu (4s)`,
`test result: ok. 34 passed, 34 total`. Nothing in the tree cited it. The
order it was about is `kernel/src/power.rs`'s `shutdown`, which says it.

Not built and not run by this commit's author: the request is in the pull
request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Mutation kill-the-child-alone, against 1a89ecb85: a_step_that_goes_quiet_is_killed_with_what_it_spawned_and_names_its_last_line must turn red on it. Not yet run; git apply --check exits 0.

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -214,5 +214,5 @@
             Err(RecvTimeoutError::Timeout) => {
                 // SAFETY: the group the child leads, which no other process can name until it is reaped.
-                let killed = match unsafe { libc::killpg(child.id() as i32, libc::SIGKILL) } {
+                let killed = match child.kill().map_or(-1, |()| 0) {
                     0 => {
                         child.wait().map_err(|e| format!("wait: {e}"))?;

…he report, and the issue has the measurement

The first request ran at 1a89ecb on an Apple-silicon Mac. The hang does not
reproduce there: `a_port_answers_as_its_declaration_says` builds and exits 0
under rustc 1.98.1 and under 1.99.0, at the tree's profile and at opt-level 0.
The issue says so, says what still differs from the job that hung (the tree,
the 19 tests beside it, the hosted image), and that libtest's line cannot tell
a test body that never ends from a test thread that never reports.

`heard` therefore ends a silent group with SIGQUIT and, 10 s later, SIGKILL for
what is left: macOS writes a report with every thread's stack of a process
SIGQUIT ends, and `portability-macos` uploads `~/Library/Logs/DiagnosticReports`
when it fails. That no tool is spawned for it is the point: `sample` is a
binary of one host OS. The test's fixtures ignore SIGQUIT, so the test runs the
escalation and leaves no report on a developer's machine; it takes 20 s.

`--ci host` at 1a89ecb was red in one step, `clippy, warnings denied`:
`clippy::zombie_processes` on the fixture's child, which is now reaped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu Japabu changed the title A step's cargo that goes silent is killed and named; the macOS nightly's hung test is filed; the power-off line's issue is closed A step's cargo that goes silent is ended and named; the macOS nightly's hung test is filed with its measurement; the power-off line's issue is closed Oct 8, 2026
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Mutation kill-the-child-alone, against 34367b5a6, replacing the one posted against 1a89ecb85. Run by the orchestrator: applied EXIT=0, cargo test --lib --no-run EXIT=0, a_step_that_goes_quiet_is_killed_with_what_it_spawned_and_names_its_last_line EXIT=101, restored EXIT=0 with the tree clean.

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -232,5 +232,5 @@
                 for (signal, name) in [(libc::SIGQUIT, "SIGQUIT"), (libc::SIGKILL, "SIGKILL")] {
                     // SAFETY: the group the child leads, which no other process can name until it is reaped.
-                    if unsafe { libc::killpg(child.id() as i32, signal) } != 0 {
+                    if unsafe { libc::kill(child.id() as i32, signal) } != 0 {
                         let why = std::io::Error::last_os_error();
                         return Err(format!("{quiet} and its process group could not be sent {name}: {why}; the last it said: {last}"));

@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 1 of #778 at 34367b5a6 (two commits on 809c33c0c), against .claude/agents/reviewer.md. Read: git log origin/main..34367b5a6, git diff origin/main...34367b5a6, src/ci.rs and nightly.yml whole, both issue files, the logs of jobs 113317209315, 113317209042, 113189702343 and 113320660798, and the orchestrator's round-1 and round-2 request logs. I ran no build and no test.

Growth: git diff --shortstat origin/main...34367b5a6 is 4 files, +283 −110. Production: src/ci.rs +96 −29, nightly.yml +14 −1. Tests: src/ci.rs +62. issues/ +111 −80.

BLOCKER

  • src/ci.rs:205 — a terminal's interrupt no longer ends a step — process_group(0) takes the step's cargo, its rustc and its test binaries out of the foreground group, and the driver neither handles nor forwards SIGINT/SIGHUP. Ctrl-C on a local cargo run -- --ci host now kills the driver and leaves the step's group running: a talking one until its next write to the dead pipe, a silent one for good, and the driver that held the 15-minute bound is gone with it. On main the interrupt reached all of them. The doc comment at :200 names this as "the price of the group"; the body's own reason for the group is that a test binary left spinning stays "on a development machine for good", and this path reintroduces exactly that for the one process this pull request is about. Under a local --ci guest the survivors are the harness and its QEMUs. An interrupt of the driver must end the group it started, with a test that fails without it.
  • src/ci.rs:232, nightly.yml:133, issues/a-port-answers-…-macos-runner.md:96 — the SIGQUIT leg, its GONE wait, the upload step and the issue's "what the next nightly settles" are built on two guesses that one cheap measurement each would settle, on the Apple-silicon Mac the round-1 request already ran on:
    1. That macOS writes a report with every thread's stack for a process SIGQUIT ends. No run has shown one: the test's fixtures ignore the signal on purpose, and the issue states the behaviour as fact. Measure it: SIGQUIT to a test binary that spins, then the .ips under ~/Library/Logs/DiagnosticReports, naming its frames, and how long after the signal the pipe ends (whether report generation holds the output past GONE, which would turn every real kill into by SIGKILL and may cut the report short).
    2. That the hang does not reproduce there. The three arms built 809c33c0c's tree with cargo test -p toyos-userbound --test firmware and ran one test with --exact. The job built 8a883a142 with the step's cargo test --workspace --exclude … and ran the binary's 20 tests on libtest's threads. The issue lists the tree and the neighbours as open differences and omits the invocation; all three close in one command: 8a883a142, rustc 1.99.0, the step's own cargo line, the firmware binary whole. Round 2's local --ci host ran that step green, under 1.98.1 and this head's tree, so the 1.99.0 cell is the one missing. If it hangs there, a debugger on that Mac answers the issue and the SIGQUIT leg, GONE's second use and the upload step are not needed; if it does not, the issue's list shrinks to the machine.
  • Evidence — the pull request's body is of 1a89ecb85 and reports cargo run -- --ci host EXIT=1; the mutation posted as a comment is the round-1 patch, which does not apply to this head. The measurements at 34367b5a6 exist (round 2: the test EXIT=0 finished in 20.02s; the mutation EXIT=101 with a process outside that group still held its output after it; --ci host EXIT=0, [ci] Host: 77 step(s), all green, a_port_answers_as_its_declaration_says ... ok inside it) and I read their logs, but a claim stands on the body at the head. Post the round-2 body and the round-2 mutation.
  • Evidence — nothing has run the guest suite through heard. --ci guest's the suite is now a process group with a silence bound around the harness and every QEMU it starts (their stderr is inherited, so they hold the pipe and are in the group). guest / suite at the head that lands is the first measurement of that and is not in the body.
  • issues/a-port-answers-…-macos-runner.md:108 — status: open, owner "the acpi claim's author (The acpi claim's mediated access and the Global Lock, and acpiserver loading the machine's tables through them #749)": no worktree of that author exists (git worktree list), and the file's own next step, reading an artifact the workflow keeps 7 days, has nobody named to do it. The issue this branch deletes carried the same owner line and needed "Assigned: the orchestrator … he reads" added before anybody read the nightly. Name who reads the first nightly's artifact, and set status: assigned.

NOTE

  • src/ci.rs:1210 — the test waits 20 s of wall clock, and 10 s of it is a wait the test knows will run out: both fixtures ignore SIGQUIT, so gone always expires. buildlock.rs's HEARTBEAT is the tree's precedent for a bound shortened under cfg(test); the same for GONE halves the test, and for QUIET removes heard's quiet parameter, which production gives one value. The 10 s of quiet is also the fixture's deadline to exec and print under whatever load the host carries; a red there under load is a defect by root CLAUDE.md, so do not shorten that one below what a loaded host starts a cached binary in.
  • src/ci.rs:234 — after a SIGQUIT that ended every member, a process outside the group can still hold the pipe (src/tether.rs's setsid children under cargo test --lib are such), and SIGKILL then goes to a group whose only member is the unreaped leader. If macOS answers that with ESRCH, the refusal reads "could not be sent SIGKILL", the "outside that group" diagnosis is lost and the child is never reaped. Not measured on either host; the mutation covers a live grandchild inside the group, not this.
  • src/ci.rs:176 — gone discards every line the group says after the first signal; none of it reaches the step's log.
  • src/ci.rs:229 — silence and "the last it said" are counted in lines, not bytes. With --test-threads 1 (src/tether.rs:246 uses it) libtest runs the test on its main thread, prints test <name> ... without a newline and never prints the 60-second line, so the last line names the test before the hung one.
  • src/ci.rs:166 — 15 minutes holds against what I could read: the longest silence inside a green step is 78 s in a macOS host job (job 113189702343), under 60 s in the cold Linux seal (113317209042) and 120 s in tcg / suite (113320660798, between a BUILD line and its BUILT; the harness prints nothing while a build runs). release::toolchain, bootstrap and install do not go through heard. One caller the body does not weigh: a step's cargo that waits on another cargo's lock in the same target prints one line and is ended after 15 minutes, where main waited.
  • issues/a-port-answers-…-macos-runner.md:106 — "until it is fixed the nightly is red on macOS": root CLAUDE.md gives a red test two outcomes, fixed or deleted with its issue recording the commit that restores it. Keeping it as the subject of one instrumented nightly is a third, and is the owner's to allow; the issue should say what happens after that one nightly.
  • The SIGQUIT leg and the upload step serve this one issue; its exit does not say whether they go with it.
  • What I verified and found true: the issue's table and log excerpt against the four jobs; firmware::port's one loop and the test's inputs; the test absent at 9ef866436 and present at b6bcb9691 and 8a883a142; the image line (macOS 26.6.2, 25G83) and rustc 1.99.0. The close of the-power-offs-global-lock-line-…: exit met by job 113320660798 (PASS acpi_mediated_access (4s), PASS acpi_lock_given_back_on_one_cpu (4s), 34 passed, 34 total), 15b8f9d25 an ancestor of 8a883a142, no citation of the slug or its fragments left in the tree at this head, the rule at kernel/src/power.rs:67. actions/upload-artifact at the commit guest.yml pins on main; the step-level if: failure() is in a job that names no cache. The root crate is already Unix-only (buildlock.rs, tether.rs), so CommandExt and killpg cost no host. The spawn sets the group atomically, the signal names only that group, and the mutation is red for its own reason.

What the checks must show

  • host at the head that lands: the build system with ci::tests::a_step_that_goes_quiet_… ... ok, every step green. That alone is not enough to land on: guest / suite at the same head must end [ci] the suite: test result: ok, because it is the only run of the harness and QEMU under the new group and bound.
  • The next nightly, for the bound to be believed: portability-macos concludes failure, not cancelled; the workspace's host members is red about 16 minutes after libtest's line with said nothing for 900s and was ended with its process group by SIGQUIT … the last it said: test a_port_answers_as_its_declaration_says has been running for over 60 seconds; the steps after it run and the driver prints its summary; no Terminate orphan process … firmware-* at the job's end; and macos-crash-reports holds a firmware-* report or the upload warns that it found none. by SIGKILL there means the report was probably cut. A workflow_dispatch of nightly on this branch shows all of this before landing.

SEND BACK

Japabu and others added 2 commits October 8, 2026 21:32
… longer depends on who started the driver, and the issue is assigned with both measurements

Answers review round 1 of #778.

The hang does not reproduce on an Apple-silicon Mac with the job's own tree,
compiler and invocation: `8a883a142`, rustc 1.99.0, the step's package
selection, the `firmware` binary whole, 20 runs on libtest's default threads
and 20 on three, 40 of 40 exit 0. The issue records it, is `status: assigned`
to the orchestrator, who dispatches `nightly.yml` on this branch and reads the
macOS job, and says what follows that one run: the cause fixed, or the test
deleted with the issue recording the commit that restores it.

`heard`:
- hands the driver's SIGINT, SIGTERM and SIGHUP to the group it started and
  then takes the signal itself. The group is out of a terminal's reach, and
  without this an interrupted driver left its step running, a silent one for
  good.
- starts the command with SIGQUIT at its default. A shell starts a background
  job ignoring SIGINT and SIGQUIT and an ignored disposition survives exec:
  the first measurement of the SIGQUIT report was void for that reason, and a
  driver started that way would have sent a signal nobody took.
- counts silence and "the last it said" in bytes, so libtest on one thread,
  which names a test before running it and ends the line after, still names
  the test that hangs; and passes on what the group says after the first
  signal.
- does not count as silent a cargo whose last line says it waits on another
  cargo's lock.
- sends SIGKILL to the child by pid as well as to the group, and reaps it
  whatever the group answered.
- takes `QUIET` from a constant, 10 s under `cfg(test)`.

The quiet step's test no longer waits out a bound it knows will expire: its
fixtures exit at SIGQUIT by a handler, which leaves no report behind. A second
test interrupts a driver and reads the end of a pipe its step and the step's
child write into.

Not built by this commit's author; the request is with the orchestrator.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
…me, and the handler's cast is the one the compiler asks for

`--ci host` at 8bb2e6d was red in one step, `clippy, warnings denied`:
`function_casts_as_integer` on the three `as libc::sighandler_t` casts, which
now go through `*const ()`.

The window between `spawn` and the store of the group's id is closed without
a signal mask, which would hold only for the thread that set it: the handler
writes the signal it took and then reads the group, `started` writes the group
and then reads the signal, so on whatever thread the handler runs one of the
two sees the other. An interrupt that finds the group still unnamed returns
and is taken by `started`, which hands it to the group and ends the driver.

The issue records the corrected SIGQUIT measurement: a spinning test started
as `heard` starts a step begins with SIGQUIT at its default, its output ends
in the millisecond of the signal, and the 11,980-byte report holds both of its
threads, libtest's main thread and the test's own. Started without the reset
from a shell's background job it begins with SIGQUIT ignored, outlives the
signal by the 30 s it was given and leaves no report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu Japabu changed the title A step's cargo that goes silent is ended and named; the macOS nightly's hung test is filed with its measurement; the power-off line's issue is closed A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hung test is filed with its measurements; the power-off line's issue is closed Oct 8, 2026
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

The two mutations of c952668c9, replacing every patch posted before. Run by the orchestrator at this head: each applied EXIT=0, cargo test --lib --no-run EXIT=0, its test EXIT=101, restored EXIT=0 with the tree clean.

signal-the-child-alone, red in a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line:

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -349,14 +349,14 @@
     };
     let ended = hung.map(|last| {
         // SAFETY: the group the child leads, which no other process can name until it is reaped.
-        unsafe { libc::killpg(group, libc::SIGQUIT) };
+        unsafe { libc::kill(group, libc::SIGQUIT) };
         let mut by = "SIGQUIT";
         let mut gone = said.ended();
         if !gone {
             by = "SIGKILL";
             // SAFETY: as above. The child by its own name too: a group whose
             // members are gone but for an unreaped leader may answer for nobody.
-            unsafe { libc::killpg(group, libc::SIGKILL) };
+            unsafe { libc::kill(group, libc::SIGKILL) };
             let _ = child.kill();
             gone = said.ended();
         }

the-interrupt-is-not-handed-on, red in an_interrupt_of_the_driver_ends_the_group_its_step_runs_in:

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -208,7 +208,7 @@
     // SAFETY: three async-signal-safe calls; the signal raised at its default ends the driver.
     unsafe {
         if group != 0 {
-            libc::killpg(group, signal);
+            let _ = group;
         }
         libc::signal(signal, libc::SIG_DFL);
         libc::raise(signal);

@Japabu

Japabu commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review round 2 of #778 at c952668c9 (four commits on 809c33c0c), against .claude/agents/reviewer.md. Read: git diff 34367b5a6 c952668c9 as new code, src/ci.rs and the issue whole at this head, src/tether.rs and tests/common/qemu.rs where the harness starts QEMU, the body, the patches of comment 6067758543, and the orchestrator's request scripts and logs of rounds 3, 4 and 5 (nightlymac-r3.log, -r4.log, -r5.log and their output directories, the crash report among them). I ran no build and no test.

Growth: git diff --shortstat origin/main...c952668c9 is 4 files, +487 −115. Production: src/ci.rs +221 −34, nightly.yml +14 −1. Tests: src/ci.rs +118. issues/ +134 −80.

Round 1's BLOCKERs

  1. An interrupt no longer ends a step: CLOSED for what round 1 named. nightlymac-r5.log, head c952668c9: the two tests EXIT=0, 2 passed … finished in 10.01s; the-interrupt-is-not-handed-on applied, built, and its test EXIT=101 with the step's group outlived its driver by 10s, the panic at the judge's deadline and nowhere else; restored, tree clean. The test asserts the driver's status is the signal's own (status.signal() == Some(SIGTERM)). The spawn window is not closed: first new BLOCKER below.
  2. The two guesses: CLOSED. nightlymac-r3.log: head 8a883a1428c7…, rustc 1.99.0 (b940084d7 2026-09-28), the step's selection, --test firmware --no-run: EXIT=0 with no fallback line, threads=default: 20 run(s) EXIT=0, hung=no, threads=3: 20 run(s) EXIT=0, hung=no; the 40 run logs each end test result: ok. 20 passed, and a_port_answers_as_its_declaration_says ... ok is in all 40. nightlymac-r4.log: reset: SIGQUIT is default, blocked=0, the output ended 0 ms after SIGQUIT, signal: 3 (SIGQUIT), a report of 11,980 bytes; inherit: SIGQUIT is ignored, not ended 30 s later, signal: 9 (SIGKILL), no report. I read the report: exception.signal is SIGQUIT, and threads has two entries with frames, main in semaphore_timedwait_trap under Thread::park_timeout, the CompletedTest channel's recv and test::console::run_tests_console, and spins in spins::spins under test::run_test and _pthread_start. The request log's own 1 thread(s) with frames is a count of lines in a one-line JSON body; the body's two is what the report holds.
  3. Body and mutation of another head: CLOSED. The body is of c952668c9; both patches in comment 6067758543 are the ones nightlymac-r5.log applied at that head.
  4. Nothing has run the guest suite through heard: OPEN. No run exists; CI is skipped on the draft. What it must show is below.
  5. The issue had no reader: CLOSED. status: assigned, Assigned: the orchestrator, with what he reads and what follows either reading.

Round 1's NOTEs

  • A wait the test knows will run out: CLOSED (finished in 10.01s; the fixtures exit at SIGQUIT by a handler; QUIET is a constant).
  • A group left with its unreaped leader: CLOSED by reading (src/ci.rs:359-361, :368); unmeasured, and nothing here turns on it.
  • Lines after the first signal dropped: CLOSED (Said::ended goes through Said::more).
  • Silence counted in lines: CLOSED by reading (:239-249, :268); a partial last line reaches the terminal as its bytes and the refusal through from_utf8_lossy over the whole log, so nothing is cut inside a sequence. No run of a hung test on one thread exists; libtest's unfinished test <name> ... is read, not measured.
  • A cargo waiting on another's lock: answered by an exemption, which is the second new BLOCKER.
  • "Until it is fixed the nightly is red": CLOSED in the issue, with one line under NOTE.
  • Whether the leg and the upload go with the issue: CLOSED (issues/a-port-answers-…:127-129).

BLOCKER

  • src/ci.rs:322 — the handshake leaves one interleaving in which an interrupt ends the driver and nobody signals the step — the argument at :201 pairs the handler with started only. The handler also reads GROUP == 0, between two steps, and that read is paired with nothing: thread B takes the signal, stores PENDING, loads GROUP as 0; the main thread stores STARTING and forks; B sets the default and raises; the process ends inside spawn, before started runs, with a child in a group of its own that no signal reaches. "One of the two always sees the other" is false of it: heard never looks. It needs a second thread to take the signal, which the driver has whenever an earlier step's reader still holds a pipe a process outside its group kept (:330, after the outside that group verdict the driver goes on), and which the driver of both tests always has; and it needs B to lose the processor between :204 and :214 for the length of a fork. Narrow, and the same orphan round 1's first BLOCKER was about, on the one claim the body rests on an argument alone. The closing is the same pair again: after GROUP.store(STARTING), load PENDING, and if it is set do not spawn (started(0)). Then either the handler saw STARTING and returned, or heard sees PENDING before it forks. :367 has the mirror image (B loads the group, the main thread stores 0 and reaps, B signals a group id that is no longer the step's) and the same reading settles whether it needs the same answer.
  • src/ci.rs:346 — the bound this branch exists for is switched off by a string match, for good — when QUIET expires and the last line begins Blocking waiting for file lock, heard waits again, with no count and no end. That is a wait on an event with no timeout (root CLAUDE.md, "Fail fast": bounded by a timeout that fails loudly), in new code, and the body says so under "Unsure of". The holder it trusts is "another cargo", which is exactly the process this branch shows can sit for 3 h 36 min, and one not started by a driver has no QUIET over it. The tree's own lock does not wait this way: buildlock.rs's take_lock_announcing says who holds it every HEARTBEAT, so its waiter is never silent. No test reaches the arm; with it deleted every test here is green. Round 1 asked that the caller be weighed, not exempted. Delete the arm and WAITING: a step that waits 15 minutes on a lock is then red by name, the last it said: Blocking waiting for file lock on build directory, which is the loud, bounded failure, in fewer lines. If the owner wants a lock waited out longer than a silent step, that is a second named bound that fails loudly, not none.
  • Evidence — guest / suite through heard at the head that lands (round 1's fourth, still open), and the instrumented nightly the issue has the orchestrator read before landing. Run 37834436612, workflow_dispatch at c952668c9, was in_progress when I read it: toolchain / build success, release skipped off main, and host, tcg / suite, portability-linux and portability-macos still running, so there is nothing to judge yet.

NOTE

  • src/ci.rs:268 — the last line said is unbounded in length: a step that writes megabytes without a newline and then hangs puts all of it in the refusal, which step prints twice and summary appends to the job summary. Bound it where it is formatted.
  • issues/a-port-answers-…-macos-runner.md:119 — "That run is the only one this test is kept red for" is false of main once this lands with the test in the tree: every scheduled nightly is red on macOS until the following pull request merges. The reading arrives before this lands, so the outcome it selects (the fix, or the test's deletion with the restoring commit recorded) can be in this pull request, and then the sentence is true.
  • Body, "Unmeasured", and round 1's fourth BLOCKER as I wrote it — "every QEMU it starts are now a process group" is false of the tree: tests/common/qemu.rs:2231 starts each QEMU through tether::spawn, which gives it a session of its own (src/tether.rs:55). A guest's QEMU is outside the step's group and holds the step's output through its inherited stderr (qemu.rs:2226). What ends it is the harness's death, whatever signal caused it, by the hangup of its terminal: SIGHUP, on which QEMU shuts down and gives back its sockets. So under --ci guest a silent suite is ended as cargo and the harness by SIGQUIT and each QEMU by SIGHUP a moment later, an interrupt of the driver likewise with its own signal in place of SIGQUIT, and a QEMU that does not go within GONE reads by SIGKILL; a process outside that group still held its output. The metal rows do not run under --ci. None of this path is measured through heard; a green suite shows only that the pass-through and the bound do not disturb a suite that ends.
  • What I read and found sound: interrupted is two sequentially consistent operations on AtomicI32, then killpg, signal and raise, with no allocation, lock, formatting or panic path, and on its STARTING return it makes no call at all; raise inside the handler stays pending until it returns and then ends the driver at the default, so its parent sees the signal's own status, as the test asserts; a second interrupt during the first signals the group again and ends the driver by the later one; a spawn that fails takes started(0); a driver started ignoring a signal keeps ignoring it and so does its step. For the pair the comment names, every interleaving on one thread and on two has one side see the other. The issue records both measurements as the logs have them. git merge-tree --write-tree of this head with origin/main (35d35859e) and with The host's toolchains live in one store outside every checkout, and every compiler is keyed #769's 9a7d15900 both exit 0 with no conflict: main moved one test's fixture text in src/ci.rs, The host's toolchains live in one store outside every checkout, and every compiler is keyed #769 touches no file this branch touches, and its src/buildlock.rs keeps tests::rerun, which these tests call. No hand resolution, so no new code from either merge.

What the checks must show

  • host at the head that lands: [ci] Host: … step(s), all green, with ci::tests::a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line ... ok and ci::tests::an_interrupt_of_the_driver_ends_the_group_its_step_runs_in ... ok in the build system.
  • guest / suite at the same head: [ci] the suite: test result: ok. N passed, N total, every step of [ci] Guest green, nothing left in $TMPDIR or /tmp among them, and no line said nothing for.
  • Run 37834436612, which is this head: host (--ci seal, a cold Linux tree through heard) green; tcg / suite green by the lines above, which is then a run of the suite through heard at this head; portability-macos concludes failure and not cancelled; its the workspace's host members is red about 16 minutes after libtest's line with said nothing for 900s and was ended with its process group by SIGQUIT; the last it said: test a_port_answers_as_its_declaration_says has been running for over 60 seconds; every later step runs and the driver prints [ci] Host: 1 of N step(s) red; no Terminate orphan process … firmware-* at the job's end; and macos-crash-reports holds a firmware-* report, or the upload warns it found none. by SIGKILL there means the group did not go at SIGQUIT and the report, if any, is suspect.
  • The orchestrator may not land on reading the checks and that run alone: the two BLOCKERs above change heard, and what changes is reviewed, with nightlymac-r5.log's gates repeated at the new head. The run in flight stays a valid reading of the runner's report and of the SIGQUIT leg, neither of which those two changes touch. Once a round finds the code clean, the Evidence BLOCKER closes on the checks' own lines at that head with no further round.

SEND BACK

…pt, the lock exemption is gone, and the last line said is bounded

Answers review round 2 of #778.

The handshake paired the handler with `started` only. The handler also reads
the group as 0, between two steps, and nothing read the pending signal before
the next fork: a second thread could take the signal, read 0, lose the
processor for the length of a fork and then end the driver with a step no
signal reaches. `heard` now reads the pending signal after it writes
`STARTING` and spawns nothing if one is set; and after it writes 0 at a
step's end, before it reaps, so a handler that read the group's id signals a
group that is still the step's. Argued, as before, and tested by nothing.

A cargo whose last line said it waited on a file lock was waited on without
end. That was an unbounded wait by string match; it is deleted, and such a
step is red after `QUIET` with that line as the last it said.

"The last it said" is at most 200 characters of the line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

The two mutations of 433d6d5fd, replacing every patch posted before. Run at this head in the branch's worktree, each as a checked patch: applied EXIT=0, cargo test --lib --no-run EXIT=0, its test EXIT=101, restored EXIT=0 with git status --porcelain empty.

signal-the-child-alone, red in a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line (ended with its process group by SIGKILL; a process outside that group still held its output after it):

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -351,14 +351,14 @@
     };
     let ended = hung.map(|last| {
         // SAFETY: the group the child leads, which no other process can name until it is reaped.
-        unsafe { libc::killpg(group, libc::SIGQUIT) };
+        unsafe { libc::kill(group, libc::SIGQUIT) };
         let mut by = "SIGQUIT";
         let mut gone = said.ended();
         if !gone {
             by = "SIGKILL";
             // SAFETY: as above. The child by its own name too: a group whose
             // members are gone but for an unreaped leader may answer for nobody.
-            unsafe { libc::killpg(group, libc::SIGKILL) };
+            unsafe { libc::kill(group, libc::SIGKILL) };
             let _ = child.kill();
             gone = said.ended();
         }

the-interrupt-is-not-handed-on, red in an_interrupt_of_the_driver_ends_the_group_its_step_runs_in (the step's group outlived its driver by 10s):

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -203,7 +203,7 @@
     // SAFETY: three async-signal-safe calls; the signal raised at its default ends the driver.
     unsafe {
         if group != 0 {
-            libc::killpg(group, signal);
+            let _ = group;
         }
         libc::signal(signal, libc::SIG_DFL);
         libc::raise(signal);

@Japabu

Japabu commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author

Review round 3 of #778 at 433d6d5fd (five commits on 809c33c0c), against .claude/agents/reviewer.md. Read: git diff c952668c9 433d6d5fd as new code, src/ci.rs:150-376 and the issue whole at this head, the body, comment 6067998816, and run 37834436612's job list at 20:02Z. I ran no build and no test.

Growth: git diff --shortstat origin/main...433d6d5fd is 4 files, +490 −115. Production: src/ci.rs +224 −34, nightly.yml +14 −1. Tests: src/ci.rs +118. issues/ +134 −80. This round: src/ci.rs +17 −14, the issue +2 −2.

Round 2's BLOCKERs

  1. The handshake's open interleaving: CLOSED, by proof; no measurement exists and the body says so. PENDING is written by the handler alone and never cleared, and every path that reads it non-zero ends the driver, so the proof is three pairs of store-then-load on two sequentially consistent cells. Handler: P.store (x), G.load (y).
    • Between two steps, against the pre-spawn read (:320, :323): G.store(STARTING) (c), P.load (d). If y < c then x < y < c < d, heard sees the signal and started(0) raises it at its default on the main thread before spawn; nothing is forked. If y > c the handler read STARTING or what followed it, which is the next pair.
    • Against started(group) (:328): the handler read STARTING and returned with its handler still installed; x < y < G.store(group) < P.load, so started signals the group and raises. A spawn that fails is started(0): no killpg, the raise alone. Where :323 already saw it, nothing was spawned.
    • Against the end-of-step write and the reap (:370, :371): G.store(0) (a), P.load (b). If y < a the handler read the step's id and x < y < a < b: the main thread raises inside started(0) and never reaches child.wait(), so the id the handler's killpg names is still the step's unreaped leader for as long as the driver has a thread to call it from. If y > a the handler read 0 and signals nobody, and the next step's read at :323 is the first pair again.
    • One thread: the handler runs whole between two instructions of heard, so it reads one of 0, STARTING or the id, and each is a case above with nothing interleaved. Two threads: the cases above are stated for any thread; raise is directed at the raising thread and at the default disposition ends the process from either.
    • A second signal: killpg precedes signal(SIG_DFL) in the handler, so a second delivery of the same signal either runs the handler again (group signalled twice) or finds the default after the group was already signalled; a different signal overwrites PENDING with another non-zero value and takes the same paths; the STARTING return leaves the handler installed.
    • interrupted called from started runs outside signal context on the main thread, where none of the three signals is blocked: raise does not return.
  2. The lock exemption: CLOSED. The arm and WAITING are deleted; git grep -e WAITING -e 'Blocking waiting' at this head finds nothing in src/, issues/ or .github/. hung's match is three arms, each reached, and Said::last has one caller left with the same meaning. Nothing is dead.
  3. Evidence: OPEN. Below.

Round 2's NOTEs

  • The last line unbounded: CLOSED. Said::last takes 200 chars of a &str that from_utf8_lossy made: the cut is on a scalar boundary by construction, and nothing in it indexes or slices, so it can neither split a sequence nor panic.
  • "The only run this test is kept red for": CLOSED (issues/a-port-answers-…:119-120), on the condition it states, which is the third item under "What closes it".
  • QEMU and the step's group: CLOSED in the body's "Unmeasured".

BLOCKER

  • Evidence — three readings are missing at the head that lands, and the third changes the diff.
    • The implementer's own gates. reviewer.md asks for a measurement in the body at the head, with its command, exit code and log; it does not say who runs it, and the orchestrator ran rounds 1 and 2 only because the implementer could not. So the two tests (EXIT=0, 2 passed, finished in 10.02s) and the two mutations (applied, built, EXIT=101 with the judge's own panic text, restored clean; patches in comment 6067998816 match src/ci.rs at this head line for line) stand as the orchestrator's did. The tree agrees with them: the worktree is clean at 433d6d5fd, and its build output was last written between the commit and the comment. What the body does not carry is a log: each gate is one line. For the tests and mutations the quoted panic lines are the verdict and I accept them. For cargo run -- --ci host the one line [ci] Host: 77 step(s), all green is exactly the grepped summary line Evidence refuses, and it is superseded anyway: the required host check at the head that lands is that log.
    • host and guest / suite: skipped on the draft; no run at this head.
    • Run 37834436612: at 20:02Z portability-macos was 15 minutes into cargo run -- --build-only (2 h 4 min in the hung run), with --ci host and the upload pending; host in --ci seal, tcg / suite in --ci guest, portability-linux in --build-only. Nothing to read yet, and no artifact.

NOTE

  • None.

Whether run 37834436612 stands for 433d6d5fd

It does, for what it is read for. The run is at c952668c9; this round's diff acts in three places: when PENDING is non-zero, which needs an interrupt of the driver and so a job that concludes cancelled, the reading that fails anyway; when a step is silent for QUIET with Blocking waiting for file lock as its last line, which is not libtest's line; and when the last line exceeds 200 characters, which libtest's 80 do not. The SIGQUIT leg, GONE, the reset disposition, the group, the refusal's text and the upload step are byte-identical between the two heads. A failure with the lines below is therefore the reading for this head; a cancelled is not a reading of either.

Its host and tcg / suite are a run of --ci seal and of the suite through heard one commit short of this head, by the same argument undisturbed, and worth reading for that; they are not the checks Evidence names.

What portability-macos must show

  • conclusion failure, not cancelled;
  • in the workspace's host members, about 16 minutes after test a_port_answers_as_its_declaration_says has been running for over 60 seconds: said nothing for 900s and was ended with its process group by SIGQUIT; the last it said: test a_port_answers_as_its_declaration_says has been running for over 60 seconds, with no a process outside that group still held its output. by SIGKILL means the group did not go at SIGQUIT and a report, if any, is suspect;
  • every later step runs, and [ci] Host: 1 of N step(s) red;
  • no Terminate orphan process … firmware-* at the job's end;
  • the upload step ran, and macos-crash-reports holds a firmware-* report or the step warns it found no files.

If a report exists, each thread's frames decide between the issue's three: a test thread inside firmware::port or the test's body (the loop, and the code the runner's CPU ran of it); a test thread in libtest or std outside the body (capture, the channel, thread start); or no test thread, with main parked in recv under run_tests_console as in measurement 2.

What closes it

  1. Run 37834436612's portability-macos read against the list above, the report read thread by thread if there is one, and the issue's table and text given that reading.
  2. This pull request then contains what the reading selects (issues/a-port-answers-…:119-129), and that diff is reviewed as new code:
    • the report names the cause: the fix at its owner, with the measurement behind it: a dispatch of nightly.yml on this branch whose portability-macos is green with the test in the tree. The issue stays until a nightly on main shows the same;
    • no report, or none that names the cause: a_port_answers_as_its_declaration_says deleted from toyos-userbound/tests/firmware.rs, the issue recording the commit that restores it and the instrument still owed; and if the artifact held no report of any process, heard's SIGQUIT leg and nightly.yml's upload step deleted with it, the body, the two tests and the mutation signal-the-child-alone restated for what is left.
  3. At the head that results, out of draft: host green, [ci] Host: … step(s), all green with both ci::tests::… named in round 2 ok in the build system; guest / suite green, [ci] the suite: test result: ok. N passed, N total, nothing left in $TMPDIR or /tmp among [ci] Guest's steps, and no line said nothing for. With The host's toolchains live in one store outside every checkout, and every compiler is keyed #769 landed first the branch is merged with or rebased on it before those run; round 2 found the merge clean, and git merge-tree --write-tree origin/main 433d6d5fd still exits 0.
  4. The two tests and both mutations repeated at that head if item 2 touched src/ci.rs.

SEND BACK

Japabu and others added 2 commits October 9, 2026 00:35
… issue names the cause, and the objcopy reports are filed

Run 37834436612, the dispatch of `nightly.yml` at c952668, gave the reading
the issue waited for. `portability-macos` concluded `failure`: `the
workspace's host members` red 15 minutes after libtest's line, `said nothing
for 900s and was ended with its process group by SIGQUIT`, the later steps
run, `[ci] Host: 1 of 77 step(s) red`, no orphan line, the artifact uploaded.

The report of the `firmware-*` binary has the test's thread at offset 0 of
the test's own function. Built here the way the runner builds it (cargo
under `CI=true` compiles without incremental state, which no earlier
measurement here did), rustc 1.99.0 makes that function one instruction, a
branch to itself, at `opt-level = 2`; at 1 and 0, and under 1.98.1, the
binary ends. The trigger among the test's cases is the dword at port 0xFFFC,
`firmware::port`'s inclusive range over 0xFFFC..=0xFFFF.

So the job installs 1.98.1 instead of `stable`, and the issue, renamed for
what it now says, holds the measurements and the exit: a later stable that
compiles the test to one that ends, and the job back on `stable`.
`firmware::port` is not rewritten around a compiler's fault.

The artifact also held 22 reports of `rust-objcopy` ended by dyld at launch,
at the two moments `--build-only` builds bootstrap: filed. Its `cargo` report
is this branch's own SIGQUIT, and its `loom_sleep` report is the control
`doorbell-kick-relaxed` reaching its verdict by an abort.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu Japabu changed the title A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hung test is filed with its measurements; the power-off line's issue is closed A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is rustc 1.99.0's and its job takes 1.98.1; the power-off line's issue is closed Oct 8, 2026
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

The two mutations of e0a61d070, the same two patches as at 433d6d5fd (git apply --check exits 0 on both after the merge of #769), run again at this head in the branch's worktree: each applied EXIT=0, cargo test --lib --no-run EXIT=0, its test EXIT=101, restored EXIT=0 with git status --porcelain empty.

signal-the-child-alone:

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -351,14 +351,14 @@
     };
     let ended = hung.map(|last| {
         // SAFETY: the group the child leads, which no other process can name until it is reaped.
-        unsafe { libc::killpg(group, libc::SIGQUIT) };
+        unsafe { libc::kill(group, libc::SIGQUIT) };
         let mut by = "SIGQUIT";
         let mut gone = said.ended();
         if !gone {
             by = "SIGKILL";
             // SAFETY: as above. The child by its own name too: a group whose
             // members are gone but for an unreaped leader may answer for nobody.
-            unsafe { libc::killpg(group, libc::SIGKILL) };
+            unsafe { libc::kill(group, libc::SIGKILL) };
             let _ = child.kill();
             gone = said.ended();
         }

the-interrupt-is-not-handed-on:

--- a/src/ci.rs
+++ b/src/ci.rs
@@ -203,7 +203,7 @@
     // SAFETY: three async-signal-safe calls; the signal raised at its default ends the driver.
     unsafe {
         if group != 0 {
-            libc::killpg(group, signal);
+            let _ = group;
         }
         libc::signal(signal, libc::SIG_DFL);
         libc::raise(signal);

@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 4 of #778 at e0a61d070, against .claude/agents/reviewer.md. Read: git diff 353b42999 e0a61d070 as new code (the merge 353b42999 of 6776714db is clean; src/ci.rs differs from 433d6d5fd only in the six test lines #773 brought), both new issues, nightly.yml, toyos-userbound/src/firmware.rs and tests/firmware.rs whole, the body, the job log of run 37834436612, its artifact's 25 reports, and the implementer's local logs for the reproduction and the gates. I ran no build and no test. One thing I did change and undid: asking rustc +1.99.0 -vV for its LLVM made rustup install that toolchain on the development Mac, where the implementer had removed it; I uninstalled it, and rustup toolchain list reads as it did before.

Growth: git diff --shortstat origin/main...e0a61d070 is 5 files, +503 −116. Production: src/ci.rs +224 −34, nightly.yml +17 −2. Tests: src/ci.rs +118. issues/ +144 −80. This round: nightly.yml +3 −1, issues/ +144 −134, no source.

Round 3's BLOCKER

Evidence: OPEN, in two of its three parts.

  • Run 37834436612 read against the list: CLOSED. The job log has libtest's line at 21:56:02Z, the refusal at 22:11:02Z word for word as round 3 asked, by SIGQUIT, no outside that group, === [ci] the kernel's library in the same second, [ci] Host: 1 of 77 step(s) red, no Terminate orphan process, and the artifact uploaded at 80,792 bytes. The firmware-* report is as the body reads it: two threads, main in semaphore_wait_trap under Thread::park, recv and run_tests_console; the test's thread with pc at image offset 5476, symbol offset 0 of the closure's call_once, lr in libtest's __rust_begin_short_backtrace, and standing (:495), firmware::port (:344) and the test's :513 inlined at that address; ended by toyos-build.
  • The implementer's gates at this head: CLOSED. The gate script's log names e0a61d070, a clean tree, the two tests EXIT=0 (2 passed, 10.07 s), each mutation applied, built, EXIT=101 with the judge's panic the body quotes, restored with an empty git status; --build-only EXIT=0; --ci host EXIT=0 with 78 === [ci] steps, both ci::tests::… ok and [ci] Host: 78 step(s), all green. The two mutation patches match src/ci.rs at this head.
  • The dispatch under the pin, and host and guest / suite at the head that lands: OPEN. Neither has run; the body says so.

Is it the compiler or the code

The code is right by Rust's semantics, and the fault is the toolchain's.

  • toyos-userbound/src/firmware.rs:340 refuses port as usize + width.bytes() as usize > IO_PORTS (0x10000, port.rs:14) and QWord before anything is added, so at :343 port + (width.bytes() as u16 - 1) is at most 0xFFFF: for (0xFFFC, DWord) it is exactly 0xFFFF. It cannot wrap, and under this profile it would panic if it did: [profile.dev] leaves debug-assertions on, from which overflow-checks follows.
  • :343 is for port in port..=…, a RangeInclusive<u16>. There is no hand-written counter anywhere in the function. RangeInclusive's iterator ends at u16::MAX by its exhausted flag; that is the case the type exists for.
  • No unsafe reaches it: toyos-userbound/src/lib.rs:30 and toyos-bootmap/src/lib.rs:29 forbid it, Width::bytes is self as u64 (toyos-abi/src/acpi.rs:145), and tests/firmware.rs contains none. Safe code with no overflow has no undefined behaviour for an optimiser to stand on.
  • The test's own loops (:505–:516) are over arrays.

So a function that is one b . is a miscompilation, and the measurements say the same thing independently: the local binary has the runner's hash and the runner's offset (0x1564 = 5476), its disassembly is the single self-branch, and the same source ends under 1.98.1 and under 1.99.0 at opt-level 1 and 0. The case edits isolate (0xFFFC, DWord). firmware::port should not be rewritten, and the pin is the right kind of answer for that job. What is not established is how far the fault reaches: that is the BLOCKER below.

BLOCKER

  • issues/rustc-1-99-0-makes-an-endless-loop-of-a-port-test-on-apple-silicon.md:89-90, body "What this branch does" — "The kernel is not built by this compiler: it takes the fork's" is a guess where an eight-second measurement tells, on the policy function of a trust boundary — the kernel calls this loop at kernel/src/arch/x86_64/acpi_mode.rs:733 with a port and width the acpi claim's holder chooses, and a loop there that does not end for a dword at 0xFFFC is a kernel hang from userland. What -vV says of the three compilers: 1.98.1 is LLVM 22.1.8; 1.99.0 is LLVM 23.1.1; the fork's is rustc 1.99.0-dev on LLVM 22.1.8, upstream's tree of 2 August. So the kernel's compiler shares its rustc generation with the faulting one and its LLVM with the good one, CI builds the guest image the way that hangs (CI set, so no incremental state; [profile.toyos] opt-level = 2), and nothing in the tree drives a dword at 0xFFFC through the kernel's instance. Measure: CARGO_INCREMENTAL=0 cargo +toyos test --locked -p toyos-userbound --test firmware (the row-2 command under the fork's toolchain), exit code and log in the issue; and the disassembly of the kernel's own firmware::port instance from a CI=true build, shown to have the loop's exit. That same run halves the cause between rustc and LLVM, which is what makes the record something to act on. If it hangs, the pin is beside the point and this comes back with the kernel's consequence stated.
  • same issue :77-79 and the body's "Unmeasured" — x86-64 under 1.99.0 is left unmeasured while 1.99.0 on x86-64 already compiles what two checks run: guest.yml:52 and nightly.yml's portability-linux install stable, and the tcg / suite log of the earlier nightly reads stable-x86_64-unknown-linux-gnu installed - rustc 1.99.0 (b940084d7 2026-09-28). ci.yml's host, the required check, takes its image's rustc and builds with CARGO_INCREMENTAL=0 (src/ci.rs:840); the macOS image was on 1.98.1 that day, and the Linux image moves when it moves. If x86-64 has the fault, that roll reds host on every pull request and in the merge queue at this test, 15 minutes in. Measure it on the development Mac without a Linux machine: the row-2 build under 1.99.0 with --target x86_64-apple-darwin --no-run and the test function's disassembly (a self-branch or not), or --target x86_64-unknown-linux-gnu with --emit=asm. If it is clean, the issue says so with the command; if not, the pin has to cover every job that runs this compiler, and the branch comes back.
  • Evidence, round 3's, still open — below, under "What the runs must show".

NOTE

  • same issue :84-86 and the body's "Unmeasured" — "the Linux jobs take the rustc of their runner's image" and "their logs name no version" are false of guest.yml and portability-linux, which install stable and log its version; true only of the two host jobs.
  • same issue — it does not point at issues/the-host-job-runs-the-toolchain-the-runner-ships.md, the open issue whose subject is exactly the decision the body leaves to the owner (pin the host toolchain once, or track what ships), and which this is a second measured case for. That is its home, not the fork-pin track: issues/the-forks-pin-is-a-file-and-a-worktree-checks-no-fork-out.md is about the fork's commit, not the host's stable. One sentence naming the file.
  • same issue :101-105 — "Owner of the pin: whoever next changes the rustup step" is nobody, and nothing prompts the exit: the pinned job is green on 1.98.1 for as long as nobody looks. Name the holder of the exit (the issue is assigned to the orchestrator; say he holds this too) and its trigger, each stable release, row 2 rerun. Two more rows cost seconds and belong in the table now: row 2 under nightly-2026-07-22 and under nightly-2026-09-25, both installed on the development Mac — the second says whether upstream already compiles it right, which is whether the exit is reachable at the next stable.
  • same issue — it does not say that the fork takes LLVM 23 at its next move to upstream, at which point the kernel's compiler is the faulting generation: one line for whoever moves the fork, naming row 2 under the moved toolchain.
  • same issue :72-75 — "the crate boundary and the test's other cases are part of what the optimiser needs" is stated as found; the reproducer differs from the test in more than that (a u16 width for the Width enum, three Standing arms for five, no write), and none was isolated. Say it is unknown. Say also that no upstream report is sent for now, so nobody reads the absence of one as an oversight.
  • issues/the-bootstraps-own-build-cannot-start-rust-objcopy-on-an-apple-host.md:30-32 — "Not known: … whether a developer's Mac does the same" is answered by a directory listing: the development Mac's own crash reports hold 18 of rust-objcopy, dated 4, 7 and 8 October, every one Library not loaded. And 8 of the artifact's 22 name rustc as the parent, which the "no run has shown" sentence can carry. By the file's own last sentence a host that still does it is a defect; with two hosts doing it, file it as one or say why not.
  • body, "Review round 3" table — its last row's "below" is true; its second row should carry the result of the dispatch once it exists.

The workflow edit

nightly.yml:120-124: one argument changed in a step the host-tools issue already declares (issues/the-build-runs-host-tools-outside-rust-and-qemu.md, the portability-macos rustup row); no new tool, no new action, the cited issue exists, runs-on is hosted. The runner's rustup already holds a stable at 1.98.1 and the old step updated it in place; whether --default-toolchain 1.98.1 over that installation leaves 1.98.1 the default the two cargo run steps take is for the dispatch to show, not for reading.

What the runs must show

The dispatch of nightly.yml on this branch, portability-macos:

  • concludes success;
  • the rustup step's log names rustc 1.98.1 (48a229cea 2026-09-01) as what it installed or kept as default, and no line moves anything to 1.99.0;
  • in the workspace's host members: test a_port_answers_as_its_declaration_says ... ok, with no has been running for over 60 seconds for it;
  • [ci] Host: N step(s), all green, no said nothing for anywhere, and the upload step skipped.
    It stands for the head that lands if that head differs from the dispatched one only under issues/.

On the pull request, out of draft, at the head that lands: host green, [ci] Host: … step(s), all green, with ci::tests::a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line and ci::tests::an_interrupt_of_the_driver_ends_the_group_its_step_runs_in ok in the build system and test a_port_answers_as_its_declaration_says ... ok; guest / suite green, [ci] the suite: test result: ok. N passed, N total, nothing left in $TMPDIR or /tmp among [ci] Guest's steps, and no said nothing for.

Whether the orchestrator may land on reading them. Yes, on three conditions together: both measurements of the first two BLOCKERs are in the issue with command, exit code and log and both say unaffected (the fork's toolchain exits 0 on row 2 and the kernel's instance has its exit; the x86-64 function is not a self-branch); the diff from e0a61d070 to the head that lands is confined to issues/; and the three runs read as above. If either measurement says affected, or anything outside issues/ changes, it comes back for a round.

SEND BACK

Japabu and others added 2 commits October 9, 2026 01:41
…s instance has its exit; the objcopy reports are a defect

Round 5's measurements of #778, on an Apple-silicon Mac at e0a61d0, each
row the issue's second (`CARGO_INCREMENTAL=0 cargo test --locked -p
toyos-userbound --test firmware`, the binary bounded at 120 s):

- the fork's toolchain (`rustc 1.99.0-dev`, LLVM 22.1.8, the store's sysroot
  as `RUSTUP_TOOLCHAIN`): build exit 0, the binary hung and was ended by PID;
  the test's function is one branch to itself.
- nightly-2026-07-22 (1.99.0-nightly, LLVM 22.1.8): hung. So the fault is
  not LLVM 23's.
- nightly-2026-09-25 (1.100.0-nightly, LLVM 23.1.1): exit 0.
- 1.99.0 with `--target x86_64-apple-darwin`: build exit 0, the test's
  function is a return, and the binary exits 0 under Rosetta.

The kernel's own instance of `firmware::port`, from `CI=true cargo run --
--build-only` (exit 0) with and without incremental state, has the loop's
exit in both: a function of its own in the first, inlined into
`acpi_mode::port` in the second.

The issue records these, says what is unknown of the fault's reach, names
which jobs take which rustc, points at the host-toolchain issue, and names
the holder and trigger of the exit. The `rust-objcopy` finding becomes a
defect: the development Mac holds 18 such reports of its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
…er has it for AArch64: one issue for the defect, one for the CI pin

A bisection of upstream nightlies, the same test binary built without
incremental state and run whole on an Apple-silicon development machine:

- last good nightly-2026-07-09 (14cae6813), first bad nightly-2026-07-10
  (af3d95584), both LLVM 22.1.8. The range holds rust-lang/rust #155114,
  which rewrote RangeInclusive's `next` onto Step::forward_overflowing.
- nightly-2026-09-25 hangs: its earlier "exit 0" was a script that ran an
  empty binary path. nightly-2026-10-08, the newest, hangs too. No upstream
  compiler has a fix.
- that `next` written by hand is miscompiled by stable 1.91.0, 1.95.0,
  1.98.0 and 1.98.1 (LLVM 21.1.2 to 22.1.8) and compiled right by 1.88.0
  (LLVM 20.1.5): the fault is LLVM's, and 1.98.1 escapes the tree's test
  only because its `core` lacks the shape.
- -opt-bisect-limit names `indvars`: it replaces the loop's exit at the
  type's maximum with `false`.
- the fork's own compiler turns a 30-line safe reproducer into a branch to
  itself for aarch64-unknown-toyos, aarch64-unknown-none-softfloat and
  aarch64-unknown-uefi, and compiles it right for the three x86-64 targets;
  0 of 40 sources went wrong on x86-64.

The issue this branch carried said nightly-2026-09-25 was good and looked
for its exit in a later stable. It is now two files. The new one is the
defect, the toolchain's: the reproducer as text with its command, the
compilers measured, the cause as far as measured and what is not
identified, the reach, and an exit a fix in the fork's LLVM meets. The
existing one keeps its slug, which `nightly.yml` cites, and is the pin's:
what the runner showed, the stable compilers' rows, what the pin is worth
now that 1.98.1 is known to carry the same LLVM, and the pin's own exit.

No test is added. A check that compiles the reproducer with the tree's
compiler for the AArch64 targets would be red today, and a red test is
fixed or deleted; it arrives with the fix, green.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu Japabu changed the title A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is rustc 1.99.0's and its job takes 1.98.1; the power-off line's issue is closed A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is an LLVM miscompile rustc 1.99.0 exposes, its job takes 1.98.1, and the fork's compiler has the fault for AArch64; the power-off line's issue is closed Oct 9, 2026
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 5 of #778 at f532a13a6, against .claude/agents/reviewer.md. Read: git diff e0a61d070 f532a13a6 (three files, all under issues/), the three issue files and nightly.yml whole at this head, the body, the bisection's report with its step logs, scripts, reproducers, sweep and pass-bisection outputs, the round-5 measurement logs (the fork's compiler on the tree's test, the two kernel disassemblies, 1.99.0 for x86_64-apple-darwin), and the root-cause report that arrived during this round with its IR controls, width sweep and saved upstream responses. I ran no build and no test.

Growth: git diff --shortstat origin/main...f532a13a6 is 6 files, +813 −116. Production: src/ci.rs +224 −34, nightly.yml +17 −2. Tests: src/ci.rs +118. issues/ +454 −80. Since e0a61d070: issues/ only, no source.

Round 4's BLOCKERs

  1. The fork's compiler and the kernel's instance: CLOSED as a measurement, and it says affected. The fork's compiler (rustc 1.99.0-dev, LLVM 22.1.8) on the tree's test at e0a61d070: BUILD EXIT=0, 20 tests ok, RUN: HUNG, ended by PID after 120 s (wait status 137). The x86-64 kernel from --build-only, both EXIT=0: acpi_mode::port at 0x13a bytes without incremental state and firmware::port at 0x11a with it; in the first the loop is dec r13w / je out at eef58–eef5c, both branches back to eef55 are conditional, and every unconditional jump goes forward or to the epilogue. Round 4 said this comes back with the consequence stated if it hung; it has, as a defect file.
  2. x86-64 under stable 1.99.0: CLOSED. --target x86_64-apple-darwin, BUILD EXIT=0, test result: ok. 21 passed, RUN EXIT=0, the test's function six instructions ending in ret.
  3. Evidence: OPEN. Run 37862964181 is in_progress: host (--ci seal) and tcg / suite success, toolchain / build success, portability-macos and portability-linux still inside cargo run -- --build-only. host and guest / suite on the pull request are skipped on the draft.

Round 4's NOTEs: all seven CLOSED at this head (the Linux jobs' sentence, the pointer at issues/the-host-job-runs-the-toolchain-the-runner-ships.md, the pin's holder and trigger, the two nightly rows, the line for whoever moves the fork, the unisolated "crate boundary" sentence gone, the objcopy file promoted to a defect with the development machine's 18 reports).

What the bisection's logs say of the two files

Checked and true: the committed reproducer is byte-identical to repro/minns.rs; the seven target rows (caller is br label to itself with no ret for aarch64-unknown-toyos, -none-softfloat, -uefi and aarch64-apple-darwin; ret i1 true and movb $1, %al / retq for the three x86-64 targets; each rustc exit 0); every row of the nightly table against its step log (rustc hash, LLVM version, RUN EXIT=0 or HUNG, the one-instruction function for 07-10, 09-25 and 10-08); last good nightly-2026-07-09 (14cae6813), first bad nightly-2026-07-10 (af3d95584), adjacent, seven steps; the compare of the two is 106 commits and 189 files, none under src/llvm-project, with #155114 as 71c64160bd0f; the hand-written iterator's rows (1.88.0 has a ret and prints true; 1.91.0, 1.95.0, 1.98.0, 1.98.1, the March nightly and the fork's have none on AArch64; 1.98.1 and the fork's are right for x86-64); pass-bisection limit 690 prints true, 691 prints false, pass 691 is indvars on the loop in probe, and the diff is the one quoted; the sweep is 39 files, 18 without an exit for each AArch64 target and none for either x86-64 target; the three fragility sentences each have a sweep variant that compiles right; rust/ is 6d6ad8c71906 with src/llvm-project at ceaf0fbb8440 on the branch named. The old claim that a later nightly is good is gone from the tree: the earlier RUN EXIT=0 for nightly-2026-09-25 was the status of an empty binary path after find: … No such file or directory, and the rerun hangs. rustup toolchain list before and after the bisection is identical.

BLOCKER

  • issues/the-forks-compiler-drops-the-exit-of-an-inclusive-range-loop-on-aarch64.md (slug, :7, :262-276) — the file's name and its exit are refuted by a measurement the root-cause report holds, and issues/README.md has a refuted slug renamed with every citation moved — the fork's compiler leaves caller with no ret for x86_64-unknown-toyos when the reproducer's integer is u128 (cls/summary.txt: u128 x86_64-unknown-toyos: rustc rc=0 caller: lines=7 ret_true=0 rets=0, and the IR is the same self-branch). So "on aarch64" is false, and the exit is wrong three ways:
    • it holds the fix to "the indvars step". The fault is ScalarEvolution's: it copies the increment's nuw onto the header phi's recurrence unconditionally, and indvars only asks the question. The controls say so: the 14-line m1.ll through opt -passes=indvars (LLVM 22.1.8) folds the exit to false under no data layout, AArch64's and x86-64's; the same without nuw (m4) and with the add in the header (m5) do not fold; and SCEV's own print gives {-3,+,1}<nuw> the range [-3,0) beside an exit count of 3. A patch to indvars that made the reproducer return would meet the exit as written and leave the fault.
    • its first measurement reads three AArch64 targets and no x86-64 one. It must read x86_64-unknown-toyos with the u128 form too, and that form's source has to be in the file.
    • its measurements are two Rust sources the file itself calls fragile. m1.ll is not: it is the fault with no front end in the way. The file should carry it as text, with what a fixed opt leaves of %done, so that the exit does not rest on a shape rustc may stop making.
  • Evidence, round 3's, still open — under "What the runs must show".

NOTE

Made false or incomplete by the root-cause report, each to be corrected in the same text round:

  • same file :9-15, :187-189 — "for aarch64-unknown-toyos, aarch64-unknown-none-softfloat and aarch64-unknown-uefi" and "Exposed: the AArch64 kernel, loader and userland" — both architectures are exposed; x86-64 is measured wrong for u128.
  • same file :154, :173-174 — "The step that goes wrong is indvars" — indvars is where it shows. The chain: CorrelatedValuePropagation marks the add nuw (sound, the wrapped value is dead), LoopRotate moves it to the latch (sound), SCEV transfers the flag to the pre-increment recurrence (the fault), indvars folds the exit. Say which of these is measured and which is read from LLVM's source; the report says the last step's path is read.
  • same file :176-183 — all three "not identified" items are answered or wrong. The pass that left nuw is CVP. "Between 20.1.5 and 21.1.2" is not where the fault is: under 1.88.0 the overflowing_add stays an intrinsic call so no add exists to flag, which is an escape by shape exactly as 1.98.1's is; the row at :143 should say so and not read as LLVM 20 being right. And upstream has a report: [SCEV] Long-standing miscompile due to absence of per-use flags in SCEV expressions llvm/llvm-project#175729, open, labelled miscompilation, with pull request #118959 open and unmerged since December 2024. Say also what is not established: the maintainer's sentence there is "probably the known issue", after "I've only glanced at it", and nobody has built an LLVM with #118959 to see that it fixes this.
  • same file :191-198 — "What is at risk is a range a..=b whose end is its integer type's maximum" and "Seen for u16 only" — one instance of a wider class (an increment whose wrap is dead, a rotation that makes it a header phi's increment, a start of known range, a SCEV client reasoning about a sibling value); u128 is seen on both architectures; i16 gave a non-constant caller nobody checked.
  • same file :204-211 — "x86-64: 0 of 40 sources miscompiled, which is not a proof" — there is now a counterexample, and a reason the forty escaped: all are u16, u32 or u64, and on x86-64 indvars rewrites a legal-width counter's exit test first, which gives the add a second use so CVP never flags it. Keep the count; put the u128 row beside it.
  • same file :237-239 — "nothing was looked for" stays true and is now the larger gap. The report names the measurement that would look: the fix behind an LLVM option, the tree built twice, a per-function diff. No statistic or remark in today's compiler counts it. Name that as the measurement owed, so the exit is not met by two reproducers while the tree's own functions go unread.
  • same file :241-250 — add what the report settles: waiting for upstream has no date, the proposed fix is 22 months unmerged.
  • same file :55-58 — the command runs every target from the sysroot key the ToyOS targets take; a worktree's target/.deps-stamp gives aarch64-unknown-none-softfloat and aarch64-unknown-uefi another key. As measured, so true; say it, because the exit's reader will otherwise use the kernel's key and get a different library set than the row was made with.
  • issues/rustc-1-99-0-makes-an-endless-loop-of-a-port-test-on-apple-silicon.md:16 — the citation moves with the renamed slug. :102-104 — "Every nightly from nightly-2026-07-10 through nightly-2026-10-08 … hangs" — five were run (07-10, 07-13, 07-22, 09-25, 10-08); the defect file's "No upstream compiler has a fix" at :106-107 is the same overreach. :133-134 — "one arrives when upstream's LLVM compiles the loop right" — or when core's shape moves again, which is how 1.98.1 passes today.
  • Body — false at this head against the root-cause report: the title's and first paragraph's "for AArch64" as the whole reach, "The fault is LLVM's indvars", "compiled right by 1.88.0 (LLVM 20.1.5)" as a bound on the fault, "Not identified … whether upstream has a report (none was searched for)", "Seen for u16 only", "0 for x86_64-unknown-none or x86_64-unknown-toyos" without the u128 row, and the exit as "a fix of the indvars step".
  • The defect is status: open with nobody holding it, which the tracker allows. It is a miscompile in the compiler of every ToyOS binary on both architectures: the orchestrator puts it before the owner by name when this lands, since work on it starts only on his go.

Is the pin still the right change

Yes, and it lands as it is. What is known now makes the pin worth less and no less proper.

  • The red was never the test's or the code's. firmware::port is right, the test is right, and round 4 walked why. "Fixed, or deleted with its issue" has no honest application to it: the fix is in LLVM, which this pull request does not own and upstream has not made, and deleting the test removes the only thing in the tree that has ever caught this fault, from every host where it is green.
  • Leaving stable keeps the job red every night on one upstream bug. A job that is always red shows nothing of macOS: a second red reads the same as the first.
  • Rewriting the loop, ignoring the test on one host or lowering the profile's optimisation would each be a workaround in the product. The pin changes one word of a workflow and nothing a program is built from.
  • What the job exists to show, that the host suite's macOS arms build and run under a stock compiler, it still shows, under a stock compiler one release back. What the pin hides is that the current stable hangs one test there, and anything else 1.99.0 alone would red. That is a recorded compromise in root CLAUDE.md's sense, with a holder, evidence and an exit, and the record does not flatter it: "1.98.1 has the faulty LLVM too … The pin keeps one test of one job green. It is no statement that 1.98.1 compiles the rest of the host suite right."
  • It is honest only beside the defect file. The pin's issue must keep pointing at it, and the defect must not close on the pin's exit: at this head neither does.
  • The root-cause report strengthens this: the fault is in every LLVM the stables carry and upstream's fix is not merged, so there is no compiler to move the job to, and no stable for the other jobs that is free of the class either. Whether every job's host compiler is declared once stays issues/the-host-job-runs-the-toolchain-the-runner-ships.md's and the owner's.

No test now

Right. The cheapest tier that reaches the defect is a host check on the tree's own compiler's output, and no type or reading sees that. It would be red on main from the commit that added it, and a check written to pass while the bug is present is a gate that guards nothing and has to be inverted by hand. It arrives with the fix, green. What that leaves owed is that the checks survive until then: they are text in the issue, so the u128 source and m1.ll go in beside minns.rs, as the BLOCKER above asks.

What the runs must show

Run 37862964181 (e52e6275b; git diff --stat e52e6275b f532a13a6 is the two issue files), portability-macos:

  • concludes success;
  • the rustup step's log names 1.98.1 as installed and default, and no line in the job installs or selects 1.99.0; where the build or the driver prints the compiler it ran, that line says 1.98.1 too. The image ships its own rustup, so the install step alone does not prove which compiler cargo took;
  • in the workspace's host members: test a_port_answers_as_its_declaration_says ... ok, and no has been running for over 60 seconds for it;
  • [ci] Host: N step(s), all green, no said nothing for anywhere, the upload step skipped.
    It stands for the head that lands if that head differs from e52e6275b only under issues/. portability-linux is not what this change is measured by; read its conclusion and report a red.

On the pull request, out of draft, at the head that lands:

  • host: green; [ci] Host: … step(s), all green; in the build system, ci::tests::a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line ... ok and ci::tests::an_interrupt_of_the_driver_ends_the_group_its_step_runs_in ... ok; test a_port_answers_as_its_declaration_says ... ok; and the gates that read issues/, sourcegate among them, green, since the last two rounds changed nothing else.
  • guest / suite: green; [ci] the suite: test result: ok. N passed, N total; nothing left in $TMPDIR or /tmp among [ci] Guest's steps; no said nothing for.

Whether the orchestrator may land on reading them. Not at this head: the first BLOCKER is in the tree. After the text round, yes, on three conditions together: that round's review finds the defect file renamed, its exit and reach corrected and nothing new false; the diff from f532a13a6 to the head that lands is confined to issues/; and the three runs read as above at that head. No further round on src/ci.rs or nightly.yml is owed unless one of them changes.

SEND BACK

…: the defect is renamed, its cause, reach and exit rewritten

Review round 5 of #778 refuted the defect file's slug and exit by the
root-cause measurements: the fork's compiler leaves `caller` with no `ret` for
`x86_64-unknown-toyos` when the reproducer's integer is `u128`, and the fault
is not `indvars`.

`issues/the-forks-compiler-drops-the-exit-of-an-inclusive-range-loop-on-aarch64.md`
becomes
`issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md`,
and its one citation, in the pin's file, moves with it. The pin's file keeps
its slug, which `nightly.yml` cites.

The defect file now carries:

- the fault as 14 lines of IR with no front end, what `opt -passes=indvars`
  makes of it and what a fixed one leaves;
- the cause as four steps (CorrelatedValuePropagation marks the increment
  `nuw`, LoopRotate makes it a header phi's, ScalarEvolution copies the flag
  onto the phi's recurrence, `indvars` folds the exit), each saying what is
  measured and what is read from LLVM's source;
- the `u128` source beside the `u16` one, the width table for both ToyOS
  targets, and the sysroot key the rows were made with;
- 1.88.0 as an escape by shape and not a bound on the fault;
- upstream: llvm/llvm-project#175729 open, pull request #118959 unmerged for
  22 months, the maintainer's hedged sentence, and that nobody has built an
  LLVM with the proposed fix;
- reach on both architectures, the class the range is one instance of, why
  forty sources escaped on x86-64 and the one that does not;
- an exit held to ScalarEvolution and to no client of it, measured by the IR,
  by both Rust forms for both ToyOS targets, by the tree's test, and by a
  per-function comparison of the tree built with the fix given and withheld.

It is `status: assigned`: the toolchain holds it, and a fix in the fork's LLVM
is being built on `wt/toyos-scevfix`.

The pin's file: five nightlies were run, not every one; its exit can also be
met by `core` moving the loop's shape again, which fixes nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu Japabu changed the title A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is an LLVM miscompile rustc 1.99.0 exposes, its job takes 1.98.1, and the fork's compiler has the fault for AArch64; the power-off line's issue is closed A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is an LLVM ScalarEvolution miscompile rustc 1.99.0 exposes, its job takes 1.98.1, and the fork's compiler has the fault on both architectures; the power-off line's issue is closed Oct 9, 2026
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Answer to review round 5, at aeb396f32. git diff f532a13a6 aeb396f32 --stat names three files, all under issues/: no source, test or workflow changed. I ran no compiler on any reproducer: every number below was checked against the root-cause report's and the bisection's files (the IR controls and their outputs, the pass traces, cls/summary.txt and the per-type caller IR, the saved upstream responses), and the cited LLVM lines against the fork's source at ceaf0fbb8440.

BLOCKER 1: the slug and the exit

  • Renamed. issues/the-forks-compiler-drops-the-exit-of-an-inclusive-range-loop-on-aarch64.md is now issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md. Its one citation, in the pin's file, moved. git grep -e the-forks-compiler-drops -e inclusive-range-loop-on-aarch64 aeb396f32 finds nothing (exit 1), the bare name included. The pin's file keeps its slug, which nightly.yml cites.
  • The exit is held to ScalarEvolution, "by a change to ScalarEvolution and to no client of it", and "What does not end it" says why a change to indvars does not.
  • m1.ll is in the file as text, with what opt -passes=indvars leaves today (br i1 false, under the three data layouts) and what a fixed one leaves of %done. It is the exit's first measurement.
  • The u128 source is in the file beside minns.rs, and the exit's second measurement reads caller for aarch64-unknown-toyos and x86_64-unknown-toyos in both.
  • The exit keeps the tree's test binary and adds the measurement over the tree's own functions as its fourth item.

BLOCKER 2: evidence

Not mine to close in a text round: the dispatch and the two checks out of draft are the orchestrator's. The body's "Unmeasured" still says so.

NOTEs

  • :9-15, :187-189, both architectures. The opening and "How far it reaches" say the kernel, the loader and userland are exposed on both.
  • :154, :173-174, the chain. "The cause" is the four steps, each with what is measured and what is read from LLVM's source. Step 4's fold is measured (the pass bisection and the controls); the path inside it is read.
  • :176-183, the three "not identified". Gone. CorrelatedValuePropagation is named; 1.88.0's row says right by shape, with the intrinsic call as the reason, and that it says nothing of LLVM 20.1.5's ScalarEvolution; a new "Upstream" section has #175729 open and #118959 an unmerged draft, and under "Not established" the maintainer's sentence quoted with "I've only glanced at it" and "probably", that nobody upstream has seen ToyOS's loop, and that nobody has built an LLVM with #118959.
  • :191-198, the class. Stated as the class with the range as one instance; the width table has u128 wrong on both and the i16 row as unchecked.
  • :204-211, x86-64. The count is kept with the u128 counterexample beside it and the reason the forty escaped. One correction to the review's wording, from the sweep's files: I could not show that every one of the forty loops over u16, u32 or u64 without reading each loop, so the file says what a count shows, that none names an integer wider than 64 bits, and that the mechanism is read from the x86-64 trace of minns.rs and inferred for the other thirty-nine.
  • :237-239, the measurement owed. "Not measured" says no statistic or remark counts it and why, and the exit's fourth item is the differential build.
  • :241-250. "Waiting for upstream" is in "What does not end it": no date, 22 months unmerged.
  • :55-58, the key. The file names the key the rows were made with, 8618c089fa736cb0, the other key the stamp gives the four kernel and loader targets, and that no log records the compiler's path: it rests on the bisection's report. One measurement of mine beside it: the core crate hash in the fork's compiler's pass traces for the two ToyOS targets is in that directory's two ToyOS libcore files.
  • The pin's file. :16 cites the new slug. :102-104 says five nightlies were run and names them. :133-134 says the exit can also be met by core moving the shape again, which fixes nothing. The defect file's "No upstream compiler has a fix" is replaced by the same five.
  • Body. Rewritten at this head: title, first paragraph, cause, upstream, the width table, reach, the exit, and "Unmeasured".
  • The holder. status: assigned: the toolchain holds it, and a fix in the fork's LLVM is being built on wt/toyos-scevfix, as the orchestrator briefed. That branch was not on the remote when I pushed.

Gates at aeb396f32

gate result
cargo test --lib sourcegate EXIT=0, 10 passed, no_tracked_file_identifies_a_machine_or_its_network ... ok
cargo test --lib -- userlandhost every_program_an_architecture_leaves_out EXIT=0, 8 passed
git status --porcelain --ignore-submodules=none empty

No other gate was rerun: the diff is issues/ only.

@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 6 of #778 at aeb396f32, against .claude/agents/reviewer.md. Narrow: git diff f532a13a6 aeb396f32 (three files, all under issues/), the two issue files whole at this head, the body, and the root-cause report with its files (the IR controls and their outputs, the analysis prints, the pass traces, the width sweep, the saved upstream responses), plus the cited LLVM lines in the fork's source at ceaf0fbb8440. src/ci.rs and nightly.yml were not re-read. I ran no build and no test.

Growth: git diff --shortstat origin/main...aeb396f32 is 6 files, +1038 −116. Production: src/ci.rs +224 −34, nightly.yml +17 −2. Tests: src/ci.rs +118. issues/ +679 −80. Since f532a13a6: issues/ only, +509 −284.

Round 5's BLOCKERs

  1. The defect file's slug and exit: CLOSED.
    • Renamed to issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md. git grep -e the-forks-compiler-drops -e inclusive-range-loop-on-aarch64 aeb396f32 exits 1; the new slug has one citation, in the pin's file at :16; the pin's slug is still cited by nightly.yml:121 and by the defect file.
    • The exit is held to ScalarEvolution "and to no client of it", and "What does not end it" refuses a change to indvars by name.
    • m1.ll, minns.rs and c_u128.rs in the file are byte-identical to the report's files. The three m1 outputs (no data layout, AArch64's, x86-64's) each have the header br i1 false, and each stderr file is empty.
    • The exit reads x86_64-unknown-toyos with the u128 form.
  2. Evidence, round 3's: OPEN. Run 37862964181 (e52e6275b) is in_progress: host, tcg / suite and toolchain / build success, portability-macos and portability-linux not concluded. host, guest and toolchain on the pull request are skipping on the draft.

Round 5's NOTEs

All CLOSED at this head, each against the report's files:

  • Both architectures in the opening and in "How far it reaches".
  • The four-step chain with measured and read told apart. The AArch64 trace has the increment as add i16 before CorrelatedValuePropagation on port and add nuw i16 after it, add nuw i16 %0, 1 after LoopRotate on probe, and the br i1 false with add nuw nsw after indvars on probe. Both analysis prints say {-3,+,1}<nuw><nsw> with U: [-3,0), and m2's says exit count for header: i16 3.
  • Every cited line is where the file says in the fork's source: LazyValueInfo.cpp:1821, ScalarEvolution.cpp:5776-5783, :5798, :5879-5909, :11490, SimplifyIndVar.cpp:275 and :46, IndVarSimplify.cpp:964. That source has no canPreservePreIncAddRecNoWrapFlags and no optimisation remark in IndVarSimplify.cpp, and src/llvm-project is one commit.
  • The controls: folded in m2, at i8, i32 and i64, and with range(i16 10, 100); not folded without nuw, with the add in the header, or with an unknown start; the same under the three layouts; of the eleven passes only indvars leaves br i1 false.
  • The three "not identified" items are gone. 1.88.0's row is an escape by shape: in its trace the loop from 0xFFFC keeps llvm.uadd.with.overflow.i16 in every dump, and 1.91.0's has add nuw i16 after CorrelatedValuePropagation on port.
  • Upstream, against the saved responses: #175729 open, opened 2026-01-13, labelled miscompilation; #118959 open, draft, unmerged, opened 2024-12-06, updated 2026-03-26, body "Test updates incomplete", ScalarEvolution.cpp +73 −37, ScalarEvolution.h +10 −4, 22 test files. Both quoted sentences are verbatim, and the trace through eliminateIVComparison, getMinusSCEV and isKnownNonZero is the reporter's own comment.
  • The class, the width table (cls/summary.txt row for row; i16 is 41 lines with one ret, i8 does not compile), the u128 counterexample beside the forty, and the implementer's correction of my wording there, which is right: what the files show is that none names an integer wider than 64 bits.
  • The measurement owed is the exit's fourth item; -stats output is 0 bytes and -debug-only=indvars is refused.
  • The keys: the worktree's target/.deps-stamp gives the two ToyOS targets 8618c089fa736cb0 and the other four dc7c468f0c07446e, and the core hash in each ToyOS trace is in one libcore under the first.
  • The pin's file: the citation moved, the five nightlies named, the exit's second way in.
  • The holder: status: assigned and a named task satisfy issues/README.md. wt/toyos-scevfix is a worktree on the development machine at 6776714db with no commit of its own in this repository yet; the body says it is not on the remote.
  • The body: title, cause, upstream, tables, reach, exit and "Unmeasured" agree with the files at this head; the growth figures add up.

BLOCKER

  • Evidence, as above. Nothing new.

NOTE

  • issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md:482-484 — exit item 2 reads caller for aarch64-unknown-toyos and x86_64-unknown-toyos only — the round-5 exit read aarch64-unknown-none-softfloat and aarch64-unknown-uefi too, the kernel's and the loader's targets, and this file's own table at :227-228 has minns.rs wrong for both. Round 5 asked for the x86-64 u128 row to be added, not for those two to go. Put them back in item 2 for minns.rs, naming the key they are read from, since the rows were made from the ToyOS targets' key and the kernel and loader take another.
  • same file :381-382 — "the trace has no add nuw anywhere" — the x86-64 trace has four add nuw nsw i32 lines, on the result's packing and not on the counter. What it lacks is nuw on the i16 increment, in every dump.
  • same file :18-19, the pin's file :105-106, and the body's second paragraph — "has a report of it open" as fact — the file's own "Not established" says that ToyOS's loop is #175729's is this file's reading, and that #175729 is the known fault is a maintainer's "probably". Say it in the lead as the section does.

What changes in the landing conditions

Round 5's first condition is met except for exit item 2 above: a change of that one item, confined to issues/, needs no further round, and the orchestrator reads its diff himself. The other two stand as written: the diff from f532a13a6 to the head that lands confined to issues/, and the three runs read as round 5's "What the runs must show" lists them, the nightly run standing for a head that differs from e52e6275b only under issues/, which aeb396f32 does (two files).

SEND BACK

…the upstream report is a reading, not a fact

Round 6's three NOTEs on the two issue files.

Exit item 2 reads `caller` for `aarch64-unknown-none-softfloat` and
`aarch64-unknown-uefi` too, in `minns.rs`, from the key the stamp gives those
two targets, and says that the table's rows for them were made from the ToyOS
targets' key.

The x86-64 trace has `add nuw nsw i32` on the result's packing, so "no `add
nuw` anywhere" was false of it: what no dump has is `nuw` on the `i16`
increment.

"Has a report of it open" stated as fact what the file's own "Not established"
section does not: that ToyOS's loop is llvm/llvm-project#175729's fault is the
file's reading, and that the report is the known fault a maintainer's
"probably". The defect's lead and the pin's file now say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

The nightly under the pin, run 37862964181 at e52e6275b (the orchestrator's reading; git diff e52e6275b 82ee7ab86 is under issues/ only, so it stands for this head). Every job concluded success: host, toolchain / build, tcg / suite, portability-linux, portability-macos (release skipped, as on a dispatch).

portability-macos, from the job's own log:

  • 1.98.1-aarch64-apple-darwin installed - rustc 1.98.1 (48a229cea 2026-09-01), and rustup's default toolchain implicitly overridden with 1.98.1-aarch64-apple-darwin by environment variable;
  • no line installs or selects stable 1.99.0: the seventeen lines naming 1.99.0 are the fork's own rustc-std-workspace-* crates and the tree's pinned bootstrap nightly;
  • test a_port_answers_as_its_declaration_says ... ok, no has been running for over 60 seconds;
  • [ci] Host: 78 step(s), all green, no said nothing for.

The round-6 text round is at 82ee7ab86; I read its diff (two issue files, +18 -9, the three NOTEs and nothing else). Marked ready; it lands on host and guest / suite read at this head.

@Japabu
Japabu marked this pull request as ready for review October 9, 2026 02:21
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

CI at 82ee7ab86, read from each job's own log (the orchestrator). host: [ci] Host: 79 step(s), all green; ci::tests::a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line ... ok; ci::tests::an_interrupt_of_the_driver_ends_the_group_its_step_runs_in ... ok; test a_port_answers_as_its_declaration_says ... ok; no said nothing for. guest / suite: test result: ok. 36 passed, 36 total, [ci] Guest: 5 step(s), all green, nothing left in $TMPDIR or /tmp, no FAIL line, no said nothing for. With the nightly's reading above, round 5's three conditions are met. Queued.

@Japabu
Japabu added this pull request to the merge queue Oct 9, 2026
Merged via the queue into main with commit 0b755c7 Oct 9, 2026
6 checks passed
@Japabu
Japabu deleted the wt/toyos-nightlymac branch October 9, 2026 02:54
Japabu added a commit that referenced this pull request Oct 9, 2026
One conflict, the track's stage-5 paragraph: this branch named datagram
senders that wait on the hop as still to build, #783 named the datagram
sockets' broadcast permission. Both stand, in that order, before the
move. The clause that ordered the first after #777 is met and is
restated as where the work lies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant