Repository navigation
A step's cargo that goes silent is ended and named, and an interrupt of the driver ends its step; the macOS nightly's hang is an LLVM ScalarEvolution miscompile rustc 1.99.0 exposes, its job takes 1.98.1, and the fork's compiler has the fault on both architectures; the power-off line's issue is closed - #778
Conversation
…wer-off issue is closed The nightly's `portability-macos` sat in `cargo run -- --ci host` until its `timeout-minutes: 350` cancelled it (run 37778826093, job 113317209315, head 8a883a1): `toyos-userbound`'s `a_port_answers_as_its_declaration_says` never ended, libtest said so once after 60 s, and the driver waited on cargo with no bound for the remaining 3 h 36 min. The test's cause is not in the log and is filed with the measurement that settles it. The unbounded wait is the build system's, and is fixed here. `src/ci.rs`: `cargo` and `cargo_logged` both run through `heard`, which reads the command's output through one pipe and ends a command that says nothing for `QUIET`, 15 minutes: the command leads a process group, the group is killed, and the step is red with the last line said. The longest silences measured in green steps: 78 s in a macOS host job (run 37740449787, a compile), and in run 37778826093 under 60 s in the Linux `host` and 120 s in `tcg / suite` (a kernel build). `cargo` used to hand its child the driver's own stdout and stderr; it now passes the lines through as `cargo_logged` always did. `nightly.yml`: the macOS `--ci host` step gets `timeout-minutes: 90`, the limit the Linux `host` job has, for a step that hangs while it talks. `issues/the-power-offs-global-lock-line-is-logged-after-the-last-console-drain.md` is deleted, its exit met: the first nightly `tcg / suite` on a `main` carrying #767 is run 37778826093 at 8a883a1 (job 113320660798), `PASS acpi_mediated_access (4s)`, `PASS acpi_lock_given_back_on_one_cpu (4s)`, `test result: ok. 34 passed, 34 total`. Nothing in the tree cited it. The order it was about is `kernel/src/power.rs`'s `shutdown`, which says it. Not built and not run by this commit's author: the request is in the pull request. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
Mutation --- a/src/ci.rs
+++ b/src/ci.rs
@@ -214,5 +214,5 @@
Err(RecvTimeoutError::Timeout) => {
// SAFETY: the group the child leads, which no other process can name until it is reaped.
- let killed = match unsafe { libc::killpg(child.id() as i32, libc::SIGKILL) } {
+ let killed = match child.kill().map_or(-1, |()| 0) {
0 => {
child.wait().map_err(|e| format!("wait: {e}"))?; |
…he report, and the issue has the measurement The first request ran at 1a89ecb on an Apple-silicon Mac. The hang does not reproduce there: `a_port_answers_as_its_declaration_says` builds and exits 0 under rustc 1.98.1 and under 1.99.0, at the tree's profile and at opt-level 0. The issue says so, says what still differs from the job that hung (the tree, the 19 tests beside it, the hosted image), and that libtest's line cannot tell a test body that never ends from a test thread that never reports. `heard` therefore ends a silent group with SIGQUIT and, 10 s later, SIGKILL for what is left: macOS writes a report with every thread's stack of a process SIGQUIT ends, and `portability-macos` uploads `~/Library/Logs/DiagnosticReports` when it fails. That no tool is spawned for it is the point: `sample` is a binary of one host OS. The test's fixtures ignore SIGQUIT, so the test runs the escalation and leaves no report on a developer's machine; it takes 20 s. `--ci host` at 1a89ecb was red in one step, `clippy, warnings denied`: `clippy::zombie_processes` on the fixture's child, which is now reaped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
Mutation --- a/src/ci.rs
+++ b/src/ci.rs
@@ -232,5 +232,5 @@
for (signal, name) in [(libc::SIGQUIT, "SIGQUIT"), (libc::SIGKILL, "SIGKILL")] {
// SAFETY: the group the child leads, which no other process can name until it is reaped.
- if unsafe { libc::killpg(child.id() as i32, signal) } != 0 {
+ if unsafe { libc::kill(child.id() as i32, signal) } != 0 {
let why = std::io::Error::last_os_error();
return Err(format!("{quiet} and its process group could not be sent {name}: {why}; the last it said: {last}")); |
|
Review round 1 of #778 at Growth: BLOCKER
NOTE
What the checks must show
SEND BACK |
… longer depends on who started the driver, and the issue is assigned with both measurements Answers review round 1 of #778. The hang does not reproduce on an Apple-silicon Mac with the job's own tree, compiler and invocation: `8a883a142`, rustc 1.99.0, the step's package selection, the `firmware` binary whole, 20 runs on libtest's default threads and 20 on three, 40 of 40 exit 0. The issue records it, is `status: assigned` to the orchestrator, who dispatches `nightly.yml` on this branch and reads the macOS job, and says what follows that one run: the cause fixed, or the test deleted with the issue recording the commit that restores it. `heard`: - hands the driver's SIGINT, SIGTERM and SIGHUP to the group it started and then takes the signal itself. The group is out of a terminal's reach, and without this an interrupted driver left its step running, a silent one for good. - starts the command with SIGQUIT at its default. A shell starts a background job ignoring SIGINT and SIGQUIT and an ignored disposition survives exec: the first measurement of the SIGQUIT report was void for that reason, and a driver started that way would have sent a signal nobody took. - counts silence and "the last it said" in bytes, so libtest on one thread, which names a test before running it and ends the line after, still names the test that hangs; and passes on what the group says after the first signal. - does not count as silent a cargo whose last line says it waits on another cargo's lock. - sends SIGKILL to the child by pid as well as to the group, and reaps it whatever the group answered. - takes `QUIET` from a constant, 10 s under `cfg(test)`. The quiet step's test no longer waits out a bound it knows will expire: its fixtures exit at SIGQUIT by a handler, which leaves no report behind. A second test interrupts a driver and reads the end of a pipe its step and the step's child write into. Not built by this commit's author; the request is with the orchestrator. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
…me, and the handler's cast is the one the compiler asks for `--ci host` at 8bb2e6d was red in one step, `clippy, warnings denied`: `function_casts_as_integer` on the three `as libc::sighandler_t` casts, which now go through `*const ()`. The window between `spawn` and the store of the group's id is closed without a signal mask, which would hold only for the thread that set it: the handler writes the signal it took and then reads the group, `started` writes the group and then reads the signal, so on whatever thread the handler runs one of the two sees the other. An interrupt that finds the group still unnamed returns and is taken by `started`, which hands it to the group and ends the driver. The issue records the corrected SIGQUIT measurement: a spinning test started as `heard` starts a step begins with SIGQUIT at its default, its output ends in the millisecond of the signal, and the 11,980-byte report holds both of its threads, libtest's main thread and the test's own. Started without the reset from a shell's background job it begins with SIGQUIT ignored, outlives the signal by the 30 s it was given and leaves no report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
The two mutations of
--- a/src/ci.rs
+++ b/src/ci.rs
@@ -349,14 +349,14 @@
};
let ended = hung.map(|last| {
// SAFETY: the group the child leads, which no other process can name until it is reaped.
- unsafe { libc::killpg(group, libc::SIGQUIT) };
+ unsafe { libc::kill(group, libc::SIGQUIT) };
let mut by = "SIGQUIT";
let mut gone = said.ended();
if !gone {
by = "SIGKILL";
// SAFETY: as above. The child by its own name too: a group whose
// members are gone but for an unreaped leader may answer for nobody.
- unsafe { libc::killpg(group, libc::SIGKILL) };
+ unsafe { libc::kill(group, libc::SIGKILL) };
let _ = child.kill();
gone = said.ended();
}
--- a/src/ci.rs
+++ b/src/ci.rs
@@ -208,7 +208,7 @@
// SAFETY: three async-signal-safe calls; the signal raised at its default ends the driver.
unsafe {
if group != 0 {
- libc::killpg(group, signal);
+ let _ = group;
}
libc::signal(signal, libc::SIG_DFL);
libc::raise(signal); |
|
Review round 2 of #778 at Growth: Round 1's BLOCKERs
Round 1's NOTEs
BLOCKER
NOTE
What the checks must show
SEND BACK |
…pt, the lock exemption is gone, and the last line said is bounded Answers review round 2 of #778. The handshake paired the handler with `started` only. The handler also reads the group as 0, between two steps, and nothing read the pending signal before the next fork: a second thread could take the signal, read 0, lose the processor for the length of a fork and then end the driver with a step no signal reaches. `heard` now reads the pending signal after it writes `STARTING` and spawns nothing if one is set; and after it writes 0 at a step's end, before it reaps, so a handler that read the group's id signals a group that is still the step's. Argued, as before, and tested by nothing. A cargo whose last line said it waited on a file lock was waited on without end. That was an unbounded wait by string match; it is deleted, and such a step is red after `QUIET` with that line as the last it said. "The last it said" is at most 200 characters of the line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
The two mutations of
--- a/src/ci.rs
+++ b/src/ci.rs
@@ -351,14 +351,14 @@
};
let ended = hung.map(|last| {
// SAFETY: the group the child leads, which no other process can name until it is reaped.
- unsafe { libc::killpg(group, libc::SIGQUIT) };
+ unsafe { libc::kill(group, libc::SIGQUIT) };
let mut by = "SIGQUIT";
let mut gone = said.ended();
if !gone {
by = "SIGKILL";
// SAFETY: as above. The child by its own name too: a group whose
// members are gone but for an unreaped leader may answer for nobody.
- unsafe { libc::killpg(group, libc::SIGKILL) };
+ unsafe { libc::kill(group, libc::SIGKILL) };
let _ = child.kill();
gone = said.ended();
}
--- a/src/ci.rs
+++ b/src/ci.rs
@@ -203,7 +203,7 @@
// SAFETY: three async-signal-safe calls; the signal raised at its default ends the driver.
unsafe {
if group != 0 {
- libc::killpg(group, signal);
+ let _ = group;
}
libc::signal(signal, libc::SIG_DFL);
libc::raise(signal); |
|
Review round 3 of #778 at Growth: Round 2's BLOCKERs
Round 2's NOTEs
BLOCKER
NOTE
Whether run 37834436612 stands for
|
… issue names the cause, and the objcopy reports are filed Run 37834436612, the dispatch of `nightly.yml` at c952668, gave the reading the issue waited for. `portability-macos` concluded `failure`: `the workspace's host members` red 15 minutes after libtest's line, `said nothing for 900s and was ended with its process group by SIGQUIT`, the later steps run, `[ci] Host: 1 of 77 step(s) red`, no orphan line, the artifact uploaded. The report of the `firmware-*` binary has the test's thread at offset 0 of the test's own function. Built here the way the runner builds it (cargo under `CI=true` compiles without incremental state, which no earlier measurement here did), rustc 1.99.0 makes that function one instruction, a branch to itself, at `opt-level = 2`; at 1 and 0, and under 1.98.1, the binary ends. The trigger among the test's cases is the dword at port 0xFFFC, `firmware::port`'s inclusive range over 0xFFFC..=0xFFFF. So the job installs 1.98.1 instead of `stable`, and the issue, renamed for what it now says, holds the measurements and the exit: a later stable that compiles the test to one that ends, and the job back on `stable`. `firmware::port` is not rewritten around a compiler's fault. The artifact also held 22 reports of `rust-objcopy` ended by dyld at launch, at the two moments `--build-only` builds bootstrap: filed. Its `cargo` report is this branch's own SIGQUIT, and its `loom_sleep` report is the control `doorbell-kick-relaxed` reaching its verdict by an abort. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
The two mutations of
--- a/src/ci.rs
+++ b/src/ci.rs
@@ -351,14 +351,14 @@
};
let ended = hung.map(|last| {
// SAFETY: the group the child leads, which no other process can name until it is reaped.
- unsafe { libc::killpg(group, libc::SIGQUIT) };
+ unsafe { libc::kill(group, libc::SIGQUIT) };
let mut by = "SIGQUIT";
let mut gone = said.ended();
if !gone {
by = "SIGKILL";
// SAFETY: as above. The child by its own name too: a group whose
// members are gone but for an unreaped leader may answer for nobody.
- unsafe { libc::killpg(group, libc::SIGKILL) };
+ unsafe { libc::kill(group, libc::SIGKILL) };
let _ = child.kill();
gone = said.ended();
}
--- a/src/ci.rs
+++ b/src/ci.rs
@@ -203,7 +203,7 @@
// SAFETY: three async-signal-safe calls; the signal raised at its default ends the driver.
unsafe {
if group != 0 {
- libc::killpg(group, signal);
+ let _ = group;
}
libc::signal(signal, libc::SIG_DFL);
libc::raise(signal); |
|
Review round 4 of #778 at Growth: Round 3's BLOCKEREvidence: OPEN, in two of its three parts.
Is it the compiler or the codeThe code is right by Rust's semantics, and the fault is the toolchain's.
So a function that is one BLOCKER
NOTE
The workflow edit
What the runs must showThe dispatch of
On the pull request, out of draft, at the head that lands: Whether the orchestrator may land on reading them. Yes, on three conditions together: both measurements of the first two BLOCKERs are in the issue with command, exit code and log and both say unaffected (the fork's toolchain exits 0 on row 2 and the kernel's instance has its exit; the x86-64 function is not a self-branch); the diff from SEND BACK |
…s instance has its exit; the objcopy reports are a defect Round 5's measurements of #778, on an Apple-silicon Mac at e0a61d0, each row the issue's second (`CARGO_INCREMENTAL=0 cargo test --locked -p toyos-userbound --test firmware`, the binary bounded at 120 s): - the fork's toolchain (`rustc 1.99.0-dev`, LLVM 22.1.8, the store's sysroot as `RUSTUP_TOOLCHAIN`): build exit 0, the binary hung and was ended by PID; the test's function is one branch to itself. - nightly-2026-07-22 (1.99.0-nightly, LLVM 22.1.8): hung. So the fault is not LLVM 23's. - nightly-2026-09-25 (1.100.0-nightly, LLVM 23.1.1): exit 0. - 1.99.0 with `--target x86_64-apple-darwin`: build exit 0, the test's function is a return, and the binary exits 0 under Rosetta. The kernel's own instance of `firmware::port`, from `CI=true cargo run -- --build-only` (exit 0) with and without incremental state, has the loop's exit in both: a function of its own in the first, inlined into `acpi_mode::port` in the second. The issue records these, says what is unknown of the fault's reach, names which jobs take which rustc, points at the host-toolchain issue, and names the holder and trigger of the exit. The `rust-objcopy` finding becomes a defect: the development Mac holds 18 such reports of its own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
…er has it for AArch64: one issue for the defect, one for the CI pin A bisection of upstream nightlies, the same test binary built without incremental state and run whole on an Apple-silicon development machine: - last good nightly-2026-07-09 (14cae6813), first bad nightly-2026-07-10 (af3d95584), both LLVM 22.1.8. The range holds rust-lang/rust #155114, which rewrote RangeInclusive's `next` onto Step::forward_overflowing. - nightly-2026-09-25 hangs: its earlier "exit 0" was a script that ran an empty binary path. nightly-2026-10-08, the newest, hangs too. No upstream compiler has a fix. - that `next` written by hand is miscompiled by stable 1.91.0, 1.95.0, 1.98.0 and 1.98.1 (LLVM 21.1.2 to 22.1.8) and compiled right by 1.88.0 (LLVM 20.1.5): the fault is LLVM's, and 1.98.1 escapes the tree's test only because its `core` lacks the shape. - -opt-bisect-limit names `indvars`: it replaces the loop's exit at the type's maximum with `false`. - the fork's own compiler turns a 30-line safe reproducer into a branch to itself for aarch64-unknown-toyos, aarch64-unknown-none-softfloat and aarch64-unknown-uefi, and compiles it right for the three x86-64 targets; 0 of 40 sources went wrong on x86-64. The issue this branch carried said nightly-2026-09-25 was good and looked for its exit in a later stable. It is now two files. The new one is the defect, the toolchain's: the reproducer as text with its command, the compilers measured, the cause as far as measured and what is not identified, the reach, and an exit a fix in the fork's LLVM meets. The existing one keeps its slug, which `nightly.yml` cites, and is the pin's: what the runner showed, the stable compilers' rows, what the pin is worth now that 1.98.1 is known to carry the same LLVM, and the pin's own exit. No test is added. A check that compiles the reproducer with the tree's compiler for the AArch64 targets would be red today, and a red test is fixed or deleted; it arrives with the fix, green. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
Review round 5 of #778 at Growth: Round 4's BLOCKERs
Round 4's NOTEs: all seven CLOSED at this head (the Linux jobs' sentence, the pointer at What the bisection's logs say of the two filesChecked and true: the committed reproducer is byte-identical to BLOCKER
NOTEMade false or incomplete by the root-cause report, each to be corrected in the same text round:
Is the pin still the right changeYes, and it lands as it is. What is known now makes the pin worth less and no less proper.
No test nowRight. The cheapest tier that reaches the defect is a host check on the tree's own compiler's output, and no type or reading sees that. It would be red on What the runs must showRun 37862964181 (
On the pull request, out of draft, at the head that lands:
Whether the orchestrator may land on reading them. Not at this head: the first BLOCKER is in the tree. After the text round, yes, on three conditions together: that round's review finds the defect file renamed, its exit and reach corrected and nothing new false; the diff from SEND BACK |
…: the defect is renamed, its cause, reach and exit rewritten Review round 5 of #778 refuted the defect file's slug and exit by the root-cause measurements: the fork's compiler leaves `caller` with no `ret` for `x86_64-unknown-toyos` when the reproducer's integer is `u128`, and the fault is not `indvars`. `issues/the-forks-compiler-drops-the-exit-of-an-inclusive-range-loop-on-aarch64.md` becomes `issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md`, and its one citation, in the pin's file, moves with it. The pin's file keeps its slug, which `nightly.yml` cites. The defect file now carries: - the fault as 14 lines of IR with no front end, what `opt -passes=indvars` makes of it and what a fixed one leaves; - the cause as four steps (CorrelatedValuePropagation marks the increment `nuw`, LoopRotate makes it a header phi's, ScalarEvolution copies the flag onto the phi's recurrence, `indvars` folds the exit), each saying what is measured and what is read from LLVM's source; - the `u128` source beside the `u16` one, the width table for both ToyOS targets, and the sysroot key the rows were made with; - 1.88.0 as an escape by shape and not a bound on the fault; - upstream: llvm/llvm-project#175729 open, pull request #118959 unmerged for 22 months, the maintainer's hedged sentence, and that nobody has built an LLVM with the proposed fix; - reach on both architectures, the class the range is one instance of, why forty sources escaped on x86-64 and the one that does not; - an exit held to ScalarEvolution and to no client of it, measured by the IR, by both Rust forms for both ToyOS targets, by the tree's test, and by a per-function comparison of the tree built with the fix given and withheld. It is `status: assigned`: the toolchain holds it, and a fix in the fork's LLVM is being built on `wt/toyos-scevfix`. The pin's file: five nightlies were run, not every one; its exit can also be met by `core` moving the loop's shape again, which fixes nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
Answer to review round 5, at BLOCKER 1: the slug and the exit
BLOCKER 2: evidenceNot mine to close in a text round: the dispatch and the two checks out of draft are the orchestrator's. The body's "Unmeasured" still says so. NOTEs
Gates at
|
| gate | result |
|---|---|
cargo test --lib sourcegate |
EXIT=0, 10 passed, no_tracked_file_identifies_a_machine_or_its_network ... ok |
cargo test --lib -- userlandhost every_program_an_architecture_leaves_out |
EXIT=0, 8 passed |
git status --porcelain --ignore-submodules=none |
empty |
No other gate was rerun: the diff is issues/ only.
|
Review round 6 of #778 at Growth: Round 5's BLOCKERs
Round 5's NOTEsAll CLOSED at this head, each against the report's files:
BLOCKER
NOTE
What changes in the landing conditionsRound 5's first condition is met except for exit item 2 above: a change of that one item, confined to SEND BACK |
…the upstream report is a reading, not a fact Round 6's three NOTEs on the two issue files. Exit item 2 reads `caller` for `aarch64-unknown-none-softfloat` and `aarch64-unknown-uefi` too, in `minns.rs`, from the key the stamp gives those two targets, and says that the table's rows for them were made from the ToyOS targets' key. The x86-64 trace has `add nuw nsw i32` on the result's packing, so "no `add nuw` anywhere" was false of it: what no dump has is `nuw` on the `i16` increment. "Has a report of it open" stated as fact what the file's own "Not established" section does not: that ToyOS's loop is llvm/llvm-project#175729's fault is the file's reading, and that the report is the known fault a maintainer's "probably". The defect's lead and the pin's file now say so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
|
The nightly under the pin, run 37862964181 at
The round-6 text round is at |
|
CI at |
One conflict, the track's stage-5 paragraph: this branch named datagram senders that wait on the hop as still to build, #783 named the datagram sockets' broadcast permission. Both stand, in that order, before the move. The clause that ordered the first after #777 is met and is restated as where the work lies. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
Head
82ee7ab86:e0a61d070(six commits and a merge oforigin/main,6776714db, #769, on809c33c0c) and four commits that change onlyissues/(git diff e0a61d070 82ee7ab86 --statnames three files, all underissues/;git diff aeb396f32 82ee7ab86 --stat, two).Open, and the toolchain's: the fork's compiler, which builds every ToyOS kernel, loader and program, turns a correct loop of safe Rust into an endless one on both architectures. Measured for a
u16range to its maximum on the three AArch64 targets ToyOS builds for, and for au128range to its maximum onaarch64-unknown-toyosandx86_64-unknown-toyos. The fault is in LLVM's ScalarEvolution and is target-independent; upstream has the fault and has merged no fix. Its open report llvm/llvm-project#175729 has the same symptom on another loop: that ToyOS's loop is that report's fault is the issue file's reading, and that the report is the known fault is a maintainer's "probably". Nothing on this branch fixes it:issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.mdis filed for it, held by the toolchain, with a fix in the fork's LLVM being built onwt/toyos-scevfix. "How far the fault reaches" below has the measurements.Every gate at
e0a61d070was run by the implementer in this branch's worktree, on an Apple-silicon development machine, macOS 27.0.1; each result is the command's own exit code. The measurements of rounds 1 to 5 were run by the orchestrator on the same machine, and the bisection (ate52e6275b) and the root-cause work by agents of his. The implementer of the last two rounds ran none of those: he read both reports and checked each number written here against their logs, traces and saved upstream responses.1. The nightly's
portability-macoshang: an LLVM miscompile that rustc 1.99.0'scoreexposesWhat the nightly showed before this branch
Run 37778826093, head
8a883a142, job 113317209315: incargo run -- --ci host's stepthe workspace's host members,toyos-userbound/tests/firmware.rs'sa_port_answers_as_its_declaration_saysnever reported, libtest said so once after 60 s, and the driver waited on cargo with no bound for 3 h 36 min until the job'stimeout-minutes: 350. The same test isokin the same run's Linuxhostjob (113317209042). It arrived with #749 and was never green on that runner.What the instrumented nightly showed
Run 37834436612, a dispatch of
nightly.ymlon this branch atc952668c9, read against review round 2's list:failure, notcancelledfailure, 19:46 to 22:23Z;host,tcg / suite,portability-linuxandtoolchain / buildsuccess22:11:02 [ci] the workspace's host members: cargo test --workspace …: said nothing for 900s and was ended with its process group by SIGQUIT; the last it said: test a_port_answers_as_its_declaration_says has been running for over 60 seconds, 15 min 0 s after libtest's line at 21:56:02; nooutside that group=== [ci] the kernel's libraryat 22:11:02 and every step after it;[ci] Host: 1 of 77 step(s) redfirmware-*Cleaning up orphan processesis followed by noTerminate orphan processlinemacos-crash-reportsuploaded, 80,792 bytes, with afirmware-*reportThe report, thread by thread. Two threads.
main, the faulting one, is parked:semaphore_wait_trapunderThread::park, theCompletedTestchannel'srecvandtest::console::run_tests_console. The thread nameda_port_answers_as_its_declaration_saysis in the test's body and nowhere else: its program counter is at offset 0 of the test's own function (the closure'sFnOnce::call_once), its link register in libtest's__rust_begin_short_backtrace, and the report lists as inlined frames at that one addressstanding(tests/firmware.rs:495),firmware::port(src/firmware.rs:344, thematchin its loop) and the test's line 513, the loop over the ports nothing declared.The same binary made here, and what it is
Cargo compiles without incremental state where
CIis set. No earlier measurement here set it, which is why 40 runs of 40 at the job's tree, compiler and package selection ended.cargo +<toolchain> test --locked -p toyos-userbound --test firmware --no-run, then the binary whole, ended after 60 s if it had not ended:CI=trueok, then libtest's line for this oneCARGO_INCREMENTAL=0CARGO_INCREMENTAL=0opt-level = 1CARGO_INCREMENTAL=0opt-level = 0CARGO_INCREMENTAL=0The hung binary's test function is at the offset the runner's report names (
0x1564, 5476), and its disassembly is one instruction, a branch to itself:The hanging build repeated with the test's cases edited, the file restored each time: without
(0xFFFC, Width::DWord)EXIT=0; with it and without(0xFFFF, Width::Byte)hung; with the byte and without the dword EXIT=0. So atopt-level = 2foraarch64-apple-darwinthe compiler endsfirmware::port's inclusive range over ports0xFFFC..=0xFFFF, inlined into the test, wrongly, and drops the rest of the function. The source ends: the crate forbidsunsafe, and every other build above runs it to its end. The cause is in "How far the fault reaches".What this branch does about it
portability-macosinstalls1.98.1where it installedstable. Under 1.98.1 the hanging build ends (the table's last row). That is a recorded compromise, and a narrower one than it looked: 1.98.1 carries the same faulty LLVM and escapes this test only because itscoredoes not yet give the loop the shape the fault needs.issues/rustc-1-99-0-makes-an-endless-loop-of-a-port-test-on-apple-silicon.mdis the pin's record,kind: tooling, assigned to the orchestrator: the runner's reading, the stable compilers' rows, what the pin is and is not worth, and its exit, the job back onstableand green in a nightly onmain. No stable that meets it exists.firmware::portis not rewritten. The loop is right; a compiler that drops a loop's exit here is not made safe by rewriting the one loop where it was seen.nightly.ymlonwt/toyos-nightlymac, at a head that differs from the one that lands only underissues/;portability-macosmust conclude success withtest a_port_answers_as_its_declaration_says ... okinthe workspace's host members.The bound, at the wait that had none
src/ci.rs'scargocalledCommand::status()andcargo_loggedread a pipe to its end. Both now run throughheard:QUIET, 15 minutes, is ended and the step is red with the last it said, counted in bytes and cut at 200 characters: libtest'stest <name> has been running for over 60 seconds, or on one thread the unfinishedtest <name> .... The driver goes on to its other steps and its summary. The runner's reading above is this, end to end.SIGQUITfirst, thenSIGKILL10 s later for what still holds the output, to the group and to the child by pid, which is always reaped. The command starts withSIGQUITat its default whatever the driver inherited: measured, a spinning test started that way ends in the millisecond of the signal and leaves a report with both its threads' stacks, and started from a shell's background job without the reset it ignores the signal for the 30 s it was given and leaves none.SIGINT,SIGTERMandSIGHUPare handed to the group and then taken at their default, unless the driver was started ignoring them. The handler writes the signal it took and then reads the group;heardreads the pending signal after every write of the group (before a spawn, after it, and before the reap), all sequentially consistent, so one of the two sees the other on whatever thread the handler runs. That rests on the argument, which review round 3 walked pair by pair, and on no test.hostjob and 120 s at most intcg / suite, both of run 37778826093. 15 minutes is 7.5 times the longest, and above the guest harness's own longest ceiling (GUEST_WEDGED, 141 s).nightly.yml,portability-macos:timeout-minutes: 90on the--ci hoststep, for a step that hangs while it talks; and~/Library/Logs/DiagnosticReportsuploaded asmacos-crash-reportswhen the job fails, byactions/upload-artifactat the commitguest.ymlpins. Both stay: the report they kept is what named the cause.What the two tests see that reading cannot
Both run a step that is this test binary, which starts a second process, prints a line and hangs by writing into a pipe its judge's side holds; both judge by a pipe's end, in a process of their own (
issues/a-pipe-made-on-macos-can-leak-into-a-sibling-spawn-and-red-the-tether-tests.mdis why).a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line: the refusal ends in the step's last line,by SIGQUIT, with no holder of the output left.an_interrupt_of_the_driver_ends_the_group_its_step_runs_in: a driver sentSIGTERMdies of it and the pipe its step and the step's child write into ends within 10 s.QUIETis 10 s undercfg(test).2. The artifact's other reports
cargo, 22:11:02Z:SIGQUIT, sent bytoyos-build, its main thread in__wait4undercargo_test::run_unit_tests. This branch's own signal to the silent step's group.loom_sleep, 22:15:40Z:SIGABRT,abort() called, undercore::panicking::panic_in_cleanup. The controldoorbell-kick-relaxedreaching its verdict: the log has the model's panic,thread caused non-unwinding panic. aborting., and[ci] control \doorbell-kick-relaxed`: 1 verdict(s) reached`. Nothing owed.rust-objcopy, 22 reports (not 25: the artifact has 25 files in all), each ended by dyld at launch inside--build-only, which went on and succeeded: 13 after the firstBuilding bootstrap(Library not loaded: @rpath/libLLVM.dylib) and 9 after the second (Symbol not found, expected in alibLLVM.dylib). The development Mac's own reports hold 18 more, of 4, 7 and 8 October, every oneLibrary not loaded. Filed as a defect, not fixed:issues/the-bootstraps-own-build-cannot-start-rust-objcopy-on-an-apple-host.md.3. The power-off line's issue is closed
issues/the-power-offs-global-lock-line-is-logged-after-the-last-console-drain.mdis deleted. Its exit's first clause,acpi_mediated_accessgreen in the nightly'stcg / suiteon amaincarrying #767: run 37778826093,nightlyonmainat8a883a142, of which15b8f9d25(#767) is an ancestor; job 113320660798, success:PASS acpi_mediated_access (4s),PASS acpi_lock_given_back_on_one_cpu (4s),test result: ok. 34 passed, 34 total (598.5s; workers: 427s building, 171s testing). The file itself records #767 meeting the other two. No citation of the slug is left; its rule is atkernel/src/power.rs'sshutdown.How far the fault reaches
A bisection on the same machine at
e52e6275b, the tree clean before and after and built only into a target directory of the bisection's own, and after it a root-cause investigation with LLVM 22.1.8'soptand the fork's compiler, which touched no ToyOS file. The issue file carries every table, both sources and the IR; what they say:Which compilers. Each step installed one nightly with rustup, built
CARGO_INCREMENTAL=0 cargo test --locked -p toyos-userbound --test firmware --no-run, ran the binary whole with a bound of 120 s, and uninstalled it (UNINSTALL EXIT=0;rustup toolchain listbyte-identical before and after).b .b .Seven steps, the five good ones before 07-09 in the issue. Five nightlies from 07-10 on were run (07-10, 07-13, 07-22, 09-25, 10-08) and each hangs; none between them was.
nightly-2026-10-08was the newest there was.What put the test's loop in the shape. Between the two adjacent nightlies are 106 commits, none under
src/llvm-project; among them is rust-lang/rust #155114, which rewroteRangeInclusive'snextontoStep::forward_overflowing. That code is correct. The samenextwritten by hand, with nocorerange, is miscompiled by stable 1.91.0, 1.95.0, 1.98.0 and 1.98.1 (LLVM 21.1.2 to 22.1.8). 1.88.0 (LLVM 20.1.5) compiles it right, which is an escape by shape and no bound on the fault: in its trace theoverflowing_addstays a call ofllvm.uadd.with.overflow.i16, so the chain below has noaddto start from. 1.98.1 passes the tree's test the same way.The cause, four steps, the first two sound:
nuw: its wrapped result is used by nothing. Measured in the pass traces; why it may, read from LLVM's source.nuwonto the phi's recurrence without a condition: the fault. Measured in its own print, which gives{-3,+,1}<nuw>the range[-3,0)beside an exit count of 3, the iteration at which it is 0; the lines that copy (ScalarEvolution.cpp:5776-5783in the fork's source), read.indvarsasks whether the exit's compare can hold and is told no. Measured:-C llvm-args=-opt-bisect-limitundernightly-2026-07-22, limit 690 printstrue, 691 printsfalse, pass 691 isindvars, and it replaces the loop's exit at0xFFFFwithfalse. The path inside is read from the source and not measured.With no front end: 14 lines of IR, in the issue as text, through
opt -passes=indvars(LLVM 22.1.8,nightly-2026-07-22's) come out with the headerbr i1 falseunder no data layout, AArch64's and x86-64's. The same with a second exit and without thenuw, or with the add still in the header, is not folded. Nooptbuilt from the fork's LLVM was run.Upstream. llvm/llvm-project#175729 is open, labelled
miscompilation, with the same symptom and the same source path on another loop. Pull request #118959, proposed as its fix, is an unmerged draft opened in December 2024. Not established: a maintainer's sentence on the issue is "I've only glanced at it, but this is probably the known issue"; nobody upstream has seen ToyOS's loop; and nobody, here or there, has built an LLVM with #118959 to see what it fixes. No rust-lang/rust report was found. None is sent.The fork's compiler (
rust/at6d6ad8c71906,rustc 1.99.0-dev, LLVM 22.1.8; it contains #155114),--emit llvm-ir,asm -C opt-level=2, each rustc exit 0, run from the sysroot key the ToyOS targets take. On theu16reproducer:caller, which returnstrueaarch64-unknown-toyosret;b .LBB0_1aarch64-unknown-none-softfloatret;b .LBB0_1aarch64-unknown-uefiret;b .LBB0_1aarch64-apple-darwinret;b LBB0_1x86_64-unknown-toyos,x86_64-unknown-none,x86_64-unknown-uefiret i1 true;movb $1, %al,retqAnd by width, the range ending at the type's maximum,
callerin the IR:aarch64-unknown-toyosx86_64-unknown-toyosu8,u32,u64ret i1 trueret i1 trueu16retret i1 trueu128retreti16gave acallerthat is not constant, on both; whether it is right was not checked.And on the tree's test, on
aarch64-apple-darwinwith the store's sysroot asRUSTUP_TOOLCHAIN: build EXIT=0, the binary hung and was ended at 120 s, the test's function one branch to itself.Reach. The kernel, the loader and userland are exposed on both architectures. Nothing wrong has been found in them, and nothing was looked for; no statistic or remark of today's compiler would count it. An inclusive range to its type's maximum is one instance of a class: an increment whose wrapped result is dead, a rotation that makes it a header phi's, a start of known range, and a client of ScalarEvolution reasoning about another value with the same expression. Of 40 sources made while reducing, 18 are miscompiled for each of the two AArch64 targets swept and 0 for
x86_64-unknown-noneorx86_64-unknown-toyos; theu128row is the counterexample that makes that no immunity. Why the forty escaped: none names an integer wider than 64 bits, and on x86-64, where 16 is a legal width,indvarsfirst rewrites the loop's exit test onto the incremented value, which gives the add a second use so step 1 never marks it. That is read from the x86-64 trace of the one reproducer and inferred for the rest. With the passes afterindvarswithheld the outcome is a wrong value and not a hang, so a silent wrong answer is possible.The x86-64 kernel's own instance of
firmware::porthas the loop's exit.CI=true cargo run -- --build-onlyate0a61d070, EXIT=0, thex86_64-unknown-nonekernel built from nothing and disassembled with the Xcode tools'objdump: with incremental state, a function of its own of 0x11a bytes; withCARGO_INCREMENTAL=0besideCI=true, EXIT=0, inlined intoacpi_mode::port, 0x13a bytes. In both the loop counts the width's bytes down in a 16-bit register and leaves at zero, every refusal jumps out, and no backward branch is unconditional. The AArch64 kernel has no instance: the one call is underarch/x86_64, by the source.Stable 1.99.0 on x86-64 compiles this test right:
--target x86_64-apple-darwin, build EXIT=0, the test's function six instructions ending inret, the binary EXIT=0 under Rosetta with21 passed.One issue became two
issues/the-forks-llvm-deletes-a-loops-exit-on-a-no-wrap-flag-scalar-evolution-gives-the-wrong-value.md, new,kind: defect,status: assigned: the toolchain holds it, and a fix in the fork's LLVM is being built onwt/toyos-scevfix. Round 5's file was named for AArch64 and for an inclusive range; theu128row refuted both, and it is renamed with its one citation moved. It is the graver of the two and outlives the pin: no job installs the fork's compiler by a version. Its exit is held to ScalarEvolution and to no client of it: a change toindvarsthat made the reproducers return would leave the fault. It is measured by four things together: the IR above through the fixed LLVM'sopt -passes=indvarswith nobr i1 falseleft;callerasret i1 trueforaarch64-unknown-toyosandx86_64-unknown-toyosin theu16and theu128source from the sysroot key the ToyOS targets take, and foraarch64-unknown-none-softfloatandaarch64-unknown-uefiin theu16source from the key the kernel and the loader take, which is not the key the table's rows for those two were made from; the tree's test binary exiting 0 on Apple silicon; and the tree's own functions read, by the fix behind an LLVM option, the tree built twice and compared function by function. The last is the measurement owed, and the only one that reads what ToyOS ships.issues/rustc-1-99-0-makes-an-endless-loop-of-a-port-test-on-apple-silicon.mdkeeps its slug, whichnightly.ymlcites and which is still true, and becomes the pin's record,kind: tooling. Its row fornightly-2026-09-25and its sentences that the nightly was good, that 1.98.1's LLVM is right and that a single file does not reproduce the fault are deleted. Its exit says a stable can meet it bycoremoving the loop's shape again, which fixes nothing.The reproducers are text in the issue, and no test arrives here
A host check that compiles the reproducer with the tree's own compiler for
aarch64-unknown-toyosand reads whethercallerreturns is the cheapest tier that reaches the defect: it needs no machine and no QEMU, and no type or reading sees what a compiler emits. It is not added now. It would be red onmainfrom the commit that added it, and a red test is a defect that is fixed or deleted; the tree has no known-red test and this is no reason to invent one. So both sources, the IR, their commands and what each must show are in the defect's file, where the exit reads them, and the check arrives with the fix, green. The tree's existing test already goes red on it, but only where the host suite is built by the fork's compiler without incremental state on Apple silicon, which no job does.Review round 3
portability-macosis green with the test in the tree, is still to be run. Round 5 found the pin does not cover the fork's compiler, which has the fault toohostandguest / suiteat the head that landsGrowth
git diff --shortstat origin/main...82ee7ab86: 6 files, 1047 insertions, 116 deletions.src/ci.rsproduction +224 −34, tests +118;nightly.yml+17 −2;issues/+688 −80.Gates at
e0a61d070cargo run -- --build-onlyBuild finished.cargo test --lib -- --exactthe two tests2 passed,finished in 10.07ssignal-the-child-alone: applied EXIT=0,cargo test --lib --no-runEXIT=0said nothing for 10s and was ended with its process group by SIGKILL; a process outside that group still held its output after itgit status --porcelainemptythe-interrupt-is-not-handed-on: applied EXIT=0, builds EXIT=0the step's group outlived its driver by 10sgit status --porcelainemptycargo run -- --ci host=== [ci] the build systemwithtest ci::tests::an_interrupt_of_the_driver_ends_the_group_its_step_runs_in ... okandtest ci::tests::a_step_that_goes_quiet_is_ended_with_what_it_spawned_and_names_its_last_line ... ok;[ci] Host: 78 step(s), all greenThe local
--ci hostruns under this machine'srustc 1.98.1with incremental compilation, so it says nothing of the hang; the table in section 1 does.At
aeb396f32, which differs frome0a61d070only underissues/(git diff f532a13a6 aeb396f32 --stat: three files, all underissues/):cargo test --lib sourcegateEXIT=0,10 passed,no_tracked_file_identifies_a_machine_or_its_network ... okamong them;cargo test --lib -- userlandhost every_program_an_architecture_leaves_outEXIT=0,8 passed, the two other gates that readissues/.git grepof the old slug's bare name at that head finds nothing. No other gate was rerun at it.At
82ee7ab86, which differs fromaeb396f32in the two issue files only (+18 −9, round 6's three NOTEs):cargo test --lib sourcegateEXIT=0,10 passed,no_tracked_file_identifies_a_machine_or_its_network ... okamong them. No other gate was rerun at it.Unmeasured
portability-macosgreen under the pin. The dispatch named above.--ci guestthroughheard, at this head. The suite's cargo and the harness are a process group under the silence bound; each guest's QEMU is not in it (the harness starts it throughtether::spawn, in a session of its own, holding the step's output through its inherited stderr) and is ended by its terminal's hangup when the harness dies. Run 37834436612'stcg / suitesucceeded atc952668c9, which is a suite that ends, throughheard, one round short of this head; a silent or interrupted suite has not been run.u128form for the kernel's and the loader's targets, andoptbuilt from the fork's own LLVM on the IR.guest.ymlandportability-linuxinstallstableand log its version (thetcg / suiteof run 37778826093:rustc 1.99.0 (b940084d7 2026-09-28)); the twohostjobs take their image's.aarch64-unknown-none,aarch64-unknown-linux-gnuandx86_64-unknown-linux-gnu: the x86-64 host build measured isx86_64-apple-darwin's.82ee7ab86other thansourcegate: the two other gates that readissues/areaeb396f32's, and the rest aree0a61d070's.Unsure of
wt/toyos-scevfixwas not on the remote when this was written. The defect names it as where the fix is being built, as the orchestrator briefed; nothing here says what that fix is or that it works.stableis. Whether the host's rustc should be declared once, as.github/qemu-versiondeclares QEMU, is the owner's to decide and is not done here; its home isissues/the-host-job-runs-the-toolchain-the-runner-ships.md.--ci hostleaves a report in that machine's~/Library/Logs/DiagnosticReports, and on a Linux host that keeps cores, a core.--ci hostloses cargo's colour and progress bar.🤖 Generated with Claude Code
https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A