Skip to content

The host gate's model controls run only the tests their verdicts name, and refuse a run that selected none - #668

Merged
Japabu merged 5 commits into
mainfrom
wt/toyos-citests
Oct 1, 2026
Merged

Japabu merged 5 commits into
mainfrom
wt/toyos-citests

Conversation

@Japabu

@Japabu Japabu commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

The host gate's model controls now run only the tests their verdicts name, and a control whose names drifted from its model is refused as that. Nothing leaves the gate, and cargo still runs every test.

What changed, per decision

  • A control's verdict names its test. Verdict is Fails(test), Passes(test) or Says { test, message }, and the run is -- --exact <those tests>. A control on a library's own tests passes --lib instead of building every target. judge_control reads the same lines and the same exit status as before.
  • A run that selected no test is refused as drift. A drifted name makes --exact select nothing, and a must-red control used to blame the model ("no teeth"). judge_control now refuses a log with the line running 0 tests as CONTROLS having drifted, before it reads the exit. A drifted test: target makes cargo error, and one drifted name among several leaves its line missing: both red as "no verdict".
  • The judge is tested on libtest's own text. a_control_is_judged_by_its_verdict_and_not_its_exit_alone feeds judge_control log text copied from run 36844759528, never Verdict::line's output:
    • serial-try-lock-then-some with both tests FAILED is accepted; with one FAILED and one ok it is refused.
    • every control, given error[E0425]: cannot find value and the exit its teeth would have, is refused.
    • doorbell-kick-relaxed's real panic text, and no-preempt-guard's ... ok on a green exit, are accepted.
    • a running 0 tests log is refused as drift.
  • The build system and the harness's checks stay two steps, as on main. Merged into one invocation, their refusal said only exited exit status: 101 and no longer which target went red. The most the merge can save is the checks' own compile, which two steps run after the lib's test build: 7.85 s in main's run 36842009422.
  • Two comments are deleted: the one calling commit-ignores-notify a double panic (under that feature both its tests print FAILED), and the one saying a double panic aborts before FAILED (run 36844759528 prints FAILED for push-fence-relaxed).
  • issues/build/a-clean-cannot-land-inside-a-build-judges-a-flat-700-ms-window.md files the flat wait in src/buildlock.rs's clean_racing_a_build. That code is outside this branch's scope.

Measured on CI

Run 36844759528 at a51bcc7b2, the previous head, which still had the merged build/check step:

  • host took 452 s, against a 944 s PR median and 870 s at round 1's head.
  • Same base: The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660's merge-queue run 36842009422 at 06788146b took 670 s. The controls took 44.6 s in run 36844759528 and 115.1 s in run 36842009422 (first === [ci] control line to the next step's). Steps this branch does not touch also moved, e.g. the host workspace 238.5 s → 184.7 s, so the job's total mixes the change with the runner's spread; the controls' 70.5 s does not rest on it.

This head's own host run is the reading of it.

Gates (exit codes)

  • cargo run -- --ci host on 991f666e0: EXIT=0 in 194 s on the dev host. All 56 steps were green, and each of the 39 controls printed N verdict(s) reached.
  • cargo test --lib ci::: EXIT=0, 12 passed.

Negative controls

Each mutation was applied with git apply, shown to build with cargo test --lib --no-run (EXIT=0), run with cargo test --lib ci::, and reverted with git apply -R (EXIT=0). git status --porcelain was empty after each.

Mutation cargo test --lib ci:: Where it reds
A (review): Says { message, .. } => message.to_string() → Says { .. } => String::new() EXIT=101 the compile-error loop: doorbell-kick-relaxed: Ok("1 verdict(s) reached")
B (review): Fails(test) => format!("{test} ... FAILED") → Fails(test) => test.to_string() EXIT=101 the one-ok log: unwrap_err() on Ok("2 verdict(s) reached")
C: the running 0 tests refusal deleted EXIT=101 the running 0 tests log reads as "no teeth"

Independent oracle: the logs the test judges are libtest's text as a real control run printed it (run 36844759528), not text produced by the code under test.

At a51bcc7b2, wake-fence-off's verdict re-pointed at exactly_one_producer_owns_a_park, a test that mutation leaves green, made cargo run -- --ci host exit 1 with RED: control `wake-fence-off`: passed with `wake-fence-off`: the model has no teeth. The control path that run took is unchanged at this head.

What I am unsure of

  • Independence of the names. A control's verdict and its filter are built from the same name, kept by hand. No independent oracle checks that the named tests are the ones whose verdicts matter; the review checked all 47 names against their source.
  • running 0 tests is libtest's wording. If libtest rewords it, the refusal stops firing and a drifted control reds as "no teeth" or "no verdict" again: still red, but blaming the model.

🤖 Generated with Claude Code

https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L

Japabu and others added 2 commits October 1, 2026 11:06
…l runs only the tests its verdicts name

`--ci host` spent most of its test time waiting on itself. Cargo's `test`
runs one executable at a time, so the host workspace's 176 test executables
ran in series after the whole build, and the build system's lib tests and the
harness checks were two more serial invocations ahead of it. Each of the 40
controls ran its whole test binary under the mutation, though its judge reads
only the verdict tests, which go red in milliseconds.

The build system, the harness's checks and the host workspace are now one
step. `cargo test --no-run --message-format=json-render-diagnostics` builds
the workspace and then the build system's `--lib` and `--test toyos-checks`
in one invocation; every test executable goes to a pool of
`available_parallelism` runners the moment cargo reports it linked, so the
workspace's tests run while the build system's two big crates compile. Each
run's two streams share one pipe and print whole when it ends, with its exit
status and wall time; any red names its source. The doctests are cargo's own
`--doc` pass, queued last because it takes the build lock.

An executable runs where cargo would run it: in its package's directory, with
that package's `CARGO_MANIFEST_DIR` and `CARGO_MANIFEST_PATH`. The
`CARGO_PKG_*` the driver inherited from `cargo run` are toyos-build's, so they
are removed rather than handed to other packages' tests. No host test reads a
`CARGO_*` variable at run time today.

A control's verdict is now a `Verdict` that names its test: `Fails`, `Passes`
or `Says { test, message }`. The run is `-- --exact <those tests>`, and a
control on a library's own tests passes `--lib` instead of building every
target. The test names of the panic-message verdicts come from the source of
each message. The comment calling `commit-ignores-notify` a double panic is
deleted: its two tests print `FAILED` and no abort.

Oracle: cargo's own serial runner and the new one, over the same tree, give
the same 176 executables, 245 result lines, 3440 passed and 7 ignored.

Measured on the dev host (M4 Pro, 14 CPUs, shared with other agents' builds,
load 26-47). Each run touched every tracked file first, as a checkout does,
and the runs alternated before, after, before, after:

  --ci host total:   281 s, 362 s before; 137 s, 175 s after
  the three steps:   158 s, 187 s before;  53 s,  62 s after
  the 40 controls:    77 s, 115 s before;  39 s,  54 s after

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu
Japabu marked this pull request as ready for review October 1, 2026 09:14
@Japabu Japabu changed the title ci: host tests run side by side as cargo links them; a control runs only its verdicts' tests ci: the host tests run side by side as cargo links them, and a control runs only the tests its verdicts name Oct 1, 2026
@Japabu

Japabu commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Review of ee2fc3fa8, round 1.

CI: host success, run 36841411417. The job took 870 s (09:14:31→09:29:01) against the 944 s PR median. Within it, cargo run -- --ci host took 692 s against an 817 s median, and the cache restore took 166 s against a 134 s median. The combined test step took 369.5 s, against 425 s for old steps 1–3 at PR medians. The controls, from kernel-loom without loom on, took 85 s against 161 s. The run covered 176 executables plus the doctests: 245 test result lines, 3445 passed, 0 failed, 7 ignored. 39 controls printed verdict(s) reached.

Net lines: 1 file, +291/−86. Production is +286/−82 (net +204); tests are +5/−4 (net +1).

BLOCKER

  • toyos-fat32/tests/common/mod.rs:32-35 with src/ci.rs:479-481 — this branch breaks an invariant the tree states. The comment says two hdiutil attaches at once are "a race in macOS", and HDIUTIL serialises attaches only within one process. This branch runs host_read, host_write and hostile as concurrent processes on macos-latest. On CI all three overlapped from 09:20:04 to 09:20:44; host_read took 123 s against a 33.9 s serial median, and host_write took 148 s against 16.9 s. The only evidence that this is safe is 6 green rounds out of 6. Closes when this head carries The host job moves to Linux: FAT32 reads committed volumes, hex input is judged by exact rounding, and winit and softbuffer get their Wayland backends #667's removal of hdiutil (rg hdiutil toyos-fat32 empty). More green rounds do not close it.
  • src/ci.rs:537 — no test can fail on which artifacts the runner runs, and a wrong filter gives a green gate that runs fewer tests. Patch || artifact["profile"]["test"] != true { → || artifact["profile"]["test"] != true || artifact["fresh"] == true {: a warm tree then runs no executable. Or add || artifact["target"]["kind"][0] == "test": every integration test is dropped. Under both, cargo test --lib ci:: and --ci host stay green. Fix: move the per-message decision into its own function. Add a_test_executable_is_a_test_profile_artifact_and_nothing_else, fed recorded cargo lines: a lib unittest, an integration test, a fresh one, an example, and a CARGO_BIN_EXE bin (the last two with profile.test false). Both patches must turn it red.
  • src/ci.rs:597-610 — run_whole duplicates cargo_logged (src/ci.rs:157-180). Each puts a command's stdout and stderr into one std::io::pipe, reads it to the end and waits. Merge them into one function that takes the Command and whether to echo live, called by run_control, guest and run_tests.
  • src/ci.rs:464-610 — about 150 production lines whose only claim is speed on the runner, and the one runner reading cannot tell their effect from runner noise. The combined step's −56 s sits inside the spread of old steps 1–3 themselves: medians of 425 s (PR), 332 s (MQ) and 317 s (nightly). Running tests beside the compile also took the workspace --no-run from a 132 s median to 230 s. The job's clear gain is the controls' −76 s, which is cut 2 and needs none of this runner. Either measure the runner on the runner it will gate on (Linux, after The host job moves to Linux: FAT32 reads committed volumes, hex input is judged by exact rounding, and winit and softbuffer get their Wayland backends #667): 3 or more runs, compared with The host job moves to Linux: FAT32 reads committed volumes, hex input is judged by exact rounding, and winit and softbuffer get their Wayland backends #667's own steps 1–3. Or land cut 2 and cargo test --lib --test toyos-checks without the runner.

NOTE

  • src/ci.rs:580-585 — a hang names nothing. Each run's output is held until it exits, so when the job hits timeout-minutes: 45, the log names no running executable. It also loses the output of every run still going, including libtest's own "has been running for over 60 seconds". Print --- <label>: started when a runner takes a run.
  • src/ci.rs:586-589 — the exit reading has no test either. if status.success() || status.code() == Some(101) { continue; } passes every gate. It was measured once, by the planted failing test. The new test can cover it with a run that exits non-zero.
  • src/ci.rs:534-547 — a ? or return inside the message loop leaves cargo unwaited. It keeps building into a closed pipe while the next build_tests waits on its lock. Kill it and wait for it on that path.
  • src/ci.rs:524, :600-601 — on macOS, std marks a pipe close-on-exec only after pipe() returns. This is the race toyos-tmpdir/tests/reclaim.rs:22-24 names. A sibling runner's child can inherit a run's write end, and that run's read_to_end then ends only when the unrelated child does. Harmless on Linux; a delay on every macOS dev host.
  • src/buildlock.rs:1051 — clean_racing_a_build(false) asserts a child finishes within a flat 700 ms window. It now runs beside up to width test binaries and a compile. Record it in issues/: the flat wait predates this branch and is outside its fence.
  • src/ci.rs:611-612 — two blank lines; delete one.
  • src/ci.rs:472 — the alternatives were weighed; no change needed. cargo test --jobs widens only the build. cargo-nextest is a new host tool that has to be fetched. It also starts only after the whole build, runs no doctests, and gives each test its own process, which would also break HDIUTIL and reclaim.rs's SPAWNING inside a binary. Conditional on the measurement BLOCKER.
  • src/ci.rs:449-462 — control scoping: no finding. Every named test exists in its stated target. A Fails/Passes line is built from the same name as the filter. A name that drifts selects nothing, and the control then fails as "no teeth" or "no verdict".

REMOVE (PR body)

  • "(cut 1)", "(cut 2)", "(cut 3: before, these were two serial single-crate builds)" — numbering from a document that is not in the tree.
  • "All of it is in src/ci.rs (+291/−86)." — the diff already says it.
  • "the 40 controls" and "All 40 controls reach their verdicts filtered" — false: CONTROLS has 39 entries, and CI printed 39 verdict lines.
  • "The runner has 14 CPUs here and 3 on macos-latest, so the gain on CI will be smaller. This PR's own host run is the measurement on the runner." — stale now that run 36841411417 exists.
  • "What I am unsure of": the "Contention on 3 cores" and "The doctest pass is the last job" bullets are stale, since run 36841411417 measured both (the doctests took 11.2 s). The "hdiutil across processes" bullet goes with its BLOCKER.

SEND BACK

Japabu and others added 2 commits October 1, 2026 11:34
…d its checks build in one invocation

Review round 1 of #668 measured the concurrent runner on CI: the combined
step took 369.5 s against 425 s for the three steps it replaced, a gain
inside the spread of those steps' own medians (425 s PR, 332 s merge
queue, 317 s nightly), and it pushed the workspace `--no-run` from a
132 s median to 230 s. About 150 lines whose only claim is speed, with
no speed shown on the 3-CPU runner, are deleted: `host_tests`,
`build_tests`, `run_tests`, `run_whole` and `TestRun`. With them go the
artifact filter no test could fail on, `run_whole`'s copy of
`cargo_logged`, and three `hdiutil` test binaries attaching at once
from separate processes.

What stays of the step is `cargo test --no-fail-fast --lib --test
toyos-checks`: one invocation, so the build system's two targets
compile side by side instead of in two serial builds. `--no-fail-fast`
keeps what the two separate steps gave: a red in the build system's own
tests still runs the checks. The host workspace is its own step again,
exactly as on main.

The control scoping stays: it is where the job's gain was (85 s against
161 s for the controls on CI).

The flat 700 ms window in `buildlock`'s `clean_racing_a_build`, which
the review noted, is filed as
issues/build/a-clean-cannot-land-inside-a-build-judges-a-flat-700-ms-window.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu Japabu changed the title ci: the host tests run side by side as cargo links them, and a control runs only the tests its verdicts name ci: a control runs only the tests its verdicts name, and the build system and its checks build in one invocation Oct 1, 2026
@Japabu

Japabu commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Review of a51bcc7b2, round 2 (non-merge change since ee2fc3fa8: a51bcc7b2 alone).

Round 1 BLOCKERs

  • 1, hdiutil binaries attaching from concurrent processes: CLOSED. rg 'host_tests|build_tests|run_tests|run_whole|TestRun' src/ is empty, and src/ci.rs:477-479 is cargo test --workspace --exclude toyos-build again, so cargo runs one test executable at a time, as on main. Run 36844759528 is green.
  • 2, the artifact filter no test could fail on: CLOSED. It was deleted with build_tests.
  • 3, run_whole duplicating cargo_logged: CLOSED. It was deleted.
  • 4, about 150 lines whose only claim was speed: CLOSED. They were deleted. The diff against main is now src/ci.rs (+139/−83) plus one issue file.

CI. host passed at a51bcc7b2 in run 36844759528.

Net lines (git diff --shortstat origin/main...HEAD): 2 files, +163/−83.

  • Production src/ci.rs: about +134/−79, net +55. This is the Verdict type and the reshaped CONTROLS, accepted for the controls' −71 s.
  • Tests: +5/−4.
  • Issue file: +24.

Control scoping. The mapping is kept by hand. Each test name, each test: target and each Says pairing of test and message is a literal copied from the model crates. Nothing derives them: every_model_control_is_run holds the features only. Within ci.rs, a Fails or Passes filter and its line come from one string. All 47 distinct names exist at this head.

Every kind of drift reds host on the PR that causes it:

  • A renamed test: --exact selects nothing and the run exits 0. A must-red control reds "no teeth"; a self-catching control reds "no verdict".
  • A moved target: cargo errors, and the control reds "no verdict".
  • A message that moves to another test: it is absent from the named test's output, and the control reds "no verdict".
  • One drifted name in a control with several names: the run stays red and that name's line is missing, so the control reds "no verdict".

No drift is silent. The silent hole is in turning a name into a line; see the BLOCKER below.

BLOCKER

  • src/ci.rs:203-209 with :869-883 — Verdict::line is tested only against its own output, so no test can fail on the claim "judge_control reads the same lines as before".
    • Why: the one test builds its log from line() itself and never feeds a Says control.
    • Patch A: Says { message, .. } => message.to_string(), → Says { .. } => String::new(),. Any log then meets every Says verdict. Seven controls fall back to the exit code alone, which verdicts' doc (:179-180) says reads a compile error as teeth: doorbell-kick-relaxed, push-fence-relaxed, commit-ignores-notify, notify-flag-load-only, gate-fence-off, fault-posted-before-it-is-set and victim-retires-mid-probe.
    • Patch B: Fails(test) => format!("{test} ... FAILED"), → Fails(test) => test.to_string(),. A test X ... ok line then meets Fails(X). In a control with several names, such as serial-try-lock-then-some, a test that stops failing goes unseen while its sibling keeps the run red.
    • Under either patch, cargo test --lib ci:: and cargo run -- --ci host stay green.
    • Fix: a test on literal libtest text, never on line():
      • For every control, judge_control(c, !c.must_red, "error[E0425]: cannot find value") is refused. This turns Patch A red.
      • For serial-try-lock-then-some, judge_control with exited_green = false on "test a_lost_try_lock_leaves_the_lock_held ... FAILED\ntest two_writers_never_overlap ... ok\n" is refused. This turns Patch B red.
    • The PR body shows both patches red.

NOTE

  • src/ci.rs:411-414 — A drifted name reds a must-red control as "the model has no teeth", which blames the model for a filter that selected nothing. Refusing a run whose log says running 0 tests, and saying so, would point at CONTROLS instead.
  • src/ci.rs:473-476 — The merged step reports two targets as one: the refusal reads ... exited exit status: 101 and no longer says whether the build system or the checks went red. --no-fail-fast was measured carrying only a planted test failure past the lib. A compile error in one target was not measured.

REMOVE

  • src/ci.rs:300-301 — "A double panic aborts before the harness prints a FAILED line" is false for push-fence-relaxed: run 36844759528 prints test a_cpu_that_halts_without_seeing_the_surplus_was_pushed ... FAILED.
  • issues/build/a-clean-cannot-land-inside-a-build-judges-a-flat-700-ms-window.md:19-20 — "Found in the review of PR The host gate's model controls run only the tests their verdicts name, and refuse a run that selected none #668, … that runner did not land." is story. "Not seen red." stays.
  • PR body — the "No custom test runner." bullet describes a head that main never receives.
  • PR body — the "Measured on CI" section measures ee2fc3fa8, the runner head, and "This head has no CI reading yet" is stale since run 36844759528.
  • PR body — "Whether the merged build/check step gains anything on the runner" is answered by run 36844759528.
  • PR body — "Each panic-message verdict's test was found from the source of its message." narrates provenance.

SEND BACK

…a run that selected no test

Review round 2's BLOCKER: the only test of `Verdict::line` built its log from
`line()` itself, so a `Says` line emptied to "" (any log meets it) or a `Fails`
line cut to the bare test name (`test X ... ok` meets it) left
`cargo test --lib ci::` and `cargo run -- --ci host` green. The test now feeds
`judge_control` literal libtest text, copied from run 36844759528's log:

- `serial-try-lock-then-some` with one test `FAILED` and its sibling `ok` is
  refused, which reds the cut `Fails` line;
- every control, given a compile error and the exit its teeth would have, is
  refused, which reds the empty `Says` line;
- the both-`FAILED` log, `doorbell-kick-relaxed`'s real panic text and
  `no-preempt-guard`'s `... ok` are accepted.

A drifted name made `--exact` select nothing, and a must-red control then
blamed the model ("no teeth"). A run whose log says `running 0 tests` is now
refused as `CONTROLS` having drifted, before the exit is read.

The build system and the harness's checks are two steps again, as on main.
Merged, the refusal said only `exited exit status: 101` and no longer which
target went red. Its gain is bounded by the checks' own compile, which runs
after the lib's test build when they are two steps: 7.85 s in main's run
36842009422 (34.56 s for the lib before it). In run 36844759528 the merged
compile took 22.99 s against those 42.41 s, but the same run's `cargo run`
build took 18.92 s against main's 31.94 s, so most of that gap is the runner.

The comment that a double panic aborts before `FAILED` is deleted: run
36844759528 prints `FAILED` for `push-fence-relaxed`. The issue file's story
sentence goes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu Japabu changed the title ci: a control runs only the tests its verdicts name, and the build system and its checks build in one invocation The host gate's model controls run only the tests their verdicts name, and refuse a run that selected none Oct 1, 2026
@Japabu
Japabu added this pull request to the merge queue Oct 1, 2026
Merged via the queue into main with commit 5620b73 Oct 1, 2026
1 check passed
@Japabu
Japabu deleted the wt/toyos-citests branch October 1, 2026 10:28
Japabu added a commit that referenced this pull request Oct 1, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Japabu added a commit that referenced this pull request Oct 1, 2026
Japabu added a commit that referenced this pull request Oct 2, 2026
…among them) into wt/toyos-winitstall

Three content conflicts, each a deletion on main's side of a block this
branch had edited:

- kernel/src/inbox/mod.rs: main deleted `Staged`, `handler-post`'s ring,
  with the actuator (#660). This branch's hunks inside it — the `owed`
  field, `Poll::new`, the `WatchFlags` direction and "fires" for
  "completes" in its doc — adapted it to the ring's new fields and go with
  it. `Inbox::complete` keeps this branch's wording and loses main's
  `raise_if_staged` call.
- kernel/src/watch.rs: main deleted the `handler_post` module, `holding`,
  `note_post` and the `raise_if_staged` call in `IrqLock::with`. This
  branch's one hunk inside it was "fires" for "completes" in the module's
  doc, which goes with it.
- userland/fsd/src/main.rs: main deleted the four test actuators
  (`--end-on`, `--end-at-read`, `--end-at-mount`, `--let-go-at-read`) and
  kept the acceptor probe; this branch deleted the probe and kept the
  actuators. Both deletions stand: no `caps_len`, no `probe` field, no
  actuator field, and `accept` is this branch's.

Two resolutions no marker asked for:

- src/ci.rs: #668 made a control's verdicts `Fails(..)` values, so
  `post-is-an-answer`'s three verdict strings become three `Fails`.
- issues/build: main filed the C++ runtime's scratch removal as an issue
  of its own (`the-cxx-runtimes-scratch-removal-dies-on-a-finder-file.md`)
  beside the sweep's, which this branch had merged into one file. Main's
  two files stand and this branch's file goes.

The `rust` gitlink is main's, `95960d6c2`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant