Repository navigation
The host gate's model controls run only the tests their verdicts name, and refuse a run that selected none - #668
Conversation
…l runs only the tests its verdicts name
`--ci host` spent most of its test time waiting on itself. Cargo's `test`
runs one executable at a time, so the host workspace's 176 test executables
ran in series after the whole build, and the build system's lib tests and the
harness checks were two more serial invocations ahead of it. Each of the 40
controls ran its whole test binary under the mutation, though its judge reads
only the verdict tests, which go red in milliseconds.
The build system, the harness's checks and the host workspace are now one
step. `cargo test --no-run --message-format=json-render-diagnostics` builds
the workspace and then the build system's `--lib` and `--test toyos-checks`
in one invocation; every test executable goes to a pool of
`available_parallelism` runners the moment cargo reports it linked, so the
workspace's tests run while the build system's two big crates compile. Each
run's two streams share one pipe and print whole when it ends, with its exit
status and wall time; any red names its source. The doctests are cargo's own
`--doc` pass, queued last because it takes the build lock.
An executable runs where cargo would run it: in its package's directory, with
that package's `CARGO_MANIFEST_DIR` and `CARGO_MANIFEST_PATH`. The
`CARGO_PKG_*` the driver inherited from `cargo run` are toyos-build's, so they
are removed rather than handed to other packages' tests. No host test reads a
`CARGO_*` variable at run time today.
A control's verdict is now a `Verdict` that names its test: `Fails`, `Passes`
or `Says { test, message }`. The run is `-- --exact <those tests>`, and a
control on a library's own tests passes `--lib` instead of building every
target. The test names of the panic-message verdicts come from the source of
each message. The comment calling `commit-ignores-notify` a double panic is
deleted: its two tests print `FAILED` and no abort.
Oracle: cargo's own serial runner and the new one, over the same tree, give
the same 176 executables, 245 result lines, 3440 passed and 7 ignored.
Measured on the dev host (M4 Pro, 14 CPUs, shared with other agents' builds,
load 26-47). Each run touched every tracked file first, as a checkout does,
and the runs alternated before, after, before, after:
--ci host total: 281 s, 362 s before; 137 s, 175 s after
the three steps: 158 s, 187 s before; 53 s, 62 s after
the 40 controls: 77 s, 115 s before; 39 s, 54 s after
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Review of CI: Net lines: 1 file, +291/−86. Production is +286/−82 (net +204); tests are +5/−4 (net +1). BLOCKER
NOTE
REMOVE (PR body)
SEND BACK |
…d its checks build in one invocation Review round 1 of #668 measured the concurrent runner on CI: the combined step took 369.5 s against 425 s for the three steps it replaced, a gain inside the spread of those steps' own medians (425 s PR, 332 s merge queue, 317 s nightly), and it pushed the workspace `--no-run` from a 132 s median to 230 s. About 150 lines whose only claim is speed, with no speed shown on the 3-CPU runner, are deleted: `host_tests`, `build_tests`, `run_tests`, `run_whole` and `TestRun`. With them go the artifact filter no test could fail on, `run_whole`'s copy of `cargo_logged`, and three `hdiutil` test binaries attaching at once from separate processes. What stays of the step is `cargo test --no-fail-fast --lib --test toyos-checks`: one invocation, so the build system's two targets compile side by side instead of in two serial builds. `--no-fail-fast` keeps what the two separate steps gave: a red in the build system's own tests still runs the checks. The host workspace is its own step again, exactly as on main. The control scoping stays: it is where the job's gain was (85 s against 161 s for the controls on CI). The flat 700 ms window in `buildlock`'s `clean_racing_a_build`, which the review noted, is filed as issues/build/a-clean-cannot-land-inside-a-build-judges-a-flat-700-ms-window.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Review of Round 1 BLOCKERs
CI.
Net lines (
Control scoping. The mapping is kept by hand. Each test name, each Every kind of drift reds
No drift is silent. The silent hole is in turning a name into a line; see the BLOCKER below. BLOCKER
NOTE
REMOVE
SEND BACK |
…a run that selected no test
Review round 2's BLOCKER: the only test of `Verdict::line` built its log from
`line()` itself, so a `Says` line emptied to "" (any log meets it) or a `Fails`
line cut to the bare test name (`test X ... ok` meets it) left
`cargo test --lib ci::` and `cargo run -- --ci host` green. The test now feeds
`judge_control` literal libtest text, copied from run 36844759528's log:
- `serial-try-lock-then-some` with one test `FAILED` and its sibling `ok` is
refused, which reds the cut `Fails` line;
- every control, given a compile error and the exit its teeth would have, is
refused, which reds the empty `Says` line;
- the both-`FAILED` log, `doorbell-kick-relaxed`'s real panic text and
`no-preempt-guard`'s `... ok` are accepted.
A drifted name made `--exact` select nothing, and a must-red control then
blamed the model ("no teeth"). A run whose log says `running 0 tests` is now
refused as `CONTROLS` having drifted, before the exit is read.
The build system and the harness's checks are two steps again, as on main.
Merged, the refusal said only `exited exit status: 101` and no longer which
target went red. Its gain is bounded by the checks' own compile, which runs
after the lib's test build when they are two steps: 7.85 s in main's run
36842009422 (34.56 s for the lib before it). In run 36844759528 the merged
compile took 22.99 s against those 42.41 s, but the same run's `cargo run`
build took 18.92 s against main's 31.94 s, so most of that gap is the runner.
The comment that a double panic aborts before `FAILED` is deleted: run
36844759528 prints `FAILED` for `push-fence-relaxed`. The issue file's story
sentence goes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…among them) into wt/toyos-winitstall Three content conflicts, each a deletion on main's side of a block this branch had edited: - kernel/src/inbox/mod.rs: main deleted `Staged`, `handler-post`'s ring, with the actuator (#660). This branch's hunks inside it — the `owed` field, `Poll::new`, the `WatchFlags` direction and "fires" for "completes" in its doc — adapted it to the ring's new fields and go with it. `Inbox::complete` keeps this branch's wording and loses main's `raise_if_staged` call. - kernel/src/watch.rs: main deleted the `handler_post` module, `holding`, `note_post` and the `raise_if_staged` call in `IrqLock::with`. This branch's one hunk inside it was "fires" for "completes" in the module's doc, which goes with it. - userland/fsd/src/main.rs: main deleted the four test actuators (`--end-on`, `--end-at-read`, `--end-at-mount`, `--let-go-at-read`) and kept the acceptor probe; this branch deleted the probe and kept the actuators. Both deletions stand: no `caps_len`, no `probe` field, no actuator field, and `accept` is this branch's. Two resolutions no marker asked for: - src/ci.rs: #668 made a control's verdicts `Fails(..)` values, so `post-is-an-answer`'s three verdict strings become three `Fails`. - issues/build: main filed the C++ runtime's scratch removal as an issue of its own (`the-cxx-runtimes-scratch-removal-dies-on-a-finder-file.md`) beside the sweep's, which this branch had merged into one file. Main's two files stand and this branch's file goes. The `rust` gitlink is main's, `95960d6c2`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
The
hostgate's model controls now run only the tests their verdicts name, and a control whose names drifted from its model is refused as that. Nothing leaves the gate, and cargo still runs every test.What changed, per decision
VerdictisFails(test),Passes(test)orSays { test, message }, and the run is-- --exact <those tests>. A control on a library's own tests passes--libinstead of building every target.judge_controlreads the same lines and the same exit status as before.--exactselect nothing, and a must-red control used to blame the model ("no teeth").judge_controlnow refuses a log with the linerunning 0 testsasCONTROLShaving drifted, before it reads the exit. A driftedtest:target makes cargo error, and one drifted name among several leaves its line missing: both red as "no verdict".a_control_is_judged_by_its_verdict_and_not_its_exit_alonefeedsjudge_controllog text copied from run 36844759528, neverVerdict::line's output:serial-try-lock-then-somewith both testsFAILEDis accepted; with oneFAILEDand oneokit is refused.error[E0425]: cannot find valueand the exit its teeth would have, is refused.doorbell-kick-relaxed's real panic text, andno-preempt-guard's... okon a green exit, are accepted.running 0 testslog is refused as drift.exited exit status: 101and no longer which target went red. The most the merge can save is the checks' own compile, which two steps run after the lib's test build: 7.85 s in main's run 36842009422.commit-ignores-notifya double panic (under that feature both its tests printFAILED), and the one saying a double panic aborts beforeFAILED(run 36844759528 printsFAILEDforpush-fence-relaxed).issues/build/a-clean-cannot-land-inside-a-build-judges-a-flat-700-ms-window.mdfiles the flat wait insrc/buildlock.rs'sclean_racing_a_build. That code is outside this branch's scope.Measured on CI
Run 36844759528 at
a51bcc7b2, the previous head, which still had the merged build/check step:hosttook 452 s, against a 944 s PR median and 870 s at round 1's head.06788146btook 670 s. The controls took 44.6 s in run 36844759528 and 115.1 s in run 36842009422 (first=== [ci] controlline to the next step's). Steps this branch does not touch also moved, e.g. the host workspace 238.5 s → 184.7 s, so the job's total mixes the change with the runner's spread; the controls' 70.5 s does not rest on it.This head's own
hostrun is the reading of it.Gates (exit codes)
cargo run -- --ci hoston991f666e0: EXIT=0 in 194 s on the dev host. All 56 steps were green, and each of the 39 controls printedN verdict(s) reached.cargo test --lib ci::: EXIT=0, 12 passed.Negative controls
Each mutation was applied with
git apply, shown to build withcargo test --lib --no-run(EXIT=0), run withcargo test --lib ci::, and reverted withgit apply -R(EXIT=0).git status --porcelainwas empty after each.cargo test --lib ci::Says { message, .. } => message.to_string()→Says { .. } => String::new()doorbell-kick-relaxed: Ok("1 verdict(s) reached")Fails(test) => format!("{test} ... FAILED")→Fails(test) => test.to_string()oklog:unwrap_err()onOk("2 verdict(s) reached")running 0 testsrefusal deletedrunning 0 testslog reads as "no teeth"Independent oracle: the logs the test judges are libtest's text as a real control run printed it (run 36844759528), not text produced by the code under test.
At
a51bcc7b2,wake-fence-off's verdict re-pointed atexactly_one_producer_owns_a_park, a test that mutation leaves green, madecargo run -- --ci hostexit 1 withRED: control `wake-fence-off`: passed with `wake-fence-off`: the model has no teeth. The control path that run took is unchanged at this head.What I am unsure of
running 0 testsis libtest's wording. If libtest rewords it, the refusal stops firing and a drifted control reds as "no teeth" or "no verdict" again: still red, but blaming the model.🤖 Generated with Claude Code
https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L