Repository navigation
Every QEMU ceiling is at most three times what its test takes; wedges end in seconds; two LAN boots ride the talking boot - #638
Conversation
…to take The rule is stated once, at `qemu::budget_smp`: a ceiling is at most three times the slowest the test it bounds was measured to take in the whole-suite runs on the dev host at the default width, whose times already carry the guests a test shares that host with. A wait is bounded by that multiple of its whole test, the one number measured. `host_scale` still widens it on a slower host and `oversubscription` for a guest wider than the host; nothing else does. The evidence is the orchestrator's six whole-suite logs (536r20, 536r21, 588r3, 633-x86, 633r2, 634: 495-499 tests each, 272-1000 s), the slowest PASS of each test across them. What moved: - The phase width no longer multiplies a ceiling. At 12 wide it made every number in the source twelve times itself; `port_poll_churn`'s 5 s became 263 s of guard (with an early host scale of 4.4) before it was called stalled. The measured times already carry the width, so `WIDTH` and `set_width` are deleted, and `round_trip`, which existed only to not be width-scaled, is `qemu.budget` now. - A guest still talking at twice its ceiling is timed out. The backstop was `ceiling.max(GUEST_WEDGED)`, so every chatty stall cost 300 s. - A boot's ceiling is 30 s, three times the slowest test that is one boot and one trivial command (10 s); it was 10 s times max(width, 2), 120 s at 12 wide. - The shared boot's members: 5 s (every member measured at 1.5 s or under), with mutual_kill 47, poll_wake_pipe 17, process_stats 11, abuse_elf_loader and toybox_file_tools 8; `disk_backtrace`'s 15 s override went (0.15 s at most). C members 2 s (0.54 s at most), from 10. - Every mapped wait: at most three times its test's slowest whole run. Cut where the literal was above that (doom_frames 300 -> 77, metal_sim_* 240 -> 74 and 14, cache_eviction 180 -> 59, blockd's 600 -> 152 and 251, ...); raised where a width-scaled wait sat under its own test's measured time and had only passed on the width's twelve-fold (toolkit_iced 30/60 -> 332, blocking_read_window 30 -> 179, mkdir_cap 60 -> 188, i8042_health 20 -> 164, metal_sim_pointer_churn 20 -> 92, kernel_log_file 10 -> 56). Waits that were never width-scaled keep any literal under the rule, since every measured run passed inside it. - Tests that wait out a deadline arm a shorter one: `boot-deadline=8000` for the wedges (15000), `boot-deadline=12000` for the hard lockup (30000, whose half stays above the probe's 3 s reach). - readdir_bound's `/home` arm lists one past the bound (16,385 files) rather than one past the old walk's doubling (32,769): the doubling is the bcachefs host test `a_tree_past_the_ceiling_is_refused_before_it_materialises`, with the allocator's peak as its instrument. Its ceiling is three times its slowest run, 442 s on #536's file server. Untouched: the redlisted tests (no measurement), `GUEST_WEDGED` and `GUEST_QUIET` (the `Liveness` waits have no per-test number; 300 s is under twice the slowest test that uses one), the QMP plumbing's own bounds, and pacing drains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…s the machine once a second Four cuts to what a full T14 run spends waiting, from the rig research's per-boot budget (three full runs' readbacks) and the latest metal logs. - A boot armed with one of `WEDGE_ARMS` (deadlinewedge, hardlockup, usbload) carries `boot-deadline=` `metal::STAGED_BOUND_MS`, 10 s, instead of the 120 s every other boot keeps for a wedge nobody staged. The T14 stages its wedge 1.505-1.511 s into the kernel (630r4), and those three boots came back in 110-166 s against a ~40-84 s ordinary boot. The lockup's bound is half of it, 5 s, which still outlasts the probe's 3 s reach to its lock; the load sweep simply stops sooner. `bound_for` is the one place the choice is made, in `tests/common/metal.rs`'s build, where the header already said a boot wanting a different bound is a change here and not a field on an arm. The three boots' `complete_ms` rows name the new bound, and the two lateness rows go from 10 s, "a twelfth of the bound", to twice their reading. - `wait` polls `ssh`'s port once a second, and asks `ssh true` only once the port accepts: coming back is still `ssh` answering. It slept 5 s before every probe, and a probe of a machine that is down cost `ConnectTimeout`'s 10 s. - A refusal after `reboot` waits for the machine to answer `ssh` again before it returns, unless the refusal was that wait itself, so the next boot's first `ssh` does not refuse a boot that never happened (the lanleasecase and lanswapcase arms of `issues/hardware/the-t14-stopped-answering-ssh-between- two-lan-boots.md`). The steps after `reboot` move into `after_the_reboot`. - A swapping boot whose own invocation was refused kills its `--swap` invocation instead of letting it dial for the whole 420 s. - The list allowances are twice the slowest per-member mean over five full runs: shared 845 -> 428 ms measured, allowance 1400 -> 860; shared-debug 1926 -> 398, 2900 -> 800; ccorpus 396 -> 127, 600 -> 260. The corpus now fits one boot (130 cases at 207 a boot), so `ccorpus-2` and its six rows go, with the two tests that pinned the old cut. `shared` stays two boots. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Three conflicts, every hunk of both sides accounted for: - tests/common/volumes.rs: #536 deleted `writeback_durability`; this branch had cut its one `run_test` ceiling (60 -> 26). The test is gone, so is the cut. - tests/toyos.rs: #536 deleted the `cache_eviction` arm; this branch had cut its ceiling (180 -> 59). Same. - tests/metal-profile.toml: #536 added `shared-3`'s rows and kept `ccorpus-2`'s; this branch's allowances (shared 860 ms, ccorpus 260 ms a member) fit #536's 77 shared members in two boots and the corpus in one, so neither chunk exists: the `ccorpus-2` rows stay deleted and every `shared-3` row (six, five of them outside the conflict) goes too. What #536 brought under the rule: - Two shared members it measured past 1.5 s get their own ceilings: fs_cache_eviction 50 s (16 s measured), fs_turns 11 s (3 s). - Its new ceilings, three times their tests' slowest runs on #536's two whole suites: blockd_serves_nothing 60 -> 29, fsd_restart 60 and 120 -> 35, fsd_claim_held 60 -> 59. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…by `NotFound` at any count
The arm created 32,769 files on `/home` and passed on any error from
`read_dir("/home")`. Before #536 that error was the listing bound's
(`out of memory` on 588r3, 633r2 and 634); on #536's two whole runs it was
`entity not found`, and so was `/home`'s own root afterwards, which an empty
`/home` answers too. The 32,769 creates through the DATA server were most of
readdir_bound's 364 s and 442 s there, against 46-181 s before.
So the `/home` arm and the `/home` half of the per-directory check are
deleted, and the bound is `/tmp`'s alone, which is still the kernel's
`vfs::MAX_LIST_ENTRIES` at its full count. The bcachefs library's own walk past
the old doubling stays covered on the host by
`a_tree_past_the_ceiling_is_refused_before_it_materialises`. Filed:
`issues/filesystem/listing-home-answers-not-found-and-its-bound-has-no-guest-test.md`,
for the `NotFound` and for the DATA server's `WINDOW_BYTES` refusal that no
guest test now reaches.
The ceiling is three times the slowest whole run on the runs where the arm
was the kernel's (181 s), an upper bound on the test that is left until a
suite measures it: 545 s.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…were four `lancase`, `lanicscase` and `lanleasecase` each flashed the stick and booted the machine once more (100-208 s a cycle in the rig's budget) to ask a question the talking boot, `lantalkcase`, now answers: - `lan_dhcp_lease` rides `lantalkcase`, which names the I219's function, so the loop pings the address it held under Ubuntu there as it did on `lancase`. netd's lines cross on the stick since each program has a log ring (#616), and on 630r4's full run `lantalkcase` carried the same MAC, link-up and lease lines `lancase` did. The judge drops `lan_hold`'s exit, which the talking boot's own judge replaces with `lan_talk_hold`'s. `lan.lancase.*` and `boot.lancase.ping_secs` are the talking boot's rows now. - `lan_message_delivery` rides it too, as `delivered_on_metal`. Its issue's exit condition was the shipping boot recording `pcidev: slot N took its first message` without the actuator, and all three LAN boots of 630r4 did, the talking one at 8.498 s. So `--provoke-message` has no question left: it goes from netd, from `toyos-i219` (`provoke_message` and the two tests of it) and from `build.rs`'s Intel-actuator gate, and `issues/hardware/the-lanicscase-boot-is-a-second-t14-flash-for-one-question.md` closes. - `lan_lease_report`'s metal row goes: netd's probe exists for a console line that could not cross, and the lease it reports is `lan_dhcp_lease`'s to judge off netd's own lines. The QEMU registration stays, since its link-flap check reads the probe's report; the issue that tracked the whole probe keeps that half under a slug that says what is left, `issues/diagnostics/netds-lease-probe-answers-a-question-its-lines-already-answer.md`, and `Readback::log_volume_file`, its only reader on the metal side, goes. Deleted with them: the three configs and their `ALL_CONFIGS` rows, their profile rows (the talking and swapping boots' rows that were derived "as lancase" now state that derivation), and `lan::CONFIG`, `BOOT`, `ICS_*`, `LEASE_CONFIG`, `LEASE_BOOT` and `JOBS`. `lan_hold` stays: `testcases-deaf` holds its boot open with it, and the metal-profile check of its window now reads that boot's allowance. The ssh issue's exit condition names the one arm it still owns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`blockd_serves_partitions` is 40-50 s on every whole run, and its `bench` role writes and reads its partition twice through blockd, one request at a time and then fifteen at once, with QEMU tracing every NVMe command to a file. What the bench's size buys is the throughput it prints, and a QEMU run judges no time. What the test asserts off the role holds at a quarter of it: every block reads back what was written; 2048 blocks are 64 requests of the ring's 32-block `MAX_REQUEST_BLOCKS`, past the fifteen in flight that make QEMU's trace see more than one command outstanding and more than one submission queue; and `write_all` ends every pass in a flush, so the trace carries a Flush whatever the arena's size. The reset role writes two of the partition's blocks and the survival test counts reissues in its span, neither of which reads its length. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`host` on `macos-latest` took 11, 13 and 14 minutes on the last three nightly runs that reached their end (36496779560, 36550208853, 36600425263) and was bounded at 90. The other nightly bounds are already within three times their slowest runs over the same four nightlies: `guest` 60 against 27, `tcg` 60 against 18, and `build` and the two `portability` jobs 350 against 184 and 166; the PR gate's `host` is 45 against 27. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…hat use them Two helper waits in `tests/common/power.rs` sat under the slowest test that reaches them and had passed only on the phase width's twelve-fold: - `WAIT`, the bound on QEMU's shutdown event and the drain after it, was 20 s (240 s at 12 wide). Its callers' slowest whole runs are 23 s (job_deadline_reboots), 25 s (watchdog_resets) and 110 s (quiesce_refuses_a_second_shutdown, on #536's first whole run): 332 s. - `CHAIN_WAIT`, what the boot after a reset has to arrive inside, was `PANIC_FAST_SECS + 60` (780 s at 12 wide). The tests that are one chain took at most 37 s (hard_lockup_ends_a_deaf_cpu), 33 s (boot_deadline_ends_a_wedge), 27 s and 25 s: 113 s, which the multi-chain usb_reset tests take per chain. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`tcg` took 11, 17 and 18 minutes on the last nightlies that finished it, so its bound is 54 minutes, not 60. `cfd4ba79a`'s message counted it among the bounds already inside the rule, and it was not. That message also named the wrong runs: the `host` times it gives are nightlies 36696295750 (11.6 min), 36709239346 (13.6), 36600425263 (14.0) and 36597513896 (12.9). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
#630 records metal timings per machine in tests/metal/<machine>.toml and deletes tests/metal-profile.toml and src/metalprofile.rs; a shared boot's list is cut by `SharedBoot::members`, a count fixed at the site. Every hunk of both sides is accounted for: - tests/metal-profile.toml, src/metalprofile.rs: deleted, as #630 does. What this branch carried in them moves or goes: - The per-member allowances (shared 860 ms, shared-debug 800 ms, ccorpus 260 ms, twice the slowest per-member mean over five full T14 runs) become `members` 62, 67 and 207 in tests/toyos.rs, where #630 had 38, 18 and 90. Both sets are the same derivation, (JOB_BOUND_MS - JOB_BOUND_MS / 10) / allowance, so #630's are the old 1400/2900/600 ms allowances and these replace them in the one place a list is sized. - The staged-bound rows (`complete_ms` and lateness at STAGED_BOUND_MS) and their constant in the declared-ceilings list go: #630 judges a deadline's lateness against one timer period and records no ceiling. `metal::bound_for`'s own test still pins the bound per arm. - `the_window_lan_hold_sleeps_is_the_window_this_file_prices` goes with the job allowances it read; #630 prices no list. - src/metal.rs: #630's `run` reads the machine and ends in `judge_and_write_readback`; this branch moved everything after `reboot` into `after_the_reboot`. The body is #630's, and `after_the_reboot` takes the `Machine` it writes into the boot file. - tests/common/lan.rs: #630 dropped the profile doc on `CONFIG`/`BOOT`; this branch deletes `CONFIG`/`BOOT` and the lanicscase/lanleasecase judges. - issues/diagnostics/the-lanleasecase-boot-...md and issues/hardware/the-lanicscase-boot-...md: #630 struck their metal-profile rows; this branch deletes both (the first renamed to netds-lease-probe-answers-a-question-its-lines-already-answer.md, which names no profile row; the second closed). The T14's record, tests/metal/lenovo-20w0003amz.toml: the rows of boots that no longer exist go, nine of them: `ccorpus-2` (the corpus fits one boot at 207), `lanicscase` and `lanleasecase` (folded into `lantalkcase`). It held none for `lancase`. The wedge boots' rows stay: `complete_ms` and the panel are read before the staged bound matters. No number is added; the next full T14 run records the new boots' rows. Filed issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md: with the profile gone, `toyos_tco::HARD_LOCKUP_BOUND_MS` has no reader but the assertion beside it, as on main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Orchestrator runs at b2ccd1c (logs
The same full T14 run on main-equivalent code (#630's re-record at 0d2dda6): 27 boots, wall 2899 s. The three T14 reds:
The run added three rows to the machine record, |
|
Round 1, head Readiness. CI Growth. +414 / −881 over 44 files:
BLOCKER
NOTE
REMOVE
SEND BACK |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…and the last 300 s waits go The backstop was `ceiling * 2`, and it ran ahead of the panic and silence arms, which both wait for `GUEST_QUIET` (15 s). Under any ceiling below 7.5 s (the 5 s shared default, the 2 s C members, `ipc_hostile_peer`'s 1 s) a kernel panic or a silent wedge was reported as a slow guest that was still talking. It is now `(ceiling * 2).max(ceiling + GUEST_QUIET)`. `ceiling_self_check` gains the review's two cases: - a panic under a 5 s ceiling, quiet for 10 s at 11 s: no verdict yet. Red at b2ccd1c's backstop (EXIT=101), green with this one. - the same panic at 21 s, quiet for `GUEST_QUIET`: named. Green at b2ccd1c too, because the pure function puts the panic arm first. It reds (EXIT=101) when the backstop is moved ahead of that arm. `a_stall_stays_red` stages a 30 s ceiling instead of 5 s, so its `past` (twice the ceiling, plus one) still clears the backstop. Four waits still bounded by `GUEST_WEDGED` are cut to about three times their slowest whole-suite run: `lan_no_lease` 99 s, `kernel_log_file`'s device poll 54 s, `usb_flush_optional` 99 s, and `guest_dies_with_its_harness` 24 s, which is now host-scaled like the other three. `readdir_bound` is 147 s, three times the 49 s it took at b2ccd1c without its `/home` arm. `toolkit_winit_loop`'s `Liveness` had a 40 s quiet guard behind a 35 s total, so the guard could never fire. The loop now runs to a host-scaled 35 s deadline. The i219 model's `ICS` write arm, read arm and register constant go. Nothing has written `ICS` since `--provoke-message` went. `build.rs`'s Intel-actuator gate holds one flag, so it is one `const` and one membership test. Its declaration test is renamed to say "flag". `STAGED_BOUND_MS` moves into `toyos-tco` beside `WEDGE_BOUND_MS`. That crate holds every bound a boot arms. The metal loop's wait is split into `wait_on` and a `port_accepts(host, port)`. A new host test drives them against a loopback listener that accepts and never speaks `ssh`. With `let answered = listening;` the build passes (EXIT=0) and `metal::tests` fails (EXIT=101) on that test alone. The harness kills a swap only when the flashing invocation did not exit 0 or 1. Exit 1 is a boot that ran and was judged red, and its swap is left to finish. The T14's record gains `boot.shared.{complete_ms,panel_max_us,panel_us}` = 1154 / 3608 / 21429. These are the first passing reading of `shared` at its 62 members. Deleted as the review asked: - the backstop's "six times" clause - the measured "slowest of which took" clauses - the staging time of the T14's wedge - `lan_hold`'s docs, which describe a user it no longer has - the restated twenty seconds - two issue sentences that name boots which no longer exist - the two "twice its ceiling" phrasings, which the new backstop makes false for short ceilings Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`coming_back_is_ssh_answering_and_not_its_port` went red in `cargo run -- --ci host` (EXIT=1). Its last case dialled the loopback listener's port after dropping the listener and found it still accepting for the whole second. Once let go, that port can be another socket's before it is dialled: another test's listener, or a self-connect from an ephemeral source port. Which of the two it was is not established. Port 0 is one nothing listens on. Three runs of `cargo test -p toyos-build --lib` after this change: EXIT=0 each. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Orchestrator run at edefd81: |
|
Round 2, head Round 1 BLOCKER
Readiness
Growth
BLOCKER
NOTE
REMOVE
SEND BACK |
…at the backstop The backstop arm of `ceiling_verdict` now names a death it holds before it calls the guest slow. A 5 s ceiling and a kernel death at 7 s read "timed out after 20s, with the guest still talking 14s ago" at edefd81, because the panic arm waits for `GUEST_QUIET` and the backstop came first; `main`'s 300 s floor named that panic at 22 s. The comment's clause that a death "is named by the arms above" was false for that case and goes. `ceiling_self_check`'s case (e) is two walks under a 5 s ceiling, a second at a time: a guest silent from any second in its budget must first read as `STALLED`, and a kernel dead at any second up to four ceilings must first be named. The old case (e) required only a backstop of 11 s or more, so a backstop of `ceiling + 6 s` passed it. The walk steps the quiet seconds and adds them to the start, where the review's form subtracted a `Duration`, which the host gate's clippy denies (`unchecked_time_subtraction`); the elapsed/quiet pairs are the same. Each arm was a checked patch, shown to build and reversed: - death walk at edefd81's verdict code: EXIT=101, on `died = 7`. - backstop `(ceiling * 2).max(ceiling + Duration::from_secs(6))`: EXIT=101, the silence walk on `since = 0` ("timed out after 11s"). - backstop `ceiling * 2`: EXIT=101, the silence walk on `since = 0` ("timed out after 10s"). - this commit: EXIT=0. `lan_hold` is `deaf_cpu_hold`: its one runner is `dump_nmi_probe`'s `testcases-deaf` boot, which runs no netd, so its hold is no longer netd's lease bound but its own `PAST_THE_DEAF_WINDOW`, at the same 20 s. `issues/diagnostics/the-cable-judge-reads-three-netd-records-that-cannot-arrive-on-the-t14.md` goes: netd's lease record reached the T14's stick at b2ccd1c, and the boot and keys it names are deleted by this branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Round 3, head Round 2 BLOCKERs
Readiness
Nightly 36753172688 at
Growth
BLOCKERNone. NOTE
REMOVE
LAND AFTER NAMED CHANGES |
|
Orchestrator runs at 1c81e3e:
|
Clean merge. Main's one new timed wait, `handler_post_without_a_pass`'s `drain_until(30 s)`, and its new shared member `cpal_drop_unreleased` were written against the width-scaled budget; the next commit brings them under this branch's ceiling rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…ll, and main's new wait is under the rule
- `ceiling_self_check`'s two walks step `since`, `died` and `quiet` by
100 ms, the read loop's poll of a silent guest, instead of by whole
seconds. At whole seconds a backstop one second short,
`(ceiling * 2).max(ceiling + GUEST_QUIET - 1 s)`, passed both walks: at
`since = 5` the walk reached 19 s at quiet 14, and 19 is not past 19. It
now fails at `since = 4.2 s` ("timed out after 19s"), EXIT=101.
- The walks' comment goes: "any point" claimed more than the walk checks,
and the rest restated the loops.
- `handler_post_without_a_pass`'s drain, from main's #634, was 30 s under
the width-scaled budget. The test took 3-4 s in #634's whole runs
(4 s in 634r6), so it is 12 s.
- `cpal_drop_unreleased`, #641's new shared member, took 111 and 138 ms in
#641's whole runs, so the 5 s shared default covers it and it needs no
row.
- `deaf_cpu_hold` keeps its fixed 20 s. A binary test-runner spawns holds
no `logread`: test-runner hands down its namespace, and `logread` is a
`SysCap` duplicate, not a namespace entry. logd serves no reader port on
`tests/testcases`. The only record the hold could read is logd's file
under `/log`, which gives no notification, so waiting on it would be a
sleep-and-reread poll.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
The orchestrator's whole run at 09d74fc (638r4) went red on toolkit_winit_loop "never finished" after 57 s. This branch had replaced main's Liveness(40 s quiet, 240 s total) with a hard `Instant::now() + qemu.budget(35 s)`, and a hard cut does not ask whether the guest is still talking. `qemu::Ceiling` reads a capture its caller keeps and appends to. It takes the guest's silence from when the capture last grew, and a kernel death from the first one appended. The decision is ceiling_verdict's, so every one of these waits ends the way run_test_paced's does: - a guest still talking past its budget waits for the backstop and reds as TIMED_OUT - one silent for GUEST_QUIET past its budget reds as STALLED - a dead kernel is named, with its report `backstop()` is ceiling_verdict's formula, pulled out so a wait with no console can sit behind another wait's ceiling. The waits this branch cut from main's far bound to a tight hard one: - toolkit_winit_loop, 35 s: Liveness(40, 240) on main. - toolkit_winit_pace, 53 s: a Liveness whose total was cut from 120 s. That total is a hard cut however loud the guest is, and it was not host-scaled. A wait it ended also went on to count presents, so a slow guest read as a wrong count. It is now host-scaled, and the ceiling's verdict is the red. - lan_no_lease, 99 s: drain_until(GUEST_WEDGED) on main. The guest is silent until netd gives up, and the ceiling allows silence inside the budget. - kernel_log_file's device poll, 54 s, and usb_flush_optional's, 99 s: budget(GUEST_WEDGED) on main. Each now drains the console between polls instead of sleeping 50 ms, so the ceiling can hear the guest. The lines it reads still reach the shutdown tail's panic check. - guest_dies_with_its_harness, 24 s: GUEST_WEDGED on main. The wait reads the owner's stdout, which carries nothing before HELD. The owner's own boot ceiling (BOOT_CEILING, host scale 1 in a fresh process) names a boot that never came up, and 24 s could cut ahead of that 30 s. The wait now sits at that ceiling's backstop. Not converted: every wait that was already a hard deadline on main and only had its number moved. Among them are the boot ceiling, toolkit_iced's three waits, metal_sim_pointer_churn's three, handler_post_without_a_pass's drain, power's WAIT and CHAIN_WAIT, logstream's FLOOD_CEILING, and the screendump_until/drain_until literals. What the 638r4 capture shows is not a slow guest. The terminal logged stage 1 at 6.534 s and nothing else from the app. Meanwhile the app opened 23 windows, the last being stage 6's at 9.448 s, and every one of those windows is created after a stage line it had printed. CLOSE-ME never arrived. From 11.4 s to 42.8 s every sched line reads ready=0 on every CPU. Under this change that guest still reds, at the backstop as TIMED_OUT instead of at 35 s. That is filed as issues/build/winit-loops-output-stopped-reaching-the-log-after-stage-1.md. metal_sim_client_death's red in the same run was not a cut. Its run_test returned on ===TEST_END with exit 0 after 153 ms. The missing line is its grandchild's, printed after the root had ended. That is filed as issues/build/client-death-ends-before-its-reaped-creators-request-is-served.md. capture_ceiling_self_check stages Ceiling on instants ahead of now: - a capture growing past its ceiling is ended only by the backstop - one that stopped growing is ended by the stall - a kernel death appended after the first read is named, with its report Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
… the detector on every boot, and the watchdog class refused by right Designed against #638 (23 boots a full T14 run, lists sized at `JOB_BOUND_MS` less a tenth over 860/800/260 ms allowances) as if landed. The session track: - The fold is six boots of #638's run: shared, shared-2, ccorpus, testcases, testcases-mkdir, testcases-readdir. Each `Boot parameter:` line in toyos-tight/target/metal (b2ccd1c) carries only root=, boot-deadline=120000, boot-slot=A and blackbox=. 220 members there: 62+18+130+8+1+1 `===TEST_END` lines. - Prices: #638's 860 ms a shipping member and 260 ms a corpus case; the other three lists at twice their slowest span over ten full T14 runs (536-metal-full, 536r21, 590r6-m1, 590r6-m2, 616, 616r3, 630-full1, 630-full2, 630r4, 638-metal-full): testcases 11.861 s (590r6-m1), mkdir 2.106 s (536r21), readdir_bound 5.043 s (638, the only run without its /home arm). - Per-job bound 14.1 s: twice mutual_kill's 7.041 s (630-full2), the slowest member of those runs once readdir_bound's /home arm (24.8 s on the two 536 runs) is set aside. - Budget: 120000 - 12000 - 19100 - 14100 = 74800 ms. 80 x 860 = 68800; 130 x 260 + 23722 + 4212 + 10086 = 71820. Two sessions for six boots: 23 -> 19. - The fence is every root `/` has. Of the 89 folded binaries, 14 name /home and hierarchy_paths writes /apps; the file-left-behind control is staged in /home. - A window's copy of the kernel's records is the exec-started runner's only: on QEMU's serial runner a copy would double every record on the console, where 4 must_be_clean_apart_from sites and 55 `.matches(..).count()` sites in 10 files count lines. - device_claim_lifetime, the one folded member that mints a claim, runs last until the deferred-release defect closes. - The QEMU e1000e session arm is restored, and the last stage waits on the detector on every boot. It answers the swap-bound issue. The watchdog track: - The owner ruled the hard-lockup detector armed on every boot, so its question file is deleted and the arm is the track's next step. All 23 boots of #638's run say `hard lockup:` (60000 ms, 5000 ms on the three staged boots); only hardlockup's loader.log carries the lockup record. - Refusal by right: a new Rights::WATCHDOG that no syscap name grants, demanded by sys_device_claim for the class. Declared as an ABI change. - The kernel feeds while no claim is held, so watchdogd's exit or swap hands the timer back; watchdogd waits on the deferred-release defect. - QEMU: q35's TCO counts QEMU_CLOCK_VIRTUAL and its second expiry calls watchdog_perform_action (QEMU v11.1.1 hw/acpi/ich9_tco.c:61-70,244); tests/qtest/tco-test.c:325-347 tests the `none` action. Every guest that stages no reset runs with -action watchdog=none. - The T14 reading moves to the loader's report pass right after the reset: Linux v6.12 drivers/watchdog/iTCO_wdt.c:545-560 clears SECOND_TO_STS in iTCO_wdt_probe. The starved boot's deadline must outlast tco-starve's 5 s plus the 9.6 s bound, which #638's 10 s STAGED_BOUND_MS does not. - The go-ahead is an owner question file of its own. The loader track: stage 6's exit asks a whole run, every --once boot included, and a loader sent over ssh to boot once; a stick booting neither loader costs a hand and diag/flash.sh; stage 7 closes the four Ubuntu-loop issues; the 53 MiB ROOT clause goes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…slowest run Five whole runs of `cargo test --test toyos-build` at bf28c1e, 12 wide on the dev host at load average 40 to 59 on 14 cores (guest-m2..m6 in the orchestrator's scratchpad): 26, 25, 26, 26 and 26 passed, in 229.2, 66.8, 37.1, 32.9 and 34.1 s. The slowest pass of each test, in seconds: iommu_virtio_platform 47 virt_first_entry 14 machine_shutdown 20 virt_fp_isolation 14 nested_nmi_is_loud 21 virt_irq_storm 12 screen_fatal_behind_a_painter 29 virt_mask_windows 8 screen_fatal_halt_composited 32 virt_off_names_the_cpus_left_on 16 screen_panic_muted 19 virt_readonly_copyout 18 virt_debug_refused 18 virt_reboot 8 virt_early_fault 2 virt_reboot_refused_without_psci 8 virt_early_panic 2 virt_smp 10 virt_el1_smp 10 virt_timer_floor 9 virt_el2_drop 7 virt_timer_preempts 29 virt_failed_ap_leaves_no_hole 15 virt_unmap_touch 18 virt_fatal_halts_the_others_first 34 virt_user_mode 10 The numbers, each three times the slowest test that waits through it: - BOOT_CEILING 63: nested_nmi_is_loud, the slowest test that is one boot and nothing after it (21). main's was 10 s times the width, 120 s. - GUEST_WEDGED 141, now paid out through `budget`: every wait through `await_guest`, the slowest of which is iommu_virtio_platform (47). `guest_liveness` had one caller and goes. - virt_selftest's drain 36, was 180: virt_irq_storm (12) and virt_timer_floor (9). - virt_user_mode's drain 30, was 180 (10). - virt_early_panic's and virt_early_fault's two waits 6 each, were 30 and 10 (2 each). - screen_fatal_behind_a_painter's second wait 87, was 20, and screen_fatal_halt_composited's first 96, was 30: each sat under its own test's slowest run (29 and 32) and passed on main only because the width multiplied it twelvefold. - Kept, at or above their test's slowest run and under three times it: screen_panic_muted 30 (19), virt_el2_drop 10 (7), the painter's first wait 30 (29), the composited test's second 40 (32). REFERENCE_BOOT_MS is 1424, the largest fastest boot of the five runs (1419, 1424, 1048, 1055, 750 ms), so the measuring host pays every ceiling at 1x. The red of the second run is virt_mask_windows, on an assertion and not a ceiling: nine census lines and eight windows lines for cpu7. The test, the kernel and its judge are main's. Filed as issues/build/virt-mask-windows-read-nine-censuses-and-eight-windows-lines-for-one-cpu.md. The first run after the merge exited 1 with 18 reds before any of these: twelve threads moved the fork checkout to the new pin at once. Filed as issues/build/a-suites-first-run-after-the-fork-pin-moves-races-its-own-checkout.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Clean. #651 deletes the retired ABI numbers and adds one shared member, `syscall_unassigned`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
|
Mutation patches for the negative controls at
--- a/kernel/src/arch/aarch64/trap.rs
+++ b/kernel/src/arch/aarch64/trap.rs
@@ -525,7 +525,7 @@
pub(super) fn tick() {
if RUNNING.load(Relaxed) {
- TICKS.fetch_add(1, Relaxed);
+ TICKS.fetch_add(0, Relaxed);
}
}
--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -977,7 +977,7 @@
/// first line, so the backlog crosses at the network's pace and not the
/// conversation's.
fn hand_back<T>(stream: &Stream, began: Instant, by: Duration, fire: impl FnOnce() -> T) -> T {
- match stream.wait_for_boot(by) {
+ match stream.wait_for_boot(Duration::ZERO.min(by)) {
Some(ms) => println!(
" talk: the stream carried `Boot: complete` ({ms} ms), {} ms after it opened",
began.elapsed().as_millis()
--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -286,7 +286,7 @@
/// This boot's `Boot: complete` in milliseconds, read as [`judge`] reads
/// it, once a line carries it, or `None` after `by`.
pub fn wait_for_boot(&self, by: Duration) -> Option<u64> {
- self.wait_until(by, |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
+ self.wait_until(by.min(Duration::ZERO), |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
}
/// The peer, once a connection has carried a line, or `None` after `by`
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -110,7 +110,7 @@
pub fn said_refusal(stderr: &str) -> Option<String> {
let lines: Vec<&str> = stderr.lines().collect();
let at = lines.iter().rposition(|line| line.starts_with(REFUSAL_HEAD))?;
- Some(lines[at..].join("\n")[REFUSAL_HEAD.len()..].trim_end().to_string())
+ Some(lines[at..=at].join("\n")[REFUSAL_HEAD.len()..].trim_end().to_string())
}
/// Every way this loop refuses, by name.
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -1509,7 +1509,7 @@
while began.elapsed().as_secs() < secs {
let next = std::time::Instant::now() + POLL;
let listening = accepts();
- let answered = listening && (!answering || answers());
+ let answered = listening;
if answered == answering {
return Ok(began.elapsed().as_secs());
}
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -852,7 +852,7 @@
/// The `boot-deadline=` bound an image armed with `armed` carries.
pub fn bound_for(armed: &[impl AsRef<str>]) -> u64 {
- if stages_a_wedge(armed) {
+ if stages_a_wedge(armed) && false {
toyos_tco::STAGED_BOUND_MS
} else {
toyos_tco::WEDGE_BOUND_MS
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
// For a guest that is stuck *and* chatty and so never trips the silence
// guard; never before a guest silent since `ceiling` has been silent for
// [`GUEST_QUIET`].
- let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+ let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET - Duration::from_secs(1));
if elapsed > backstop {
if let Some(line) = dying {
return Some(kernel_died_here(line));
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
// For a guest that is stuck *and* chatty and so never trips the silence
// guard; never before a guest silent since `ceiling` has been silent for
// [`GUEST_QUIET`].
- let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+ let backstop = (ceiling * 2).max(ceiling + Duration::from_secs(6));
if elapsed > backstop {
if let Some(line) = dying {
return Some(kernel_died_here(line));
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
// For a guest that is stuck *and* chatty and so never trips the silence
// guard; never before a guest silent since `ceiling` has been silent for
// [`GUEST_QUIET`].
- let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+ let backstop = ceiling * 2;
if elapsed > backstop {
if let Some(line) = dying {
return Some(kernel_died_here(line));
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
// For a guest that is stuck *and* chatty and so never trips the silence
// guard; never before a guest silent since `ceiling` has been silent for
// [`GUEST_QUIET`].
- let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+ let backstop = ceiling.max(Duration::from_secs(300));
if elapsed > backstop {
if let Some(line) = dying {
return Some(kernel_died_here(line));
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -519,7 +519,7 @@
// [`GUEST_QUIET`].
let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
if elapsed > backstop {
- if let Some(line) = dying {
+ if let Some(line) = dying.filter(|_| false) {
return Some(kernel_died_here(line));
}
return Some(format!( |
|
T14 at
Not run, and owed: the five rows the request left out for the owner's OK ( Readbacks and judge logs: |
|
Round 4, head Earlier BLOCKERs: none open. Round 3 closed round 2's two at Evidence at
Growth: +648 / −677 over 36 files, net −29.
Rulings on the brief
The T14 readings this head owes None of these ran at
BLOCKER
NOTE
REMOVE
SEND BACK |
…the shared boots' allowances are named Round 4's review of 4859e8d (#638): - metal: `wait_on` dials the port within CONNECT_SECS when it watches the machine go down, what `ssh`'s own ConnectTimeout gives, and within one POLL coming back. A single lost SYN while Ubuntu is still up no longer reads as down. `coming_back_is_ssh_answering_and_not_its_port` records the dial it is handed going down. - metal: the wait for `ssh` after a refusal that followed `reboot`, and `after_the_reboot`'s split, are deleted. No recorded run has the case it was for: the T14 run the body cited timed out on the boot after a pass, and the post-reboot refusals in the kept T14 logs are a machine that never came back, a stick, a ping and a mount, each with the machine already answering or already waited out. `run` is main's again past `bootnext`. - toyos-tco: `RUST_MEMBER_MS` (860) and `C_MEMBER_MS` (260) beside `JOB_BOUND_MS`; `metal::members_fitting` derives `shared`'s and `ccorpus`'s member counts from them, 62 and 207 as before. `shared-debug` keeps main's 18: on a five-member list the count moves nothing. - qemu: `VERDICT_POLL` is the read loop's poll, and the ceiling walk in `tests/checks/qemu.rs` steps by it. - metaltalk: two comments deleted. - issues: owners for the fork-pin race (the tooling track's toolchain stage), the hard-lockup bound's reader (the T14 unattended track, which holds the detector) and the I219 transmit burst (the LAN track's stage 2); the `virt_mask_windows` red moves to `issues/kernel/`; filed: `run_test` and `ceiling_verdict` serve only `--debug`'s `run`, and `host_scale` reads host speed off boots that stop at different markers (491 ms for `virt_early_panic` alone, 3318 ms for `machine_shutdown` alone, same host, same minute). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
|
Negative controls at
|
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
|
Negative controls at
The unmutated head: |
|
The wedge's third arm at |
…red in one session The filing at 4b009e7 set an x86-64 boot under TCG beside an AArch64 boot under HVF and named only the marker. At e5daa61, one run after the other: virt_early_panic (Virt, HVF, to `EARLY PANIC:`) 479 ms, virt_smp (VirtEl2, TCG, to `SCTLR_EL1=`) 1180 ms, machine_shutdown (x86-64, TCG, to `===READY===`) 3316 ms and 2.33x. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
#680 stamps every statement on stderr with the time of day, the driver's refusal among them, so `said_refusal` reads a line through `printer::unstamped`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
|
Negative controls at
The unmutated head:
|
|
T14 at
The unfiltered run in detail:
Logs and readbacks: |
|
Round 5, head Round 4's BLOCKERs
Rulings on the brief
Evidence at
Growth: +740 −671 over 38 files, net +69.
BLOCKER
NOTE
REMOVE
SEND BACK |
… issue closes Round 5 of #638's review. `coming_back_is_ssh_answering_and_not_its_port` wrapped `accepts` to record the bound `wait_on` hands it going down and asserted that bound was `CONNECT_SECS`. That checks one line of `wait_on` a reader of the diff checks, so the wrapper, its assertion and the doc clause that named it go, and `accepts` is passed directly. `issues/build/the-talking-boots-reboot-outruns-its-log-stream.md` is deleted: its exit, two consecutive T14 runs passing `lan_talk`, is met by the runs at `4859e8d71` (PR comment 5962753432) and at `75dc056f1` (PR comment 5966698095), both with `metaltalk::hand_back` waiting on the stream's `Boot: complete`. The rule it carried is already `hand_back`'s doc. Its one citation, in `issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md`, goes with it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
|
Round 6, head Round 5's BLOCKER
Round 5's NOTE and REMOVE
Whether the readings at
Evidence at
Growth.
BLOCKERNone. NOTENone. REMOVENone. LAND |
…hers) into wt/toyos-resident One conflict: issues/hardware/the-t14-boots-toyos-unattended.md, which this branch deletes and main modified. Every hunk of main's side since the merge base 4f2bea1 is accounted for: - e9f67e7 (#639) appended one line: "`smp_failed_ap_leaves_no_hole` is deleted; issues/build/smp-ap-hole-and-log-reserve-window-red-under-a-loaded-host.md records the commit that restores it." It qualified the sentinel's finding 2, which named that test as the roster's gate. No file replacing the track plans a sentinel or names the test; the gap a sentinel was for (the span before clock::init) is stated in kernel/src/deadline.rs's header, and the restore the line pointed at stays where it was recorded, in smp-ap-hole-and-log-reserve-window-red-under-a-loaded-host.md ("`git revert aedcf17` brings it back"). So the file stays deleted and the line has nothing left to qualify. The deleted track's one other citation, from #638: issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md named it as owner. Its owner is now issues/hardware/a-frozen-toyos-waits-for-a-hand-on-the-power-button.md, the track #638's review named. Its two exits disagreed: that track's first step gives HARD_LOCKUP_BOUND_MS a reader (the detector's bound on a boot that names no boot-deadline=), and the issue's exit deleted the constant. The reader stands: the issue now closes on that step, with the constant's doc saying what it bounds, and says deleting it is not the exit, since the step would declare it again. issues/boot-media/the-loader-does-only-what-must-precede-the-handover.md merged cleanly; main's two hunks (create_boot_image's per-image GUIDs, and root_read_ticks) are in it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
…today Each citation the branch's files make was checked against 4adc1c0. Every name a track or issue cites exists there (toyos_tco::STAGED_BOUND_MS, RUST_MEMBER_MS, C_MEMBER_MS, SharedBoot's `members`, metal::judge_arms, panic_console::hold_the_panel, arch::watchdog's `arm` and its TCO2_STS read, bootlog::split_listing, toyos::log::LogTail, vfs::ROOT_ENTRIES, Armed's (0, _) arm, every boot and row name), but for these, which this commit fixes: - `device_claim_lifetime` is gone (51cc87f, #639). `endowment_denied` is now the one session member that mints a claim (`git grep -l device_claim` over the Rust tests, test-runner and the C corpus finds it alone), so the track says it runs last in its session rather than that the two end one each. - `log_poll_outlives_a_close` is gone (ad6dc07, #639). Of the judges on the six folded boots, `syscall_cost` alone reads the job's own output; the audio judges read soundd's lines between the kernel's `spawn:` records, which reach the stick. - The guest-suite track does not itself leave the shared boot to the T14: #660 did, and `shared_metal` and `c_corpus_metal` are its only runs outside `--debug`. The track now says the first and cites the issue for the rows it adds. - The prices name #638's allowances, toyos_tco::RUST_MEMBER_MS and C_MEMBER_MS, rather than "as SharedBoot::members prices them": `members` is a count derived from them. - Line citations into test-runner and toyos/src/syscap.rs moved: `--bound-ms=` is main.rs:78-83, run_one's duplicate 248-250, the namespace's inheritance 62-68, SysCap::duplicate 63-70 and SysCap::narrowed 133-139. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Head
e3094a94a, on75dc056f1, which mergesorigin/mainat73282fa93(#648, #680).git diff --shortstat origin/main...HEAD: 38 files, +717 −704. Shipping code (netd,toyos-i219'slib.rsandregs.rs,toyos-tco) +23 −63;toyos-i219's model and tests +2 −57;src/production +112 −14 and its tests +109 −58;tests/+169 −351;issues/+302 −161.Between
75dc056f1and this headgit diff --stat 75dc056f1 e3094a94a:Test code and issues only. Both hunks in
src/metal.rslie inside its#[cfg(test)] mod tests, incoming_back_is_ssh_answering_and_not_its_port. Nothing the loop does at run time changed, nor anything the image or the guest suite builds, so the T14 readings, the image and the guest suite at75dc056f1stand for this head.Every QEMU ceiling is at most three times what its test takes
The rule is stated once, at
tests/common/qemu.rs'sbudget_smp: a ceiling is at most three times the slowest its test was measured to take, and it is a guard against a wedge, never a verdict on duration.WIDTH,set_widthand the freebudgetare deleted. At the default 12 wide every number in the source was twelve times itself:virt_irq_storm's 180 s was 36 minutes.cargo test --test toyos-buildatbf28c1e38, 12 wide, on the dev host at load average 40 to 59 on 14 cores (guest-m2.logtoguest-m6.log): 26, 25, 26, 26 and 26 passed in 229.2, 66.8, 37.1, 32.9 and 34.1 s. A wait is bounded by three times its whole test, the one number the suite prints.BOOT_CEILINGGUEST_WEDGED,await_guest's total, now paid out throughQemuInstance::budgetawait_guestvirt_selftest's drainREFERENCE_BOOT_MSis 1424, the largest fastest boot of the five runs (1419, 1424, 1048, 1055, 750 ms), so the measuring host pays every ceiling at 1×.ceiling_verdictis(ceiling * 2).max(ceiling + GUEST_QUIET), and a kernel that died is named at it. Onmainit wasceiling.max(GUEST_WEDGED), which gaveGUEST_WEDGEDtwo readers: this floor andawait_guest's total.GUEST_WEDGEDisawait_guest's alone now. On today's suite only--debugreachesceiling_verdict, throughrun_test; their deletion is filed (Issues, below).VERDICT_POLL(100 ms) is the read loop's poll, and the walk intests/checks/qemu.rsthat holds the backstop steps by it, so the walk moves with the poll.GUEST_QUIET(15 s): the silence that makes a wait a stall, the wait for QEMU's own stop after a guest's last word, and virt_fatal_halts_the_others_first's wait for the other vCPUs to halt. Also the QMP plumbing's bounds, pacing drains and--debug's 60 s.guest / suiteruns (37064140924, 37063415370, 37061831501; 1 wide, 4 cores), each at or under the dev host's slowest. virt_timer_preempts printed 76 to 100 s there: its time includes the build of the job's binary, which no ceiling covers.Wedges end in seconds
QEMU, measured. The wedge is
wedge-irq-storm.patch(in the negative-control comment): theirq-stormselftest loses every tick, so its flood never ends and it says nothing. Its test's verdict is the ceiling alone.75dc056f1, alonecargo test --test toyos-build -- virt_irq_stormFAIL virt_irq_storm: irq-storm never reported, 37 s75dc056f1, whole suite 12 widecargo test --test toyos-buildorigin/mainat73282fa93, alone (638-r5/at-head/mutations/wedge-main.log)cargo test --test toyos-build -- virt_irq_stormAlone the width is 1. In a whole run
mainmultiplies the 180 s by 12: 2160 s (derived, not run).The T14.
toyos_tco::STAGED_BOUND_MSis 10 s, andmetal::bound_forarms the three images that stage their own wedge with it (deadlinewedge,hardlockup,usbload) where every other image keepsWEDGE_BOUND_MS, 120 s. The lockup bound is half of it, 5 s. The full T14 run atb2ccd1c84passed all three under this bound.Two LAN boots ride the talking boot
lan_dhcp_leaseandlan_message_deliveryridetests/lantalkcase, which now names the I219's PCI function, so the loop still pings the address Ubuntu held.tests/lancaseandtests/lanicscaseare deleted with theirALL_CONFIGSrows andlanicscase's machine-record rows.--provoke-messageis deleted from netd, fromtoyos-i219(provoke_message, its two tests, the model'sICSregister) and fromsrc/build.rs's gate, which is one flag now. Atb2ccd1c84the T14 recordedpcidev: slot 0 took its first message on vector 0x28onlantalkcasewith no actuator armed, which was the exit of thelanicscaseissue.tests/lanleasecasestays. netd keeps--exit-with-leasefor it, and deleting the probe takesuserland/netd/src/report.rs,toyos_i219::leaseandtoyos_i219::phy::Outcomewith the driver tests that assert its codes: the I219 driver crate, past this brief.issues/diagnostics/netds-lease-probe-answers-a-question-its-lines-already-answer.mdnames the deletion as its exit.lan_holdkeeps its name: it still holdslanleasecaseanddump_nmi_probe's boot.The T14 loop
wait_onpollsssh's port once a second and asksssh trueonly once the port accepts. It slept 5 s before every probe.CONNECT_SECS(10 s), whatssh's ownConnectTimeoutgives, so a SYN lost while Ubuntu is still up does not read as down; coming back, a dial gets one poll.SharedBoot::membersfollows the runner's bound.toyos_tco::RUST_MEMBER_MS(860) andC_MEMBER_MS(260), besideJOB_BOUND_MS, are what one shipping Rust member and one C case are allowed of it, andmetal::members_fittinggivessharedandccorpusas many as fit inJOB_BOUND_MSless a tenth: 62 and 207, wheremainhas 38 and 90. Each allowance is twice the slowest per-member mean of five full T14 runs taken before The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660 changed the lists.shared-debugkeepsmain's 18: its list is five members.readdir_bound's/homearm is deleted: on the file serversread_dir("/home")answersNotFoundat any count, so 32,769 creates proved nothing. Filed asissues/filesystem/listing-home-answers-not-found-and-its-bound-has-no-guest-test.md.boot.shared.*(1154 / 3608 / 21429, the T14's reading atb2ccd1c84) and loses the rows ofccorpus-2andlanicscase.Carried from #609
metaltalk::hand_backfiresrebootonce the stream has carried this boot'sBoot: complete, or afterCARRIED_WAIT(30 s). This is the exit ofissues/build/the-talking-boots-reboot-outruns-its-log-stream.md, deleted here (Issues, below).metal::said_refusal: a failedtoyos-metalrun's verdict carries the refusal it ended on. Since A test and a build each say when they start and how long they took, the suite's last line splits building from testing, and every statement on stderr opens with the UTC time of day #680 the driver's statement opens with the printer's time-of-day stamp, so the reader findsREFUSAL_HEADpast it (printer::unstamped) on the statement's first line, the one the stamp is on.issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md(Issues, below).Issues
the-lanicscase-boot-is-a-second-t14-flash-for-one-question.md(its exit, above) andthe-cable-judge-reads-three-netd-records-that-cannot-arrive-on-the-t14.md(netd's lease record reached the stick atb2ccd1c84), andthe-talking-boots-reboot-outruns-its-log-stream.md: its exit, two consecutive T14 runs passinglan_talk, is the runs at4859e8d71(comment 5962753432) and75dc056f1(comment 5966698095), both withhand_back. Its rule ishand_back's doc, and its one citation, inissues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md, goes with it;git grepfinds neither its path nor its slug.lanleasecaseissue, whose slug said a third flash.issues/build/a-suites-first-run-after-the-fork-pin-moves-races-its-own-checkout.md(the first run after the merge: EXIT=1, 18 of 26 red,guest-m1.log), held by the toolchain stage ofissues/build/the-tooling-is-a-review-prompt-and-three-workflows.md. A test and a build each say when they start and how long they took, the suite's last line splits building from testing, and every statement on stderr opens with the UTC time of day #680 filed the same unlockedfork_checkoutreddening a new worktree's first run,issues/build/a-worktrees-first-wide-run-reds-on-the-fork-checkout-it-is-still-making.md, which this one names.issues/kernel/virt-mask-windows-read-nine-censuses-and-eight-windows-lines-for-one-cpu.md(the one red of the five measuring runs, onmain's own test, at load average 43 to 59), held by the windows instrument's author.issues/build/run-test-and-its-ceiling-verdict-serve-only-debugs-run.md: The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660's cut leftrun_test,run_test_hooked,run_test_pacedandceiling_verdictone caller,--debug'srun; held byissues/build/the-guest-suite-runs-only-what-no-cheaper-tier-reaches.md.issues/build/host-scale-reads-host-speed-off-boots-that-stop-at-different-markers.md: on the dev host, one run after the other ate5daa61ff,virt_early_panicalone (HVF, toEARLY PANIC:) read a fastest boot of 479 ms andvirt_smpalone (TCG, toSCTLR_EL1=) 1180 ms, both paid at 1.00×, andmachine_shutdownalone (x86-64, TCG, to===READY===) 3316 ms, paid at 2.33×: the scale reads which boots a run held; held by the tooling track.issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md, held by stage 2 ofissues/hardware/the-lan-is-not-yet-production-grade.md.issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md, held byissues/hardware/the-t14-boots-toyos-unattended.md, the track that holds the detector and its bound.Gates
e3094a94acargo run -- --ci host638-r6/host.loge3094a94acargo test --test toyos-checks638-r6/checks.log75dc056f1cargo run -- --build-only638-r5/at-head/build-only.log75dc056f1cargo test --test toyos-build638-r5/at-head/guest.logThe last two stand for this head (Between
75dc056f1and this head, above). Logs are under/Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/;4859e8d71's and the measuring runs' under…/orch/638-recut/.Negative controls
Each is a checked patch, built (
--no-run, EXIT=0), run, and reversed, with the tree clean after it (638-r5/at-head/mutations/run.log). The patches and their run at75dc056f1are in comment 5963627444. The one whose test this round changed,wait_onwithlet answered = listening;, ran again ate3094a94a(638-r6/mutations/run.log): unmutated EXIT=0, built EXIT=0, mutated EXIT=101 at the first assertion (left: Ok(0),right: Err(Silent { what: "come back", secs: 1 })), tree clean before and after.hand_backwithout its waitthe_hand_back_waits_for_the_boot_record_the_judge_readswait_for_bootanswering at oncesaid_refusalkeeping its first line onlya_refusal_is_read_back_off_the_drivers_stderrsaid_refusallooking forREFUSAL_HEADat the start of the stamped linewait_onwithlet answered = listening;coming_back_is_ssh_answering_and_not_its_portbound_forarming every image with 120 severy_arm_that_stops_this_machine_is_cleared_and_judged_as_oneceiling + GUEST_QUIET - 1 sserial_vocabulary: "silent from 4.2s under a 5s ceiling"ceiling + 6 sserial_vocabulary: "silent from 0ns"ceiling * 2serial_vocabulary: "silent from 0ns"ceiling.max(300 s),main'sa_stall_stays_redserial_vocabulary: "a kernel death at 5.2s"Oracles
hand_back: the T14's own records at4859e8d71(below) and atb2ccd1c84.75dc056f1(below).port_acceptsthe loop dials the T14 with.said_refusal: the T14's own stderr from metal: the LAN lease rides the talking boot, and two T14 flashes go #609's run at7d2e15c0, with the printer's stamp A test and a build each say when they start and how long they took, the suite's last line splits building from testing, and every statement on stderr opens with the UTC time of day #680 puts on the driver's statement.The T14
Read at
4859e8d71(comment 5962753432)The orchestrator ran three boots from the staged request, each image's sha256 checked before it was flashed, each
toyos-metalexit 0, then the two judges from the clean worktree at that head. No row armed the TCO watchdog, a staged wedge or the hard-lockup probe.toyos-metallantalkcase(--nic 0000:00:1f.6 --talk …)c01bdbdda51f554281b3729a9519689e020d57094720699a25a5634c67682814talk: the stream carried `Boot: complete` (1153 ms), 822 ms after it openedlanleasecase(--nic 0000:00:1f.6)5efcceb3e0455de666fe68617cf8f6b4b72fbb4a88069b926ea57acf101970a7testcases-readdirb51844becf2b8f8704067c1afde7dc4969082dc0f1f00092ba8e9cc7a61e62aacargo test --test toyos-build -- --metal --metal-readback <dir> lan_, judged: exit 1,3 passed, 1 failed, 2 boot(s).PASS lan_talk(415 lines over the cable,Boot: completeamong them in the stick's order;echoanswered byte for byte, status 0, 73 ms after the stream opened;reboottaken).PASS lan_message_delivery(slot 0 on vector 0x28 at 1.193 s, its first message at 8.485 s).PASS lan_lease_report(netd exit 83, leased 192.168.1.48/24, 5 sent, 14 received).FAIL lan_dhcp_lease,main's red, both findings the oneissues/hardware/the-benchs-router-leases-toyos-another-address-than-ubuntu.mdrecords: the boot leased .48, the host pinged .46, and that ping was answered on the far side of the reset.… readdir_bound, judged: exit 0,PASS readdir_bound,1 passed, 0 failed, 1 boot(s).Read at
75dc056f1, Drive mode (comment 5966698095)The owner approved these, watchdog rows included. The orchestrator ran each as
cargo test --test toyos-build -- --metal [filter]from the clean worktree at the head, without--metal-readback, the stick's log partition saved before each run's first flash; no judging left a row intests/metal.boot_deadline_ends_a_wedgethe boot deadline expired: a bound of 10000 ms, reached at 10062 ms; back in 51 s;PASShard_lockup_ends_a_deaf_cpucpu7 has taken no interrupt for 5000 ms, with IF clear at every sample in that span. Its bound is 5000 ms.; back in 59 s;PASSusb_reset_records_the_phase_it_cutthe controller had that device's data endpoint Running with 252 TRB(s) it had not reached on the ring; back in 54 s;PASS270 passed, 2 failed, 25 boot(s)in 2273 sshared62 (61 ms per member),shared-219 (26 ms),shared-debug5 (31 ms),ccorpus136 (25 ms), noccorpus-2.PASSin the unfiltered run: the three rows above again,loader_watchdog_arms,watchdog_fed,lan_talk,lan_message_delivery,lan_lease_report,readdir_bound,hda_tone.main's and recorded, and there is no other:lan_dhcp_lease(issues/hardware/the-benchs-router-leases-toyos-another-address-than-ubuntu.md) andhda_client_stall(issues/audio/hda-client-stall-reads-one-resume-where-its-judge-wants-two.md).Unsure
host_scale, which reads the run's fastest boot; filed above, that is an early-panic boot on nearly any host, so a whole run pays 1× and an AArch64 guest on a 4-core runner is widened only byvcpus/cores. The three CI runs above say the margin is there today.BOOT_CEILINGhave the least room: screen_fatal_behind_a_painter's first wait is 30 s against a 29 s test.75dc056f1readsharedat 62 members in 61 ms per member; one boot is all that says so.toyos-tco, beside the bound they divide, and only the harness reads them.HARD_LOCKUP_BOUND_MS's reader, so its merge moves the citation.lan_dhcp_leaseismain's red on the T14, so after the fold no green row says the T14 leased an address except throughlan_talkandlan_lease_report.guest / suitebound is 60 minutes against 18 m 11 s for the slowest of six of CI's gate runs, 3.3 times. It is untouched: the job builds the tree, and its TCG arm has no run to measure.🤖 Generated with Claude Code
https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm