Skip to content

Every QEMU ceiling is at most three times what its test takes; wedges end in seconds; two LAN boots ride the talking boot - #638

Merged
Japabu merged 27 commits into
mainfrom
wt/toyos-tight
Oct 3, 2026
Merged

Japabu merged 27 commits into
mainfrom
wt/toyos-tight

Conversation

@Japabu

@Japabu Japabu commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Head e3094a94a, on 75dc056f1, which merges origin/main at 73282fa93 (#648, #680). git diff --shortstat origin/main...HEAD: 38 files, +717 −704. Shipping code (netd, toyos-i219's lib.rs and regs.rs, toyos-tco) +23 −63; toyos-i219's model and tests +2 −57; src/ production +112 −14 and its tests +109 −58; tests/ +169 −351; issues/ +302 −161.

Between 75dc056f1 and this head

git diff --stat 75dc056f1 e3094a94a:

 ...-talking-boots-reboot-outruns-its-log-stream.md | 45 ----------------------
 ...ds-i219-drops-a-transmit-burst-past-its-ring.md |  5 +--
 src/metal.rs                                       | 14 +------
 3 files changed, 4 insertions(+), 60 deletions(-)

Test code and issues only. Both hunks in src/metal.rs lie inside its #[cfg(test)] mod tests, in coming_back_is_ssh_answering_and_not_its_port. Nothing the loop does at run time changed, nor anything the image or the guest suite builds, so the T14 readings, the image and the guest suite at 75dc056f1 stand for this head.

Every QEMU ceiling is at most three times what its test takes

The rule is stated once, at tests/common/qemu.rs's budget_smp: a ceiling is at most three times the slowest its test was measured to take, and it is a guard against a wedge, never a verdict on duration.

  • The width no longer multiplies a ceiling. WIDTH, set_width and the free budget are deleted. At the default 12 wide every number in the source was twelve times itself: virt_irq_storm's 180 s was 36 minutes.
  • The measurement is five whole runs of cargo test --test toyos-build at bf28c1e38, 12 wide, on the dev host at load average 40 to 59 on 14 cores (guest-m2.log to guest-m6.log): 26, 25, 26, 26 and 26 passed in 229.2, 66.8, 37.1, 32.9 and 34.1 s. A wait is bounded by three times its whole test, the one number the suite prints.
ceiling was is the tests that wait through it, slowest pass in s
BOOT_CEILING 10 s × width, 120 s 63 nested_nmi_is_loud 21, the slowest test that is one boot and nothing after it
GUEST_WEDGED, await_guest's total, now paid out through QemuInstance::budget 300 141 iommu_virtio_platform 47, the slowest of the sixteen tests that wait through await_guest
virt_selftest's drain 180 36 virt_irq_storm 12, virt_timer_floor 9
virt_user_mode's drain 180 30 10
virt_early_panic's two waits 30, 10 6, 6 2
virt_early_fault's two waits 30, 10 6, 6 2
screen_fatal_behind_a_painter's second wait 20 87 29
screen_fatal_halt_composited's first wait 30 96 32
screen_panic_muted's wait 30 30 19
virt_el2_drop's drain 10 10 7
screen_fatal_behind_a_painter's first wait 30 30 29
screen_fatal_halt_composited's second wait 40 40 32
  • The two raised waits were each under their own test's slowest run (20 s against 29 s, 30 s against 32 s), so how much room the wait itself had is unmeasured; they are at the rule's bound. Every measured run passed inside the old numbers too.
  • The four kept ones are at or above their test's slowest run and under three times it, and every measured run passed inside them.
  • REFERENCE_BOOT_MS is 1424, the largest fastest boot of the five runs (1419, 1424, 1048, 1055, 750 ms), so the measuring host pays every ceiling at 1×.
  • The backstop of ceiling_verdict is (ceiling * 2).max(ceiling + GUEST_QUIET), and a kernel that died is named at it. On main it was ceiling.max(GUEST_WEDGED), which gave GUEST_WEDGED two readers: this floor and await_guest's total. GUEST_WEDGED is await_guest's alone now. On today's suite only --debug reaches ceiling_verdict, through run_test; their deletion is filed (Issues, below).
  • VERDICT_POLL (100 ms) is the read loop's poll, and the walk in tests/checks/qemu.rs that holds the backstop steps by it, so the walk moves with the poll.
  • Unchanged: GUEST_QUIET (15 s): the silence that makes a wait a stall, the wait for QEMU's own stop after a guest's last word, and virt_fatal_halts_the_others_first's wait for the other vCPUs to halt. Also the QMP plumbing's bounds, pacing drains and --debug's 60 s.
  • On CI's runners the same tests took 1 to 9 s in three of CI's guest / suite runs (37064140924, 37063415370, 37061831501; 1 wide, 4 cores), each at or under the dev host's slowest. virt_timer_preempts printed 76 to 100 s there: its time includes the build of the job's binary, which no ceiling covers.

Wedges end in seconds

  • QEMU, measured. The wedge is wedge-irq-storm.patch (in the negative-control comment): the irq-storm selftest loses every tick, so its flood never ends and it says nothing. Its test's verdict is the ceiling alone.

    arm command result
    75dc056f1, alone cargo test --test toyos-build -- virt_irq_storm EXIT=1, FAIL virt_irq_storm: irq-storm never reported, 37 s
    75dc056f1, whole suite 12 wide cargo test --test toyos-build EXIT=1, 25 passed and that one red at 38 s, 45.4 s in all
    the whole change reverted onto origin/main at 73282fa93, alone (638-r5/at-head/mutations/wedge-main.log) cargo test --test toyos-build -- virt_irq_storm EXIT=1, the same red at 181 s

    Alone the width is 1. In a whole run main multiplies the 180 s by 12: 2160 s (derived, not run).

  • The T14. toyos_tco::STAGED_BOUND_MS is 10 s, and metal::bound_for arms the three images that stage their own wedge with it (deadlinewedge, hardlockup, usbload) where every other image keeps WEDGE_BOUND_MS, 120 s. The lockup bound is half of it, 5 s. The full T14 run at b2ccd1c84 passed all three under this bound.

Two LAN boots ride the talking boot

  • lan_dhcp_lease and lan_message_delivery ride tests/lantalkcase, which now names the I219's PCI function, so the loop still pings the address Ubuntu held. tests/lancase and tests/lanicscase are deleted with their ALL_CONFIGS rows and lanicscase's machine-record rows.
  • --provoke-message is deleted from netd, from toyos-i219 (provoke_message, its two tests, the model's ICS register) and from src/build.rs's gate, which is one flag now. At b2ccd1c84 the T14 recorded pcidev: slot 0 took its first message on vector 0x28 on lantalkcase with no actuator armed, which was the exit of the lanicscase issue.
  • tests/lanleasecase stays. netd keeps --exit-with-lease for it, and deleting the probe takes userland/netd/src/report.rs, toyos_i219::lease and toyos_i219::phy::Outcome with the driver tests that assert its codes: the I219 driver crate, past this brief. issues/diagnostics/netds-lease-probe-answers-a-question-its-lines-already-answer.md names the deletion as its exit.
  • So lan_hold keeps its name: it still holds lanleasecase and dump_nmi_probe's boot.

The T14 loop

  • wait_on polls ssh's port once a second and asks ssh true only once the port accepts. It slept 5 s before every probe.
  • Going down is a dial given CONNECT_SECS (10 s), what ssh's own ConnectTimeout gives, so a SYN lost while Ubuntu is still up does not read as down; coming back, a dial gets one poll.
  • SharedBoot::members follows the runner's bound. toyos_tco::RUST_MEMBER_MS (860) and C_MEMBER_MS (260), beside JOB_BOUND_MS, are what one shipping Rust member and one C case are allowed of it, and metal::members_fitting gives shared and ccorpus as many as fit in JOB_BOUND_MS less a tenth: 62 and 207, where main has 38 and 90. Each allowance is twice the slowest per-member mean of five full T14 runs taken before The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660 changed the lists. shared-debug keeps main's 18: its list is five members.
  • readdir_bound's /home arm is deleted: on the file servers read_dir("/home") answers NotFound at any count, so 32,769 creates proved nothing. Filed as issues/filesystem/listing-home-answers-not-found-and-its-bound-has-no-guest-test.md.
  • The machine record gains boot.shared.* (1154 / 3608 / 21429, the T14's reading at b2ccd1c84) and loses the rows of ccorpus-2 and lanicscase.

Carried from #609

Issues

  • Closed: the-lanicscase-boot-is-a-second-t14-flash-for-one-question.md (its exit, above) and the-cable-judge-reads-three-netd-records-that-cannot-arrive-on-the-t14.md (netd's lease record reached the stick at b2ccd1c84), and the-talking-boots-reboot-outruns-its-log-stream.md: its exit, two consecutive T14 runs passing lan_talk, is the runs at 4859e8d71 (comment 5962753432) and 75dc056f1 (comment 5966698095), both with hand_back. Its rule is hand_back's doc, and its one citation, in issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md, goes with it; git grep finds neither its path nor its slug.
  • Renamed: the lanleasecase issue, whose slug said a third flash.
  • Filed here, each with its owner:
    • issues/build/a-suites-first-run-after-the-fork-pin-moves-races-its-own-checkout.md (the first run after the merge: EXIT=1, 18 of 26 red, guest-m1.log), held by the toolchain stage of issues/build/the-tooling-is-a-review-prompt-and-three-workflows.md. A test and a build each say when they start and how long they took, the suite's last line splits building from testing, and every statement on stderr opens with the UTC time of day #680 filed the same unlocked fork_checkout reddening a new worktree's first run, issues/build/a-worktrees-first-wide-run-reds-on-the-fork-checkout-it-is-still-making.md, which this one names.
    • issues/kernel/virt-mask-windows-read-nine-censuses-and-eight-windows-lines-for-one-cpu.md (the one red of the five measuring runs, on main's own test, at load average 43 to 59), held by the windows instrument's author.
    • issues/build/run-test-and-its-ceiling-verdict-serve-only-debugs-run.md: The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660's cut left run_test, run_test_hooked, run_test_paced and ceiling_verdict one caller, --debug's run; held by issues/build/the-guest-suite-runs-only-what-no-cheaper-tier-reaches.md.
    • issues/build/host-scale-reads-host-speed-off-boots-that-stop-at-different-markers.md: on the dev host, one run after the other at e5daa61ff, virt_early_panic alone (HVF, to EARLY PANIC:) read a fastest boot of 479 ms and virt_smp alone (TCG, to SCTLR_EL1=) 1180 ms, both paid at 1.00×, and machine_shutdown alone (x86-64, TCG, to ===READY===) 3316 ms, paid at 2.33×: the scale reads which boots a run held; held by the tooling track.
    • issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md, held by stage 2 of issues/hardware/the-lan-is-not-yet-production-grade.md.
    • issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md, held by issues/hardware/the-t14-boots-toyos-unattended.md, the track that holds the detector and its bound.

Gates

head command exit log
e3094a94a cargo run -- --ci host 0, "Host: 67 step(s), all green" 638-r6/host.log
e3094a94a cargo test --test toyos-checks 0, 31 passed 638-r6/checks.log
75dc056f1 cargo run -- --build-only 0 638-r5/at-head/build-only.log
75dc056f1 cargo test --test toyos-build 0, 26 passed in 44.5 s, 12 wide from load average 3.8 638-r5/at-head/guest.log

The last two stand for this head (Between 75dc056f1 and this head, above). Logs are under /Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/; 4859e8d71's and the measuring runs' under …/orch/638-recut/.

Negative controls

Each is a checked patch, built (--no-run, EXIT=0), run, and reversed, with the tree clean after it (638-r5/at-head/mutations/run.log). The patches and their run at 75dc056f1 are in comment 5963627444. The one whose test this round changed, wait_on with let answered = listening;, ran again at e3094a94a (638-r6/mutations/run.log): unmutated EXIT=0, built EXIT=0, mutated EXIT=101 at the first assertion (left: Ok(0), right: Err(Silent { what: "come back", secs: 1 })), tree clean before and after.

patch test it reds exit
hand_back without its wait the_hand_back_waits_for_the_boot_record_the_judge_reads 101
wait_for_boot answering at once the same 101
said_refusal keeping its first line only a_refusal_is_read_back_off_the_drivers_stderr 101
said_refusal looking for REFUSAL_HEAD at the start of the stamped line the same 101
wait_on with let answered = listening; coming_back_is_ssh_answering_and_not_its_port 101
bound_for arming every image with 120 s every_arm_that_stops_this_machine_is_cleared_and_judged_as_one 101
backstop ceiling + GUEST_QUIET - 1 s serial_vocabulary: "silent from 4.2s under a 5s ceiling" 101
backstop ceiling + 6 s serial_vocabulary: "silent from 0ns" 101
backstop ceiling * 2 serial_vocabulary: "silent from 0ns" 101
backstop ceiling.max(300 s), main's a_stall_stays_red 101
no death named at the backstop serial_vocabulary: "a kernel death at 5.2s" 101

Oracles

The T14

Read at 4859e8d71 (comment 5962753432)

The orchestrator ran three boots from the staged request, each image's sha256 checked before it was flashed, each toyos-metal exit 0, then the two judges from the clean worktree at that head. No row armed the TCO watchdog, a staged wedge or the hard-lockup probe.

boot image sha256 toyos-metal
lantalkcase (--nic 0000:00:1f.6 --talk …) c01bdbdda51f554281b3729a9519689e020d57094720699a25a5634c67682814 exit 0; talk: the stream carried `Boot: complete` (1153 ms), 822 ms after it opened
lanleasecase (--nic 0000:00:1f.6) 5efcceb3e0455de666fe68617cf8f6b4b72fbb4a88069b926ea57acf101970a7 exit 0
testcases-readdir b51844becf2b8f8704067c1afde7dc4969082dc0f1f00092ba8e9cc7a61e62aa exit 0
  • cargo test --test toyos-build -- --metal --metal-readback <dir> lan_, judged: exit 1, 3 passed, 1 failed, 2 boot(s). PASS lan_talk (415 lines over the cable, Boot: complete among them in the stick's order; echo answered byte for byte, status 0, 73 ms after the stream opened; reboot taken). PASS lan_message_delivery (slot 0 on vector 0x28 at 1.193 s, its first message at 8.485 s). PASS lan_lease_report (netd exit 83, leased 192.168.1.48/24, 5 sent, 14 received). FAIL lan_dhcp_lease, main's red, both findings the one issues/hardware/the-benchs-router-leases-toyos-another-address-than-ubuntu.md records: the boot leased .48, the host pinged .46, and that ping was answered on the far side of the reset.
  • … readdir_bound, judged: exit 0, PASS readdir_bound, 1 passed, 0 failed, 1 boot(s).

Read at 75dc056f1, Drive mode (comment 5966698095)

The owner approved these, watchdog rows included. The orchestrator ran each as cargo test --test toyos-build -- --metal [filter] from the clean worktree at the head, without --metal-readback, the stick's log partition saved before each run's first flash; no judging left a row in tests/metal.

run exit read
boot_deadline_ends_a_wedge 0 the boot deadline expired: a bound of 10000 ms, reached at 10062 ms; back in 51 s; PASS
hard_lockup_ends_a_deaf_cpu 0 cpu7 has taken no interrupt for 5000 ms, with IF clear at every sample in that span. Its bound is 5000 ms.; back in 59 s; PASS
usb_reset_records_the_phase_it_cut 0 deadline at 10066 ms; the controller had that device's data endpoint Running with 252 TRB(s) it had not reached on the ring; back in 54 s; PASS
unfiltered 1 270 passed, 2 failed, 25 boot(s) in 2273 s
  • The shared boots at this head's sizes, each one boot with every member run: shared 62 (61 ms per member), shared-2 19 (26 ms), shared-debug 5 (31 ms), ccorpus 136 (25 ms), no ccorpus-2.
  • PASS in the unfiltered run: the three rows above again, loader_watchdog_arms, watchdog_fed, lan_talk, lan_message_delivery, lan_lease_report, readdir_bound, hda_tone.
  • The two reds are main's and recorded, and there is no other: lan_dhcp_lease (issues/hardware/the-benchs-router-leases-toyos-another-address-than-ubuntu.md) and hda_client_stall (issues/audio/hda-client-stall-reads-one-resume-where-its-judge-wants-two.md).

Unsure

  • The ceilings were measured on the dev host, where x86-64 guests and the EL2 guests are emulated. CI pays them out through host_scale, which reads the run's fastest boot; filed above, that is an early-panic boot on nearly any host, so a whole run pays 1× and an AArch64 guest on a 4-core runner is widened only by vcpus/cores. The three CI runs above say the margin is there today.
  • A wait is bounded by three times its whole test because nothing measures the wait alone. The four kept waits and BOOT_CEILING have the least room: screen_fatal_behind_a_painter's first wait is 30 s against a 29 s test.
  • The shared boots' allowances were measured on lists The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660 has since changed. One unfiltered run at 75dc056f1 read shared at 62 members in 61 ms per member; one boot is all that says so.
  • The allowances sit in toyos-tco, beside the bound they divide, and only the harness reads them.
  • The hard-lockup issue is held by the in-tree track that holds the detector. Tracks: the T14 runs its tests in sessions of a booted ToyOS, and every machine ships a watchdog and a hard-lockup detector #631, not landed, deletes that track for one that plans HARD_LOCKUP_BOUND_MS's reader, so its merge moves the citation.
  • lan_dhcp_lease is main's red on the T14, so after the fold no green row says the T14 leased an address except through lan_talk and lan_lease_report.
  • CI's guest / suite bound is 60 minutes against 18 m 11 s for the slowest of six of CI's gate runs, 3.3 times. It is untouched: the job builds the tree, and its TCG arm has no run to measure.

🤖 Generated with Claude Code

https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm

Japabu and others added 9 commits September 30, 2026 16:18
…to take

The rule is stated once, at `qemu::budget_smp`: a ceiling is at most three
times the slowest the test it bounds was measured to take in the whole-suite
runs on the dev host at the default width, whose times already carry the guests
a test shares that host with. A wait is bounded by that multiple of its whole
test, the one number measured. `host_scale` still widens it on a slower host and
`oversubscription` for a guest wider than the host; nothing else does.

The evidence is the orchestrator's six whole-suite logs (536r20, 536r21, 588r3,
633-x86, 633r2, 634: 495-499 tests each, 272-1000 s), the slowest PASS of each
test across them.

What moved:

- The phase width no longer multiplies a ceiling. At 12 wide it made every
  number in the source twelve times itself; `port_poll_churn`'s 5 s became
  263 s of guard (with an early host scale of 4.4) before it was called
  stalled. The measured times already carry the width, so `WIDTH` and
  `set_width` are deleted, and `round_trip`, which existed only to not be
  width-scaled, is `qemu.budget` now.
- A guest still talking at twice its ceiling is timed out. The backstop was
  `ceiling.max(GUEST_WEDGED)`, so every chatty stall cost 300 s.
- A boot's ceiling is 30 s, three times the slowest test that is one boot and
  one trivial command (10 s); it was 10 s times max(width, 2), 120 s at 12 wide.
- The shared boot's members: 5 s (every member measured at 1.5 s or under),
  with mutual_kill 47, poll_wake_pipe 17, process_stats 11, abuse_elf_loader
  and toybox_file_tools 8; `disk_backtrace`'s 15 s override went (0.15 s at
  most). C members 2 s (0.54 s at most), from 10.
- Every mapped wait: at most three times its test's slowest whole run. Cut
  where the literal was above that (doom_frames 300 -> 77, metal_sim_* 240 ->
  74 and 14, cache_eviction 180 -> 59, blockd's 600 -> 152 and 251, ...);
  raised where a width-scaled wait sat under its own test's measured time and
  had only passed on the width's twelve-fold (toolkit_iced 30/60 -> 332,
  blocking_read_window 30 -> 179, mkdir_cap 60 -> 188, i8042_health 20 -> 164,
  metal_sim_pointer_churn 20 -> 92, kernel_log_file 10 -> 56). Waits that were
  never width-scaled keep any literal under the rule, since every measured run
  passed inside it.
- Tests that wait out a deadline arm a shorter one: `boot-deadline=8000` for
  the wedges (15000), `boot-deadline=12000` for the hard lockup (30000, whose
  half stays above the probe's 3 s reach).
- readdir_bound's `/home` arm lists one past the bound (16,385 files) rather
  than one past the old walk's doubling (32,769): the doubling is the bcachefs
  host test `a_tree_past_the_ceiling_is_refused_before_it_materialises`, with
  the allocator's peak as its instrument. Its ceiling is three times its
  slowest run, 442 s on #536's file server.

Untouched: the redlisted tests (no measurement), `GUEST_WEDGED` and
`GUEST_QUIET` (the `Liveness` waits have no per-test number; 300 s is under
twice the slowest test that uses one), the QMP plumbing's own bounds, and
pacing drains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…s the machine once a second

Four cuts to what a full T14 run spends waiting, from the rig research's
per-boot budget (three full runs' readbacks) and the latest metal logs.

- A boot armed with one of `WEDGE_ARMS` (deadlinewedge, hardlockup, usbload)
  carries `boot-deadline=` `metal::STAGED_BOUND_MS`, 10 s, instead of the
  120 s every other boot keeps for a wedge nobody staged. The T14 stages its
  wedge 1.505-1.511 s into the kernel (630r4), and those three boots came back
  in 110-166 s against a ~40-84 s ordinary boot. The lockup's bound is half of
  it, 5 s, which still outlasts the probe's 3 s reach to its lock; the load
  sweep simply stops sooner. `bound_for` is the one place the choice is made,
  in `tests/common/metal.rs`'s build, where the header already said a boot
  wanting a different bound is a change here and not a field on an arm. The
  three boots' `complete_ms` rows name the new bound, and the two lateness
  rows go from 10 s, "a twelfth of the bound", to twice their reading.
- `wait` polls `ssh`'s port once a second, and asks `ssh true` only once the
  port accepts: coming back is still `ssh` answering. It slept 5 s before every
  probe, and a probe of a machine that is down cost `ConnectTimeout`'s 10 s.
- A refusal after `reboot` waits for the machine to answer `ssh` again before
  it returns, unless the refusal was that wait itself, so the next boot's first
  `ssh` does not refuse a boot that never happened (the lanleasecase and
  lanswapcase arms of `issues/hardware/the-t14-stopped-answering-ssh-between-
  two-lan-boots.md`). The steps after `reboot` move into `after_the_reboot`.
- A swapping boot whose own invocation was refused kills its `--swap`
  invocation instead of letting it dial for the whole 420 s.
- The list allowances are twice the slowest per-member mean over five full
  runs: shared 845 -> 428 ms measured, allowance 1400 -> 860; shared-debug
  1926 -> 398, 2900 -> 800; ccorpus 396 -> 127, 600 -> 260. The corpus now
  fits one boot (130 cases at 207 a boot), so `ccorpus-2` and its six rows
  go, with the two tests that pinned the old cut. `shared` stays two boots.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Three conflicts, every hunk of both sides accounted for:

- tests/common/volumes.rs: #536 deleted `writeback_durability`; this branch
  had cut its one `run_test` ceiling (60 -> 26). The test is gone, so is the
  cut.
- tests/toyos.rs: #536 deleted the `cache_eviction` arm; this branch had cut
  its ceiling (180 -> 59). Same.
- tests/metal-profile.toml: #536 added `shared-3`'s rows and kept
  `ccorpus-2`'s; this branch's allowances (shared 860 ms, ccorpus 260 ms a
  member) fit #536's 77 shared members in two boots and the corpus in one, so
  neither chunk exists: the `ccorpus-2` rows stay deleted and every
  `shared-3` row (six, five of them outside the conflict) goes too.

What #536 brought under the rule:

- Two shared members it measured past 1.5 s get their own ceilings:
  fs_cache_eviction 50 s (16 s measured), fs_turns 11 s (3 s).
- Its new ceilings, three times their tests' slowest runs on #536's two whole
  suites: blockd_serves_nothing 60 -> 29, fsd_restart 60 and 120 -> 35,
  fsd_claim_held 60 -> 59.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…by `NotFound` at any count

The arm created 32,769 files on `/home` and passed on any error from
`read_dir("/home")`. Before #536 that error was the listing bound's
(`out of memory` on 588r3, 633r2 and 634); on #536's two whole runs it was
`entity not found`, and so was `/home`'s own root afterwards, which an empty
`/home` answers too. The 32,769 creates through the DATA server were most of
readdir_bound's 364 s and 442 s there, against 46-181 s before.

So the `/home` arm and the `/home` half of the per-directory check are
deleted, and the bound is `/tmp`'s alone, which is still the kernel's
`vfs::MAX_LIST_ENTRIES` at its full count. The bcachefs library's own walk past
the old doubling stays covered on the host by
`a_tree_past_the_ceiling_is_refused_before_it_materialises`. Filed:
`issues/filesystem/listing-home-answers-not-found-and-its-bound-has-no-guest-test.md`,
for the `NotFound` and for the DATA server's `WINDOW_BYTES` refusal that no
guest test now reaches.

The ceiling is three times the slowest whole run on the runs where the arm
was the kernel's (181 s), an upper bound on the test that is left until a
suite measures it: 545 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…were four

`lancase`, `lanicscase` and `lanleasecase` each flashed the stick and booted
the machine once more (100-208 s a cycle in the rig's budget) to ask a question
the talking boot, `lantalkcase`, now answers:

- `lan_dhcp_lease` rides `lantalkcase`, which names the I219's function, so the
  loop pings the address it held under Ubuntu there as it did on `lancase`.
  netd's lines cross on the stick since each program has a log ring (#616), and
  on 630r4's full run `lantalkcase` carried the same MAC, link-up and lease
  lines `lancase` did. The judge drops `lan_hold`'s exit, which the talking
  boot's own judge replaces with `lan_talk_hold`'s. `lan.lancase.*` and
  `boot.lancase.ping_secs` are the talking boot's rows now.
- `lan_message_delivery` rides it too, as `delivered_on_metal`. Its issue's exit
  condition was the shipping boot recording `pcidev: slot N took its first
  message` without the actuator, and all three LAN boots of 630r4 did, the
  talking one at 8.498 s. So `--provoke-message` has no question left: it goes
  from netd, from `toyos-i219` (`provoke_message` and the two tests of it) and
  from `build.rs`'s Intel-actuator gate, and
  `issues/hardware/the-lanicscase-boot-is-a-second-t14-flash-for-one-question.md`
  closes.
- `lan_lease_report`'s metal row goes: netd's probe exists for a console line
  that could not cross, and the lease it reports is `lan_dhcp_lease`'s to judge
  off netd's own lines. The QEMU registration stays, since its link-flap check
  reads the probe's report; the issue that tracked the whole probe keeps that
  half under a slug that says what is left,
  `issues/diagnostics/netds-lease-probe-answers-a-question-its-lines-already-answer.md`,
  and `Readback::log_volume_file`, its only reader on the metal side, goes.

Deleted with them: the three configs and their `ALL_CONFIGS` rows, their
profile rows (the talking and swapping boots' rows that were derived "as
lancase" now state that derivation), and `lan::CONFIG`, `BOOT`, `ICS_*`,
`LEASE_CONFIG`, `LEASE_BOOT` and `JOBS`. `lan_hold` stays: `testcases-deaf`
holds its boot open with it, and the metal-profile check of its window now
reads that boot's allowance. The ssh issue's exit condition names the one arm
it still owns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`blockd_serves_partitions` is 40-50 s on every whole run, and its `bench` role
writes and reads its partition twice through blockd, one request at a time and
then fifteen at once, with QEMU tracing every NVMe command to a file. What the
bench's size buys is the throughput it prints, and a QEMU run judges no time.
What the test asserts off the role holds at a quarter of it: every block reads
back what was written; 2048 blocks are 64 requests of the ring's 32-block
`MAX_REQUEST_BLOCKS`, past the fifteen in flight that make QEMU's trace see
more than one command outstanding and more than one submission queue; and
`write_all` ends every pass in a flush, so the trace carries a Flush whatever
the arena's size. The reset role writes two of the partition's blocks and the
survival test counts reissues in its span, neither of which reads its length.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`host` on `macos-latest` took 11, 13 and 14 minutes on the last three nightly
runs that reached their end (36496779560, 36550208853, 36600425263) and was
bounded at 90. The other nightly bounds are already within three times their
slowest runs over the same four nightlies: `guest` 60 against 27, `tcg` 60
against 18, and `build` and the two `portability` jobs 350 against 184 and 166;
the PR gate's `host` is 45 against 27.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…hat use them

Two helper waits in `tests/common/power.rs` sat under the slowest test that
reaches them and had passed only on the phase width's twelve-fold:

- `WAIT`, the bound on QEMU's shutdown event and the drain after it, was 20 s
  (240 s at 12 wide). Its callers' slowest whole runs are 23 s
  (job_deadline_reboots), 25 s (watchdog_resets) and 110 s
  (quiesce_refuses_a_second_shutdown, on #536's first whole run): 332 s.
- `CHAIN_WAIT`, what the boot after a reset has to arrive inside, was
  `PANIC_FAST_SECS + 60` (780 s at 12 wide). The tests that are one chain took
  at most 37 s (hard_lockup_ends_a_deaf_cpu), 33 s (boot_deadline_ends_a_wedge),
  27 s and 25 s: 113 s, which the multi-chain usb_reset tests take per chain.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`tcg` took 11, 17 and 18 minutes on the last nightlies that finished it, so its
bound is 54 minutes, not 60. `cfd4ba79a`'s message counted it among the bounds
already inside the rule, and it was not. That message also named the wrong
runs: the `host` times it gives are nightlies 36696295750 (11.6 min),
36709239346 (13.6), 36600425263 (14.0) and 36597513896 (12.9).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu
Japabu marked this pull request as ready for review September 30, 2026 15:00
@Japabu Japabu changed the title Every ceiling is at most three times what its test takes; wedges end in seconds Every QEMU ceiling is at most three times what its test takes; wedges end in seconds; the LAN boots ride the talking boot Sep 30, 2026
#630 records metal timings per machine in tests/metal/<machine>.toml and
deletes tests/metal-profile.toml and src/metalprofile.rs; a shared boot's
list is cut by `SharedBoot::members`, a count fixed at the site. Every hunk
of both sides is accounted for:

- tests/metal-profile.toml, src/metalprofile.rs: deleted, as #630 does. What
  this branch carried in them moves or goes:
  - The per-member allowances (shared 860 ms, shared-debug 800 ms, ccorpus
    260 ms, twice the slowest per-member mean over five full T14 runs) become
    `members` 62, 67 and 207 in tests/toyos.rs, where #630 had 38, 18 and 90.
    Both sets are the same derivation, (JOB_BOUND_MS - JOB_BOUND_MS / 10) /
    allowance, so #630's are the old 1400/2900/600 ms allowances and these
    replace them in the one place a list is sized.
  - The staged-bound rows (`complete_ms` and lateness at STAGED_BOUND_MS) and
    their constant in the declared-ceilings list go: #630 judges a deadline's
    lateness against one timer period and records no ceiling.
    `metal::bound_for`'s own test still pins the bound per arm.
  - `the_window_lan_hold_sleeps_is_the_window_this_file_prices` goes with the
    job allowances it read; #630 prices no list.
- src/metal.rs: #630's `run` reads the machine and ends in
  `judge_and_write_readback`; this branch moved everything after `reboot`
  into `after_the_reboot`. The body is #630's, and `after_the_reboot` takes
  the `Machine` it writes into the boot file.
- tests/common/lan.rs: #630 dropped the profile doc on `CONFIG`/`BOOT`; this
  branch deletes `CONFIG`/`BOOT` and the lanicscase/lanleasecase judges.
- issues/diagnostics/the-lanleasecase-boot-...md and
  issues/hardware/the-lanicscase-boot-...md: #630 struck their
  metal-profile rows; this branch deletes both (the first renamed to
  netds-lease-probe-answers-a-question-its-lines-already-answer.md, which
  names no profile row; the second closed).

The T14's record, tests/metal/lenovo-20w0003amz.toml: the rows of boots that
no longer exist go, nine of them: `ccorpus-2` (the corpus fits one boot at
207), `lanicscase` and `lanleasecase` (folded into `lantalkcase`). It held
none for `lancase`. The wedge boots' rows stay: `complete_ms` and the panel
are read before the staged bound matters. No number is added; the next full
T14 run records the new boots' rows.

Filed issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md:
with the profile gone, `toyos_tco::HARD_LOCKUP_BOUND_MS` has no reader but
the assertion beside it, as on main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu

Japabu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Orchestrator runs at b2ccd1c (logs orch/logs/638-*.log):

job command exit
whole cargo test --test toyos-build -- 0
T14 full cargo test --test toyos-build -- --metal 1 — [metal] 260 passed, 3 failed, 23 boot(s), wall 2101 s

The same full T14 run on main-equivalent code (#630's re-record at 0d2dda6): 27 boots, wall 2899 s.

The three T14 reds:

The run added three rows to the machine record, boot.shared.{complete_ms,panel_max_us,panel_us} = 1154 / 3608 / 21429. They are not committed yet; the patch is kept by the orchestrator.

@Japabu

Japabu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Round 1, head b2ccd1c84.

Readiness. CI host passed at this head (run 36739037448). The whole guest suite passed 498 of 498 (638-whole.log). The full T14 run gave 260 passed, 3 failed, over 23 boots.

Growth. +414 / −881 over 44 files:

  • Shipping code (netd, toyos-i219's lib.rs): +11 / −67.
  • The metal loop (src/metal.rs): +75 / −12. The port poll, bound_for and the wait after a refusal account for it, and I accept that growth.
  • The rest is the harness, configs and issues.

BLOCKER

  • tests/common/qemu.rs:564 — the backstop is now ceiling * 2, and it fires before the kernel-death arm and the silence arm, which both wait for GUEST_QUIET (15 s). That happens whenever the guest went quiet later than 2×ceiling − 15 s, and so at once for every ceiling under 7.5 s: the 5 s shared-member default, the 2 s C members, ipc_hostile_peer's 1 s and locale_detect_unrecognized's 2 s.
    • What goes wrong: a kernel panic or a silent wedge in any of them now reports timed out after 10s, with the guest still talking 10s ago … it was working and did not finish, and the summary counts it under "N of those reds are the ceiling". The panic is not named.
    • Before this branch: the 300 s floor named the panic first, unless it came after 285 s.
    • Patch: let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
    • Test: add to tests/checks/qemu.rs ceiling_self_check. It is red at b2ccd1c84 and green with the patch:
      • ceiling_verdict(Some(KERNEL), Duration::from_secs(11), Duration::from_secs(5), Duration::from_secs(10), 40) must be None;
      • ceiling_verdict(Some(KERNEL), Duration::from_secs(21), Duration::from_secs(5), GUEST_QUIET, 40) must contain kernel panic.
    • With the patch, tests/CLAUDE.md:18 ("a wedge verdict needs both the budget spent and the guest gone quiet … only a far backstop") is true again.

NOTE

  • hda_tone (question 1): not this branch.

    • The testcases boot is the same on both runs: 8 jobs in the same order, and no shared or LAN work rides it.
    • The tone's own window is clean here too: underruns=0 deferred=0 starve_max=0 at 4.456 s.
    • The 104 are 39 + 65 from hda_client_stall's staged session. Its client 1 connected at 5.657 s, before soundd removed the tone's client 0 at 5.661 s. So the tone never flushed a clients=0 line, and job_window (tests/common/audio.rs:13) ran on into the next job.
    • On 630r4 the removal (6.234 s) came before the next connect (6.241 s). This is the same 39 + 65 the issue file records at f5c80264f: a race in the judge's window, not starvation. hda_client_stall's red here is the same race.
  • Four waits still bounded by GUEST_WEDGED (300 s) (question 2). Until they are cut, the rule stated at tests/common/qemu.rs:303 and the PR title are false. Every other unchanged literal I checked is within three times its test.

    site test slowest of the six logs cut to about
    tests/common/lan.rs:408 lan_no_lease 33 s 99 s
    tests/common/volumes.rs:563 kernel_log_file 18 s 54 s
    tests/common/usb.rs:821 usb_flush_optional 33 s 99 s
    tests/common/orphan.rs:42 guest_dies_with_its_harness 8 s 24 s
  • tests/toyos.rs:11132 — readdir_bound's 545 s is three times the old shape. At this head the test took 49 s (638-whole.log), so re-derive it now.

  • tests/toyos.rs:7800 — Liveness::new(40 s quiet, 35 s total): the quiet guard can never end the wait first, so the 40 s has no effect.

  • toyos-i219/src/stub.rs:1165, :1044 and toyos-i219/src/regs.rs:35 — nothing writes ICS since provoke_message went. Delete the model's write arm, its read arm and the constant, so the --provoke-message deletion is whole (question 3).

  • src/build.rs:3350 — INTEL_ACTUATORS: [&str; 1] keeps array-and-find machinery for one flag. One const and one contains would do.

  • src/metal.rs:874 — STAGED_BOUND_MS is a bound a boot arms. toyos-tco's description says it holds "every bound a boot arms", so it belongs there beside WEDGE_BOUND_MS.

  • tests/common/metal.rs:1038 — unbooted is true for any failed exit of the flashing invocation. That includes a boot that ran and was judged red, although the comment says "refused". In that case SIGKILL can cut a --swap that is still writing its readback.

  • src/metal.rs:1450 — no host test reaches wait or port_accepts. The mutation let answered = listening; (no ssh true on come-back) survives every host test. The 23 T14 returns are its only evidence.

  • Deletions (question 3).

    • The /home arm of readdir_bound is deleted whole.
    • netd's --exit-with-lease still ships for one live QEMU check, lan_lease_report's flap. It is tracked in issues/diagnostics/netds-lease-probe-answers-a-question-its-lines-already-answer.md with an owner and an exit condition, so it is not kept just in case.
  • Prompts (question 4): no CLAUDE.md and no .claude/agents/*.md names --provoke-message, lancase, lanicscase, lanleasecase, set_width, round_trip or POLL_SECS. Nothing is owed.

  • Machine record (question 5): commit the three boot.shared.* rows (1154 / 3608 / 21429) in this PR.

    • They are this head's first passing reading of shared, a boot whose list this branch moves from 38 to 62 members.
    • metaltimings::judge wrote them only because shared passed and had no row, and the harness says "commit it".
    • Landing without them leaves a record that this head's own run says is behind.
  • CI's guest shards were not run under the cut ceilings (--jobs 1, KVM; question 2). nightly.yml has workflow_dispatch, so one dispatch on wt/toyos-tight measures it before landing.

REMOVE

  • tests/common/qemu.rs:562-563 — "at most six times the slowest this test was ever measured": false for readdir_bound (1090 s against 49 s), and it restates the rule.
  • tests/common/qemu.rs:4668 — "the slowest of which took 10 s": where a measurement came from goes in the commit, not the source.
  • tests/common/power.rs:1071 — "the slowest of which took 37 s": same.
  • src/metal.rs:870 — "the T14 stages its wedge 1.5 s into the kernel": a measured number that will rot.
  • tests/toyos-rust-tests/src/bin/lan_hold.rs:1-4 and :9-10 — the cable, netd's records and the lease bound are all false for its one remaining user, testcases-deaf.
  • tests/toyos.rs:261 — the rewritten comment restates LEASE_BOUND_MS as "twenty seconds".
  • issues/hardware/the-t14-stopped-answering-ssh-between-two-lan-boots.md:9-10 — "are the two arms the metal suite flashes after lanicscase" is false.
  • issues/hardware/the-cable-judge-spends-two-premises-nothing-has-measured.md:26-27 — "Closed by a lancase run" names a boot that no longer exists.
  • PR body: "Runs owed", "Not run: any guest test, and the T14" and "No number is added; the next full T14 run records every boot the record lacks" are stale now that both runs happened.
  • PR body: "Corrections to this branch's own record" is not main's record.
  • PR body: the 20-slowest table's "kept / Not cut blind" rows carry nothing.

SEND BACK

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Japabu and others added 2 commits September 30, 2026 19:29
…and the last 300 s waits go

The backstop was `ceiling * 2`, and it ran ahead of the panic and silence
arms, which both wait for `GUEST_QUIET` (15 s). Under any ceiling below
7.5 s (the 5 s shared default, the 2 s C members, `ipc_hostile_peer`'s
1 s) a kernel panic or a silent wedge was reported as a slow guest that
was still talking. It is now `(ceiling * 2).max(ceiling + GUEST_QUIET)`.
`ceiling_self_check` gains the review's two cases:
- a panic under a 5 s ceiling, quiet for 10 s at 11 s: no verdict yet.
  Red at b2ccd1c's backstop (EXIT=101), green with this one.
- the same panic at 21 s, quiet for `GUEST_QUIET`: named. Green at
  b2ccd1c too, because the pure function puts the panic arm first. It
  reds (EXIT=101) when the backstop is moved ahead of that arm.
`a_stall_stays_red` stages a 30 s ceiling instead of 5 s, so its `past`
(twice the ceiling, plus one) still clears the backstop.

Four waits still bounded by `GUEST_WEDGED` are cut to about three times
their slowest whole-suite run: `lan_no_lease` 99 s, `kernel_log_file`'s
device poll 54 s, `usb_flush_optional` 99 s, and
`guest_dies_with_its_harness` 24 s, which is now host-scaled like the
other three. `readdir_bound` is 147 s, three times the 49 s it took at
b2ccd1c without its `/home` arm.

`toolkit_winit_loop`'s `Liveness` had a 40 s quiet guard behind a 35 s
total, so the guard could never fire. The loop now runs to a host-scaled
35 s deadline.

The i219 model's `ICS` write arm, read arm and register constant go.
Nothing has written `ICS` since `--provoke-message` went.

`build.rs`'s Intel-actuator gate holds one flag, so it is one `const`
and one membership test. Its declaration test is renamed to say "flag".

`STAGED_BOUND_MS` moves into `toyos-tco` beside `WEDGE_BOUND_MS`. That
crate holds every bound a boot arms.

The metal loop's wait is split into `wait_on` and a `port_accepts(host,
port)`. A new host test drives them against a loopback listener that
accepts and never speaks `ssh`. With `let answered = listening;` the
build passes (EXIT=0) and `metal::tests` fails (EXIT=101) on that test
alone.

The harness kills a swap only when the flashing invocation did not exit
0 or 1. Exit 1 is a boot that ran and was judged red, and its swap is
left to finish.

The T14's record gains `boot.shared.{complete_ms,panel_max_us,panel_us}`
= 1154 / 3608 / 21429. These are the first passing reading of `shared`
at its 62 members.

Deleted as the review asked:
- the backstop's "six times" clause
- the measured "slowest of which took" clauses
- the staging time of the T14's wedge
- `lan_hold`'s docs, which describe a user it no longer has
- the restated twenty seconds
- two issue sentences that name boots which no longer exist
- the two "twice its ceiling" phrasings, which the new backstop makes
  false for short ceilings

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`coming_back_is_ssh_answering_and_not_its_port` went red in
`cargo run -- --ci host` (EXIT=1). Its last case dialled the loopback
listener's port after dropping the listener and found it still
accepting for the whole second. Once let go, that port can be another
socket's before it is dialled: another test's listener, or a
self-connect from an ephemeral source port. Which of the two it was is
not established. Port 0 is one nothing listens on. Three runs of
`cargo test -p toyos-build --lib` after this change: EXIT=0 each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu

Japabu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Orchestrator run at edefd81: cargo test --test toyos-build -- EXIT=0 — test result: ok. 502 passed, 502 total (309.7s) (log orch/logs/638r2-638r2-whole.log). Nightly on the branch: https://github.com/ToyOSOrg/ToyOS/actions/runs/36753172688 (running).

@Japabu

Japabu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Round 2, head edefd8193.

Round 1 BLOCKER

  • tests/common/qemu.rs:564 (the backstop) is OPEN, but narrower.
    • The patch I prescribed is at :563. Case (e)1 fails at b2ccd1c84's ceiling * 2 (EXIT=101) and passes with the patch (EXIT=0), so a kernel death before the ceiling is now named.
    • A death after the ceiling is still not named. My round-1 patch did not cover the whole defect it described; BLOCKER 1 below is the rest.

Readiness

  • CI host passed at edefd8193 (run 36753095781, 15m14s).
  • The whole guest suite passed at edefd8193: EXIT=0, 502 passed (orchestrator comment).
  • The T14 reading is from b2ccd1c84. Since then the metal loop has changed only in a refactor of wait, a same-value move of STAGED_BOUND_MS, and the swap's failure path.

Growth

  • Whole branch: +520 / −948.
  • Shipping code (netd, toyos-i219 lib.rs/regs.rs, toyos-tco): +18 / −69.
  • Since b2ccd1c84: +161 / −122 over 20 files.

BLOCKER

  • tests/common/qemu.rs:563-570 — a kernel that dies after its ceiling is still reported as a slow guest, and it is counted as a ceiling red.
    • Example: a 5 s ceiling and a death at 7 s gives "timed out after 20s, with the guest still talking 13s ago … it was working and did not finish". On main, the 300 s floor named that panic at 22 s.
    • The comment at :561-562 says "so a death is named by the arms above". That is false for this case.
    • Patch: in the backstop arm, before the TIMED_OUT format, add if let Some(line) = dying { return Some(kernel_died_here(line)); }.
    • Test: the death walk below. It fails at edefd8193 on died = 7 and passes with the patch.
  • tests/checks/qemu.rs:201-220 — case (e) cannot fail on the claim :563 makes.
    • To the brief's question: case (e)1 alone does not prove the fix. It only requires a backstop of at least 11 s under a 5 s ceiling. The claim is that the backstop is at least ceiling + GUEST_QUIET.
    • The implementer is right about (e)2: it tests the order of the arms, not the backstop.
    • Mutation that passes today: let backstop = (ceiling * 2).max(ceiling + Duration::from_secs(6));. It passes:
      • (e)1, because 11 s is not past 11 s;
      • (e)2, because the panic arm is checked first;
      • every other ceiling_self_check case, whose ceilings are 153 s and 380 s;
      • a_stall_stays_red, where 30 s gives a 60 s backstop.
    • Under that mutation, a guest silent from 0 s under a 5 s ceiling reads "timed out after 11s" instead of STALLED at 15 s. That is round 1's defect back.
    • Replace (e) with two walks:
      const SHORT: Duration = Duration::from_secs(5);
      let first = |dying: Option<&str>, from: u64| {
          (from..).map(Duration::from_secs).find_map(|elapsed| {
              ceiling_verdict(dying, elapsed, SHORT, elapsed - Duration::from_secs(from), 40)
          })
      };
      for since in 0..=SHORT.as_secs() {
          let got = first(None, since);
          if !got.as_deref().is_some_and(|v| v.starts_with(STALLED)) {
              return Err(format!("silent from {since}s under a {SHORT:?} ceiling: {got:?}"));
          }
      }
      for died in 0..=4 * SHORT.as_secs() {
          let got = first(Some(KERNEL), died);
          if !got.as_deref().is_some_and(|v| v.contains("kernel panic")) {
              return Err(format!("a kernel death at {died}s under a {SHORT:?} ceiling: {got:?}"));
          }
      }
    • What each walk catches:
      • The silence walk passes at edefd8193. It fails under the mutation above and under b2ccd1c84's ceiling * 2, both on since = 0.
      • The death walk fails at edefd8193 on died = 7, and passes with BLOCKER 1's patch.
      • Together they also catch the backstop moved ahead of the panic arm, which is what (e)2 guarded.

NOTE

  • Nightly 36753172688 at edefd8193: guest (1)–(12), tcg and host must be green before landing.
    • The guest shards are the only run of the cut ceilings on CI's runners, paid out through host_scale.
    • tcg runs process_stats's new 11 s ceiling under TCG, and its own bound drops to 54 min.
    • host is the new 42 min bound, and it runs the new loopback test on macos-latest.
    • build is already green. portability-linux and portability-macos are build-only, and nothing in this branch moves them.
  • tests/common/metal.rs:1040 — the implementer's named gap is not a defect.
    • run returns an exit-2 refusal after the reboot only after its own come-back wait (src/metal.rs:2090-2096), so the ToyOS boot the swap dials is already over.
    • The label's refused entry is the loop's either way: the booted insert at :1062 replaces the swap's at :1052.
    • Child::kill on a swap that has already exited keeps its real status.
  • tests/toyos-rust-tests/src/bin/lan_hold.rs:4 — its only user now is dump_nmi_probe's testcases-deaf (tests/toyos.rs:1726), which runs no netd.
    • Its name and its LEASE_BOUND_MS are netd's.
    • The hold should be named for what dump-deaf-cpu needs, and say so in its constant.
  • PR body "Runs" carries only b2ccd1c84's readings. The head that lands has its own: CI host run 36753095781, and the whole suite at EXIT=0 with 502 passed. They belong in main's record.

REMOVE

  • issues/diagnostics/the-cable-judge-reads-three-netd-records-that-cannot-arrive-on-the-t14.md — delete the whole file.
    • Its premise is false: netd's lease record reached the T14's stick at b2ccd1c84 ("this boot leased 192.168.1.49").
    • The boot and the keys it names, tests/lancase and lan.lancase.*, are deleted by this branch.

SEND BACK

…at the backstop

The backstop arm of `ceiling_verdict` now names a death it holds before it
calls the guest slow. A 5 s ceiling and a kernel death at 7 s read "timed out
after 20s, with the guest still talking 14s ago" at edefd81, because the
panic arm waits for `GUEST_QUIET` and the backstop came first; `main`'s 300 s
floor named that panic at 22 s. The comment's clause that a death "is named by
the arms above" was false for that case and goes.

`ceiling_self_check`'s case (e) is two walks under a 5 s ceiling, a second at a
time: a guest silent from any second in its budget must first read as
`STALLED`, and a kernel dead at any second up to four ceilings must first be
named. The old case (e) required only a backstop of 11 s or more, so a backstop
of `ceiling + 6 s` passed it. The walk steps the quiet seconds and adds them to
the start, where the review's form subtracted a `Duration`, which the host
gate's clippy denies (`unchecked_time_subtraction`); the elapsed/quiet pairs
are the same. Each arm was a checked patch, shown to build and reversed:
- death walk at edefd81's verdict code: EXIT=101, on `died = 7`.
- backstop `(ceiling * 2).max(ceiling + Duration::from_secs(6))`: EXIT=101,
  the silence walk on `since = 0` ("timed out after 11s").
- backstop `ceiling * 2`: EXIT=101, the silence walk on `since = 0` ("timed
  out after 10s").
- this commit: EXIT=0.

`lan_hold` is `deaf_cpu_hold`: its one runner is `dump_nmi_probe`'s
`testcases-deaf` boot, which runs no netd, so its hold is no longer netd's
lease bound but its own `PAST_THE_DEAF_WINDOW`, at the same 20 s.

`issues/diagnostics/the-cable-judge-reads-three-netd-records-that-cannot-arrive-on-the-t14.md`
goes: netd's lease record reached the T14's stick at b2ccd1c, and the boot and
keys it names are deleted by this branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu

Japabu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Round 3, head 1c81e3e66.

Round 2 BLOCKERs

  • tests/common/qemu.rs:563-570, a kernel death after the ceiling reported as a slow guest: CLOSED.
    • The backstop arm now names dying (:565-567).
    • The PR body's control: edefd81's verdict code fails the death walk on died = 7 ("timed out after 20s, with the guest still talking 14s ago", EXIT=101), and 1c81e3e passes it (EXIT=0).
    • Checked by hand under a 5 s ceiling, where the backstop is 20 s: 7 s is the first death that fails at edefd81. A death at 6 s reaches the panic arm at 21 s, before the backstop fires.
  • tests/checks/qemu.rs:201-221, case (e) unable to fail on the backstop claim: CLOSED.
    • The silence walk fails at since = 0 under ceiling + 6 s ("timed out after 11s") and under ceiling * 2 ("timed out after 10s"), each EXIT=101.
    • It passes at the head (EXIT=0).

Readiness

  • CI host at 1c81e3e66: run 36756594838 succeeded. It printed "Host: 55 step(s), all green", checks::serial_vocabulary ... ok and checks::a_stall_stays_red ... ok.
  • T14 at 1c81e3e66, --metal dump_nmi_probe: "[metal] 1 passed, 0 failed, 1 boot(s)".
    • test_rs_deaf_cpu_hold was spawned at 1.200 s. The dump ran at 11.251 s, cpu7 rejoined at 11.651 s, and the hold exited 0 at 21.201 s.
    • Source: the orchestrator's 638r3-T14-dump_nmi_probe.log. It is not on the PR.
  • The whole guest suite last ran at edefd8193: 502 passed. The only change since then that a guest sees is the hold's new name, and rg lan_hold finds no reference left.

Nightly 36753172688 at edefd8193: neither red is this branch's.

  • guest (10), screen_loader_lines: "the panel carried 26 rows at the GOP query and 41 at the loader's last line, a growth of 15 …". Main's scheduled nightly 36696295750, guest (8) (job 109873180506), failed with the same words, and so did wt/toyos-castore's 36736950844, guest (10).
  • guest (7), blockd_survives_its_death: "role crash exited Some(1)". The guest's own blockd_io: FAIL /AFTER.BIN after the restart: Io came at 23.909 s, 2.7 s into the role.
    • No bound this branch cut shortened anything it waits on. The crash role's host ceiling is 251 s × host_scale (was 600) and was never approached. The boot passes no kernel parameter (boot(…, "blockd-death", &[])), so boot-deadline=8000/12000 and STAGED_BOUND_MS never reach it. blockd_io's crash path waits with no deadline.
    • The branch changes no line of blockd, fsd, the kernel or the crash role.
    • Its one change to that boot is BENCH_BLOCKS 8192 → 2048. That shrinks the bench partition, which lies after the FS partition. In the failing log the crash role's B2D4F6A8… at block 35072, 16384 blocks sits where it does on main.
    • The test passed elsewhere: at b2ccd1c84 (68 s) and at edefd8193 (61 s) in the orchestrator's whole runs, and on castore's nightly (58 s).

Growth

  • Round 3: +36 / −66. Tests are +36 / −28 and issues −38.
  • Branch: +540 / −998.
    • Shipping code (netd, toyos-i219, toyos-tco): +20 / −126.
    • src/: +145 / −76.
    • Tests: +270 / −665.
    • Issues: +103 / −129.

BLOCKER

None.

NOTE

  • tests/checks/qemu.rs:205-209 — the walks step whole seconds, so a backstop one second short passes both.
    • Mutation: let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET - Duration::from_secs(1));. At since = 5 the walk reaches 19 s at quiet 14, and 19 is not past 19. At quiet 15 the STALLED arm comes first.
    • The read loop polls every 100 ms (qemu.rs:3357). Under that mutation, a guest silent from any point in (4 s, 5 s] would read "timed out after 19s".
    • Patch: step since, died and quiet by Duration::from_millis(100). The silence walk then fails under the mutation from since = 4.2 s.
  • tests/toyos-rust-tests/src/bin/deaf_cpu_hold.rs:7 — the boot is held 20 s for a dump that finished at 11.651 s on the T14. The judge reads the dump's own record, so this is not a flat wait on a verdict, but every full run pays about 9.5 s past the event.
  • PR — the T14 reading at 1c81e3e66 and CI host run 36756594838 belong to the head that lands. Neither is in the body or in a comment.
  • issues/filesystem/blockd-survives-its-death-reds-on-a-replacement-that-reorders-acknowledged-writes.md — it records the trace-order verdict, not "/AFTER.BIN after the restart: Io". That signature is recorded nowhere. Recording it falls to main, not to this branch.
  • origin/main has moved past the merge base 6e1f1661a (issues: child processes — end events, a parent takes its children down, sessions and quit, with the owner's fourteen questions #617, Stage 6 step 1: a device's handler posts its watch; irq_ring's UserDev and Audio go #634, hda_tone: the judge refuses multi-stream sessions, and a dropped cpal stream waits a bounded time on soundd and then lets go #641). merge-tree is clean.
    • main's tests/toyos.rs:12182 (drain_until(Duration::from_secs(30), handler_post_said)) and the new shared member test_rs_cpal_drop_unreleased were written against the width-scaled budget.
    • Neither has a measurement under this branch's rule. No run has measured the merged head.

REMOVE

  • PR body, "Unsure": "The nightly shards have run the whole suite since Delete the test tiers: a run is the whole suite or its filter #625, but none has run under these ceilings yet." — false: 36753172688 ran all twelve guest shards at edefd8193.
  • PR body, "Unsure": "No T14 run has booted deaf_cpu_hold under its new name yet." — false: the orchestrator's run at 1c81e3e66 booted it.
  • tests/checks/qemu.rs:201-203 — "from any point in its budget" and "at any point" claim more than a whole-second walk checks. The rest restates the loops below it.

LAND AFTER NAMED CHANGES

@Japabu

Japabu commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Orchestrator runs at 1c81e3e:

  • T14 cargo test --test toyos-build -- --metal dump_nmi_probe: EXIT=0 — [metal] 1 passed, 0 failed, 1 boot(s) (log orch/logs/638r3-T14-dump_nmi_probe.log).
  • CI host run 36756594838: success.

Japabu and others added 3 commits September 30, 2026 22:52
Clean merge. Main's one new timed wait, `handler_post_without_a_pass`'s
`drain_until(30 s)`, and its new shared member `cpal_drop_unreleased` were
written against the width-scaled budget; the next commit brings them under
this branch's ceiling rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…ll, and main's new wait is under the rule

- `ceiling_self_check`'s two walks step `since`, `died` and `quiet` by
  100 ms, the read loop's poll of a silent guest, instead of by whole
  seconds. At whole seconds a backstop one second short,
  `(ceiling * 2).max(ceiling + GUEST_QUIET - 1 s)`, passed both walks: at
  `since = 5` the walk reached 19 s at quiet 14, and 19 is not past 19. It
  now fails at `since = 4.2 s` ("timed out after 19s"), EXIT=101.
- The walks' comment goes: "any point" claimed more than the walk checks,
  and the rest restated the loops.
- `handler_post_without_a_pass`'s drain, from main's #634, was 30 s under
  the width-scaled budget. The test took 3-4 s in #634's whole runs
  (4 s in 634r6), so it is 12 s.
- `cpal_drop_unreleased`, #641's new shared member, took 111 and 138 ms in
  #641's whole runs, so the 5 s shared default covers it and it needs no
  row.
- `deaf_cpu_hold` keeps its fixed 20 s. A binary test-runner spawns holds
  no `logread`: test-runner hands down its namespace, and `logread` is a
  `SysCap` duplicate, not a namespace entry. logd serves no reader port on
  `tests/testcases`. The only record the hold could read is logd's file
  under `/log`, which gives no notification, so waiting on it would be a
  sleep-and-reread poll.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
The orchestrator's whole run at 09d74fc (638r4) went red on
toolkit_winit_loop "never finished" after 57 s. This branch had replaced
main's Liveness(40 s quiet, 240 s total) with a hard
`Instant::now() + qemu.budget(35 s)`, and a hard cut does not ask whether the
guest is still talking.

`qemu::Ceiling` reads a capture its caller keeps and appends to. It takes the
guest's silence from when the capture last grew, and a kernel death from the
first one appended. The decision is ceiling_verdict's, so every one of these
waits ends the way run_test_paced's does:
- a guest still talking past its budget waits for the backstop and reds as
  TIMED_OUT
- one silent for GUEST_QUIET past its budget reds as STALLED
- a dead kernel is named, with its report

`backstop()` is ceiling_verdict's formula, pulled out so a wait with no
console can sit behind another wait's ceiling.

The waits this branch cut from main's far bound to a tight hard one:
- toolkit_winit_loop, 35 s: Liveness(40, 240) on main.
- toolkit_winit_pace, 53 s: a Liveness whose total was cut from 120 s. That
  total is a hard cut however loud the guest is, and it was not host-scaled.
  A wait it ended also went on to count presents, so a slow guest read as a
  wrong count. It is now host-scaled, and the ceiling's verdict is the red.
- lan_no_lease, 99 s: drain_until(GUEST_WEDGED) on main. The guest is silent
  until netd gives up, and the ceiling allows silence inside the budget.
- kernel_log_file's device poll, 54 s, and usb_flush_optional's, 99 s:
  budget(GUEST_WEDGED) on main. Each now drains the console between polls
  instead of sleeping 50 ms, so the ceiling can hear the guest. The lines it
  reads still reach the shutdown tail's panic check.
- guest_dies_with_its_harness, 24 s: GUEST_WEDGED on main. The wait reads
  the owner's stdout, which carries nothing before HELD. The owner's own boot
  ceiling (BOOT_CEILING, host scale 1 in a fresh process) names a boot that
  never came up, and 24 s could cut ahead of that 30 s. The wait now sits at
  that ceiling's backstop.

Not converted: every wait that was already a hard deadline on main and only
had its number moved. Among them are the boot ceiling, toolkit_iced's three
waits, metal_sim_pointer_churn's three, handler_post_without_a_pass's drain,
power's WAIT and CHAIN_WAIT, logstream's FLOOD_CEILING, and the
screendump_until/drain_until literals.

What the 638r4 capture shows is not a slow guest. The terminal logged stage 1
at 6.534 s and nothing else from the app. Meanwhile the app opened 23 windows,
the last being stage 6's at 9.448 s, and every one of those windows is created
after a stage line it had printed. CLOSE-ME never arrived. From 11.4 s to
42.8 s every sched line reads ready=0 on every CPU. Under this change that
guest still reds, at the backstop as TIMED_OUT instead of at 35 s. That is
filed as issues/build/winit-loops-output-stopped-reaching-the-log-after-stage-1.md.

metal_sim_client_death's red in the same run was not a cut. Its run_test
returned on ===TEST_END with exit 0 after 153 ms. The missing line is its
grandchild's, printed after the root had ended. That is filed as
issues/build/client-death-ends-before-its-reaped-creators-request-is-served.md.

capture_ceiling_self_check stages Ceiling on instants ahead of now:
- a capture growing past its ceiling is ended only by the backstop
- one that stopped growing is ended by the stall
- a kernel death appended after the first read is named, with its report

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Japabu added a commit that referenced this pull request Oct 1, 2026
… the detector on every boot, and the watchdog class refused by right

Designed against #638 (23 boots a full T14 run, lists sized at
`JOB_BOUND_MS` less a tenth over 860/800/260 ms allowances) as if landed.

The session track:
- The fold is six boots of #638's run: shared, shared-2, ccorpus,
  testcases, testcases-mkdir, testcases-readdir. Each `Boot parameter:`
  line in toyos-tight/target/metal (b2ccd1c) carries only root=,
  boot-deadline=120000, boot-slot=A and blackbox=. 220 members there:
  62+18+130+8+1+1 `===TEST_END` lines.
- Prices: #638's 860 ms a shipping member and 260 ms a corpus case;
  the other three lists at twice their slowest span over ten full T14
  runs (536-metal-full, 536r21, 590r6-m1, 590r6-m2, 616, 616r3,
  630-full1, 630-full2, 630r4, 638-metal-full): testcases 11.861 s
  (590r6-m1), mkdir 2.106 s (536r21), readdir_bound 5.043 s (638, the
  only run without its /home arm).
- Per-job bound 14.1 s: twice mutual_kill's 7.041 s (630-full2), the
  slowest member of those runs once readdir_bound's /home arm (24.8 s on
  the two 536 runs) is set aside.
- Budget: 120000 - 12000 - 19100 - 14100 = 74800 ms. 80 x 860 = 68800;
  130 x 260 + 23722 + 4212 + 10086 = 71820. Two sessions for six boots:
  23 -> 19.
- The fence is every root `/` has. Of the 89 folded binaries, 14 name
  /home and hierarchy_paths writes /apps; the file-left-behind control is
  staged in /home.
- A window's copy of the kernel's records is the exec-started runner's
  only: on QEMU's serial runner a copy would double every record on the
  console, where 4 must_be_clean_apart_from sites and 55
  `.matches(..).count()` sites in 10 files count lines.
- device_claim_lifetime, the one folded member that mints a claim, runs
  last until the deferred-release defect closes.
- The QEMU e1000e session arm is restored, and the last stage waits on
  the detector on every boot. It answers the swap-bound issue.

The watchdog track:
- The owner ruled the hard-lockup detector armed on every boot, so its
  question file is deleted and the arm is the track's next step. All 23
  boots of #638's run say `hard lockup:` (60000 ms, 5000 ms on the three
  staged boots); only hardlockup's loader.log carries the lockup record.
- Refusal by right: a new Rights::WATCHDOG that no syscap name grants,
  demanded by sys_device_claim for the class. Declared as an ABI change.
- The kernel feeds while no claim is held, so watchdogd's exit or swap
  hands the timer back; watchdogd waits on the deferred-release defect.
- QEMU: q35's TCO counts QEMU_CLOCK_VIRTUAL and its second expiry calls
  watchdog_perform_action (QEMU v11.1.1 hw/acpi/ich9_tco.c:61-70,244);
  tests/qtest/tco-test.c:325-347 tests the `none` action. Every guest
  that stages no reset runs with -action watchdog=none.
- The T14 reading moves to the loader's report pass right after the
  reset: Linux v6.12 drivers/watchdog/iTCO_wdt.c:545-560 clears
  SECOND_TO_STS in iTCO_wdt_probe. The starved boot's deadline must
  outlast tco-starve's 5 s plus the 9.6 s bound, which #638's 10 s
  STAGED_BOUND_MS does not.
- The go-ahead is an owner question file of its own.

The loader track: stage 6's exit asks a whole run, every --once boot
included, and a loader sent over ssh to boot once; a stick booting
neither loader costs a hand and diag/flash.sh; stage 7 closes the four
Ubuntu-loop issues; the 53 MiB ROOT clause goes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Japabu and others added 3 commits October 3, 2026 00:04
…slowest run

Five whole runs of `cargo test --test toyos-build` at bf28c1e, 12 wide
on the dev host at load average 40 to 59 on 14 cores (guest-m2..m6 in the
orchestrator's scratchpad): 26, 25, 26, 26 and 26 passed, in 229.2, 66.8,
37.1, 32.9 and 34.1 s. The slowest pass of each test, in seconds:

  iommu_virtio_platform 47         virt_first_entry 14
  machine_shutdown 20              virt_fp_isolation 14
  nested_nmi_is_loud 21            virt_irq_storm 12
  screen_fatal_behind_a_painter 29 virt_mask_windows 8
  screen_fatal_halt_composited 32  virt_off_names_the_cpus_left_on 16
  screen_panic_muted 19            virt_readonly_copyout 18
  virt_debug_refused 18            virt_reboot 8
  virt_early_fault 2               virt_reboot_refused_without_psci 8
  virt_early_panic 2               virt_smp 10
  virt_el1_smp 10                  virt_timer_floor 9
  virt_el2_drop 7                  virt_timer_preempts 29
  virt_failed_ap_leaves_no_hole 15 virt_unmap_touch 18
  virt_fatal_halts_the_others_first 34   virt_user_mode 10

The numbers, each three times the slowest test that waits through it:
- BOOT_CEILING 63: nested_nmi_is_loud, the slowest test that is one boot
  and nothing after it (21). main's was 10 s times the width, 120 s.
- GUEST_WEDGED 141, now paid out through `budget`: every wait through
  `await_guest`, the slowest of which is iommu_virtio_platform (47).
  `guest_liveness` had one caller and goes.
- virt_selftest's drain 36, was 180: virt_irq_storm (12) and
  virt_timer_floor (9).
- virt_user_mode's drain 30, was 180 (10).
- virt_early_panic's and virt_early_fault's two waits 6 each, were 30
  and 10 (2 each).
- screen_fatal_behind_a_painter's second wait 87, was 20, and
  screen_fatal_halt_composited's first 96, was 30: each sat under its
  own test's slowest run (29 and 32) and passed on main only because
  the width multiplied it twelvefold.
- Kept, at or above their test's slowest run and under three times it:
  screen_panic_muted 30 (19), virt_el2_drop 10 (7), the painter's
  first wait 30 (29), the composited test's second 40 (32).

REFERENCE_BOOT_MS is 1424, the largest fastest boot of the five runs
(1419, 1424, 1048, 1055, 750 ms), so the measuring host pays every
ceiling at 1x.

The red of the second run is virt_mask_windows, on an assertion and not
a ceiling: nine census lines and eight windows lines for cpu7. The
test, the kernel and its judge are main's. Filed as
issues/build/virt-mask-windows-read-nine-censuses-and-eight-windows-lines-for-one-cpu.md.

The first run after the merge exited 1 with 18 reds before any of
these: twelve threads moved the fork checkout to the new pin at once.
Filed as
issues/build/a-suites-first-run-after-the-fork-pin-moves-races-its-own-checkout.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Clean. #651 deletes the retired ABI numbers and adds one shared member,
`syscall_unassigned`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
@Japabu Japabu changed the title Every QEMU ceiling is at most three times what its test takes; wedges end in seconds; the LAN boots ride the talking boot Every QEMU ceiling is at most three times what its test takes; wedges end in seconds; two LAN boots ride the talking boot Oct 2, 2026
@Japabu

Japabu commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Mutation patches for the negative controls at 4859e8d71, each applied with git apply, built, run and reversed (run.sh, wedge.sh). The wedge's third arm reverts the whole change with git diff --binary HEAD origin/main (a93067fc2) before applying wedge-irq-storm.patch.

wedge-irq-storm.patch

--- a/kernel/src/arch/aarch64/trap.rs
+++ b/kernel/src/arch/aarch64/trap.rs
@@ -525,7 +525,7 @@
 
     pub(super) fn tick() {
         if RUNNING.load(Relaxed) {
-            TICKS.fetch_add(1, Relaxed);
+            TICKS.fetch_add(0, Relaxed);
         }
     }
 

hand-back-no-wait.patch

--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -977,7 +977,7 @@
 /// first line, so the backlog crosses at the network's pace and not the
 /// conversation's.
 fn hand_back<T>(stream: &Stream, began: Instant, by: Duration, fire: impl FnOnce() -> T) -> T {
-    match stream.wait_for_boot(by) {
+    match stream.wait_for_boot(Duration::ZERO.min(by)) {
         Some(ms) => println!(
             "  talk: the stream carried `Boot: complete` ({ms} ms), {} ms after it opened",
             began.elapsed().as_millis()

wait-for-boot-at-once.patch

--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -286,7 +286,7 @@
     /// This boot's `Boot: complete` in milliseconds, read as [`judge`] reads
     /// it, once a line carries it, or `None` after `by`.
     pub fn wait_for_boot(&self, by: Duration) -> Option<u64> {
-        self.wait_until(by, |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
+        self.wait_until(by.min(Duration::ZERO), |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
     }
 
     /// The peer, once a connection has carried a line, or `None` after `by`

said-refusal-first-line.patch

--- a/src/metal.rs
+++ b/src/metal.rs
@@ -110,7 +110,7 @@
 pub fn said_refusal(stderr: &str) -> Option<String> {
     let lines: Vec<&str> = stderr.lines().collect();
     let at = lines.iter().rposition(|line| line.starts_with(REFUSAL_HEAD))?;
-    Some(lines[at..].join("\n")[REFUSAL_HEAD.len()..].trim_end().to_string())
+    Some(lines[at..=at].join("\n")[REFUSAL_HEAD.len()..].trim_end().to_string())
 }
 
 /// Every way this loop refuses, by name.

wait-on-port-alone.patch

--- a/src/metal.rs
+++ b/src/metal.rs
@@ -1509,7 +1509,7 @@
     while began.elapsed().as_secs() < secs {
         let next = std::time::Instant::now() + POLL;
         let listening = accepts();
-        let answered = listening && (!answering || answers());
+        let answered = listening;
         if answered == answering {
             return Ok(began.elapsed().as_secs());
         }

bound-for-always-wedge.patch

--- a/src/metal.rs
+++ b/src/metal.rs
@@ -852,7 +852,7 @@
 
 /// The `boot-deadline=` bound an image armed with `armed` carries.
 pub fn bound_for(armed: &[impl AsRef<str>]) -> u64 {
-    if stages_a_wedge(armed) {
+    if stages_a_wedge(armed) && false {
         toyos_tco::STAGED_BOUND_MS
     } else {
         toyos_tco::WEDGE_BOUND_MS

backstop-one-second-short.patch

--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET - Duration::from_secs(1));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));

backstop-plus-six.patch

--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = (ceiling * 2).max(ceiling + Duration::from_secs(6));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));

backstop-twice-only.patch

--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = ceiling * 2;
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));

backstop-mains.patch

--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = ceiling.max(Duration::from_secs(300));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));

backstop-no-death.patch

--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -519,7 +519,7 @@
     // [`GUEST_QUIET`].
     let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
     if elapsed > backstop {
-        if let Some(line) = dying {
+        if let Some(line) = dying.filter(|_| false) {
             return Some(kernel_died_here(line));
         }
         return Some(format!(

@Japabu

Japabu commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

T14 at 4859e8d71, run by the orchestrator from 638-recut/metal/request.txt: three boots, each image's sha256 checked against the request before it was flashed, each toyos-metal exit 0; then the two judges from the clean worktree at the head. No row armed the TCO watchdog, a staged wedge or the hard-lockup probe.

boot image sha256 toyos-metal
lantalkcase (--nic 0000:00:1f.6 --talk …) c01bdbdda51f554281b3729a9519689e020d57094720699a25a5634c67682814 exit 0; the loop printed talk: the stream carried Boot: complete (1153 ms), 822 ms after it opened
lanleasecase (--nic 0000:00:1f.6) 5efcceb3e0455de666fe68617cf8f6b4b72fbb4a88069b926ea57acf101970a7 exit 0
testcases-readdir b51844becf2b8f8704067c1afde7dc4969082dc0f1f00092ba8e9cc7a61e62aa exit 0
  • lan_: exit 1, by main's known red alone, as the request predicted. PASS lan_talk (415 lines arrived over the cable, Boot: complete among them, in the stick's order; echo answered byte for byte with status 0, 73 ms after the stream opened; reboot taken), PASS lan_message_delivery (slot 0 handed over on vector 0x28 at 1.193 s, its first message at 8.485 s), PASS lan_lease_report (netd exit 83, leased 192.168.1.48/24, 5 sent and 14 received). FAIL lan_dhcp_lease with two findings, both the one issues/hardware/the-benchs-router-leases-toyos-another-address-than-ubuntu.md records: this boot leased 192.168.1.48 and the host pinged 192.168.1.46, and that ping was answered on the far side of the reset. 3 passed, 1 failed, 2 boot(s).
  • readdir_bound: exit 0. PASS readdir_bound; 1 passed, 0 failed, 1 boot(s).

Not run, and owed: the five rows the request left out for the owner's OK (boot_deadline_ends_a_wedge, hard_lockup_ends_a_deaf_cpu, usb_reset_records_the_phase_it_cut, loader_watchdog_arms, watchdog_fed), and with them the unfiltered run that judges the shared boots at their new sizes.

Readbacks and judge logs: /Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/638-recut/ (metal/…, judge-lan.log, judge-readdir.log).

@Japabu

Japabu commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Round 4, head 4859e8d71 (the re-cut), reviewed as a first review of this head.

Earlier BLOCKERs: none open. Round 3 closed round 2's two at 1c81e3e66. Round 3's NOTE is closed too: the walks at tests/checks/qemu.rs:199-217 now step 100 ms, and a backstop one second short reds serial_vocabulary (EXIT=101, mutations/run.log).

Evidence at 4859e8d71

  • cargo run -- --ci host: EXIT=0, "Host: 67 step(s), all green" (host2.log, headed 4859e8d71).
  • cargo test --test toyos-build: EXIT=0, 26 passed in 146.9 s (guest-head2.log, headed 4859e8d71).
  • cargo test --test toyos-checks: EXIT=0, 31 passed.
  • Ten negative-control patches: each EXIT=101 on the test the body names.
  • The wedge: 37 s on the head alone, 47.6 s in the whole suite, 181 s on main's tree.
  • T14 (comment 5962753432): lan_talk, lan_message_delivery, lan_lease_report and readdir_bound PASS.
  • lan_dhcp_lease fails only on the router-lease red. Its first finding is the issue's words. Its second follows from it on a boot the host ends: .46 answered only after the reset.
  • CI: the PR is a draft, so every job was skipped.

Growth: +648 / −677 over 36 files, net −29.

  • Shipping (netd, toyos-i219 lib/regs, toyos-tco): +18 / −63.
  • toyos-i219 model and tests: +2 / −57.
  • src/ production: +136 / −20. src/ tests: +107 / −58.
  • tests/: +160 / −351.
  • issues/: +225 / −128.
  • I accept the production growth for hand_back: the T14 carried Boot: complete 822 ms after the stream opened, and lan_talk passed. I accept it for bound_for and for the port poll. I do not accept it for the refusal-wait (NOTE).

Rulings on the brief

  • Ceilings: the claim holds.
    • Each number checks against m2–m6: 21→63, 47→141, 12→36, 10→30, 2→6, 29→87 and 32→96.
    • The kept waits against their test's slowest run: 30 over 19, 10 over 7, 30 over 29, 40 over 32.
    • None became a duration verdict. An expiry only ends a wait whose red is a thing the guest never said (never reported, never paged, never took the screen).
    • Every kept or raised wait sits at or above its test's slowest whole run.
    • The measuring commit bf28c1e38 already ran without the width.
  • lanleasecase stays: accepted. Folding it deletes --exit-with-lease, toyos_i219::lease (262 lines) and phy::Outcome with their tests. That is a high-risk driver change, and the renamed issue carries an owner and an exit.
  • Tracks: the T14 runs its tests in sessions of a booted ToyOS, and every machine ships a watchdog and a hard-lockup detector #631's three citations: two hold.
    • toyos_tco::STAGED_BOUND_MS (toyos-tco/src/lib.rs:149) and the hard-lockup issue exist.
    • The allowances do not. The tree holds only their quotients, 62/67/207, as literals; 860/800/260 ms exist in this body alone.
  • Off-task filings.
    • Both are filed in the present tense, with their evidence.
    • virt_mask_windows is a defect with an owner by role.
    • The fork race is tooling, which README's "the build system" allows. It names no owner.

The T14 readings this head owes

None of these ran at 4859e8d71.

  1. Changed by this diff, and reset rows, so the owner's call:
    • The rows: boot_deadline_ends_a_wedge (deadlinewedge), hard_lockup_ends_a_deaf_cpu (hardlockup), usb_reset_records_the_phase_it_cut (usbload).
    • Their boot-deadline= goes from 120000 to 10000, and the lockup bound from 60 s to 5 s.
    • They were last read at b2ccd1c84, before main's kernel merges.
    • Each is one --metal <name> run of one boot; no shared member matches those names.
  2. Changed by this diff, with no TCO and no staged wedge:
    • The boots: shared (62 of today's 81 shipping members), shared-2 (19), and ccorpus (every C case, up to 160, in one boot where main boots two).
    • shared at 62 has never booted on today's list. ccorpus in one chunk was read only at b2ccd1c84, under that day's kernel and libc.
    • The harness judges a member only in a run whose filter keeps it. So only an unfiltered --metal run judges these, and that run also boots everything in 1, plus loader_watchdog_arms and watchdog_fed.
  3. Not owed by this diff:
    • loader_watchdog_arms and watchdog_fed: their image keeps 120000, as src/metal.rs:2851 asserts.
    • shared-debug: its 5 members make one boot at 18 or at 67.
    • Every other row: its image is unchanged, and the three boots at this head crossed the new wait_on both ways.
  4. The owner's other route:
    • Reverting 62/67/207 to 38/18/90 removes 2.
    • Dropping STAGED_BOUND_MS and bound_for removes 1.
  5. Whichever run closes 1 or 2 runs in Drive mode (--metal without --metal-readback). The three boots at this head were flashed by hand from request.txt, so drive (tests/common/metal.rs:805) has met no machine.

BLOCKER

  • toyos-tco/src/lib.rs:149, src/metal.rs:854, tests/common/metal.rs:743 — the three staged-wedge images now end at 10 s with lockup at 5 s, and no T14 reading at this head judges them (T14 list, 1). This is a hardware claim QEMU cannot reach: hard_lockup_ends_a_deaf_cpu holds only while staging, plus 5 s, plus one sample still comes before the 10 s deadline on today's kernel. It closes with those three rows green on the T14 at 4859e8d71, or with the bound dropped.
  • tests/toyos.rs:763, :853 — the T14's shared and ccorpus now run 62 and up to 160 members under the runner's 60 s JOB_BOUND_MS. The counts are priced from per-member means of lists The guest suite keeps the 21 tests only a booted machine answers; the rest are metal, host or tracked #660 has since replaced, and shared at 62 has never booted on today's list (T14 list, 2). It closes with shared, shared-2 and ccorpus green on the T14 at 4859e8d71, or with main's counts kept.

NOTE

  • src/metal.rs:1512, :1523-1529 — "go down" is now one failed 1 s connect. A single lost SYN while Ubuntu is still up reads as down, then "come back" answers at once, and the loop reads a stick ToyOS never booted. That is exactly what ride_the_reboot's doc (:1362-1364) exists to prevent. Main's go-down was an ssh with ConnectTimeout=10; dial with CONNECT_SECS when going down.
  • src/metal.rs:2025-2034 — the refusal-wait has no measured case. The body cites issues/hardware/the-t14-stopped-answering-ssh-between-two-lan-boots.md, but that file's recorded timeout came after lanicscase passed, a path this wait never takes. Name the refusal it is for (the reboot job's ssh dying with the machine) with a reading, or delete it along with after_the_reboot's split.
  • tests/toyos.rs:763, :782, :853 — Tracks: the T14 runs its tests in sessions of a booted ToyOS, and every machine ships a watchdog and a hard-lockup detector #631 prices sessions from allowances that exist only in this body. Declare them beside JOB_BOUND_MS and derive members from them, so the counts follow the bound and Tracks: the T14 runs its tests in sessions of a booted ToyOS, and every machine ships a watchdog and a hard-lockup detector #631 has a name to cite. Separately, shared-debug's 18→67 moves nothing on a 5-member list.
  • issues/build/a-suites-first-run-after-the-fork-pin-moves-races-its-own-checkout.md:29 — "Whoever next changes sysroot::fork_checkout" names nobody. The race reds every agent's first run after a pin moves (18 of 26 here), so it needs a holder now.
  • issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md:22 — "Whoever next changes toyos-tco" is this branch (toyos-tco/src/lib.rs:144-149), which leaves the issue open. Name Tracks: the T14 runs its tests in sessions of a booted ToyOS, and every machine ships a watchdog and a hard-lockup detector #631's track, which plans the reader.
  • issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md — no owner.
  • issues/build/virt-mask-windows-read-nine-censuses-and-eight-windows-lines-for-one-cpu.md — its subject is the kernel's windows instrument, and its owner's track is in issues/kernel/; move it there.
  • tests/checks/qemu.rs:201 — STEP restates the read loop's 100 ms poll (tests/common/qemu.rs:1571). Make it one const read by both, or the walk goes blind once the poll moves.
  • tests/common/qemu.rs:493-532, :1474-1528 — on today's suite, ceiling_verdict, run_test_paced and run_test_hooked have one caller: --debug's run (tests/toyos.rs:3418). The branch reworks a backstop and its walks for that one path. File their deletion.
  • tests/common/qemu.rs:172-198, :276-278 — the run's fastest boot is a virt_early_* boot whose ready marker EARLY PANIC: comes before userland (head2: 502 ms; virt_early_fault's whole test took 731 ms). So host_scale pays 1× on nearly any host, and with the width gone it is the only host correction budget_smp's doc promises. File it.
  • PR body — "No T14 run has measured this head: not the fold, not hand_back, not the staged bound" is false. The body also carries none of the readings at this head from comment 5962753432. Keep the record true.

REMOVE

  • src/metaltalk.rs:55 — "the stream stalls only while netd retransmits what its NIC dropped".
  • src/metaltalk.rs:1475-1477 — "The lines are the T14's: … Boot: complete was 72 lines further on."
  • PR body — the opening paragraph, "Re-cut onto main at a93067fc2. …".
  • PR body — the section "Dropped in the re-cut".
  • PR body — ", and the branch had deleted it".
  • PR body — "metal: the LAN lease rides the talking boot, and two T14 flashes go #609's Driver and its reader thread existed for the swap beside it; main deleted the swap, so it is one function."

SEND BACK

…the shared boots' allowances are named

Round 4's review of 4859e8d (#638):

- metal: `wait_on` dials the port within CONNECT_SECS when it watches
  the machine go down, what `ssh`'s own ConnectTimeout gives, and within
  one POLL coming back. A single lost SYN while Ubuntu is still up no
  longer reads as down. `coming_back_is_ssh_answering_and_not_its_port`
  records the dial it is handed going down.
- metal: the wait for `ssh` after a refusal that followed `reboot`, and
  `after_the_reboot`'s split, are deleted. No recorded run has the case
  it was for: the T14 run the body cited timed out on the boot after a
  pass, and the post-reboot refusals in the kept T14 logs are a machine
  that never came back, a stick, a ping and a mount, each with the
  machine already answering or already waited out. `run` is main's
  again past `bootnext`.
- toyos-tco: `RUST_MEMBER_MS` (860) and `C_MEMBER_MS` (260) beside
  `JOB_BOUND_MS`; `metal::members_fitting` derives `shared`'s and
  `ccorpus`'s member counts from them, 62 and 207 as before.
  `shared-debug` keeps main's 18: on a five-member list the count moves
  nothing.
- qemu: `VERDICT_POLL` is the read loop's poll, and the ceiling walk in
  `tests/checks/qemu.rs` steps by it.
- metaltalk: two comments deleted.
- issues: owners for the fork-pin race (the tooling track's toolchain
  stage), the hard-lockup bound's reader (the T14 unattended track,
  which holds the detector) and the I219 transmit burst (the LAN
  track's stage 2); the `virt_mask_windows` red moves to
  `issues/kernel/`; filed: `run_test` and `ceiling_verdict` serve only
  `--debug`'s `run`, and `host_scale` reads host speed off boots that
  stop at different markers (491 ms for `virt_early_panic` alone, 3318
  ms for `machine_shutdown` alone, same host, same minute).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Negative controls at 4b009e712. Each patch is checked with git apply --check, applied, built (cargo test <target> --no-run, EXIT=0 each), run, and reversed, with the tree clean after each (mutations/run.sh and run.log under /Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/638-r5/). The wedge patch is run by mutations/wedge.sh.

patch command test it reds exit
hand-back-no-wait cargo test --lib -- the_hand_back_waits_for_the_boot_record_the_judge_reads that test 101
wait-for-boot-at-once the same that test 101
said-refusal-first-line cargo test --lib -- a_refusal_is_read_back_off_the_drivers_stderr that test 101
wait-on-port-alone cargo test --lib -- coming_back_is_ssh_answering_and_not_its_port that test 101
go-down-dial-one-poll (new) the same that test: "going down dialled within [1s], where ssh's own connect is given 10s" 101
bound-for-always-wedge cargo test --lib -- every_arm_that_stops_this_machine_is_cleared_and_judged_as_one that test 101
backstop-one-second-short cargo test --test toyos-checks serial_vocabulary: "silent from 4.2s under a 5s ceiling" 101
backstop-plus-six the same serial_vocabulary: "silent from 0ns" 101
backstop-twice-only the same serial_vocabulary: "silent from 0ns" 101
backstop-mains the same a_stall_stays_red 101
backstop-no-death the same serial_vocabulary: "a kernel death at 5.2s" 101
wedge-irq-storm cargo test --test toyos-build -- virt_irq_storm; then cargo test --test toyos-build virt_irq_storm: "irq-storm never reported", at 37 s alone and 38 s in a 45.7 s suite 1; 1
hand-back-no-wait.patch
--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -977,7 +977,7 @@
 /// first line, so the backlog crosses at the network's pace and not the
 /// conversation's.
 fn hand_back<T>(stream: &Stream, began: Instant, by: Duration, fire: impl FnOnce() -> T) -> T {
-    match stream.wait_for_boot(by) {
+    match stream.wait_for_boot(Duration::ZERO.min(by)) {
         Some(ms) => println!(
             "  talk: the stream carried `Boot: complete` ({ms} ms), {} ms after it opened",
             began.elapsed().as_millis()
wait-for-boot-at-once.patch
--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -286,7 +286,7 @@
     /// This boot's `Boot: complete` in milliseconds, read as [`judge`] reads
     /// it, once a line carries it, or `None` after `by`.
     pub fn wait_for_boot(&self, by: Duration) -> Option<u64> {
-        self.wait_until(by, |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
+        self.wait_until(by.min(Duration::ZERO), |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
     }
 
     /// The peer, once a connection has carried a line, or `None` after `by`
said-refusal-first-line.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -110,7 +110,7 @@
 pub fn said_refusal(stderr: &str) -> Option<String> {
     let lines: Vec<&str> = stderr.lines().collect();
     let at = lines.iter().rposition(|line| line.starts_with(REFUSAL_HEAD))?;
-    Some(lines[at..].join("\n")[REFUSAL_HEAD.len()..].trim_end().to_string())
+    Some(lines[at..=at].join("\n")[REFUSAL_HEAD.len()..].trim_end().to_string())
 }
 
 /// Every way this loop refuses, by name.
wait-on-port-alone.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -1513,7 +1513,7 @@
     while began.elapsed().as_secs() < secs {
         let next = std::time::Instant::now() + POLL;
         let listening = accepts(dial);
-        let answered = listening && (!answering || answers());
+        let answered = listening;
         if answered == answering {
             return Ok(began.elapsed().as_secs());
         }
go-down-dial-one-poll.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -1508,7 +1508,7 @@
     mut accepts: impl FnMut(std::time::Duration) -> bool,
     mut answers: impl FnMut() -> bool,
 ) -> Result<u64, Refusal> {
-    let dial = if answering { POLL } else { std::time::Duration::from_secs(CONNECT_SECS) };
+    let dial = POLL;
     let began = std::time::Instant::now();
     while began.elapsed().as_secs() < secs {
         let next = std::time::Instant::now() + POLL;
bound-for-always-wedge.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -852,7 +852,7 @@
 
 /// The `boot-deadline=` bound an image armed with `armed` carries.
 pub fn bound_for(armed: &[impl AsRef<str>]) -> u64 {
-    if stages_a_wedge(armed) {
+    if stages_a_wedge(armed) && false {
         toyos_tco::STAGED_BOUND_MS
     } else {
         toyos_tco::WEDGE_BOUND_MS
backstop-one-second-short.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET - Duration::from_secs(1));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-plus-six.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = (ceiling * 2).max(ceiling + Duration::from_secs(6));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-twice-only.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = ceiling * 2;
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-mains.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = ceiling.max(Duration::from_secs(300));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-no-death.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -519,7 +519,7 @@
     // [`GUEST_QUIET`].
     let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
     if elapsed > backstop {
-        if let Some(line) = dying {
+        if let Some(line) = dying.filter(|_| false) {
             return Some(kernel_died_here(line));
         }
         return Some(format!(
wedge-irq-storm.patch
--- a/kernel/src/arch/aarch64/trap.rs
+++ b/kernel/src/arch/aarch64/trap.rs
@@ -525,7 +525,7 @@
 
     pub(super) fn tick() {
         if RUNNING.load(Relaxed) {
-            TICKS.fetch_add(1, Relaxed);
+            TICKS.fetch_add(0, Relaxed);
         }
     }
 

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Negative controls at e5daa61ff (4b009e712 with origin/main at 322085f23, #648, merged in). The patches are the ones in comment 5963363857, unchanged; each is checked with git apply --check, applied, built (cargo test <target> --no-run, EXIT=0 each), run, and reversed, with the tree clean after each (at-merge/mutations/run.sh and run.log under /Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/638-r5/).

patch command test it reds exit
hand-back-no-wait cargo test --lib -- the_hand_back_waits_for_the_boot_record_the_judge_reads that test 101
wait-for-boot-at-once the same that test 101
said-refusal-first-line cargo test --lib -- a_refusal_is_read_back_off_the_drivers_stderr that test 101
wait-on-port-alone cargo test --lib -- coming_back_is_ssh_answering_and_not_its_port that test 101
go-down-dial-one-poll the same that test: "going down dialled within [1s], where ssh's own connect is given 10s" 101
bound-for-always-wedge cargo test --lib -- every_arm_that_stops_this_machine_is_cleared_and_judged_as_one that test 101
backstop-one-second-short cargo test --test toyos-checks serial_vocabulary: "silent from 4.2s under a 5s ceiling" 101
backstop-plus-six the same serial_vocabulary: "silent from 0ns" 101
backstop-twice-only the same serial_vocabulary: "silent from 0ns" 101
backstop-mains the same a_stall_stays_red 101
backstop-no-death the same serial_vocabulary: "a kernel death at 5.2s" 101
wedge-irq-storm cargo test --test toyos-build -- virt_irq_storm; then cargo test --test toyos-build virt_irq_storm: "irq-storm never reported", at 37 s alone and 37 s in a 44.3 s suite 1; 1

The unmutated head: cargo test --lib over the four tests EXIT=0, cargo test --test toyos-checks EXIT=0 (31 passed).

@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

The wedge's third arm at e5daa61ff: the whole change reverted onto origin/main at 322085f23 with git diff --binary HEAD origin/main, then wedge-irq-storm.patch (comment 5963363857) applied, built (EXIT=0), cargo test --test toyos-build -- virt_irq_storm run, and both reversed with the tree clean after (at-merge/mutations/wedge-main.sh, wedge-main.log): EXIT=1, FAIL virt_irq_storm: irq-storm never reported at 181 s, against 37 s at the head.

Japabu and others added 2 commits October 3, 2026 02:16
…red in one session

The filing at 4b009e7 set an x86-64 boot under TCG beside an AArch64
boot under HVF and named only the marker. At e5daa61, one run after
the other: virt_early_panic (Virt, HVF, to `EARLY PANIC:`) 479 ms,
virt_smp (VirtEl2, TCG, to `SCTLR_EL1=`) 1180 ms, machine_shutdown
(x86-64, TCG, to `===READY===`) 3316 ms and 2.33x.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
#680 stamps every statement on stderr with the time of day, the
driver's refusal among them, so `said_refusal` reads a line through
`printer::unstamped`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Negative controls at 75dc056f1 (1ba6b43e3 with origin/main at 73282fa93, #680, merged in; #648 came in at e5daa61ff). These patches replace the ones in comments 5963363857 and 5963458288: said-refusal-first-line is cut against the reader the #680 merge changed, and said-refusal-stamped is new. Each is checked with git apply --check, applied, built (cargo test <target> --no-run, EXIT=0 each), run, and reversed, with the tree clean after each (at-head/mutations/run.sh, wedge.sh, wedge-main.sh and their logs under /Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/638-r5/).

patch command test it reds exit
hand-back-no-wait cargo test --lib -- the_hand_back_waits_for_the_boot_record_the_judge_reads that test 101
wait-for-boot-at-once the same that test 101
said-refusal-first-line cargo test --lib -- a_refusal_is_read_back_off_the_drivers_stderr that test 101
said-refusal-stamped (new) the same that test: left: None 101
wait-on-port-alone cargo test --lib -- coming_back_is_ssh_answering_and_not_its_port that test 101
go-down-dial-one-poll (new) the same that test: "going down dialled within [1s], where ssh's own connect is given 10s" 101
bound-for-always-wedge cargo test --lib -- every_arm_that_stops_this_machine_is_cleared_and_judged_as_one that test 101
backstop-one-second-short cargo test --test toyos-checks serial_vocabulary: "silent from 4.2s under a 5s ceiling" 101
backstop-plus-six the same serial_vocabulary: "silent from 0ns" 101
backstop-twice-only the same serial_vocabulary: "silent from 0ns" 101
backstop-mains the same a_stall_stays_red 101
backstop-no-death the same serial_vocabulary: "a kernel death at 5.2s" 101
wedge-irq-storm, at the head cargo test --test toyos-build -- virt_irq_storm; then cargo test --test toyos-build virt_irq_storm: "irq-storm never reported", at 37 s alone and 38 s in a 45.4 s suite 1; 1
wedge-irq-storm on the whole change reverted onto origin/main at 73282fa93 (git diff --binary HEAD origin/main) cargo test --test toyos-build -- virt_irq_storm the same red at 181 s 1

The unmutated head: cargo test --lib over the four tests EXIT=0, cargo test --test toyos-checks EXIT=0 (31 passed).

hand-back-no-wait.patch
--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -977,7 +977,7 @@
 /// first line, so the backlog crosses at the network's pace and not the
 /// conversation's.
 fn hand_back<T>(stream: &Stream, began: Instant, by: Duration, fire: impl FnOnce() -> T) -> T {
-    match stream.wait_for_boot(by) {
+    match stream.wait_for_boot(Duration::ZERO.min(by)) {
         Some(ms) => println!(
             "  talk: the stream carried `Boot: complete` ({ms} ms), {} ms after it opened",
             began.elapsed().as_millis()
wait-for-boot-at-once.patch
--- a/src/metaltalk.rs
+++ b/src/metaltalk.rs
@@ -286,7 +286,7 @@
     /// This boot's `Boot: complete` in milliseconds, read as [`judge`] reads
     /// it, once a line carries it, or `None` after `by`.
     pub fn wait_for_boot(&self, by: Duration) -> Option<u64> {
-        self.wait_until(by, |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
+        self.wait_until(by.min(Duration::ZERO), |lines| lines.iter().find_map(|line| crate::bootlog::boot_millis(line)))
     }
 
     /// The peer, once a connection has carried a line, or `None` after `by`
said-refusal-first-line.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -115,6 +115,6 @@
         .iter()
         .rposition(|line| crate::printer::unstamped(line).starts_with(REFUSAL_HEAD))?;
     let mut said = vec![&crate::printer::unstamped(lines[at])[REFUSAL_HEAD.len()..]];
-    said.extend(&lines[at + 1..]);
+    said.extend(&lines[at + 1..at + 1]);
     Some(said.join("\n").trim_end().to_string())
 }
said-refusal-stamped.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -113,7 +113,7 @@
     let lines: Vec<&str> = stderr.lines().collect();
     let at = lines
         .iter()
-        .rposition(|line| crate::printer::unstamped(line).starts_with(REFUSAL_HEAD))?;
+        .rposition(|line| line.starts_with(REFUSAL_HEAD))?;
     let mut said = vec![&crate::printer::unstamped(lines[at])[REFUSAL_HEAD.len()..]];
     said.extend(&lines[at + 1..]);
     Some(said.join("\n").trim_end().to_string())
wait-on-port-alone.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -1513,7 +1513,7 @@
     while began.elapsed().as_secs() < secs {
         let next = std::time::Instant::now() + POLL;
         let listening = accepts(dial);
-        let answered = listening && (!answering || answers());
+        let answered = listening;
         if answered == answering {
             return Ok(began.elapsed().as_secs());
         }
go-down-dial-one-poll.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -1508,7 +1508,7 @@
     mut accepts: impl FnMut(std::time::Duration) -> bool,
     mut answers: impl FnMut() -> bool,
 ) -> Result<u64, Refusal> {
-    let dial = if answering { POLL } else { std::time::Duration::from_secs(CONNECT_SECS) };
+    let dial = POLL;
     let began = std::time::Instant::now();
     while began.elapsed().as_secs() < secs {
         let next = std::time::Instant::now() + POLL;
bound-for-always-wedge.patch
--- a/src/metal.rs
+++ b/src/metal.rs
@@ -852,7 +852,7 @@
 
 /// The `boot-deadline=` bound an image armed with `armed` carries.
 pub fn bound_for(armed: &[impl AsRef<str>]) -> u64 {
-    if stages_a_wedge(armed) {
+    if stages_a_wedge(armed) && false {
         toyos_tco::STAGED_BOUND_MS
     } else {
         toyos_tco::WEDGE_BOUND_MS
backstop-one-second-short.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET - Duration::from_secs(1));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-plus-six.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = (ceiling * 2).max(ceiling + Duration::from_secs(6));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-twice-only.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = ceiling * 2;
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-mains.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -517,7 +517,7 @@
     // For a guest that is stuck *and* chatty and so never trips the silence
     // guard; never before a guest silent since `ceiling` has been silent for
     // [`GUEST_QUIET`].
-    let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
+    let backstop = ceiling.max(Duration::from_secs(300));
     if elapsed > backstop {
         if let Some(line) = dying {
             return Some(kernel_died_here(line));
backstop-no-death.patch
--- a/tests/common/qemu.rs
+++ b/tests/common/qemu.rs
@@ -519,7 +519,7 @@
     // [`GUEST_QUIET`].
     let backstop = (ceiling * 2).max(ceiling + GUEST_QUIET);
     if elapsed > backstop {
-        if let Some(line) = dying {
+        if let Some(line) = dying.filter(|_| false) {
             return Some(kernel_died_here(line));
         }
         return Some(format!(
wedge-irq-storm.patch
--- a/kernel/src/arch/aarch64/trap.rs
+++ b/kernel/src/arch/aarch64/trap.rs
@@ -525,7 +525,7 @@
 
     pub(super) fn tick() {
         if RUNNING.load(Relaxed) {
-            TICKS.fetch_add(1, Relaxed);
+            TICKS.fetch_add(0, Relaxed);
         }
     }
 

@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

T14 at 75dc056f1, Drive mode, run by the orchestrator from 638-r5/metal/request.txt, with the owner's approval for the watchdog and reset rows. Each run was cargo test --test toyos-build -- --metal [filter] from the clean worktree at the head, without --metal-readback; the stick's log partition was saved before each run's first flash; no judging left a row in tests/metal (each saved patch is empty) and the worktree was clean after each.

run exit what it read
boot_deadline_ends_a_wedge 0 (118 s) armed ["wedge-before-reset", "boot-deadline=10000"]; the boot deadline expired: a bound of 10000 ms, reached at 10062 ms; back in 51 s; PASS, 1 passed, 0 failed, 1 boot(s)
hard_lockup_ends_a_deaf_cpu 0 (114 s) armed ["hard-lockup-probe", "boot-deadline=10000"]; a cpu locked up with interrupts off: cpu7 has taken no interrupt for 5000 ms, with IF clear at every sample in that span. Its bound is 5000 ms.; back in 59 s; PASS
usb_reset_records_the_phase_it_cut 0 (113 s) armed ["usb-reset-under-load", "boot-deadline=10000"]; deadline at 10066 ms; a Bulk-Only command was open in its data phase on slot 1 after 2000 ms; the controller had that device's data endpoint Running with 252 TRB(s) it had not reached on the ring; back in 54 s; PASS
unfiltered 1 (2273 s) 50 registration(s) and 222 shared member(s) over 25 boot(s); 270 passed, 2 failed, 25 boot(s)

The unfiltered run in detail:

  • The shared boots at this head's sizes, each in one boot and every member run: shared: 62 member(s) (61 ms per member over the 62 that ran), shared-2: 19 member(s) (26 ms per member), shared-debug: 5 member(s) (31 ms per member), ccorpus: 136 member(s) (25 ms per member over the 136 that ran), no ccorpus-2.
  • PASS among the rest: boot_deadline_ends_a_wedge, hard_lockup_ends_a_deaf_cpu, usb_reset_records_the_phase_it_cut, loader_watchdog_arms, watchdog_fed, lan_talk, lan_message_delivery, lan_lease_report, readdir_bound, hda_tone.
  • The two reds, both main's and recorded, and no other: lan_dhcp_lease (issues/hardware/the-benchs-router-leases-toyos-another-address-than-ubuntu.md) and hda_client_stall (soundd resumed 1 time(s) — the second stream did not find a suspended daemon; issues/audio/hda-client-stall-reads-one-resume-where-its-judge-wants-two.md). The third allowed red, hda_tone, passed in this run.

Logs and readbacks: /Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/638-r5/metal/ (<run>.log, <run>-readback/, <run>.record-rows.patch).

@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Round 5, head 75dc056f1. This round reads round 4's BLOCKERs and what changed since 4859e8d71: 4b009e712, 1ba6b43e3, and the merges of #648 and #680.

Round 4's BLOCKERs

  • toyos-tco/src/lib.rs:154, the three staged-wedge rows at 10 s: CLOSED. Comment 5966698095 ran Drive mode at 75dc056f1. Each row exited 0 alone and passed again in the unfiltered run.
    • boot_deadline_ends_a_wedge: armed ["wedge-before-reset", "boot-deadline=10000"]. The deadline was reached at 10062 ms (10063 ms in the unfiltered run). PASS.
    • hard_lockup_ends_a_deaf_cpu: armed ["hard-lockup-probe", "boot-deadline=10000"]. The detector sealed the record, not the deadline: cpu7 has taken no interrupt for 5000 ms … Its bound is 5000 ms. The wedge was staged at 1.210 s and the ring's last line is cpu7's LOCK CONTENTION at 7.082 s, so the deadline had about 2.9 s to spare. PASS both times.
    • usb_reset_records_the_phase_it_cut: the deadline was reached at 10066 ms, with 252 TRBs (119 in the unfiltered run). PASS.
  • tests/toyos.rs:766, :856, the shared boots at their new sizes: CLOSED. The unfiltered --metal run exited 1 with 270 passed, 2 failed, 25 boot(s) in 2273 s.
    • shared: 62 member(s), at 61 ms per member over the 62 that ran.
    • shared-2: 19 and shared-debug: 5. ccorpus: 136 ran in one boot, with no ccorpus-2.
    • 50 registrations and 222 members make 272, which is 270 passed and 2 failed. Both reds are registrations, so every member passed.

Rulings on the brief

Evidence at 75dc056f1

  • cargo run -- --ci host: EXIT=0, "Host: 67 step(s), all green".
  • cargo test --test toyos-build: EXIT=0, 26 passed in 44.5 s.
  • toyos-checks: EXIT=0, 31 passed.
  • The twelve negative controls: EXIT=101 each.
  • The wedge: 37 s alone, 38 s inside a 45.4 s suite, and 181 s on main's tree.
  • CI: the PR is a draft, so every job was skipped.

Growth: +740 −671 over 38 files, net +69.

  • Code alone: +425 −543, net −118.
    • Shipping (netd, toyos-i219 lib/regs, toyos-tco): +23 −63.
    • src/ production: +112 −14. src/ tests: +119 −58.
    • toyos-i219 model and tests: +2 −57.
    • tests/: +169 −351.
  • issues/: +315 −128.
  • Production has shrunk since round 4, because the refusal-wait and after_the_reboot are gone. Accepted.

BLOCKER

  • src/metal.rs:2608-2612, :2615-2619 — the dialled wrapper and its assertion check only that wait_on hands accepts CONNECT_SECS when going down, which is the one line at :1517. The go-down-dial-one-poll control is a defect any reader of the diff catches, and the body's Unsure admits no test stages a lost SYN. A test that guards what a reader can check goes. Delete both spans and pass accepts at :2613. Cut :2590 back to "a machine going down is its port refusing". Drop the body's control row and its line "No test stages a lost SYN…".

NOTE

  • issues/build/the-talking-boots-reboot-outruns-its-log-stream.md:45 — its exit, "Two consecutive T14 runs pass lan_talk", is met: PASS lan_talk at 4859e8d71 (comment 5962753432) and at 75dc056f1 (comment 5966698095), both runs with hand_back. Delete the file, its citation at issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md:27, and the body's "which stays until two T14 runs pass lan_talk".

REMOVE

  • PR body, the "The T14." bullet — "no run has read them since main's kernel merges. They are owed at this head (The T14, below)." This has been false since comment 5966698095.

SEND BACK

… issue closes

Round 5 of #638's review.

`coming_back_is_ssh_answering_and_not_its_port` wrapped `accepts` to record
the bound `wait_on` hands it going down and asserted that bound was
`CONNECT_SECS`. That checks one line of `wait_on` a reader of the diff
checks, so the wrapper, its assertion and the doc clause that named it go,
and `accepts` is passed directly.

`issues/build/the-talking-boots-reboot-outruns-its-log-stream.md` is
deleted: its exit, two consecutive T14 runs passing `lan_talk`, is met by
the runs at `4859e8d71` (PR comment 5962753432) and at `75dc056f1` (PR
comment 5966698095), both with `metaltalk::hand_back` waiting on the
stream's `Boot: complete`. The rule it carried is already `hand_back`'s
doc. Its one citation, in
`issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md`, goes
with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
@Japabu

Japabu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Round 6, head e3094a94a. This round reads round 5's findings and what changed since 75dc056f1, which is the one commit e3094a94a. origin/main is still 73282fa93, the tip that 75dc056f1 merged.

Round 5's BLOCKER

  • src/metal.rs:2608-2612, :2615-2619, the test-only dialled wrapper: CLOSED.
    • Both spans are deleted, and accepts is passed directly at :2608. The doc at :2590 now reads "a machine going down is its port refusing". The body's control row and its "No test stages a lost SYN" line are both gone.
    • At this head, 638-r6/host.log shows metal::tests::coming_back_is_ssh_answering_and_not_its_port ... ok, and the run ends Host: 67 step(s), all green (host.exit EXIT=0). The run started at 07:33:27 UTC, 16 s after the commit.
    • The remaining control was run again at this head (638-r6/mutations/run.log). The patch was let answered = listening;. Unmutated EXIT=0, build EXIT=0, mutated EXIT=101 at :2605 with left: Ok(0), right: Err(Silent { what: "come back", secs: 1 }). The tree was [] before and after.

Round 5's NOTE and REMOVE

  • The NOTE on issues/build/the-talking-boots-reboot-outruns-its-log-stream.md: CLOSED.
    • The file is deleted, and so is its citation in issues/hardware/netds-i219-drops-a-transmit-burst-past-its-ring.md. The body's "which stays until…" is replaced.
    • git grep for the slug at e3094a94a exits 1.
    • Its durable rule is already the doc of hand_back (src/metaltalk.rs:974-978), which is what issues/README.md's closing procedure asks. The commit message carries the evidence.
  • The REMOVE of the T14 bullet's "no run has read them…": CLOSED, the sentence is gone from the body.

Whether the readings at 75dc056f1 stand for this head

  • They do. git diff -U0 75dc056f1 e3094a94a -- src/metal.rs has three hunks, at :2590, :2608 and :2615. All three lie inside #[cfg(test)] mod tests, which opens at :2578.
    • Only cargo test --lib compiles that module.
    • The toyos-metal loop, the image and the toyos-build integration harness link the library without it.
  • The other two files are under issues/. No include_str! or include_bytes! in the tree reads issues/, and the host steps that read it ran green at this head.
  • So the T14 Drive-mode readings, build-only.log and guest.log at 75dc056f1 cover this head. toyos-checks was run again at this head: EXIT=0, 31 passed.

Evidence at e3094a94a

  • cargo run -- --ci host: EXIT=0. The FAILED lines in that log belong to the control steps, which are required to fail, for example === [ci] control \wake-fence-off``.
  • cargo test --test toyos-checks: EXIT=0, 31 passed.
  • CI: the PR is a draft, so host, toolchain and guest were all SKIPPED. The merge queue runs them.

Growth. git diff --shortstat origin/main...e3094a94a: 38 files, +717 −704, net +13.

  • Production: src/ +112 −14, and netd, toyos-i219 lib and regs and toyos-tco +23 −63.
  • Tests: src/ test modules +109 −58, tests/ +169 −351, and toyos-i219's model and tests +2 −57.
  • Issues: +302 −161.
  • Since round 5 the branch is −56 lines, all in test code and issues. Production has not moved since round 5, when it was accepted.

BLOCKER

None.

NOTE

None.

REMOVE

None.

LAND

@Japabu
Japabu marked this pull request as ready for review October 3, 2026 07:48
@Japabu
Japabu enabled auto-merge October 3, 2026 07:48
@Japabu
Japabu added this pull request to the merge queue Oct 3, 2026
Merged via the queue into main with commit 4adc1c0 Oct 3, 2026
6 checks passed
@Japabu
Japabu deleted the wt/toyos-tight branch October 3, 2026 08:24
Japabu added a commit that referenced this pull request Oct 3, 2026
…hers) into wt/toyos-resident

One conflict: issues/hardware/the-t14-boots-toyos-unattended.md, which this
branch deletes and main modified. Every hunk of main's side since the merge
base 4f2bea1 is accounted for:

- e9f67e7 (#639) appended one line: "`smp_failed_ap_leaves_no_hole` is
  deleted; issues/build/smp-ap-hole-and-log-reserve-window-red-under-a-loaded-host.md
  records the commit that restores it." It qualified the sentinel's finding 2,
  which named that test as the roster's gate. No file replacing the track
  plans a sentinel or names the test; the gap a sentinel was for (the span
  before clock::init) is stated in kernel/src/deadline.rs's header, and the
  restore the line pointed at stays where it was recorded, in
  smp-ap-hole-and-log-reserve-window-red-under-a-loaded-host.md
  ("`git revert aedcf17` brings it back"). So the file stays deleted and
  the line has nothing left to qualify.

The deleted track's one other citation, from #638:
issues/build/hard-lockup-bound-ms-is-read-by-nothing-but-its-own-assertion.md
named it as owner. Its owner is now
issues/hardware/a-frozen-toyos-waits-for-a-hand-on-the-power-button.md, the
track #638's review named. Its two exits disagreed: that track's first step
gives HARD_LOCKUP_BOUND_MS a reader (the detector's bound on a boot that names
no boot-deadline=), and the issue's exit deleted the constant. The reader
stands: the issue now closes on that step, with the constant's doc saying what
it bounds, and says deleting it is not the exit, since the step would declare
it again.

issues/boot-media/the-loader-does-only-what-must-precede-the-handover.md
merged cleanly; main's two hunks (create_boot_image's per-image GUIDs, and
root_read_ticks) are in it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Japabu added a commit that referenced this pull request Oct 3, 2026
…today

Each citation the branch's files make was checked against 4adc1c0. Every
name a track or issue cites exists there (toyos_tco::STAGED_BOUND_MS,
RUST_MEMBER_MS, C_MEMBER_MS, SharedBoot's `members`, metal::judge_arms,
panic_console::hold_the_panel, arch::watchdog's `arm` and its TCO2_STS read,
bootlog::split_listing, toyos::log::LogTail, vfs::ROOT_ENTRIES, Armed's
(0, _) arm, every boot and row name), but for these, which this commit fixes:

- `device_claim_lifetime` is gone (51cc87f, #639). `endowment_denied` is now
  the one session member that mints a claim (`git grep -l device_claim` over
  the Rust tests, test-runner and the C corpus finds it alone), so the track
  says it runs last in its session rather than that the two end one each.
- `log_poll_outlives_a_close` is gone (ad6dc07, #639). Of the judges on the
  six folded boots, `syscall_cost` alone reads the job's own output; the audio
  judges read soundd's lines between the kernel's `spawn:` records, which reach
  the stick.
- The guest-suite track does not itself leave the shared boot to the T14: #660
  did, and `shared_metal` and `c_corpus_metal` are its only runs outside
  `--debug`. The track now says the first and cites the issue for the rows it
  adds.
- The prices name #638's allowances, toyos_tco::RUST_MEMBER_MS and
  C_MEMBER_MS, rather than "as SharedBoot::members prices them": `members` is
  a count derived from them.
- Line citations into test-runner and toyos/src/syscap.rs moved:
  `--bound-ms=` is main.rs:78-83, run_one's duplicate 248-250, the namespace's
  inheritance 62-68, SysCap::duplicate 63-70 and SysCap::narrowed 133-139.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant