Repository navigation
Conversation
The orchestrator stops re-running reds on a quiet host. Every red below is on main's code, so each is disabled now, and its issue says what it is and what ends it. Defects already on main: - process_stats: a connection park charged 0 ns to ipc and to pipe, at process_stats.rs:280 on wt/toyos-loaderlines 6e0d7da (fastest boot 480 ms) and at :263 on wt/toyos-proclife1 60ec86d. Neither branch touches the test or the charge. The existing issue takes both sightings and becomes expected-red. - lan_dhcp_lease: the test awaits the lease and asserts netd's ready line after it, but the capture ends at test-runner's ===READY===, which can come first (wt/toyos-tight 59940c4, ceilings at 1.00x). - root_candidate_malformed: wait_for_ready reads a 16550 ready marker off QEMU's log file while the loader is still writing the line, and the test got "Slot A: REFUSED, its root is" (same run). - redirty_mid_flush: the guest went silent after spawning its child (wt/toyos-proclife 6e9d6a4). The branch's own kernel change is two retired-syscall lines; the test passed at the branch's previous head and in every other kept whole-suite log. Verdicts the host's speed decided: - update_boots_the_new_kernel, update_falls_back_from_a_dying_kernel, update_floor_is_the_images_own, update_refusals_boot_the_other_slot: 15 s of silence across the firmware reset reads as a stall at load average 84, on #648's run and again on #642's. - guest_dies_with_its_harness: the owner's boot is judged at an unscaled 20 s on the host's clock. - syscall_window_nmi_controls: the spinner stops after ten seconds of the guest's clock and the storm waits for a million syscalls; starved, it made 950000. - virt_smp: an AP has 100 ms to echo, and a starved vCPU did not. src/redlist.rs's whole-name test probed `<row>_controls` against None; with syscall_window_nmi and syscall_window_nmi_controls both disabled that probe finds the second row, so it now asserts only that the probe does not find the row it extends. The old assertion reds on this table (checked: EXIT=101), the new one passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…uest's own clock, and the kernel's waits on another CPU are bounds on a dead one Seven reds were the host's speed, not the guest's behaviour. Each timing assumption is removed where it lives. The harness. Every wait on a guest measured the host's wall clock, so a TCG guest whose QEMU sat runnable and unserved read as one that stopped (the update tests' 15 s of silence across a firmware reset; the owner's boot in guest_dies_with_its_harness killed at 20 s with its loader still printing). The opposite error had the same cause: `budget` multiplied a ceiling by the phase width and by a fastest-boot `host_scale` taken early in a loaded run, so redirty_mid_flush's stopped guest was waited on for 4474 s against its 120 s budget. - tests/common/steal.rs is a guest's own clock: the wall clock less the host's steal, read from QEMU's process accounting (macOS proc_pid_rusage's run and runnable times; Linux's per-thread schedstat). Demand under one thread's worth loses the moments it waited; demand over it has the share the host served. A guest that wants nothing, idle or stopped, has the wall's time, so a stopped guest is called stopped on time however loaded the host is. - Every wait on a guest reads it: the boot wait, run_test's ceiling, the drains, Liveness, the screendump waits, await_exit, the update tests' wait, the QMP event waits (now bounded on the clock with a 200 ms read poll), the segment's frame wait, and the tests' own deadline loops. - WIDTH, set_width and round_trip go: sharing the host is the clock's to take out, and the width was a stand-in for it. The boot ceiling is the twenty seconds a phase of one guest gave it. `host_scale` reads boots on the guest's clock, so it measures the host's speed and not its load. The kernel. `time::AP_START` gave an AP 100 ms to echo, and a vCPU its host had not scheduled for that long booted the machine one CPU short (virt_smp, cpu7 of 8). It is now DEAF_CPU's span, the bound this kernel already holds a live CPU to; Linux waits 5 s for the same echo on arm64 (`__cpu_up`) and 10 s in its generic bring-up (`cpuhp_wait_for_sync_state`). The NMI storm. syscall_window_nmi_controls' spinner stopped after ten seconds of the guest's clock while the storm waits for a million syscalls on one CPU; starved, it made 950000 and the storm never fired. The spinner now spins until the harness ends its boot, and the storm's three waits on its victim (`HOLD_ACK_NS` twice, `HELD_DELIVERY_NS`) are DEAF_CPU's span rather than 100 ms; the hold's budget of turns grows to 2^31 to outlast it at the 3 ns-a-turn estimate its assertion states. The seven rows #654 disabled for these, and their three issues, go: update_boots_the_new_kernel, update_falls_back_from_a_dying_kernel, update_floor_is_the_images_own, update_refusals_boot_the_other_slot, guest_dies_with_its_harness, syscall_window_nmi_controls, virt_smp. -icount is not taken: it is TCG's alone, so the HVF and KVM guests keep the kernel's exposure to steal and only TCG's is masked; QEMU refuses MTTCG with it ("No MTTCG when icount is enabled", accel/tcg/tcg-all.c), and round-robin switches vCPUs on a 100 ms kick of the virtual clock (TCG_KICK_PERIOD), which ISB-based spin loops on AArch64 do not yield before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…only branch's run Both went red in 650-libcllvm-whole.log on wt/toyos-libcllvm, whose diff is userland/libc alone, at load average 66-74. - iommu_virtio_platform is a defect on main. It reads netd's claim line and netd's feature negotiation off a boot log that ends at test-runner's ===READY===, and netd starts before test-runner but speaks after it. The same race reds in two wordings on 634r2, 637r2 and 641f-r3, each on a branch that does not touch it. #639 carries a fix (d773a43). - loader_watchdog_arms timed out its boot at 132 s of wall clock, which is 10 s x 12 wide x host_scale. The console stood at "BdsDxe: starting Boot0001". Its 50-odd other runs in the kept logs passed in 6-34 s. It goes under the harness-wait issue, which now carries this sighting and 650's two update STALLs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
The redlist branch added iommu_virtio_platform, which stays disabled here, and loader_watchdog_arms under the harness-wait issue, which this branch closes. Its row goes with the issue. The issue's modified side carried 650's sightings: two update STALLs and loader_watchdog_arms' 132 s boot timeout at "BdsDxe: starting Boot0001", all on wall-clock ceilings this branch puts on the guest's own clock. They are recorded here and in the pull request rather than in a file this merge deletes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…d borrow Clippy's needless_borrow, red in the host gate's clippy step at 330c76a: the guest is already a reference there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
650-libcllvm-whole.log, wt/toyos-libcllvm, whose diff is userland/libc alone; load average 66-87, with the run reporting ceilings at 1.01x. - metal_sim_window_drag and usb_boot_stick_pulled timed out their boots at 21 and 22 s in the serial tail, the 20 s a phase of one guest gets. Each had the loader still printing its segments. They go under the harness-wait issue, which now also says why the run's 1.01x could not see the load. - blockd_serves_partitions: the bench passed and its process exited at 35.978 s, then no ===TEST_END=== came for 3559 s, while the kernel kept logging with every CPU idle. Either test-runner's wait on the job missed the exit's post, or logd stopped forwarding. The new issue says which evidence the next sighting needs, and names #655's poller shape as the logd reading. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
The redlist branch added blockd_serves_partitions, which stays disabled here. It also put metal_sim_window_drag and usb_boot_stick_pulled under the harness-wait issue, which this branch closes, so their rows go with it. The issue's modified side carried 650's serial-tail sightings: two boots cut at 21 and 22 s of wall clock with the loader still printing, and a run that called its ceilings 1.01x at load average 66-87. Both are what the guest's clock and the end of the width take out. They are recorded here and in the pull request rather than in a file this merge deletes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…without schedstat is refused #650's run called its ceilings "1.01x" at a load average of 66-87. Its gauge is the fastest boot, taken at the run's quietest moment, so it sees the host's speed and never its load. Every guest's clock now adds its read spans to a run-wide tally, and the summary says what share of its waits' wall clock the guests had. That is the load, measured where it fell. On Linux the clock reads each thread's /proc schedstat. A kernel built without CONFIG_SCHED_INFO has no such file, so a starved guest would read as one that wanted nothing. demand() now refuses such a kernel by name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
blockd_serves_partitions' guest printed "ready=0 ... parked=4" every ten seconds for an hour. The coordinator asked that such a guest be ended within seconds. The line cannot carry that verdict: a guest waiting out a timer prints the same line. The issue records what would carry it, a count of the parks with a deadline, and what it would end early. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
Japabu
marked this pull request as ready for review
October 1, 2026 03:58
This was referenced Oct 1, 2026
Collaborator
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The host's load no longer decides a guest test's verdict for the ten reds it decided. Every wait the harness puts on a guest reads that guest's own clock: the wall clock with the host's steal taken out. The kernel's waits on another CPU's progress are bounds on a dead CPU, not on one the hypervisor has not scheduled yet. The ten rows #654 disabled for these reds go, along with their three issues.
Triage
issues/diagnostics/an-idle-guest-and-a-hung-one-print-the-same-sched-line.md)Decisions
The guest's clock (
tests/common/steal.rs). A wait on a guest is a hang detector, and a hang is a guest that does not progress while it has the host. The clock reads QEMU's process accounting: run and runnable time fromproc_pid_rusageon macOS, and each thread'sschedstaton Linux. For each span:A guest that wants nothing gets the wall's time, so an idle guest and a stopped one are judged on time however loaded the host is. Only a starved guest's clock slows. On a quiet host the clock is the wall, and every ceiling is the number it was before. The independent oracle is steal time as hypervisors account it to a vCPU: Linux's
/proc/statsteal column and KVM's steal-time MSR are the same subtraction.Every wait on a guest reads it:
run_test's ceiling,drain_serial/drain_until,wait_for_console,Liveness,await_guest/await_resetawait_exitawait_machineQmpShutdown,QmpResets,QmpHold, now bounded on the clock with a 200 ms read poll)WIDTH,
set_widthandround_tripare deleted. The width was a stand-in for contention, and contention is now the clock's to take out. Keeping both double-counts it, which is how a stoppedredirty_mid_flushwas waited on for 4474 s.host_scalenow times boots on the guest's clock, so it measures the host's speed and not its load. The boot ceiling is the 20 s a phase of one guest already gave it.The run reports its load. #650 called its ceilings "1.01x" at load average 66-87. Its gauge was the fastest boot, which shows the host's speed at its quietest moment and never its load. Every guest's clock now adds its spans to a run-wide tally, and the summary prints
host: the guests had N% of the … their waits read across. On Linux, a kernel withoutschedstat(CONFIG_SCHED_INFO) is refused by name: without that file a starved guest would read as one that wanted nothing.AP_STARTis now DEAF_CPU's 5 s, up from 100 ms. The degraded answer on expiry is booting one CPU short for good. A vCPU whose host has not scheduled it is slow, not dead, and a live CPU is already held to DEAF_CPU's span elsewhere in this kernel. Independent oracle: Linux waits 5000 ms for the same echo on arm64 (arch/arm64/kernel/smp.c,__cpu_up) and 10 s in its generic bring-up (kernel/cpu.c,cpuhp_wait_for_sync_state). Cost on metal: a CPU that never answers costs 5 s of boot once, because the loop stops at the first one. Thesmp-skip-apboots now wait that 5 s.The NMI storm.
HOLD_ACK_NSandHELD_DELIVERY_NSare DEAF_CPU's 5 s, up from 100 ms.hold::TURNSis now 2^31, so the hold outlasts that wait at the 3 ns-a-turn estimate its own assertion states.The redlisted
syscall_window_nmiuses the same storm, and its own issue (the sprayed ratio) is unchanged.-icountis not taken.accel/tcg/tcg-all.c, v11.1.1). Every guest's vCPUs would share one thread in round-robin, switched on aTCG_KICK_PERIODof 100 ms of virtual time (accel/tcg/tcg-accel-ops-rr.c).core::hint::spin_loopisISB SY, which does not yield before that kick, so each cross-CPU spin would pay up to 100 ms.The cost in suite time is unmeasured. An optional job pair below measures it.
Gates
cargo run -- --ci hostat 1478bd4: EXIT=0, "56 step(s), all green". The head, f1ad5d1, adds one issue file after it.needless_borrowattests/common/power.rs:777its one red step. 262dd97 fixed it, and the gate there was EXIT=0, 56 green.cargo test --test toyos-checks -- steal_clock: EXIT=0.cargo run -- --build-only: EXIT=0, at 330c76a.cargo run -- --kernel-feature boot-actuators --build-only: EXIT=0. This is the kernel that compilesnmi_gate, withTURNS = 2^31and its assertion.nmi_window_spinbuilt forx86_64-unknown-toyoswith the guest toolchain and profile the harness uses: EXIT=0, no warnings.Negative controls and the oracle
servedreturning the wall span (mut-served-is-wall.patch) built and ran: EXIT=101, "light demand lost the moments it waited: … gave 1s, not 800ms". Losing the runnable time (mut-no-runnable.patch) built and ran: EXIT=101, "28 spinners on 14 cores ran 2.55s and waited 0ns".negctl-revert-the-change.patch; its tree builds at 1478bd4,cargo test --test toyos-build --no-runEXIT=0, and so does the-icountmeasurement's). It is requested under the same load as the green arm. Main's own record already holds four sightings: 648 (update ×4, guest_dies, nmi_controls), 642r2 (update_boots), and 636r2 (virt_smp).steal_clockcheck output).Unsure
wallclock.rs's host-time reads, which are the subject of those tests.issues/kernel/a-shootdown-panicked-on-a-cpu-the-host-starved.md). That issue owns the question.wt/toyos-tight) also deletes WIDTH, and it rewritesceiling_verdict, the boot ceiling and several waits converted here. Whichever lands second has a real merge.The long tests, toward ten minutes
Quiet-host whole runs: 642 (725.6 s), 638L (612.9 s), 641f-r3 (339.6 s). Each figure is that run's time.
#638 already cuts readdir_bound's
/homearm, which was most of its time.Guest runs owed
None has run on this head yet. The request is
/Users/jan/.claude/jobs/2280e09e/tmp/scratchpad/orch/loadfree/jobs.txt:std_pair with and without-icount.🤖 Generated with Claude Code
https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L