Repository navigation
The boot CPU waits for an AP's echo as long as a dead CPU is: 5 s, not 100 ms - #670
Conversation
…r as long as a dead CPU is The 21-test guest suite is meant to become a required check on GitHub's shared Linux runners, so the host's speed and load may decide none of its verdicts. On main every wait the harness puts on a guest read the host's wall clock: the boot wait, `Liveness`'s silence and backstop (every `await_guest`), `drain_until`, `screendump_while`, `run_test_paced`'s ceiling and the fatal-halt test's wait for QMP to show the other vCPUs halted. A loaded host stretches a TCG guest's work by however long QEMU's threads sat runnable and unserved, and a wall-clock wait reads that stretch as a hang. `tests/common/steal.rs` is the guest's clock, carried from #658 (`wt/toyos-loadfree`, `f1ad5d1ed`): the wall clock less what the host withheld from the guest's QEMU. Demand under one thread's worth loses the moments it waited; demand over it has the share the host served. A guest that wants nothing gets the wall's time, so a stopped guest is still called stopped on time. The clock never runs faster than the wall. Two corrections to #658's reading: - macOS: `ri_runnable_time` counts a thread running as well as waiting to. Measured on this host: one spinner alone for 500 ms ran 500.05 ms and was runnable 503.24 ms; 28 on 14 cores ran 6.93 s and were runnable 14.00 s. #658 read the whole of it as the wait, so a busy guest on a quiet host had half the wall. The wait is now what it holds past the run. - Linux: `/proc/<pid>/task/*/schedstat` is read per tid, and a span counts each thread's growth since the last reading. A thread that exits now loses only its last span, where #658's sum lost the thread's whole life and read the span as no demand. `WIDTH` and `set_width` go: the width was a stand-in for contention, and contention is now the clock's to take out, so keeping both double-counts it. `host_scale` times boots on the guest's clock, so it measures the host's speed and not its load. The boot wait is 20 s of the guest's clock, the number a run of one guest always had. The suite's summary says what share of the wall its guests had. `time::AP_START` becomes `DEAF_CPU`'s 5 s, up from 100 ms. Its expiry boots the machine one CPU short for good, and a vCPU its host has not run yet is slow, not dead. Linux waits 5000 ms for the same echo on arm64 (`arch/arm64/kernel/smp.c`, `__cpu_up`). `virt_smp` was booted 7 of 8 on a loaded host for this (`636r2-virt.log`), so `issues/kernel/an-ap-its-host-has-not-scheduled-for-100-ms-is-booted-without.md` goes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Orchestrator guest runs at
The negative control is green under a heavier load than the arm it controls: on the 21-test suite, main's wall-clock waits, |
|
Review round 1 at Gate. CI: Size. The branch adds 486 lines and removes 134, net +352:
After the cut below, it is kernel +4 −2 and issues −30, net −28. BLOCKER
NOTE
REMOVE
SEND BACK |
The harness goes back to main's, byte for byte: `tests/common/steal.rs`, `tests/checks/steal.rs` and `steal_clock` are deleted, and every wait that read the guest's clock reads the wall again. `WIDTH`, `set_width`, `host_scale`'s wall-clock boots and the boot wait's `10 s x max(width, 2)` stand as main has them. The clock fixed no measured failure. With main's wall-clock waits, `WIDTH` and the 100 ms `AP_START` reverted in, the 21-test suite went 21/21 under 28 spinners on the 14-core dev host at a load average of 86 to 98 (`loadfree2-negctl-whole-LOAD.log`). Its Linux reading had never run, and its `unimplemented!` would have panicked every wait on any other host. Deleting `WIDTH` was justified only by the clock taking contention out; without the clock it is unmeasured, and #638 deletes it on its own terms. What stays is `time::AP_START` at `DEAF_CPU`'s 5 s, from 100 ms, and the deletion of the issue it closes. Linux arm64 waits 5000 ms for the same echo (`arch/arm64/kernel/smp.c`, `__cpu_up`). `virt_smp` booted 7 of 8 with "cpu7 mpidr=0x7 did not echo within 100ms" on a loaded host (`636r2-virt.log`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Orchestrator guest run at
|
…illisecond `off` gave every other CPU 100 ms to take `SGI_OFF` and be off by `AFFINITY_INFO`, on a private `Budget` whose number was a choice. A vCPU its host serves late is slow and not dead: #670 moved `time::AP_START` off the same 100 ms on a recorded red. The wait now ends at `time::DEAF_CPU`'s span, the bound the tree already has for a CPU that does not take an interrupt, and declares nothing of its own. Its end stays a degraded answer and no panic: what keeps a CPU on may be firmware's refusal of `CPU_OFF`, and the machine was asked to end. The poll was a bare spin of firmware calls. At 100 ms `virt_off_names_the_cpus_left_on`'s trace was 6199064 bytes, 50812 calls; at five seconds that loop is fifty times it. `AFFINITY_INFO` is now asked again once per `ASK_AGAIN`, a 1 ms `Cadence`, as the xHCI port poll is paced: the same row's trace is 610488 bytes, 5004 calls, and the row says its count. Review of 17e6363, the rest: - `psci::Error::code` wrote `Error::of`'s table a second time for callers that print the number. `cpu_off`, `system_off`, `system_reset` and `affinity_info`'s refusal answer the `i32` the call returned. - `psci::init` runs first in `interrupts`, as x86-64's `init_reset` does, so a panic in the per-CPU block or the GIC's bring-up has a reset. - `virt_fatal_halts_the_others_first` is held to the armed line as a machine with a reset and no panic key says it, `panic: rebooting in 60 s, timed by`, where it took the held line too. The issue that records what stays unheld says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
time::AP_STARTis how long the boot CPU waits for an AP it started to echo its token. It is nowDEAF_CPU's 5 s, up from 100 ms. The issue it closes is deleted with it. The harness is main's, byte for byte.This supersedes #658 (
wt/toyos-loadfree,f1ad5d1ed).What changed, by decision
AP_STARTisDEAF_CPU's span (kernel/src/time.rs: one declaration, read by both architectures' bring-up loops).DEAF_CPUbefore it calls that CPU deaf.issues/kernel/an-ap-its-host-has-not-scheduled-for-100-ms-is-booted-without.mdis deleted. Its exit has two halves:virt_smpgreen in a whole-suite run at a load above the dev host's 14 cores, which is the guest run owed below.What it costs. A CPU that never echoes costs 5 s of boot, once, because the loop stops at the first one.
virt_failed_ap_leaves_no_hole(smp-skip-ap) meets it.after_consoleprints it (kernel/src/main.rs:281) beforestart_other_cpusruns (:438).judge_virt_job'sawait_marker, as 5 s of silence againstGUEST_QUIET's 15 s.sched/driver.rs:512). The T14 logs putBoot: completeat 1195–1276 ms, so a dead first AP reaches the first pass near 6.2 s.Dropped, and why
The whole harness half of round 1 and of #658 is dropped:
tests/common/steal.rs, its host checks, and every wait moved onto it. Nothing measured needs it.WIDTHand the 100 msAP_START. Under 28 spinners, at load 86 → 98, it went 21/21.unimplemented!would have panicked every wait on any host but macOS and Linux.WIDTH's deletion, and with ithost_scaleon the clock and the 20 s boot wait.WIDTHunder its own ceiling design.a-harness-wait-reads-a-starved-guest-as-a-stopped-one.mdstays unfiled, because the 21 tests give it no sighting.Gates
cargo run -- --ci hostat8b94f8438: EXIT=0, "Host: 56 step(s), all green".git diff origin/mainat this head names onlykernel/src/time.rsand the deleted issue.3f7081381's, wherecargo run -- --build-onlyexited 0.High risk: SMP bring-up
Independent oracle. Linux waits 5000 ms for the same echo on arm64:
arch/arm64/kernel/smp.c,__cpu_up,wait_for_completion_timeout(&cpu_running, msecs_to_jiffies(5000)), read at torvalds/linux551c722f4080.Recorded real failure at 100 ms.
virt_smpwent red at8f142ba9eon a loaded host, withSMP: cpu7 mpidr=0x7 did not echo within 100ms …andSMP: 7 of 8 MADT CPUs online.Negative control. No host test can reach the decision:
time.rsand botharch/*/smp.rscompile only for the kernel's targets.kernel-loomcarriessmp_roster.rsto the host, but itsawait_echotakes the expiry as a predicate and never sees the bound.The control is therefore the guest arm:
virt_smpunder load withAP_STARTat 100 ms. It goes red only when the host leaves a vCPU unrun for more than 100 ms after itsCPU_ON.636r2-virt.logrecorded that once. Round 1's negative-control run, at load 86 → 98, did not: it went 21/21.Unsure
virt_smpnames that red.time::DEAF_CPUitself can still panic a TCG guest whose host starves it for 5 s.issues/kernel/a-shootdown-panicked-on-a-cpu-the-host-starved.mdowns that.Owed
virt_smpandvirt_failed_ap_leaves_no_holegreen on main's wall-clock waits.🤖 Generated with Claude Code
https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L