Repository navigation
Fatal paths write the console UART only under its registers, whose lock knows which CPU's fatal path holds it - #675
Conversation
… burst never lands inside it `nested_nmi` wrote its report through `serial::panic_raw`, which takes no lock, before `halt_all_cpus` stops any CPU. Under KVM in CI the report on cpu0 and klogd's records on cpu1 went onto the 16550 into each other a byte at a time, and `nested_nmi_is_loud` went red twice: on #671's `guest` check (run 36863809437) it timed out with `NESTED NMI` never whole, and on main's nightly `guest` lane at 0678814 (run 36843762360) the capture stopped short of the report's tail, just before a `—` the splice could cut apart (the harness's reader stops at a line that is not UTF-8), and the boot waited out the panic bound until QEMU exited. The first spliced line of the first run is exactly a byte interleaving of `[kernel 0.385 cpu1] CPU 1: joining scheduler` and `[nmi] NESTED NMI on `: a shuffle check over the captured bytes says so, and says no with either string one byte off. The console's contract was the kernel's to keep, not the harness's to read around: the registers' lock is held for one burst, klogd writes every UART byte under it, and the panic path takes the registers alone and bypasses them only once they stay held. `panic_flush` did that; the report did not. A raw report carries no tag a reader could sort by, and a line spliced byte by byte is nothing a reader can undo. - `serial::panic_registers` is `panic_flush`'s bounded wait for the registers, lifted out unchanged: `None` once the bound runs out. `panic_flush` calls it. - `nested_nmi` holds what it returns for the whole report, and drops it before `halt_all_cpus`, whose flush takes the registers again. A hold the context it interrupted left on its own CPU delays the report by the bound and does not stop it. - `nested_nmi_is_loud` waits for the report's last line and asserts all three lines whole and back to back. It asserted only that `NESTED NMI` appeared, and the first red carried the last line whole and the first spliced, which it took a 39-second timeout to say. Filed, not fixed here: - issues/panic-path/a-fatal-report-written-raw-can-splice-into-another-cpus-console-burst.md: `panic::last_words` and the double fault's `[ist1]` line still write raw while another CPU can write, and `panic_registers` as it stands is not their fix. - issues/panic-path/a-panic-write-between-two-console-bursts-can-split-a-character.md: a burst can end inside a multibyte character, and the panic path writes between bursts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Readiness: CI Net: +106/−13. Production +20/−7, tests +28/−6, issues +58. BLOCKER
NOTE
REMOVE
SEND BACK |
…nd re-enters its own CPU's hold Review round 2 of #675. The type. `serial::PanicUart` is the one way a fatal path writes the console UART, and `serial::panic_registers` is its one constructor: `panic_raw`, `panic_raw_hex` and `panic_raw_dec` are gone. `panic::last_words` (both dead ends), `percpu::ist1_report`, `panic_reboot::arm`'s three raw lines and `panic_reboot::reboot_now` now hold the registers for their lines, beside `nested_nmi` and `panic_flush`. A write that does not ask no longer compiles. The lock. The bounded acquire moved into `serial_lock.rs` as `BackendLock::seize`, where kernel-loom drives it. The word now says who holds it: free, a burst holder (`LIVE`), or a fatal path's CPU. A fatal path entered on top of its own CPU's fatal hold re-enters at once instead of waiting for a frame that never runs again; that was a reentry inside `panic_flush`'s drain, and the report-then-halt sequence. Another holder is waited out for `PANIC_LOCK_SPIN_LIMIT` tries, and the expiry now says itself in one raw line. `within` is the console's one bounded wait, so `flush_final`'s copy of the loop goes. `lock()` is built on `try_lock()`, which lost its last other caller. `panic_flush` drains through virtio-console only under registers it took clean. Under its own CPU's fatal hold the outer path may have stopped inside a virtqueue publish, so it says so and drains raw, as it does past the bound. `nested_nmi` lets its report's registers go before the halt for that reason. The test. `nested_nmi_is_loud` now reads to the halt's flush, ending on the reboot's arm line, and refuses both lines the registers say when they were not clean. Its verdict is `faults::report`, which `tests/checks.rs` runs on a whole report, on the splice CI's KVM lane recorded before the report held the registers (run 36863809437, job 110375742604), and on each refused line, holding the refused words to `kernel/src/drivers/serial.rs`. Mutations, each a checked patch, built, run and restored: - `seize` takes the lock unbounded (the review's `Some(BackendGuard::lock())` at its new home): build 0, kernel-loom `serial_lock` 101. - `within` loses its bound: build 0, 101. - `seize` loses its re-entry arm: build 0, 101. - the kernel's re-entry line drifts from the harness's: the check 101. Issues: `a-fatal-report-written-raw-can-splice-into-another-cpus-console-burst` closes, since every raw report now holds the registers; what it left is filed as `a-fatal-path-inside-its-own-cpus-console-burst-waits-out-the-bound`: a burst holder names no CPU. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…nd a vouched commit stand behind, and one definition per lane Security (B1). No job a pull request, the merge queue or the nightly runs holds a token that writes: ci.yml and nightly.yml give their jobs `contents: read` and `actions: read` at the top, and the two workflows they call ask for nothing of their own. The one `contents: write` in any workflow is publish.yml's `release`, and `cargo run -- --ci release` refuses by name before it reads anything unless it runs as publish.yml on main, pushed or dispatched. The nightly's `build`, which published from any branch it was dispatched on, is gone: runs 36709239346 and 36600425263 published wt/toyos-castore's and wt/toyos-notiers' toolchains that way. CI no longer installs a release. toolchain.yml uploads each toolchain it bootstraps whole (upload-artifact v7, `archive: false`), so GitHub's recorded SHA-256 of the artifact is the tarball's own. `release::install` takes the newest build of its tree's tag made by main's publisher; failing that, it takes the newest made by a run of a commit its tree vouches for: its first-parent chain, and the head each merge on that chain took in. It refuses any other, and any download whose bytes hash to anything but GitHub's digest. An artifact is rewritten by nobody. The tag now hashes the build system that builds the toolchain: every module src/toolchain.rs and src/release.rs reach through `crate::`, 18 files, held to the sources by a test that recomputes the closure. Over main's last 100 first-parent landings that moves the tag on 32 where the old trees moved it on 20. The three `SOURCE` constants that existed only for the tag go. Timing (B2). Main publishes on every push (publish.yml's `toolchain` and `release`), not at 03:00. A tree no build answers for bootstraps in its own run's `toolchain` job and publishes nothing: 2h19m to 3h08m in the nightly `build` jobs that bootstrapped since 2026-09-29. Its merge group and every tree in main's window before main's own build lands install that pull request head's build, so the queue never bootstraps unless main moved the toolchain's inputs after the head's last run. One definition (B3). guest.yml is the guest lane, with KVM and the cache save as inputs; toolchain.yml is the toolchain job. ci.yml, nightly.yml and publish.yml call them. The shared-digest test and the cross-file comments go. B4: the guest lanes' gate holds `guest`'s `if:` whole, as `!cancelled()` around `host`'s condition, so the `needs.toolchain.result` mutation reds it. B5: CLAUDE.md and implementer.md say the plain suite runs in CI's `guest` check and the orchestrator runs every guest mutation; orchestrator.md is main's again. The two issue files this branch filed go: #675 and #676, batched with it, close them. The release-tag and install-digest issues close here; the release asset's remaining mutability is filed. The REMOVEd prose is deleted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Round 2 at Net Round 1
BLOCKER
NOTE
REMOVE
SEND BACK |
…t under its fatal hold; the UART moves a byte only with Registers
The round-2 review sent back six findings. Each is answered here.
The wait inside the bound now has a model. `a_seize_takes_a_burst_let_go_inside_the_bound`
holds a burst when the seize begins and lets it go on another thread. Under
`feature = "loom"`, `within`'s pause is `loom::hint::spin_loop`, imported beside
the loom atomics, so the holder runs between two looks.
The review asked for `seize(_, 2)` with `Taken` in every execution, and under
loom's default exploration that cannot hold for any count of tries. Measured:
with `LOOM_MAX_PREEMPTIONS=0` it is green, and with 1 or 2 it is red. Loom's
trace of the red execution shows the seize looking, yielding to the holder,
and then running again before the holder's store executes. That is the
holder preempted by the waiter, and each such preemption costs the seize one
try. So the model runs at preemption bounds 0, 1 and 2, with k + 2 tries at
bound k.
The console lock reads which CPU it is on through
`crate::arch::cpu::hardware_id`. kernel-loom supplies that function as a
per-model-thread shim (`arch::cpu::become_cpu`). With this, no call site
passes a CPU id, and the read the review mutated now lives in the file the
models compile. Mutated to 0 there, it fails two models, which put two CPUs
on one fatal hold.
`BackendLock::lock` refuses, by assertion, a burst on a CPU whose own fatal
path holds the registers. Holding `PanicUart` across a `log!` before klogd
runs therefore panics into the reentry guard's report instead of spinning
for ever on the CPU's own FATAL word. The patch the review named, braces
dropped in `last_words`, still compiles. Under it the hang becomes that loud
stop. The refusal is held by
`a_burst_beneath_its_own_cpus_fatal_path_is_refused`. The other CPUs' side is
held by `a_burst_waits_out_another_cpus_fatal_path`.
`arch::console_uart::{write_byte, read_byte}` take a `&mut serial::Registers`.
Only serial.rs builds one, for a `BackendGuard` or a `PanicUart`. virtio-console's
three `*_locked` functions take the `BackendGuard` itself, which closes the
review's NOTE by the same mechanism.
`nested_nmi_verdict` now also reads another CPU's burst cut into each of the
report's three lines, so `whole` is held on the host.
The false sentence in kernel-loom/tests/serial_lock.rs's header is deleted.
Filed:
- issues/kernel/a-safe-mmio-window-reaches-any-physical-address.md
- issues/panic-path/a-panic-reentry-inside-a-fatal-paths-console-hold-halts-with-it-held.md
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…e, so each of whole's clauses is held alone The burst the previous commit cut into each report line carried its own line end. That pushed the next line into the cut line's place, and a later clause of `whole` refused it. So the check passed with the first-line clause removed and also with the registers clause removed (each run exit 0, measured as mutations m4b and m4c). A burst with no line end leaves every line where it was, and only the cut line's own clause can refuse it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…ives up fails the test instead of aborting it Under mutations m2a (`Err(_) => Some(Seized::Expired)`) and m2c (`within`'s pause as `core::hint::spin_loop`), the seize gave up before the holder thread had run. The assertion then unwound while loom still had the holder's closure queued. That closure owns the burst, whose drop stores to a loom atomic outside the model. That is a panic in a destructor during cleanup, and the binary died on SIGABRT. The run was red (exit 101), but its other tests were cut short with it. Joining the holder first runs that drop inside the model, so the failure is the assertion's own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Guest runs at
Round 1's BLOCKER 2 now has its red run, for the reason the review predicted: the halt's flush finds this CPU's own |
|
Round 3 at Net
Round 3 alone: production +88/−49, tests +123/−15, issues +40. The production growth is the Round 2's BLOCKERs
BLOCKER
NOTE
REMOVE
SEND BACK |
…ring-up, and every UART touch takes the registers - kernel-loom compiles `serial_lock` and its `arch::cpu` shim under `loom` alone. Neither not-loom test (`log_zeroed_init`, `log_body_words`) names them, so the shim's not-loom arm, which no clippy shape built, is deleted rather than given a shape. - `within`'s count has a model: an attempt that never answers is asked exactly `tries` times, at 0, 1 and 100. With `0..tries.min(4)` every model passed before this one. - `arch::console_uart`'s `init`, `rx_ready` and `tx_ready` take `&mut serial::Registers` on both architectures, as `read_byte` and `write_byte` do. `serial::init` takes the registers' lock for the call. What the architecture logs under it is not drained there: `drain_inline` returns while `has_console` is false, which it is until `UART_PRESENT` is stored from the call's answer. AArch64's `init` takes them for the call it shares with the 16550's and touches no register. - `smp::echo` and `smp::commit`: a started AP says the hardware id it reads as its own with its echo, and the boot CPU refuses to commit it under any other id. x86-64's `boot_aps` holds the boot CPU's own read against its LAPIC id before the roster takes it. On AArch64 the boot CPU's roster slot is its read. Each refusal is the boot CPU's panic, before `smp::set_ready`, so a wrong read stops the boot. - Filed issues/kernel/a-safe-port-read-reaches-any-io-port.md: `inb(0x3f8)` takes a received console byte with no `Registers` and no `unsafe`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
`smp::commit` holds every AP's read against the MPIDR it was started by, and x86-64 holds its boot CPU's against the LAPIC id. On AArch64 the boot CPU's read is its roster slot and picks its redistributor, so no second id holds it: a read naming another CPU brings that CPU's redistributor up as the boot CPU's, and `CPU_ON` to the boot CPU is refused, which stops every CPU after it in the MADT. Recorded rather than fixed: the ruling for this round was to show what the read feeds there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
The review's cheapest red asserted on each CPU itself. An AP's own panic before `smp::set_ready` stops no other CPU: `stop_other_cpus` sends nothing before the release, on both architectures, and the boot CPU's `await_echo` runs out and boots on with the CPUs that came up. So the refusal is the boot CPU's, at the commit, where a wrong read stops the boot. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Guest runs at
The id → 0 pair shows the new bring-up assertion is what separates a wrong id read from a green boot. |
|
Round 4 at
The KVM arm is the batch's. Net
Round 4 alone: production +67/−26, tests and shim +26/−10, issues +45. The production growth is the Round 3's BLOCKER
Rulings
The handshake. Sound by reading:
What holds that order is the BLOCKER. Issue files. Both are accurate and have exits. The AArch64 file's two consequences check out:
BLOCKER
NOTE
REMOVE
SEND BACK |
…ly meets their exits Round 2 deleted both because #675 and #676 were to land in one batch with this branch and close them. The owner has since ruled that #675 and #676 land on their own reviews, with a nightly run on main to confirm them, so nothing closes these two by this branch's landing. They are restored as round 2 found them, at 90a989e^: when this branch lands, each goes only if main's nightly has met its exit, and otherwise stays, corrected. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…oster::commit refuses a mismatch Round 4 kept the AP's read in a second word, `smp::ECHOED_ID`, stored before the echo and loaded after `await_echo`. Its order rested on a comment that no model drove, and `ROSTER.commit` stayed a public way around the check. - `Roster::echo(token, read)` stores the token and the read as one `AtomicU64`, token in the high half, so the read cannot land after the token. `Roster::echoed` compares the token half. - `Roster::commit` loads that word and refuses a read other than the id it commits the AP under, with round 4's message. Only the boot CPU commits, before the release, so the boot CPU is the one that refuses. - `ECHOED_ID`, `smp::echo` and `smp::commit` go. Both architectures echo `cpu::hardware_id()` with the token and commit through `ROSTER.commit`, so `kernel/src/smp.rs` is main's again. - `kernel-loom/tests/smp_bringup.rs`, `an_ap_that_reads_another_id_is_refused`: an AP echoes 5, the boot CPU awaits that echo and commits the AP as 7, and the model must panic with the refusal. The count model echoes before it commits, since `commit` now reads the echo. - The review's two REMOVEs: the `within` clause in `serial_lock.rs`'s header, and the clause on the read in x86-64 `ap_entry`'s comment. - The AArch64 boot-CPU issue names `Roster::commit`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Guest runs at
Each pair shows the check is what separates a wrong id read from a green boot, for the boot CPU and for an AP. |
|
Round 5 at
The run was in progress when this round began and concluded before this comment. Guest runs at this head are the orchestrator's TCG runs (comment 5937395775). The KVM arm is a nightly on main after landing, by the owner's decision. Net
Round 5 alone: production +33/−45, tests +39/−5, issues +1/−1. Production shrinks by 12. Accepted. Round 4's BLOCKER
Round 4's NOTE on the boot CPU's check is closed: The echo is sound on both architectures, by reading:
The loom model reaches the refusal.
BLOCKER
NOTE
REMOVE
LAND AFTER NAMED CHANGES |
`Roster::echo(token, read)` took the AP's id read from its caller, so two mistakes at either architecture's call site passed every QEMU boot, where cpu N's token and hardware id are both N: the two arguments swapped, and x86-64 echoing `apic::id()` instead of `cpu::hardware_id()`, which would leave `Roster::commit`'s refusal holding nothing. `echo(token)` now reads `crate::arch::cpu::hardware_id()` itself, as the console lock's `fatal_here` does, so neither can be written: there is one argument, and the read sits in code both architectures share. kernel-loom supplies that function as its per-model-thread `arch::cpu` shim, which exists under `loom` alone, so `smp_roster` is now compiled there under `loom` alone, as `serial_lock` is; neither not-loom test names it. The two `smp_bringup` models that echo say which CPU they are with `become_cpu` first: the AP that misreads is CPU 5, and the slot model's echo is CPU 7. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
|
Mutation patches behind the body's rows m1, w1, w2 and g1 (host, at m1 (sha256 diff --git a/kernel/src/smp_roster.rs b/kernel/src/smp_roster.rs
index 7436176..097464f 100644
--- a/kernel/src/smp_roster.rs
+++ b/kernel/src/smp_roster.rs
@@ -140,14 +140,6 @@ impl Roster {
/// the release an AP's own panic stops no other CPU.
pub fn commit(&self, at: Attempt, hardware_id: u32) {
debug_assert!(at.id == self.count.load(Ordering::Relaxed));
- // Relaxed: the read is the echo's own word, which `await_echo` acquired.
- let read = self.echoed.load(Ordering::Relaxed) as u32;
- assert_eq!(
- read,
- hardware_id,
- "smp: cpu{} reads its own hardware id as {read:#x}, and its roster slot and every IPI name it {hardware_id:#x}",
- at.id
- );
self.hardware_ids[at.id as usize].store(hardware_id, Ordering::Relaxed);
// Release: the slot store above lands before the count exposes it; the
// control drops it to relaxed and the model finds the unfilled slot.w1 (sha256 diff --git a/kernel/src/arch/x86_64/smp.rs b/kernel/src/arch/x86_64/smp.rs
index 721f93d..48a9ce9 100644
--- a/kernel/src/arch/x86_64/smp.rs
+++ b/kernel/src/arch/x86_64/smp.rs
@@ -294,7 +294,7 @@ extern "C" fn ap_entry() -> ! {
apic::init_ap();
// Echo this attempt's token, so the BSP counts this AP for its own attempt.
- ROSTER.echo(percpu::ap_token());
+ ROSTER.echo(cpu::hardware_id(), percpu::ap_token());
process::ap_idle();
}w2 (sha256 diff --git a/kernel/src/smp_roster.rs b/kernel/src/smp_roster.rs
index 7436176..db63043 100644
--- a/kernel/src/smp_roster.rs
+++ b/kernel/src/smp_roster.rs
@@ -107,7 +107,7 @@ impl Roster {
/// first interrupt: the token of the attempt that started it, and the
/// hardware id this CPU reads as its own.
pub fn echo(&self, token: u32) {
- let read = crate::arch::cpu::hardware_id();
+ let read = crate::arch::apic::id();
self.echoed.store((u64::from(token) << 32) | u64::from(read), Ordering::Release);
}
g1 (sha256 diff --git a/kernel-loom/src/lib.rs b/kernel-loom/src/lib.rs
index afcade9..fe3cd2e 100644
--- a/kernel-loom/src/lib.rs
+++ b/kernel-loom/src/lib.rs
@@ -170,7 +170,6 @@ pub mod shootdown;
/// The CPU roster and the release/answer word, driven by `tests/smp_bringup.rs`;
/// under `loom` alone, as is the [`arch::cpu`] shim it names.
-#[cfg(feature = "loom")]
#[path = "../../kernel/src/smp_roster.rs"]
pub mod smp_roster;
7a (sha256 diff --git a/kernel/src/arch/x86_64/cpu.rs b/kernel/src/arch/x86_64/cpu.rs
--- a/kernel/src/arch/x86_64/cpu.rs
+++ b/kernel/src/arch/x86_64/cpu.rs
@@ -469,7 +469,7 @@
// SDM Vol. 2A, CPUID leaf 0BH: EBX[15:0] == 0 means unimplemented,
// not id 0, so the leaf-1 fallback below must still run.
if ebx & 0xFFFF != 0 {
- return edx;
+ return edx & 0;
}
}
}7b (sha256 diff --git a/kernel/src/arch/aarch64/cpu.rs b/kernel/src/arch/aarch64/cpu.rs
--- a/kernel/src/arch/aarch64/cpu.rs
+++ b/kernel/src/arch/aarch64/cpu.rs
@@ -122,5 +122,5 @@
let mpidr: u64;
// SAFETY: reads an ID register.
unsafe { asm!("mrs {}, mpidr_el1", out(reg) mpidr, options(nomem, nostack, preserves_flags)) };
- toyos_gicv3::packed_affinity(mpidr)
+ toyos_gicv3::packed_affinity(mpidr) & 0
}7d (sha256 diff --git a/kernel/src/arch/x86_64/cpu.rs b/kernel/src/arch/x86_64/cpu.rs
--- a/kernel/src/arch/x86_64/cpu.rs
+++ b/kernel/src/arch/x86_64/cpu.rs
@@ -469,7 +469,7 @@
// SDM Vol. 2A, CPUID leaf 0BH: EBX[15:0] == 0 means unimplemented,
// not id 0, so the leaf-1 fallback below must still run.
if ebx & 0xFFFF != 0 {
- return edx;
+ return if edx == 0 { 1 } else { edx };
}
}
}7e (sha256 diff --git a/kernel/src/arch/x86_64/cpu.rs b/kernel/src/arch/x86_64/cpu.rs
--- a/kernel/src/arch/x86_64/cpu.rs
+++ b/kernel/src/arch/x86_64/cpu.rs
@@ -469,7 +469,7 @@
// SDM Vol. 2A, CPUID leaf 0BH: EBX[15:0] == 0 means unimplemented,
// not id 0, so the leaf-1 fallback below must still run.
if ebx & 0xFFFF != 0 {
- return edx;
+ return if edx == 0 { 1 } else { edx };
}
}
}
diff --git a/kernel/src/arch/x86_64/smp.rs b/kernel/src/arch/x86_64/smp.rs
--- a/kernel/src/arch/x86_64/smp.rs
+++ b/kernel/src/arch/x86_64/smp.rs
@@ -165,12 +165,6 @@
/// Boot all Application Processors found in the MADT, using `boot_cr3` — the bootloader's identity+high-half PML4 — until each AP switches to the kernel PML4 in `ap_entry`.
pub fn boot_aps(madt: &MadtInfo, boot_cr3: u64) {
let bsp_id = apic::id();
- let read = cpu::hardware_id();
- assert_eq!(
- read,
- bsp_id,
- "smp: the boot CPU reads its own hardware id as {read:#x}, and its roster slot and every IPI name it {bsp_id:#x}"
- );
ROSTER.set_bsp(bsp_id);
copy_trampoline();
The four guest patches are byte-identical to round 5's. |
|
Guest runs at
All as predicted. The named changes of review round 5 are applied: |
No file changed on both sides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…s since #675 The cause the issue named is past: `nested_nmi` wrote through `serial::panic_raw` until #675 (`bc68e5d78`). The exit stays open until CI's KVM `guest` check shows `nested_nmi_is_loud` green. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
…ernel issues CI met Filed, each with its owner, evidence and exit: - a runner image that moves cc, c++ or CMake moves every LLVM key, so every pull request and merge group builds the cold path until main saves the new keys; - a toolchain job that ends inside `--ci bootstrap` saves no layer it built: run 36913380100 built all four, reded in that step and saved none, and run 36934214557 built the LLVM again. Closed on job 110649069031 (run 36934214557, `guest / suite` under KVM on the merge of f1ccb0b into 76d0d93: "test result: ok. 21 passed, 21 total"), which meets both exits as written: - every-aarch64-guest-dies-at-the-kernels-entry-on-cis-firmware: the sixteen `virt_*` tests are green in CI's `guest` check, since #677 enters the AArch64 kernel with the MMU off; - the-nested-nmi-report-interleaves-with-another-cpus-console-line: `nested_nmi_is_loud` is green there, since #675 writes the report under the console's registers, which `nested_nmi`'s own comment says. Neither was cited outside the other. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
Twenty-six landings since 08381fb. Seven files conflicted. Three of them are modify/delete in substance: #660 cut the guest suite to the tests only a booted machine answers and deleted what only the cut tests used, and this branch had modified what it deleted. kernel/src/actuator.rs: main's list, which #660 took 68 actuators out of, `power-refused-once` among them, plus this branch's two, `power-off-spares-the-last-two` and `psci-withheld`. kernel/src/arch/aarch64/irqchip.rs: `Intid::Off` follows `Halt`, and `LogNest`, which #660 deleted, is gone. `Storm` follows it unnumbered, and nothing reads its number. kernel/src/panic_reboot.rs: this branch's lines, through #675's `serial::panic_registers()` in place of `serial::panic_raw`. The arm line's three writes share one hold of the registers. kernel/src/syscall/machine.rs: main's deletion of `refused_once`, and this branch's `power::can_reboot()` and its refusal line. kernel/src/arch/aarch64/power.rs did not conflict and did not build: `serial::panic_raw` and `panic_raw_dec` are gone since #675. Each line is now written under one `panic_registers()` hold, as every fatal path on main writes, and the hold is dropped before a halt, because a CPU halted holding the registers costs every later line the whole bound. `reset` takes the registers only to say a refusal; the call itself still takes no lock. tests/common/power.rs: main's file. Every hunk this branch had there landed in code #660 deleted: - `machine_reboot` generalised into `machine_stops`: main cut the QEMU half of `machine_reboot` for its metal row, so the generalisation has one caller left, `machine_shutdown`, and is written as that. - the `init_power` story removed from `died_and_reset`, the `NO_RESET_HELD` refusal in `panic_before_peripherals_reboots`, and `acpi::reboot` renamed in `usb_reset_hands_devices_back`'s prose: all three tests are cut on main, and the track names them. What is added back is `machine_shutdown`, and `ended`, the wait both it and the virt tests make: the last word, QEMU's `SHUTDOWN` reason, then QEMU's exit. `SHUTTING_DOWN` is declared there once. tests/common/qemu.rs: main's file, plus `BootOptions::psci_trace`. Its refusal of a second `-D` log goes: main deleted `nvme_trace`, the only other one. Three things #660 deleted because only cut tests read them come back, because this branch's tests read them: `QmpShutdown` with `shutdown_reason`, `QemuInstance::await_exit` and `QemuInstance::budget`. tests/toyos.rs: main's file, plus this branch's hunks: - the three `virt_` registrations, without `Sched`, which is gone; - `machine_shutdown` in `MACHINE_TESTS`, with why only QEMU can be asked; `machine_reboot`'s QEMU registration and its `machine_stops` arm stay cut; - `judge_virt_job` by reference, `boot_virt_smp` by `BootOptions`, `virt_smp`'s power-off, `ended_through_psci` over `power::ended`, `psci_calls`, `psci_powered_off`, and the three new virt tests; - the `acpi::shutdown()` comment this branch renamed sat in a test main cut. Everything else merged without conflict. Built at this tree: `cargo run -- --build-only` EXIT=0, the same with `--arch aarch64` EXIT=0 and as the test kernel EXIT=0, and `cargo test --test toyos-build --no-run` EXIT=0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
922 commits of main since 8ee3c51. What main deleted takes the branch's hunks on it with it, and what main changed under the branch is taken as main wrote it. Conflicts, per file: - tests/toyos.rs, tests/common/power.rs, tests/common/faults.rs, tests/common/qemu.rs: main's as written (#660 cut the guest suite to what only a booted machine answers, #625 deleted the tiers). The branch's hunks there had no home: its five `isa_` guest tests and `crash_report_reads_no_kernel_memory` with their helpers (`isa_ports`, `isa_verdict`, `isa_status_byte`, `wait_for_obf`), which stood on the tier tables, `Profile::MetalNoUsb`, `BootOptions::i8042`, `qmp_send_keys`, `logstream::stage` and the `i8042-fault` and `i8042-budget-expired` actuators, all gone; `panic_ignores_keys` and its edits to `panicked()`, `PANIC_REBOOTING` and `await_reset`, whose `panic_reboots` family main cut; and its deletions of `screen_pager_keys`, `screen_paged_scrollback` and `check_no_stale_cells`, which main had already deleted. The commit after this one puts the dropped tests' behaviours on main's tiers. `tests/toyos-rust-tests/src/bin/isa_grant.rs` goes with the harness that gave it its roles, and comes back there. - tests/test-durations: deleted by #639; the branch's two removed rows have no file. - src/sourcegate.rs: `ARCH_RULES` went with #624; the branch's row has no table. - issues/build/assembly-outside-an-arch-module-in-userland-and-guest-probes.md: main rewrote the one line the branch reworded. - kernel/src/sched/driver.rs: `irq_ring`'s `UserDev` arm went with #634, and `isa::drain_pending` with it. The row's handler posts its claim's watch itself (`IrqWatch::post_in_place`), and the first interrupt is announced at the holder's read, both as `pcidev` does on main. - kernel/src/arch/{x86_64,aarch64}/pio.rs: #647 deleted both, their `EXISTS` and `out*` being dead. They come back holding only what the `isa` claim needs of an architecture. - kernel/src/panic_reboot.rs: main's `PANIC_KEYS` goes with the key the panic path no longer reads, and the line it kept for a machine that reads none is the one line now ("the bound is over"); the reset is `power::reset_now` (#647) and the raw writes are under the console's registers (#675). - kernel/src/arch/x86_64/i8042/mod.rs, aarch64/keyboard_controller.rs: main's `poll_byte` and `PANIC_KEYS` go with the pager. - kernel/src/arch/x86_64/idt/mod.rs: `log_nest` went on main. - kernel/src/actuator.rs: main's cut of the i8042's actuators stands. - issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md: main's stage 6 with the branch's stage 7 after it, and main's three "#592's i8042 stage" read "stage 7.2". Stage 7.2 loses its clause about `i8042-trace` and `shell_type_once`, which main deleted. What the merge had to change beside them, to build and to stay true on main: - `kernel/src/device.rs`: `Isa` joins the classes whose `Claimed::Class` drop cancels nothing, a match main added. - The straddle actuator is the i8042 driver's own module (`i8042::straddle`), reached from `isa::claim` through one arch function, and it raises the driver's flood itself at each answered claim: main deleted `i8042-fault`, the flood the branch's test was staged with, and the guest's keys with it. `i8042-withheld` leaves the controller unprobed, where the branch used `i8042-budget-expired`. - `panic-reboot-fast` comes back, which main deleted with the tests that armed it: `screen_fatal_behind_a_painter` proved its CPU watches the reset bound by the pager's second page, and with no pager the reset is the proof. It runs on the profile whose keys reach the i8042 and presses one inside the bound, which is what `panic_ignores_keys` asserted. `Profile::Gop`, which nothing else boots, goes. - `screen_panic_muted` reads the arm line as it is now written. - main's `report_line` and `heartbeat` are gone, so nothing reads the i8042's status port off `IRQ_CPU` and the quarantine's doc says so. - The port grant's pure half is `toyos_userbound::port`: the bitmap as the processor reads it, which rows a switch opens and closes, and the decode of a refused `in` or `out`. A grantable port is a `u8`, so one past the bitmap is no row. `percpu` embeds the bitmap in the TSS and `pio` is what is left of the architecture. - A mint no longer clears the row's record: its release cleared it, and an interrupt taken with no claim on the row is the next holder's to read. - `toyos-tco`'s doc of the panic bound and the guest-suite track's two pager lines named what this branch deletes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UDZQ6fSKw14e4w2TKTRfm
nested_nmi_is_loudfailed under KVM in CI. The cause is in the kernel:nested_nmiwrote its report throughserial::panic_raw, which took no lock, while klogd on another CPU was writing records to the same 16550. On this branch:serial::PanicUart, which it gets only by asking for the console's registers.arch::console_uartfunction that touches the UART takes aserial::Registers, which onlyserial.rsbuilds.Root cause
0edf266ab:59052827fis the merge base of The guest suite is every pull request'sguestcheck, on toolchain stores CI caches by the build system's own keys; a landing during a nightly does not red its release #671's run head09bda2f3c, and the other is06788146b.git log <base>..0edf266abtouches none ofserial.rs,serial_lock.rs,console_uart.rs,nmi.rs,log/,panic.rs,sched/driver.rs, or the harness'sfaults.rs,qemu.rsandserial.rs. The guest suite is every pull request'sguestcheck, on toolchain stores CI caches by the build system's own keys; a landing during a nightly does not red its release #671's own diff touches none of them either.guestcheck, on toolchain stores CI caches by the build system's own keys; a landing during a nightly does not red its release #671'sguestcheck, run 36863809437, job 110375742604, timed out after 39 s.NESTED NMInever arrived whole.guestlane at06788146b, run 36843762360, job 110374194368, ended withQEMU died before NESTED NMI (status 0)after 63 s.[kernel 0.385 cpu1] CPU 1: joining schedulerand the report's[nmi] NESTED NMI on(64 = 44 + 20 bytes). This was checked by an interleaving test against the line as that job's log carries it.drivers/serial.rsholds the registers' lock for one burst, and klogd writes every UART byte under it. The panic path "takes the registers alone, and bypasses them once they stay held".panic_flushdid that. The report did not, and neither didpanic::last_words,percpu::ist1_report,panic_reboot::arm's raw lines orpanic_reboot::reboot_now.Change
The lock (
kernel/src/drivers/serial_lock.rs, which kernel-loom compiles underloom). Its word is free,LIVE(a burst holder), orFATALplus the CPU whose fatal path holds it. The lock reads that CPU itself, throughcrate::arch::cpu::hardware_id. kernel-loom supplies that function as a per-model-thread shim (arch::cpu::become_cpu). No caller passes a CPU id.seize(tries)is the fatal paths' acquire. It answers one of three ways:Taken: the lock was free, or let go inside the bound.Reentered: this CPU's own fatal path already holds it.Expired: another holder kept it through every try.within, the one bounded wait for a console lock (seize, andflush_final's wait for the wire). It asks an attempt that never answers exactlytriestimes. Underfeature = "loom"its pause isloom::hint::spin_loop, imported beside the loom atomics.lock(), a burst's acquire, asserts that the word is not this CPU'sFATAL. A burst beneath its own CPU's fatal hold would wait for a frame that cannot let go while it waits. The review'slast_wordspatch produced exactly that hang: the hold was kept across the record, and before klogd runs the record's inline drain is a burst on the same CPU. Now the assertion's panic ends in the reentry guard's raw report, which names it.kernel-loom builds the lock, the roster and the CPU shim both read under
loomalone. Neither not-loom test (log_zeroed_init,log_body_words) names them. So the shim's not-loom arm, which nosrc/clippy.rsshape built, is deleted rather than given a shape, and without its gate the roster does not build there (g1). The not-loom arms of the lock and the roster are the kernel's, which the kernel's ten shapes lint.The registers' token.
arch::console_uart'sinit,rx_ready,tx_ready,read_byteandwrite_bytetake&mut serial::Registers, on both architectures. AArch64'sinittouches no register; it takes them for the call it shares with the 16550's. AArch64'sframe()answers the frame's address and touches no register.serial::initasks for them withBackendGuard::lock. What the architecture'sinitlogs under it is not drained there:drain_inlinereturns whilehas_consoleis false, which it is untilUART_PRESENTis stored from that call's answer.Registers(())has a private field, so onlyserial.rsbuilds one: atBackendGuard::lock, and inpanic_registers.write_bytes_locked,try_read_byte_lockedandhas_data_lockedtake theBackendGuarditself.Each CPU's id read, at its bring-up (
kernel/src/smp_roster.rs). The lock and the panic path's per-CPU slots (panic.rs) name a CPU byarch::cpu::hardware_id. On x86-64 that is CPUID's answer, while the roster and every IPI use LAPIC ids: the boot CPU's own (apic::id), and the MADT's for the APs.Roster::echo(token)runs on the AP and reads that CPU'sarch::cpu::hardware_iditself, as the lock does. The read and the token are oneAtomicU64store, so the read cannot land after the token, and no caller picks which read the echo carries: the review's swapped arguments do not build (w1), andapic::id()in place of the read builds for x86-64 alone, while the AArch64 kernel and kernel-loom, both compiled by the host gate, refuse it (w2).Roster::commitrefuses an AP whose echoed read is not the id it commits the AP under: the one its roster slot holds and every IPI names it by, the MADT's LAPIC id, or on AArch64 the affinity of the MADT's MPIDR. Both architectures commit only through it.boot_apsholds the boot CPU's own read against its LAPIC id before the roster takes it.ROSTER.set_bsp(me)), which also picks its GIC redistributor (irqchip.rs:184), so nothing independent holds it. That is filed.smp::set_ready, so a wrong read stops the boot. It is not the AP's own: before the release an AP's panic stops no other CPU (stop_other_cpussends nothing then), and the boot CPU boots on without it onceawait_echoruns out.PanicUart(write,hex,dec):panic_registers()is its only constructor. OnExpiredit writes one raw line before anything else is written over the holder.panic_flushdrains through virtio-console only under registers it took clean.Reentered, the outer path may have stopped inside a virtqueue publish, so the flush says so in one raw line and drains raw, as it does afterExpired.nested_nmilets its report's registers go beforehalt_all_cpus, andlast_wordslets them go before its record.nested_nmi_is_loudreads through the halt's flush to the reboot's arm line (panic: rebooting in), its ready marker.faults::UNCLEAN).faults::report.tests/checks.rsnested_nmi_verdictruns that verdict on the host:It also holds the refused words to
kernel/src/drivers/serial.rs.Issues.
issues/panic-path/a-fatal-path-inside-its-own-cpus-console-burst-waits-out-the-bound.mdissues/panic-path/a-panic-write-between-two-console-bursts-can-split-a-character.mdissues/kernel/a-safe-mmio-window-reaches-any-physical-address.mdissues/kernel/a-safe-port-read-reaches-any-io-port.mdissues/panic-path/a-panic-reentry-inside-a-fatal-paths-console-hold-halts-with-it-held.mdissues/kernel/the-aarch64-boot-cpus-id-read-is-held-against-nothing.mdnested_nmi_is_loudin its KVM lane (guest, job 110374194368). Its capture stops in the line before the—of cpu1'si8042:record, and that—lies inside one 16-byte burst: the unlocked report could split it there, and the report on this branch cannot, so the branch closes that instance.issues/panic-path/a-panic-write-between-two-console-bursts-can-split-a-character.md, filed here and left open, is another way fornested_nmi_is_loudto fail under KVM: a report taken between two bursts can still split a character that straddles them.Where each property is held
Unrepresentable. Each of these fails to build (mutations 3a to 3g):
arch::console_uartthat touches the UART without the registers:init,rx_ready,tx_ready,read_byteorwrite_byte, on either architecture;Registers;*_lockedcall made without aBackendGuard.The type does not reach a module that touches the registers by itself. Such a module could use
unsafe { outb(0x3f8, …) }, which breaksoutb's contract that the caller owns the port. It could also use a safeinb(0x3f8), which takes a received byte, or on AArch64 a safeMmio::newoverconsole_uart::frame(). Both are filed.Loom (
kernel-loom/tests/serial_lock.rs, nine models) holds:within's count;seize's bound, and the wait inside it;seize's re-entry, and its exclusion against a burst;The CPU id read is the one inside the lock, and the models drive it from two CPUs.
Loom, bring-up (
kernel-loom/tests/smp_bringup.rs,an_ap_that_reads_another_id_is_refused): an AP whose CPU reads 5 (become_cpu(5)) echoes on its own thread, the boot CPU awaits that echo withRoster::await_echoand commits the AP as 7, and the commit must panic with the refusal. m1 reds it. The read and the token are one store, so there is no second store left to swap.The bounded wait at three preemption bounds. The model of the wait inside the bound runs at preemption bounds 0, 1 and 2, with k + 2 tries at bound k.
seize(_, 2)withTakenin every execution. Under loom's default, unbounded exploration, no try count makesTakenhold in every execution. Measured with two tries:LOOM_MAX_PREEMPTIONS=0exit 0,=1exit 101,=2exit 101.PANIC_LOCK_SPIN_LIMIT, a count and not a time, isissues/panic-path/the-panic-lock-spin-limit-is-a-count-its-comment-calls-a-second.md.Host check:
nested_nmi_verdictholds:whole's clauses alone (4a to 4d);Every boot: each AP's id read against its roster id (7a to 7c), and x86-64's boot CPU's against its LAPIC id (7d, 7e).
Guest:
nmi_entry.halt_all_cpustakes by value must reach all eleven of its callers, and a mint they can all reach,nested_nmican reach a second time, so keeping the registers into the halt would still compile.Metal: none. The T14 has no 16550.
Mutations
Each mutation went through one script:
git apply --check.Where each row was measured:
3b961d5bf: m1, w1, w2 and g1, and the builds of 7a, 7b, 7d and 7e. Every other build and host run atd31faab90;git diff d31faab90 3b961d5bftouches onlyRoster::echoand its two callers, kernel-loom's gate on the roster, the two models that echo, and their doc comments.d31faab90, and row 5 at1a9380fed. Row 5's red stands for the head:git diff 1a9380fed 3b961d5bftouches neitherkernel/src/arch/x86_64/idt/nmi.rsnorkernel/src/drivers/serial.rs.The commands were:
cargo test -p kernel-loom --test serial_lock, and--test smp_bringupfor m1;cargo build -p kernel-loom --lib, with--no-default-featureswhere a row says so, andcargo test -p kernel-loom --test smp_bringup --no-run;cargo test --test toyos-checks nested_nmi_verdict;cargo build --target x86_64-unknown-none --features boot-actuatorsinkernel/, andaarch64-unknown-none-softfloatwhere a row says so; w1 and w2 with default features."(review)" marks the patch a review named.
compile_error!planted in kernel-loom'sarch::cpushimcargo build -p kernel-loom --no-default-features --lib: 0. Without--no-default-features: 101, "planted in the cpu shim"assert!(true);planted in the shim'sbecome_cpucargo clippy -p kernel-loom --all-targetswith the workspace shape's lint arguments: 101,clippy::assertions_on_constantsseize'sErr(_) => None→Err(_) => Some(Seized::Expired)(review)a_seize_takes_a_burst_let_go_inside_the_boundwithin's0..tries→0..tries.min(1)(review)within_asks_a_silent_attempt_exactly_its_tries("a wait bounded at 100 tries asked 1 times")within's pause iscore::hint::spin_loop, not loom'sa_seize_takes_a_burst_let_go_inside_the_boundwithin's0..tries→0..tries.min(4)(review)within_asks_a_silent_attempt_exactly_its_tries("a wait bounded at 100 tries asked 4 times"). With that model skipped, the other eight pass: 0crate::arch::console_uart::write_byte(b'!');at the head ofnested_nmierror[E0061], "argument #1 of type&mut serial::Registersis missing"&mut crate::drivers::serial::Registers(())error[E0603], "tuple struct constructorRegistersis private"crate::drivers::virtio_console::write_bytes_locked(b"!");thereerror[E0061], "argument #1 of type&mut BackendGuardis missing"crate::arch::console_uart::rx_ready();thereerror[E0061], "argument #1 of type&mut serial::Registersis missing"crate::arch::console_uart::tx_ready();therecrate::arch::console_uart::init(0);therecrate::arch::console_uart::tx_ready();at the head of AArch64'sap_entry, built foraarch64-unknown-none-softfloatwhole's body →report.len() == 3(review)whole's first-line clause droppedwhole's registers clause droppedwhole's last-line clause droppedlock()'s refusal deleteda_burst_beneath_its_own_cpus_fatal_path_is_refused, with loom's "Model exceeded maximum number of branches" in place of the refusallast_words' block braces dropped, so its registers live across the record (review)panic.rs, and no test boots a DOUBLE PANIC. By reading, the hang it made is now 5a's refusal:alert!→log::emit→drain_inline→write_wire→BackendGuard::lock→BackendLock::lockon this CPU'sFATALword.crate::arch::cpu::hardware_id()→0(review)a_seize_reenters_its_own_cpus_hold("a fatal path took another CPU's hold") anda_burst_waits_out_another_cpus_fatal_path(the burst is refused on CPU 2)cpu::hardware_id'sreturn edx;→return edx & 0;(review; its literalreturn 0;fails-Dwarningson the unusededx, build 101)nested_nmi_is_loud: 0/1, the boot CPU's panic, "assertionleft == rightfailed: smp: cpu1 reads its own hardware id as 0x0, and its roster slot and every IPI name it 0x1"cpu::hardware_id'spacked_affinity(mpidr)→packed_affinity(mpidr) & 0, built foraarch64-unknown-none-softfloatvirt_smp: 0/1, the same panicRoster::commitat the headnested_nmi_is_loud: 1/1 greencpu::hardware_id'sreturn edx;→return if edx == 0 { 1 } else { edx };, so QEMU's boot CPU alone reads 1 (review)nested_nmi_is_loud: 0/1, the boot CPU's panic inboot_aps, "assertionleft == rightfailed: smp: the boot CPU reads its own hardware id as 0x1, and its roster slot and every IPI name it 0x0"boot_aps(review)nested_nmi_is_loud: 1/1 greenRoster::commit(review)an_ap_that_reads_another_id_is_refused, "test did not panic as expected"; the other twosmp_bringupmodels passap_entryechoes with the review's swapped arguments,ROSTER.echo(cpu::hardware_id(), percpu::ap_token())error[E0061], "this method takes 1 argument but 2 arguments were supplied"Roster::echoreadscrate::arch::apic::id(), the review's other mistakesmp_bringup: 101, eacherror[E0433], "cannot findapicinarch"smp_rosterwithout its#[cfg(feature = "loom")]--no-default-features: 101,error[E0433], "cannot findcpuinarch"; default features: 0seize→return Seized::Taken(self.lock());(round 2's row)a_burst_beneath_its_own_cpus_fatal_path_is_refused,a_seize_gives_up_on_a_holder_that_never_lets_go,a_seize_reenters_its_own_cpus_holdwithin'sfor _ in 0..tries→loopa_seize_gives_up_on_a_holder_that_never_lets_go,a_seize_reenters_its_own_cpus_hold,within_asks_a_silent_attempt_exactly_its_triesseizeloses itsReenteredarma_seize_reenters_its_own_cpus_hold("a fatal path waited out its own CPU's hold")DRAINED_RAWdrifts (owndropped)serial.rs"writes no" the refused linenested_nmi's block braces dropped, so its registers live intohalt_all_cpus(sha256f718ce25…)nested_nmi_is_loud: 0/1, "the report's registers were not clean: [serial] this cpu's own fatal path held the console registers; drained raw"Gates
At
3b961d5bf:cargo run -- --ci host: EXIT=0,[ci] Host: 56 step(s), all green.aarch64-unknown-none-softfloat.serial_lock(9 passed) andsmp_bringup(3 passed).toyos-checks:nested_nmi_verdictok.log_zeroed_init,log_body_words): 3 passed.serial-try-lock-then-somecontrol (2 verdicts reached), andsmp_bringup'sroster-commit-relaxedandsmp-ready-split(1 verdict each).cargo run -- --clippy: EXIT=0,clippy: 17 invocations clean.cargo run -- --build-only: EXIT=0; with--arch aarch64: EXIT=0.cargo test -p kernel-loom, every model: EXIT=0.Guest, the orchestrator's runs under TCG: at
d31faab90, the whole suite 21/21; 7a 0/1, 7b 0/1 and 7d 0/1, each on its predicted line; 7c 1/1 and 7e 1/1. At1a9380fed, row 5 0/1 on its predicted line.Device driver and a lock on the panic path: the two checks
Roster::commit.seize,lockand a burst within the bounds stated above.arch::cpu::hardware_iddiffering on every CPU. On x86-64 that is the x2APIC id CPUID leaf 0BH or 1FH reports in EDX for the current logical processor, which the SDM gives as the one the local APIC's ID register holds (leaf 1's initial APIC id on a CPU without those leaves); on AArch64, MPIDR's affinity. Each AP's read is now held against the id it was started by, and on x86-64 the boot CPU's against its LAPIC id. No test puts two CPUs on fatal paths at once.tcglane. The fix's KVM arm is a nightly the orchestrator dispatches on main after this lands.Unsure
toyos-cpuvuln/fixtures/t14/cpuinfo.txt: processors 0 to 7 giveapicidandinitial apicidalike as 0, 2, 4, 6, 1, 3, 5, 7, the T14's MADT order (MADT cpus=[0, 2, 4, 6, 1, 3, 5, 7],issues/hardware/a-metal-session-runs-a-pre-flash-gate-first.md:62).last_wordspatch now does is established by reading, not by a run: no test boots a DOUBLE PANIC, and I added none. A DOUBLE PANIC needs a guest. The refusal such a panic would end in is held by loom (5a).panic_registerson that path waits out the bound again. This is filed.halt_all_cpushas stopped the other CPUs, halts its CPU with the hold kept. This is filed. The refusal is one more way in: alog!under a hold before klogd runs.🤖 Generated with Claude Code
https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L