From 24334563feabc88ddb5a9d7164b58671a299ccab Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 22:47:16 +0200 Subject: [PATCH 1/6] AArch64 stage 5: every CPU the MADT names starts through PSCI CPU_ON Stage 5 of issues/kernel/toyos-runs-on-arm64.md, its CPU half. It is the stage the small-kernel ordering leaves open: it ports no device interrupt (the kernel takes SGIs and the timer's PPI only), so nothing of the relay small-kernel stage 6 deletes, and SMP is what the track's exit, LLVM timed against a Linux guest, cannot do without. AArch64: - psci.rs: the conduit the FADT's ARM_BOOT_ARCH names (SMC, or HVC under a hypervisor), PSCI_VERSION refused below 0.2, and the SMC64 CPU_ON. - boot.rs: the declaration is one routine, apply_declaration, that the boot CPU's _start and every AP's ap_start branch to with a root in x3 and a link-address continuation in x22; the EL2 drop and the EL1 path are the same instructions they were, so every CPU applies the declaration the boot CPU does. - smp.rs: the boot CPU starts each enabled GIC CPU interface in MADT order, each AP given a leaked start block (cleaned to the point of coherency, since it reads it with its MMU off), a bring-up stack, its per-CPU block and its redistributor frame; an AP that does not echo within AP_START ends the bring-up, as on x86-64. ap_entry installs TPIDR_EL1 first, then the kernel's tables, the declaration's check, its GIC and timer, and echoes. - paging.rs: bringup_root, the kernel root's direct map again at identity, for the few instructions an AP runs at its physical address after its MMU is on; join, the one switch onto the kernel's tables, used by the boot CPU and every AP. - irqchip.rs: init is the distributor and answers the MADT's CPUs and redistributor ranges; init_cpu is one CPU's redistributor, CPU interface and timer. kick_cpu targets any online CPU through the roster; the halt SGI goes to every other CPU with IRM, and its handler never ends it, so the running priority it keeps masks everything and the WFI stays. - percpu.rs: alloc builds any CPU's block on the boot CPU, install puts it in TPIDR_EL1. - control_regs.rs: check takes the level entered at and counts; report says how many of the roster hold the declaration. Shared, one copy for both architectures: - smp_roster.rs: hardware ids rather than LAPIC ids, and the echo word and its budgeted wait, which were x86-64's AP_STARTED. - process::ap_idle: the wait for release, the halt of an uncommitted AP, the TLB join, the hard-lockup arm and the scheduler join, which were x86-64's. - log::shard_for: an AP's shard, allocated and published, which was x86-64's. - time::AP_START: the budget both wait with. Pure crates: toyos-acpi decodes ARM_BOOT_ARCH only from a FADT of 5.1 or later, with a crafted corpus and the q35 fixture (revision 3, undefined); toyos-gicv3 encodes ICC_SGI1R_EL1 with IRM. Tests: virt_smp boots tests/virtsmpcase on eight CPUs under the EL2 profile (PSCI through SMC) and asserts every CPU online, the declaration held by 8 of 8, every AP joining the scheduler, and the toybox applet unmap_seen: a page unmapped while a second thread reads it must end the process each of four times. virt_job pins one CPU: preempt and fp_isolation see a sibling run only when it took theirs. The track loses its Parked line, which the owner's running tracks in parallel supersede, gains stage 5's state and what it still owes, and its exit says the LLVM timing is one measurement taken by hand. Co-Authored-By: Claude Opus 5.5 --- issues/kernel/toyos-runs-on-arm64.md | 24 +++- kernel-loom/tests/smp_bringup.rs | 12 +- kernel/src/arch/aarch64/boot.rs | 149 +++++++++++++++-------- kernel/src/arch/aarch64/control_regs.rs | 21 +++- kernel/src/arch/aarch64/irqchip.rs | 154 ++++++++++++++---------- kernel/src/arch/aarch64/mod.rs | 1 + kernel/src/arch/aarch64/paging.rs | 49 ++++++-- kernel/src/arch/aarch64/percpu.rs | 32 +++-- kernel/src/arch/aarch64/psci.rs | 145 ++++++++++++++++++++++ kernel/src/arch/aarch64/smp.rs | 102 +++++++++++++++- kernel/src/arch/aarch64/tlb.rs | 4 + kernel/src/arch/aarch64/trap.rs | 3 + kernel/src/arch/x86_64/percpu.rs | 19 +-- kernel/src/arch/x86_64/smp.rs | 57 +-------- kernel/src/log/mod.rs | 24 +++- kernel/src/process.rs | 25 +++- kernel/src/smp_roster.rs | 57 +++++++-- kernel/src/time.rs | 6 + src/build.rs | 1 + tests/toyos.rs | 63 +++++++++- tests/virtsmpcase/system.toml | 19 +++ toyos-acpi/src/fadt.rs | 34 ++++++ toyos-acpi/src/lib.rs | 4 +- toyos-acpi/tests/corpus.rs | 46 ++++++- toyos-acpi/tests/fixtures.rs | 12 +- toyos-gicv3/src/lib.rs | 7 ++ toyos-gicv3/tests/gicv3.rs | 10 +- userland/toybox/src/main.rs | 3 +- userland/toybox/src/unmap_seen.rs | 89 ++++++++++++++ 29 files changed, 921 insertions(+), 251 deletions(-) create mode 100644 kernel/src/arch/aarch64/psci.rs create mode 100644 tests/virtsmpcase/system.toml create mode 100644 userland/toybox/src/unmap_seen.rs diff --git a/issues/kernel/toyos-runs-on-arm64.md b/issues/kernel/toyos-runs-on-arm64.md index 6ca1f8f3195..44531bf91a3 100644 --- a/issues/kernel/toyos-runs-on-arm64.md +++ b/issues/kernel/toyos-runs-on-arm64.md @@ -28,9 +28,6 @@ Every x86 guest on this host runs under TCG emulation instead — there is no ## Owner rulings, 2026-09-26 -- **Parked** (owner ruling, 2026-09-29). No later stage starts now. The Exit - (LLVM under emulation) needs stages 4-7; until they are - picked up it is unmet. - **Hardware discovery is ACPI only** (edk2 MADT, GTDT and SPCR under QEMU); no devicetree. - **Start.** Stage 0 (shared groundwork on x86) and stages 1-3 (toolchain, loader, kernel reaching serial on QEMU `virt`) start now. Stage 0 waits for @@ -295,6 +292,22 @@ Each stage names its exit; "measured" means a number from a run. `dlopen` test; until it does the kernel refuses `R_AARCH64_TLSDESC` by name (`toyos_elf::rela::ExeRefusal::TlsDescriptor` for an executable, `toyos_elf::RelocError::TlsDescriptor` for a library). + **Every CPU starts, ahead of small-kernel stage 6 as stage 4 did:** it + ports no device interrupt, so nothing of the relay. The boot CPU starts + each GIC CPU interface the MADT enables with `CPU_ON`, through the conduit + the FADT's `ARM_BOOT_ARCH` names; each AP applies the declaration through + the boot CPU's own routine, installs its per-CPU block, redistributor and + timer, and echoes its attempt's token into the roster both architectures + share before it joins the scheduler. SGIs are the kick and the halt. + `virt_smp` judges eight CPUs, each holding the declaration, and a page + unmapped beside a thread still reading it. Owed before the exit holds: + `SYSTEM_RESET`, `SYSTEM_OFF` and `CPU_OFF` behind a reset and power-off + seam that takes x86-64's reset register and PM1a out of + `drivers/acpi.rs`, and the stop shown on eight CPUs ending in that + power-off; a test of the halt SGI; the TLS-descriptor resolver; + `issues/kernel/the-crash-evidence-records-x86-fault-registers.md`; and, + for the first HVF run, the clean of an AP's start block to the point of + coherency, which TCG cannot fail on. 6. **Virtio on `virt`.** virtio-pci (ECAM from MCFG) for blk, net, gpu, sound, input and rng. virtio-input replaces the i8042 as the @@ -330,8 +343,9 @@ ARM is done when LLVM compiles under emulation on the development Mac (owner ruling, 2026-09-29). One recipe is timed three ways: macOS natively, a Linux arm64 guest under QEMU with HVF, and ToyOS arm64 under QEMU with HVF. All three times are recorded, and it passes when ToyOS is at least as fast as the -Linux guest. ToyOS compiling LLVM is `issues/build/toyos-builds-itself.md`'s -work on AArch64. +Linux guest. The timing is one measurement, taken by hand and not automated +(owner ruling, 2026-09-29). ToyOS compiling LLVM is +`issues/build/toyos-builds-itself.md`'s work on AArch64. ## Interactions with other tracks diff --git a/kernel-loom/tests/smp_bringup.rs b/kernel-loom/tests/smp_bringup.rs index 28b99566442..d89aa87a63f 100644 --- a/kernel-loom/tests/smp_bringup.rs +++ b/kernel-loom/tests/smp_bringup.rs @@ -12,8 +12,8 @@ use std::sync::atomic::{AtomicBool, Ordering::SeqCst}; use kernel_loom::smp_roster::Roster; use loom::sync::Arc; -/// A LAPIC id no uncommitted slot holds (the roster fills a slot with `u32::MAX`). -const LAPIC: u32 = 7; +/// A hardware id no uncommitted slot holds (the roster fills a slot with `u32::MAX`). +const HARDWARE_ID: u32 = 7; /// Bounded for loom (an unbounded spin never finishes); each turn a scheduling point. const POLLS: usize = 4; @@ -32,16 +32,16 @@ fn a_committed_count_never_outruns_its_slot() { let committer = { let r = r.clone(); - loom::thread::spawn(move || r.commit(attempt, LAPIC)) + loom::thread::spawn(move || r.commit(attempt, HARDWARE_ID)) }; let reader = loom::thread::spawn(move || { if r.count() >= 2 { - let got = r.apic_id(1); + let got = r.hardware_id(1); SAW.store(true, SeqCst); assert_eq!( - got, LAPIC, - "a reader saw cpu1 counted while its LAPIC slot was still unfilled", + got, HARDWARE_ID, + "a reader saw cpu1 counted while its hardware-id slot was still unfilled", ); } }); diff --git a/kernel/src/arch/aarch64/boot.rs b/kernel/src/arch/aarch64/boot.rs index cf6dd35e31a..bdd84b1aa01 100644 --- a/kernel/src/arch/aarch64/boot.rs +++ b/kernel/src/arch/aarch64/boot.rs @@ -1,18 +1,21 @@ -//! The AArch64 steps of the boot: the entry the loader jumps to, and what -//! `kernel_main` asks of this architecture at the points where one differs -//! from another. +//! The AArch64 steps of the boot: the entries firmware and the loader jump +//! to, and what `kernel_main` asks of this architecture at the points where +//! one differs from another. //! -//! **The entry.** The loader leaves the CPU as firmware ran it — at EL2 or -//! EL1, on firmware's identity tables — cleans the kernel image and its own +//! **The entries.** The loader leaves the boot CPU as firmware ran it — at EL2 +//! or EL1, on firmware's identity tables — cleans the kernel image and its own //! tables to the point of coherency, and jumps to [`_start`]'s physical -//! address with `x0 = &KernelArgs`. `_start` writes the +//! address with `x0 = &KernelArgs`. PSCI starts every other CPU at +//! [`ap_start`]'s physical address, MMU off, at the level firmware gives an +//! operating system, with `x0` its [`ApStart`]'s physical address. Each names +//! a root table and branches to [`apply_declaration`], which writes the //! [`control_regs`](super::control_regs) declaration whole: at EL2 it writes //! `HCR_EL2` first, halts in a named refusal unless it reads back as declared, -//! then programs EL1's registers with the MMU already on and drops with `ERET`; at EL1 it -//! turns the MMU off first, so no translation register changes under a live -//! walk. Either way it arrives at the kernel's link address in the view at -//! `PHYS_OFFSET`, on the kernel's own stack, with the vectors installed, and -//! calls `kernel_main`. +//! then programs EL1's registers with the MMU already on and drops with +//! `ERET`; at EL1 it turns the MMU off first, so no translation register +//! changes under a live walk. Either way the entry resumes at its link address +//! in the view at `PHYS_OFFSET`, installs the vectors on its own stack, and +//! calls `kernel_main` or `smp::ap_entry`. use core::mem::offset_of; @@ -20,6 +23,7 @@ use toyos_abi::boot::{KernelArgs, MemoryMapEntry}; use toyos_acpi::MadtEntry; use super::control_regs as regs; +use super::smp::ApStart; use crate::drivers::acpi::direct_phys; use crate::log; use crate::mm::Region; @@ -39,12 +43,80 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! { "add x20, x20, x1", "ldr x1, [x19, #{stack_size}]", "add x20, x20, x1", - // x2 = TCR_EL1 whole, x3 = the loader's L0 table, x4 = MAIR_EL1. + "ldr x3, [x19, #{root}]", + "ldr x22, =1f", + "b {declare}", + "1:", + "ldr x1, ={phys_offset}", + "add x20, x20, x1", + "mov sp, x20", + "add x19, x19, x1", + "adrp x1, {entry_el}", + "str x21, [x1, :lo12:{entry_el}]", + "bl {install}", + "mov x0, x19", + "mov x29, xzr", + "mov x30, xzr", + "bl {kernel_main}", + kernel_memory = const offset_of!(KernelArgs, kernel_memory_addr), + stack_offset = const offset_of!(KernelArgs, kernel_stack_addr), + stack_size = const offset_of!(KernelArgs, kernel_stack_size), + root = const offset_of!(KernelArgs, boot_pml4_addr), + declare = sym apply_declaration, + phys_offset = const crate::PHYS_OFFSET, + entry_el = sym regs::ENTRY_EL, + install = sym super::trap::install, + kernel_main = sym crate::kernel_main, + ); +} + +/// Where PSCI `CPU_ON` starts every other CPU, at its physical address with +/// its MMU off and `x0` the physical address of its [`ApStart`]. +/// # Safety +/// Only firmware may enter this, as `super::smp::start` asked it to. +#[unsafe(naked)] +pub(super) unsafe extern "C" fn ap_start() -> ! { + core::arch::naked_asm!( + "mov x19, x0", + "ldr x20, [x19, #{stack_top}]", + "ldr x3, [x19, #{root}]", + "ldr x22, =1f", + "b {declare}", + "1:", + "mov sp, x20", + "ldr x1, ={phys_offset}", + "add x19, x19, x1", + "bl {install}", + "mov x0, x19", + "mov x1, x21", + "mov x29, xzr", + "mov x30, xzr", + "bl {ap_entry}", + stack_top = const offset_of!(ApStart, stack_top), + root = const offset_of!(ApStart, root), + declare = sym apply_declaration, + phys_offset = const crate::PHYS_OFFSET, + install = sym super::trap::install, + ap_entry = sym super::smp::ap_entry, + ); +} + +/// The declaration written on this CPU and its MMU turned on under the root +/// in `x3`, whichever level it was entered at; then a branch to the link +/// address in `x22` at EL1 on `SP_EL1`, with `x21` the level entered at. +/// Keeps `x19`, `x20` and `x22`. Entered by a branch from an entry running at +/// its physical address, and never returns. +/// # Safety +/// Only [`_start`] and [`ap_start`] branch here, with `x3` a root that maps the +/// kernel image at both its physical and its link address. +#[unsafe(naked)] +unsafe extern "C" fn apply_declaration() -> ! { + core::arch::naked_asm!( + // x2 = TCR_EL1 whole, x4 = MAIR_EL1, x5 = SCTLR_EL1. "ldr x2, ={tcr}", "mrs x1, id_aa64mmfr0_el1", "and x1, x1, #0xf", "orr x2, x2, x1, lsl #{ips}", - "ldr x3, [x19, #{root}]", "ldr x4, ={mair}", "ldr x5, ={sctlr}", "mrs x21, CurrentEL", @@ -70,8 +142,7 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! { "isb", "msr sctlr_el1, x5", "isb", - "ldr x1, =3f", - "br x1", + "br x22", // EL2: `HCR_EL2` first, since with `E2H` set every `_el1` name // below is an EL2 register; refused unless it reads back as declared. // Then EL1's registers with the MMU on, EL2's others, and the drop. @@ -83,7 +154,7 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! { "cmp x6, x1", "b.ne {refuse_hcr}", // `ICC_SRE_EL2`, whose `Enable` lets EL1 write its own `ICC_SRE_EL1` - // (`super::irqchip::init`). Skipped where `ID_AA64PFR0_EL1.GIC` names + // (`super::irqchip::init_cpu`). Skipped where `ID_AA64PFR0_EL1.GIC` names // no system-register interface: there the write is an undefined // instruction under firmware's vectors and ends the boot silently, // while EL1's own access fails under this kernel's, which report it. @@ -112,31 +183,13 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! { "dsb nsh", "mov x1, #{spsr}", "msr spsr_el2, x1", - "ldr x1, =3f", - "msr elr_el2, x1", + "msr elr_el2, x22", "isb", "eret", - // At the link address, EL1, MMU on. - "3:", - "ldr x1, ={phys_offset}", - "add x20, x20, x1", - "mov sp, x20", - "add x19, x19, x1", - "adrp x1, {entry_el}", - "str x21, [x1, :lo12:{entry_el}]", - "bl {install}", - "mov x0, x19", - "mov x29, xzr", - "mov x30, xzr", - "bl {kernel_main}", // EL3, or anything else: nothing here may run there. "9:", "wfe", "b 9b", - kernel_memory = const offset_of!(KernelArgs, kernel_memory_addr), - stack_offset = const offset_of!(KernelArgs, kernel_stack_addr), - stack_size = const offset_of!(KernelArgs, kernel_stack_size), - root = const offset_of!(KernelArgs, boot_pml4_addr), tcr = const regs::TCR, ips = const regs::TCR_IPS_SHIFT, mair = const regs::MAIR, @@ -150,10 +203,6 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! { cntkctl = const regs::CNTKCTL, icc_sre_el2 = const regs::ICC_SRE_EL2, spsr = const regs::SPSR_EL2_TO_EL1, - phys_offset = const crate::PHYS_OFFSET, - entry_el = sym regs::ENTRY_EL, - install = sym super::trap::install, - kernel_main = sym crate::kernel_main, ); } @@ -183,7 +232,7 @@ pub fn before_panel() {} /// and what this boot found — the memory map and the tables the rest of the /// port reads. pub fn after_console(args: &KernelArgs, maps: &[MemoryMapEntry]) { - regs::check(); + regs::check(regs::ENTRY_EL.load(core::sync::atomic::Ordering::Relaxed)); for entry in maps { log!("memory: {:#014x}..{:#014x} uefi type {}", entry.start, entry.end, entry.uefi_type); } @@ -259,17 +308,20 @@ pub fn reserved() -> Region { Region { start: 0, end: 0 } } -/// What the boot learns bringing interrupts up and hands later steps: nothing -/// yet, since the other CPUs the MADT names are the port's stage 5's to read. -pub struct Platform; +/// What the boot learns bringing interrupts up and hands later steps: the +/// other CPUs the MADT names, and how to start them. +pub struct Platform { + gic: super::irqchip::Gic, + psci: Option, +} /// Interrupt delivery: this CPU's per-CPU block, the GIC and the timer's /// interrupt, and interrupts unmasked. The syscall gate is the vectors' own. pub fn interrupts(rsdp_addr: u64) -> Platform { super::percpu::init_bsp(); - super::irqchip::init(rsdp_addr); + let gic = super::irqchip::init(rsdp_addr); super::cpu::enable_interrupts(); - Platform + Platform { gic, psci: super::psci::init(rsdp_addr) } } /// The clock: the generic timer's count, at the rate firmware states in @@ -295,10 +347,9 @@ pub fn timer() { /// drives on an ACPI Arm machine. pub fn platform_devices(_rsdp_addr: u64) {} -/// Every other CPU, running: the port's stage 5, which starts each with PSCI -/// `CPU_ON`. Until then the boot CPU runs alone. -pub fn start_other_cpus(_platform: &Platform, _args: &KernelArgs) { - log!("smp: the boot CPU runs alone; the other CPUs are the port's stage 5 (PSCI CPU_ON)"); +/// Every other CPU, running. +pub fn start_other_cpus(platform: &Platform, _args: &KernelArgs) { + super::smp::start(&platform.gic, platform.psci); } /// The interrupt-controller selftests an actuator asks for. diff --git a/kernel/src/arch/aarch64/control_regs.rs b/kernel/src/arch/aarch64/control_regs.rs index 6251ba1bcdd..2940a56d268 100644 --- a/kernel/src/arch/aarch64/control_regs.rs +++ b/kernel/src/arch/aarch64/control_regs.rs @@ -9,7 +9,7 @@ //! and the GIC architecture specification (IHI 0069H), chapter 12, for the //! `ICC_` registers. -use core::sync::atomic::{AtomicU64, Ordering}; +use core::sync::atomic::{AtomicU32, AtomicU64, Ordering}; use crate::log; @@ -102,9 +102,12 @@ pub fn tcr() -> u64 { TCR | (read!("id_aa64mmfr0_el1") & 0xF) << TCR_IPS_SHIFT } +/// CPUs whose registers [`check`] found as declared. +static CHECKED: AtomicU32 = AtomicU32::new(0); + /// Read every EL1 register the declaration names back, and refuse a CPU that -/// holds anything else; then say what it holds. -pub fn check() { +/// holds anything else; then say what it holds, and that it was entered at `el`. +pub fn check(el: u64) { // `ID_AA64MMFR0_EL1.ASIDBits` = 0b0010: without 16-bit ASIDs `TCR_EL1.AS` // is RES0, and the read-back below would name the register, not the reason. let asid_bits = read!("id_aa64mmfr0_el1") >> 4 & 0xF; @@ -120,9 +123,9 @@ pub fn check() { assert_eq!(live, value, "control registers: {name} holds {live:#x}, and the declaration says {value:#x}"); } // What the drop from EL2 left, or the EL1 entry kept: EL1, on `SP_EL1`. - let (el, spsel) = (read!("CurrentEL") >> 2 & 3, read!("SPSel") & 1); - assert_eq!((el, spsel), (1, 1), "control registers: running at EL{el} on SP_EL{spsel}, not EL1 on SP_EL1"); - let el = ENTRY_EL.load(Ordering::Relaxed); + let (now, spsel) = (read!("CurrentEL") >> 2 & 3, read!("SPSel") & 1); + assert_eq!((now, spsel), (1, 1), "control registers: running at EL{now} on SP_EL{spsel}, not EL1 on SP_EL1"); + CHECKED.fetch_add(1, Ordering::Relaxed); log!( "control registers: SCTLR_EL1={SCTLR:#x} TCR_EL1={:#x} MAIR_EL1={:#x} CPACR_EL1={CPACR:#x} \ CNTKCTL_EL1={CNTKCTL:#x}, as declared; entered at EL{el}{}", @@ -135,3 +138,9 @@ pub fn check() { }, ); } + +/// How many CPUs hold the declaration, said once after the last was started: a +/// CPU that never reached [`check`] shows as a count below the roster's. +pub fn report(cpus: u32) { + log!("control registers: {} of {cpus} CPUs hold the declaration", CHECKED.load(Ordering::Relaxed)); +} diff --git a/kernel/src/arch/aarch64/irqchip.rs b/kernel/src/arch/aarch64/irqchip.rs index 676f0e8050d..a7731e738d3 100644 --- a/kernel/src/arch/aarch64/irqchip.rs +++ b/kernel/src/arch/aarch64/irqchip.rs @@ -16,9 +16,11 @@ //! zero by the entry from EL2 the two count alike. It is level-triggered, so a //! handler that neither re-arms nor stops it takes it again at once. -use core::sync::atomic::{AtomicU32, AtomicU64, Ordering::Relaxed}; +use core::sync::atomic::{AtomicBool, AtomicU32, Ordering::Relaxed}; -use toyos_acpi::MadtEntry; +use alloc::vec::Vec; + +use toyos_acpi::{Gicc, MadtEntry}; use toyos_gicv3::{packed_affinity, FRAME}; use super::{cpu, percpu}; @@ -36,6 +38,8 @@ pub(super) enum Intid { /// Asks a CPU for a scheduler pass: x86-64's kick, which rides the /// timer's vector there. Kick = 0, + /// Stops a CPU for good: [`stop_other_cpus`]'s. + Halt, /// `crate::log::nested`'s delivery, which [`send_self`] raises. LogNest, /// What `irq-storm` floods this CPU with. @@ -45,6 +49,7 @@ pub(super) enum Intid { } pub(super) const SGI_KICK: u32 = Intid::Kick as u32; +pub(super) const SGI_HALT: u32 = Intid::Halt as u32; #[cfg(feature = "boot-actuators")] pub(super) const SGI_STORM: u32 = Intid::Storm as u32; @@ -82,10 +87,10 @@ const PRIORITY_MASK: u64 = 0xF0; /// What the GIC answers `ICC_IAR1_EL1` with when nothing is pending for it. const SPURIOUS: u32 = 1023; -/// The timer's PPI, as the GTDT names it; zero until [`init`]. +/// The timer's PPI, as the GTDT names it, and whether it is edge-triggered; +/// zero until [`init`]. static TIMER_INTID: AtomicU32 = AtomicU32::new(0); -/// The distributor's frame, for [`log_state`]; zero until [`init`]. -static GICD: AtomicU64 = AtomicU64::new(0); +static TIMER_EDGE: AtomicBool = AtomicBool::new(false); /// Wait until `done`, for at most `1 / per_second` of a second counted at the /// rate firmware states: the boot has no calibrated clock yet. @@ -105,46 +110,58 @@ fn read_sysreg_pmr() -> u64 { v } -/// Bring the distributor, this CPU's redistributor, its CPU interface and its -/// timer up: every SGI and the timer's PPI enabled at [`PRIORITY`], the timer -/// stopped. Interrupts stay masked at `DAIF`; the caller unmasks them. -pub fn init(rsdp_addr: u64) { +/// What the MADT says about the machine's GIC that another CPU's bring-up needs. +pub struct Gic { + /// Every GIC CPU interface the MADT names and enables, in its order. + pub cpus: Vec, + /// Each redistributor range's physical base, and the range, mapped whole. + ranges: Vec<(u64, Mmio)>, +} + +impl Gic { + /// The physical frame of the redistributor serving the CPU whose packed + /// affinity is `affinity`, mapped: the one its GIC CPU interface names, or + /// the one in a range whose `GICR_TYPER` says it is that CPU's. + pub fn redistributor(&self, affinity: u32) -> u64 { + let named = self.cpus.iter().find(|gicc| packed_affinity(gicc.mpidr) == affinity && gicc.gicr_base != 0); + if let Some(gicc) = named { + crate::mm::paging::map_mmio(gicc.gicr_base, 2 * FRAME, MmioPolicy::Uncacheable); + return gicc.gicr_base; + } + self.ranges + .iter() + .find_map(|&(base, range)| { + toyos_gicv3::find_redistributor(range.size(), affinity, |at| range.read_u64(at + GICR_TYPER)) + .map(|offset| base + offset) + }) + .unwrap_or_else(|| panic!("GIC: no redistributor answers for MPIDR affinity {affinity:#x}")) + } +} + +/// Bring the distributor up, then this CPU's side of the GIC ([`init_cpu`]), +/// and answer what the other CPUs' bring-up needs. Interrupts stay masked at +/// `DAIF`; the caller unmasks them. +pub fn init(rsdp_addr: u64) -> Gic { let madt = toyos_acpi::find_table(direct_phys(), rsdp_addr, b"APIC", toyos_acpi::MADT_ENTRIES) .unwrap_or_else(|e| panic!("GIC: the MADT is unusable: {e:?}")); - let me = cpu::hardware_id(); - let (mut gicd, mut own_frame, mut ranges) = (None, None, [(0u64, 0u32); 4]); - let mut range_count = 0; + let (mut gicd, mut ranges, mut cpus) = (None, Vec::new(), Vec::new()); for entry in toyos_acpi::madt_entries(&madt) { match entry { Ok(MadtEntry::Gicd { base, .. }) => gicd = Some(base), - Ok(MadtEntry::Gicr { base, length }) => { - assert!(range_count < ranges.len(), "GIC: the MADT names more redistributor ranges than {}", ranges.len()); - ranges[range_count] = (base, length); - range_count += 1; - } - Ok(MadtEntry::Gicc(gicc)) if packed_affinity(gicc.mpidr) == me && gicc.gicr_base != 0 => { - own_frame = Some(gicc.gicr_base) - } + Ok(MadtEntry::Gicr { base, length }) => ranges.push(( + base, + crate::mm::paging::map_mmio(base, u64::from(length), MmioPolicy::Uncacheable), + )), + Ok(MadtEntry::Gicc(gicc)) if gicc.enabled => cpus.push(gicc), Ok(_) => {} Err(halt) => panic!("GIC: a MADT entry at +{} declares {} bytes of a {}-byte list", halt.at, halt.declared, halt.list_len), } } let gicd = gicd.expect("GIC: the MADT names no distributor"); let distributor = crate::mm::paging::map_mmio(gicd, FRAME, MmioPolicy::Uncacheable); - GICD.store(gicd, Relaxed); let revision = distributor.read_u32(GICD_PIDR2) >> 4 & 0xF; assert!(revision >= 3, "GIC: GICD_PIDR2.ArchRev is {revision}, and this kernel drives a GICv3 or later"); - // The frame is found in the ranges when the GICC entry does not name it: - // each redistributor says whose it is in `GICR_TYPER`'s affinity. - let redistributor = match own_frame { - Some(frame) => crate::mm::paging::map_mmio(frame, 2 * FRAME, MmioPolicy::Uncacheable), - None => ranges[..range_count] - .iter() - .find_map(|&(base, length)| find_redistributor(base, u64::from(length), me)) - .unwrap_or_else(|| panic!("GIC: no redistributor answers for MPIDR affinity {me:#x}")), - }; - // The distributor: affinity routing on, and Group 1 on. distributor.write_u32(GICD_CTLR, CTLR_ARE | CTLR_GRP1); settles(100, "GICD_CTLR", || distributor.read_u32(GICD_CTLR) & RWP == 0); @@ -153,6 +170,30 @@ pub fn init(rsdp_addr: u64) { "GIC: GICD_CTLR.ARE reads clear, so the distributor stays in legacy mode and no redistributor is used" ); + let timer = toyos_acpi::gtdt(direct_phys(), rsdp_addr) + .unwrap_or_else(|e| panic!("GIC: the GTDT is unusable: {e:?}")) + .virtual_el1; + assert!((16..32).contains(&timer.gsiv), "GIC: the GTDT puts the virtual timer at INTID {}, which is not a PPI", timer.gsiv); + TIMER_INTID.store(timer.gsiv, Relaxed); + TIMER_EDGE.store(timer.edge(), Relaxed); + log!( + "GIC: v{revision} distributor at {gicd:#x}; SGIs and the virtual timer's PPI {} ({}) taken at priority {PRIORITY:#x}", + timer.gsiv, + if timer.edge() { "edge" } else { "level" }, + ); + + let gic = Gic { cpus, ranges }; + init_cpu(gic.redistributor(cpu::hardware_id())); + gic +} + +/// Bring this CPU's redistributor — the frame at `frame`, which +/// [`Gic::redistributor`] mapped — its CPU interface and its timer up: every +/// SGI and the timer's PPI enabled at [`PRIORITY`], the timer stopped. +/// Interrupts stay masked at `DAIF`. +pub fn init_cpu(frame: u64) { + let redistributor = Mmio::new(DirectMap::from_phys(frame), 2 * FRAME); + // Awake, then every SGI and PPI off and cleared of what firmware left. let waker = redistributor.read_u32(GICR_WAKER); redistributor.write_u32(GICR_WAKER, waker & !WAKER_PROCESSOR_SLEEP); @@ -168,17 +209,13 @@ pub fn init(rsdp_addr: u64) { redistributor.write_u32(GICR_IPRIORITYR + word * 4, priorities); } - let timer = toyos_acpi::gtdt(direct_phys(), rsdp_addr) - .unwrap_or_else(|e| panic!("GIC: the GTDT is unusable: {e:?}")) - .virtual_el1; - assert!((16..32).contains(&timer.gsiv), "GIC: the GTDT puts the virtual timer at INTID {}, which is not a PPI", timer.gsiv); - TIMER_INTID.store(timer.gsiv, Relaxed); + let timer = TIMER_INTID.load(Relaxed); // Each PPI's two bits in `GICR_ICFGR1`: 0b10 edge, 0b00 level. - let edge = u32::from(timer.edge()) << (2 * (timer.gsiv - 16) + 1); + let edge = u32::from(TIMER_EDGE.load(Relaxed)) << (2 * (timer - 16) + 1); redistributor.write_u32(GICR_ICFGR1, edge); stop_timer_hardware(); - let enabled = 1 << SGI_KICK | 1 << timer.gsiv; + let enabled = 1 << SGI_KICK | 1 << SGI_HALT | 1 << timer; #[cfg(feature = "boot-actuators")] let enabled = enabled | 1 << Intid::LogNest as u32 | 1 << SGI_STORM; redistributor.write_u32(GICR_ISENABLER0, enabled); @@ -216,20 +253,7 @@ pub fn init(rsdp_addr: u64) { ); } assert_eq!(read_sysreg_pmr(), PRIORITY_MASK, "GIC: ICC_PMR_EL1 did not take {PRIORITY_MASK:#x}"); - log!( - "GIC: v{revision} distributor at {gicd:#x}, this CPU's redistributor at {:#x}; SGIs and the virtual \ - timer's PPI {} ({}) enabled at priority {PRIORITY:#x}", - redistributor.addr() - crate::mm::PHYS_OFFSET, - timer.gsiv, - if timer.edge() { "edge" } else { "level" }, - ); -} - -/// The redistributor frame in `[base, base + length)` whose affinity is `me`. -fn find_redistributor(base: u64, length: u64, me: u32) -> Option { - let range = crate::mm::paging::map_mmio(base, length, MmioPolicy::Uncacheable); - let offset = toyos_gicv3::find_redistributor(length, me, |at| range.read_u64(at + GICR_TYPER))?; - Some(Mmio::new(DirectMap::from_phys(base + offset), 2 * FRAME)) + log!("GIC: this CPU's redistributor at {frame:#x}, its SGIs and timer enabled"); } /// The INTID the GIC hands this CPU, or `None` when it answered spurious. @@ -253,19 +277,21 @@ pub(super) fn timer_intid() -> u32 { TIMER_INTID.load(Relaxed) } -/// Raise SGI `intid` on the CPU whose packed affinity is `target`. -fn sgi(intid: u32, target: u32) { - let value = toyos_gicv3::sgi1r(intid, target); +/// Write `ICC_SGI1R_EL1` whole: the SGI it names is raised. +fn raise(value: u64) { // SAFETY: writes `ICC_SGI1R_EL1`, which raises an SGI and touches no // memory; the `ISB` sends it before whatever follows. unsafe { core::arch::asm!("msr S3_0_C12_C11_5, {}", "isb", in(reg) value, options(nomem, nostack, preserves_flags)) }; } -/// Wake `cpu` so it runs a scheduler pass. Before the port's stage 5 the boot -/// CPU is the only one, so the only `cpu` there is is this one. +/// Raise SGI `intid` on the CPU whose packed affinity is `target`. +fn sgi(intid: u32, target: u32) { + raise(toyos_gicv3::sgi1r(intid, target)); +} + +/// Wake `cpu` so it runs a scheduler pass. pub fn kick_cpu(cpu: u32) { - assert_eq!(cpu, percpu::cpu_id(), "irqchip: a kick for cpu{cpu}, and other CPUs are the port's stage 5"); - sgi(SGI_KICK, cpu::hardware_id()); + sgi(SGI_KICK, super::smp::hardware_id_of(cpu)); } pub fn kick_all_but_self() { @@ -286,8 +312,14 @@ pub fn send_nmi(_cpu: u32) { owed!("a pseudo-NMI", "no stage yet") } -/// Nothing to stop: before the port's stage 5 the boot CPU is the only one running. -pub fn stop_other_cpus() {} +/// Halt every other CPU for good, once the machine has released them: before +/// that each is waiting for the release with interrupts masked, and the boot +/// CPU's own interface may not be up to raise anything. +pub fn stop_other_cpus() { + if super::smp::is_ready() { + raise(toyos_gicv3::sgi1r_others(SGI_HALT)); + } +} /// `CNTV_CTL_EL0.ENABLE`; `IMASK` stays clear. const TIMER_ENABLE: u64 = 1; diff --git a/kernel/src/arch/aarch64/mod.rs b/kernel/src/arch/aarch64/mod.rs index 0b8952c634a..8df6a4c33f2 100644 --- a/kernel/src/arch/aarch64/mod.rs +++ b/kernel/src/arch/aarch64/mod.rs @@ -28,6 +28,7 @@ pub mod paging; pub mod percpu; pub mod pio; pub mod pmu; +pub mod psci; pub mod rtc; pub mod smp; pub mod switch; diff --git a/kernel/src/arch/aarch64/paging.rs b/kernel/src/arch/aarch64/paging.rs index df990ec4c3b..06d954e9b75 100644 --- a/kernel/src/arch/aarch64/paging.rs +++ b/kernel/src/arch/aarch64/paging.rs @@ -587,6 +587,9 @@ static KERNEL: core::sync::atomic::AtomicPtr /// Its root, cached for lock-free access from panic and crash paths. static KERNEL_TTBR0: core::sync::atomic::AtomicU64 = core::sync::atomic::AtomicU64::new(0); +/// The kernel's own root, which every CPU's `TTBR1_EL1` holds from [`join`] on. +static KERNEL_TTBR1: core::sync::atomic::AtomicU64 = core::sync::atomic::AtomicU64::new(0); + pub fn kernel() -> &'static alloc::sync::Arc> { let ptr = KERNEL.load(core::sync::atomic::Ordering::Acquire); assert!(!ptr.is_null(), "paging not initialized"); @@ -655,19 +658,29 @@ pub(crate) fn init(memory_map: &[MemoryMapEntry], scanout: Option<(u64, u64)>) - tables.map_direct(scanout, size, kernel_leaf(CachePolicy::WriteCombining)); } - let ttbr1 = tables.root.phys(); + KERNEL_TTBR1.store(tables.root.phys(), core::sync::atomic::Ordering::Release); *high() = Some(tables); let space = AddressSpace { tables: Tables::new(), regions: vma::Regions::default(), asid: Asid::Kernel }; - let ttbr0 = space.root(); - KERNEL_TTBR0.store(ttbr0.0, core::sync::atomic::Ordering::Release); + KERNEL_TTBR0.store(space.root().0, core::sync::atomic::Ordering::Release); let published: &'static alloc::sync::Arc> = Box::leak(Box::new(alloc::sync::Arc::new(Lock::new(space)))); KERNEL.store(published as *const _ as *mut _, core::sync::atomic::Ordering::Release); - // SAFETY: both tables are built and published above, and the new - // `TTBR1_EL1` maps the code, stack and data this runs on exactly as the - // loader's did; the loader's identity view goes with the old `TTBR0_EL1`, - // and the local `TLBI` drops every translation either left behind. + join(); + crate::log!("paging: the direct map holds memory below {end:#x} in {blocks} 2 MiB blocks and {pages} 4 KiB pages"); + extent +} + +/// Put this CPU on the kernel's tables — `TTBR1_EL1` the kernel's root, +/// `TTBR0_EL1` the kernel space's empty user half — from the loader's or from +/// [`bringup_root`], either of which maps the code, stack and data this runs +/// on exactly as the kernel's root does. The local `TLBI` then drops every +/// translation the old tables left, global ones among them. +pub(crate) fn join() { + let ttbr1 = KERNEL_TTBR1.load(core::sync::atomic::Ordering::Acquire); + let ttbr0 = KERNEL_TTBR0.load(core::sync::atomic::Ordering::Acquire); + // SAFETY: both roots are published by `init` and live forever, and the + // tables this CPU ran on map everything this function touches as they do. unsafe { core::arch::asm!( "dsb ishst", @@ -678,12 +691,28 @@ pub(crate) fn init(memory_map: &[MemoryMapEntry], scanout: Option<(u64, u64)>) - "dsb nsh", "isb", ttbr1 = in(reg) ttbr1, - ttbr0 = in(reg) ttbr0.0, + ttbr0 = in(reg) ttbr0, options(nostack, preserves_flags), ); } - crate::log!("paging: the direct map holds memory below {end:#x} in {blocks} 2 MiB blocks and {pages} 4 KiB pages"); - extent +} + +/// A root for a CPU to turn its MMU on under, running at physical addresses: +/// the kernel root's direct map, and the same entries again at identity below +/// it, so one root serves both `TTBR`s as the loader's does. Its tables below +/// the root are the kernel's own. Leaked: a CPU that answered too late to be +/// counted may walk it still. +pub fn bringup_root() -> u64 { + let high = high(); + let kernel = &high.as_ref().expect("bringup_root: before the kernel's tables exist").root; + let mut root = Table::new(); + let from = index(PHYS_OFFSET, 0); + for i in from..512 { + root.0[i - from] = kernel.0[i]; + root.0[i] = kernel.0[i]; + } + published(); + Box::leak(root).phys() } /// Nothing to seal: every space runs under the one `TTBR1_EL1`, so no space diff --git a/kernel/src/arch/aarch64/percpu.rs b/kernel/src/arch/aarch64/percpu.rs index 5d7a8b8f306..3963d2a9564 100644 --- a/kernel/src/arch/aarch64/percpu.rs +++ b/kernel/src/arch/aarch64/percpu.rs @@ -1,5 +1,5 @@ //! Per-CPU state, reached through `TPIDR_EL1`, which holds this CPU's -//! [`PerCpu`] from [`init_bsp`] on and which nothing else writes. EL0 cannot +//! [`PerCpu`] from [`install`] on and which nothing else writes. EL0 cannot //! read it, so a thread learns nothing of the kernel's layout from it. //! //! Every field another context on the same CPU can write — an interrupt, a @@ -64,20 +64,20 @@ pub struct PerCpu { #[inline] fn this() -> &'static PerCpu { let block: u64; - // SAFETY: reads `TPIDR_EL1`, which [`init_bsp`] points at a leaked - // `PerCpu` before any accessor here runs (`log::PERCPU_READY` gates them). + // SAFETY: reads `TPIDR_EL1`, which [`install`] points at a leaked `PerCpu` + // before any accessor here runs on this CPU: `log::PERCPU_READY` gates them + // on the boot CPU, and an AP installs its block before it does anything else. unsafe { core::arch::asm!("mrs {}, tpidr_el1", out(reg) block, options(nomem, nostack, preserves_flags)); &*(block as *const PerCpu) } } -/// Build the boot CPU's block and publish it through `TPIDR_EL1`; after it, -/// every accessor here answers. The thread registers EL0 can read are cleared, -/// so a thread's first read of one finds nothing firmware left. -pub fn init_bsp() { +/// Build CPU `cpu_id`'s block, on the boot CPU and before that CPU runs: its +/// log shard and interrupt counters are published here, so no reader misses them. +pub fn alloc(cpu_id: u32) -> &'static PerCpu { let block: &'static PerCpu = Box::leak(Box::new(PerCpu { - cpu_id: 0, + cpu_id, current_tid: AtomicU32::new(u32::MAX), current_pid: AtomicU32::new(u32::MAX), preempt_count: AtomicU32::new(0), @@ -93,10 +93,17 @@ pub fn init_bsp() { syscall_fp: AtomicU64::new(0), syscall_sp: AtomicU64::new(0), syscall_task: AtomicU64::new(NO_SYSCALL), - log_shard: &log::BOOT_SHARD, + log_shard: log::shard_for(cpu_id), irq_counts: [const { AtomicU64::new(0) }; crate::irq_census::SLOTS], })); - crate::irq_census::publish(0, block.irq_counts.as_ptr()); + crate::irq_census::publish(cpu_id, block.irq_counts.as_ptr()); + block +} + +/// Make `block` this CPU's through `TPIDR_EL1`; after it, every accessor here +/// answers on this CPU. The thread registers EL0 can read are cleared, so a +/// thread's first read of one finds nothing firmware left. +pub fn install(block: &'static PerCpu) { // SAFETY: the block lives for the machine's life; `TPIDRRO_EL0` and // `TPIDR_EL0` are EL0's to read and hold nothing of the kernel's. unsafe { @@ -108,6 +115,11 @@ pub fn init_bsp() { options(nostack, preserves_flags), ); } +} + +/// The boot CPU's block, installed. +pub fn init_bsp() { + install(alloc(0)); crate::log::PERCPU_READY.store(true, core::sync::atomic::Ordering::Release); log!("percpu: BSP cpu_id=0 mpidr={:#x}", super::cpu::hardware_id()); } diff --git a/kernel/src/arch/aarch64/psci.rs b/kernel/src/arch/aarch64/psci.rs new file mode 100644 index 00000000000..4abbfd8dc21 --- /dev/null +++ b/kernel/src/arch/aarch64/psci.rs @@ -0,0 +1,145 @@ +//! The Power State Coordination Interface (Arm DEN0022), through the conduit +//! the FADT's `ARM_BOOT_ARCH` names (ACPI 6.5 §5.2.9): `SMC` to the secure monitor, or `HVC` +//! where a hypervisor stands in for it. Every call is SMCCC's (Arm DEN0028D): +//! the function in `x0`, its arguments in `x1`–`x3`, the answer in `x0`. + +use toyos_acpi::Psci; + +use crate::drivers::acpi::direct_phys; +use crate::log; + +/// The function IDs of `PSCI_VERSION` and the SMC64 `CPU_ON`, as PSCI +/// numbers them from 0.2 on. +const PSCI_VERSION: u32 = 0x8400_0000; +const CPU_ON: u32 = 0xC400_0003; + +/// `CPU_ON`'s `target_cpu`: `MPIDR_EL1`'s affinity fields where they stand, +/// Aff3 in bits 39:32, every other bit zero. +const TARGET_CPU: u64 = 0xFF_00FF_FFFF; + +/// How this machine reaches PSCI. +#[derive(Clone, Copy)] +pub enum Conduit { + Smc, + Hvc, +} + +/// A PSCI call's refusal, as the specification numbers its return codes. +#[derive(Clone, Copy, Debug)] +pub enum Error { + NotSupported, + InvalidParameters, + Denied, + AlreadyOn, + OnPending, + InternalFailure, + NotPresent, + Disabled, + InvalidAddress, + Unnamed(i32), +} + +impl Error { + fn of(code: i32) -> Self { + match code { + -1 => Self::NotSupported, + -2 => Self::InvalidParameters, + -3 => Self::Denied, + -4 => Self::AlreadyOn, + -5 => Self::OnPending, + -6 => Self::InternalFailure, + -7 => Self::NotPresent, + -8 => Self::Disabled, + -9 => Self::InvalidAddress, + other => Self::Unnamed(other), + } + } +} + +impl Conduit { + /// `function` with three arguments; the answer's low 32 bits, which is + /// where every function this kernel calls puts it. `x4`–`x17` are the + /// callee's to clobber under SMCCC 1.0, so they are declared so. + fn call(self, function: u32, a1: u64, a2: u64, a3: u64) -> i32 { + let answer: u64; + // SAFETY: an SMCCC call into firmware or a hypervisor, which returns + // to the next instruction; not `nomem`, since a callee may read what + // this CPU stored before it. + unsafe { + match self { + Self::Smc => core::arch::asm!( + "smc #0", + inout("x0") u64::from(function) => answer, + inout("x1") a1 => _, inout("x2") a2 => _, inout("x3") a3 => _, + out("x4") _, out("x5") _, out("x6") _, out("x7") _, out("x8") _, out("x9") _, + out("x10") _, out("x11") _, out("x12") _, out("x13") _, out("x14") _, + out("x15") _, out("x16") _, out("x17") _, + options(nostack), + ), + Self::Hvc => core::arch::asm!( + "hvc #0", + inout("x0") u64::from(function) => answer, + inout("x1") a1 => _, inout("x2") a2 => _, inout("x3") a3 => _, + out("x4") _, out("x5") _, out("x6") _, out("x7") _, out("x8") _, out("x9") _, + out("x10") _, out("x11") _, out("x12") _, out("x13") _, out("x14") _, + out("x15") _, out("x16") _, out("x17") _, + options(nostack), + ), + } + } + answer as i32 + } + + /// Start the CPU whose `MPIDR_EL1` is `mpidr` at the physical address + /// `entry`, with `context` in its `x0`, at the highest exception level + /// this CPU's firmware gives an operating system. + pub fn cpu_on(self, mpidr: u64, entry: u64, context: u64) -> Result<(), Error> { + match self.call(CPU_ON, mpidr & TARGET_CPU, entry, context) { + 0 => Ok(()), + code => Err(Error::of(code)), + } + } + + fn name(self) -> &'static str { + match self { + Self::Smc => "SMC", + Self::Hvc => "HVC", + } + } +} + +/// PSCI as the FADT names it, asked for its version: `None` where the machine +/// offers none this kernel can call, which the log says. +pub fn init(rsdp_addr: u64) -> Option { + let fadt = match toyos_acpi::find_table(direct_phys(), rsdp_addr, b"FACP", toyos_acpi::SDT_HEADER_LEN) { + Ok(fadt) => fadt, + Err(e) => { + log!("PSCI: none, because the FADT is unusable: {e:?}"); + return None; + } + }; + let conduit = match toyos_acpi::psci(&fadt) { + Psci::Smc => Conduit::Smc, + Psci::Hvc => Conduit::Hvc, + Psci::Absent => { + log!("PSCI: none: the FADT's ARM_BOOT_ARCH does not set PSCI_COMPLIANT"); + return None; + } + Psci::Undefined { revision, minor } => { + log!("PSCI: none: the FADT is version {revision}.{minor}, and ARM_BOOT_ARCH begins at 5.1"); + return None; + } + }; + let version = conduit.call(PSCI_VERSION, 0, 0, 0); + if version < 0 { + log!("PSCI: PSCI_VERSION through {} answered {:?}; none used", conduit.name(), Error::of(version)); + return None; + } + let (major, minor) = (version >> 16, version & 0xFFFF); + if major == 0 && minor < 2 { + log!("PSCI: {major}.{minor} through {}, which numbers no SMC64 CPU_ON; none used", conduit.name()); + return None; + } + log!("PSCI: {major}.{minor} through {}", conduit.name()); + Some(conduit) +} diff --git a/kernel/src/arch/aarch64/smp.rs b/kernel/src/arch/aarch64/smp.rs index a46829cd401..dff23623ca0 100644 --- a/kernel/src/arch/aarch64/smp.rs +++ b/kernel/src/arch/aarch64/smp.rs @@ -1,15 +1,46 @@ -//! Other CPUs: PSCI `CPU_ON` for each GICC the MADT names, the port's stage 5. -//! Until then the boot CPU is the only one the roster holds. +//! Other CPUs: every GIC CPU interface the MADT enables, started one at a time +//! with PSCI's `CPU_ON` (Arm DEN0022) at [`super::boot::ap_start`], and +//! committed to the roster once it has echoed. +use core::mem::size_of; + +use alloc::boxed::Box; +use toyos_gicv3::packed_affinity; + +use super::percpu::{self, PerCpu}; +use super::{cache, control_regs, cpu, irqchip, paging, psci}; +use crate::mm::DirectMap; use crate::smp_roster::Roster; +use crate::time::AP_START; +use crate::{clock, log}; static ROSTER: Roster = Roster::new(); +const _: () = assert!(crate::smp_roster::MAX_CPUS == crate::scheduler::MAX_CPUS); + +/// What an AP's entry reads with its MMU off, at the physical address `CPU_ON` +/// hands it in `x0`, and what its first Rust reads after. +#[repr(C)] +pub struct ApStart { + /// The root it turns its MMU on under: [`paging::bringup_root`]. + pub(super) root: u64, + /// The top of the stack it runs on until the idle loop, a kernel address. + pub(super) stack_top: u64, + percpu: &'static PerCpu, + /// Its redistributor's physical frame, which the boot CPU mapped. + redistributor: u64, + token: u32, +} + +/// A stack for an AP's bring-up, never freed: the AP runs on it. +#[repr(C, align(16))] +struct Stack([u8; crate::process::KERNEL_STACK_SIZE]); + pub fn cpu_count() -> u32 { ROSTER.count() } -/// Release the machine to the scheduler: with one CPU, nothing waits on it. +/// Release the APs into the scheduler. pub fn set_ready() { ROSTER.release(); } @@ -18,3 +49,68 @@ pub fn set_ready() { pub fn is_ready() -> bool { ROSTER.released() } + +/// `cpu`'s packed MPIDR affinity; panics if `cpu` is not online. +pub fn hardware_id_of(cpu: u32) -> u32 { + assert!(cpu < cpu_count(), "smp: cpu{cpu} is not online"); + ROSTER.hardware_id(cpu) +} + +/// Start every other CPU `gic` names, in the MADT's order, until one does not +/// echo within [`AP_START`] or the roster is full. +pub fn start(gic: &irqchip::Gic, psci: Option) { + let me = cpu::hardware_id(); + ROSTER.set_bsp(me); + let others = gic.cpus.iter().filter(|gicc| packed_affinity(gicc.mpidr) != me); + let Some(psci) = psci else { + log!("SMP: no PSCI to start the other {} CPUs with; the boot CPU runs alone", others.count()); + control_regs::report(cpu_count()); + return; + }; + let root = paging::bringup_root(); + let entry = DirectMap::phys_of(super::boot::ap_start as *const u8); + for gicc in others { + let Some(attempt) = ROSTER.begin_attempt() else { + log!("SMP: roster full at {} CPUs; ignoring further MADT entries", cpu_count()); + break; + }; + let stack = Box::leak(Box::::new_uninit()); + let start: &'static ApStart = Box::leak(Box::new(ApStart { + root, + stack_top: stack.as_ptr() as u64 + size_of::() as u64, + percpu: percpu::alloc(attempt.id()), + redistributor: gic.redistributor(packed_affinity(gicc.mpidr)), + token: attempt.token(), + })); + // Read with the MMU off, so from memory and never from this CPU's cache. + cache::write_back(start as *const ApStart as u64, size_of::()); + let context = DirectMap::phys_of(start as *const ApStart); + if let Err(refused) = psci.cpu_on(gicc.mpidr, entry, context) { + log!("SMP: CPU_ON refused cpu{} mpidr={:#x} ({refused:?}); the rest stay off", attempt.id(), gicc.mpidr); + break; + } + let deadline = clock::nanos_since_boot() + AP_START.nanos(); + // Committed only on this attempt's own token, so `0..cpu_count()` stays dense. + if !ROSTER.await_echo(attempt, || clock::nanos_since_boot() >= deadline) { + log!("SMP: cpu{} mpidr={:#x} did not echo within {AP_START}; the rest stay off", attempt.id(), gicc.mpidr); + break; + } + ROSTER.commit(attempt, packed_affinity(gicc.mpidr)); + log!("SMP: cpu{} mpidr={:#x} online", attempt.id(), gicc.mpidr); + } + let online = cpu_count(); + log!("SMP: {online} of {} MADT CPUs online", gic.cpus.len()); + control_regs::report(online); +} + +/// An AP's first Rust: at EL1, on its bring-up stack with the vectors +/// installed, under [`paging::bringup_root`], entered at `el`. +pub(super) extern "C" fn ap_entry(start: &'static ApStart, el: u64) -> ! { + // First: every log line, a fault's report among them, reads it. + percpu::install(start.percpu); + paging::join(); + control_regs::check(el); + irqchip::init_cpu(start.redistributor); + ROSTER.echo(start.token); + crate::process::ap_idle(); +} diff --git a/kernel/src/arch/aarch64/tlb.rs b/kernel/src/arch/aarch64/tlb.rs index d46f19b02c8..8b6fcfd500c 100644 --- a/kernel/src/arch/aarch64/tlb.rs +++ b/kernel/src/arch/aarch64/tlb.rs @@ -68,6 +68,10 @@ pub fn shootdown(_origin: Origin) {} /// Nothing to answer: no CPU waits on another's acknowledgement here. pub fn poll() {} +/// Nothing to settle: every broadcast invalidation reached a CPU that was +/// waiting for the machine's release. +pub fn join() {} + /// One `tlb:` line when the counts moved, at process exit. pub fn log_census() { let mut counts = [0u64; Origin::COUNT]; diff --git a/kernel/src/arch/aarch64/trap.rs b/kernel/src/arch/aarch64/trap.rs index a8c433f5aa9..8b510ae42ee 100644 --- a/kernel/src/arch/aarch64/trap.rs +++ b/kernel/src/arch/aarch64/trap.rs @@ -163,6 +163,9 @@ fn irq(from_el0: bool) { return; } match intid { + // Never ended: the running priority it keeps is every interrupt's + // own, so the interface signals this CPU nothing and the halt stays. + irqchip::SGI_HALT => cpu::halt(), irqchip::SGI_KICK => { percpu::irq_took(Source::Timer); irqchip::end(intid); diff --git a/kernel/src/arch/x86_64/percpu.rs b/kernel/src/arch/x86_64/percpu.rs index c9050d340c9..8e5df63dd82 100644 --- a/kernel/src/arch/x86_64/percpu.rs +++ b/kernel/src/arch/x86_64/percpu.rs @@ -356,7 +356,7 @@ fn alloc_percpu(cpu_id: u32) -> *mut PerCpu { fault_state: CpuFaultState::Normal as u8, _pad_after_fault_state: [0; 3], last_armed_ticks: AtomicU32::new(0), - log_shard: alloc_log_shard(cpu_id), + log_shard: log::shard_for(cpu_id) as *const log::Shard as u64, nmi_active: 0, ap_token: 0, #[cfg(feature = "boot-actuators")] @@ -377,23 +377,6 @@ fn alloc_percpu(cpu_id: u32) -> *mut PerCpu { ptr } -/// This CPU's log shard: cpu0's is the boot shard, every other fresh; allocated here rather than in [`init_ap`], which already logs. -fn alloc_log_shard(cpu_id: u32) -> u64 { - if cpu_id == 0 { - return &raw const log::BOOT_SHARD as u64; - } - let layout = Layout::from_size_align(size_of::(), 64).unwrap(); - // SAFETY: `alloc_percpu`'s argument — non-zero size, power-of-two alignment, never freed. - let ptr = unsafe { alloc_zeroed(layout) } as *mut log::Shard; - assert!(!ptr.is_null(), "percpu: log shard alloc failed for cpu{cpu_id}"); - // SAFETY: fresh, zeroed, 64-byte-aligned, and not yet published. - unsafe { log::Shard::initialize_zeroed(ptr) }; - // Published before the CPU executes an instruction: a reader can't find it through `gs:` otherwise. - // SAFETY: the allocation is live for the machine's life and initialised. - unsafe { log::publish_ap_shard(cpu_id, ptr) }; - ptr as u64 -} - /// This CPU's shard, its identity, and one sequence number out of that shard. /// The `xadd` has no `lock` prefix, sound only while the live [`crate::arch::IrqGuard`] proves ownership; the four reads are one `asm!` block, not four [`gs`] calls, so the absent `nomem` keeps shard selection inside the guard's barrier. pub fn reserve_log_slot( diff --git a/kernel/src/arch/x86_64/smp.rs b/kernel/src/arch/x86_64/smp.rs index 4331ff0e107..a38c792f584 100644 --- a/kernel/src/arch/x86_64/smp.rs +++ b/kernel/src/arch/x86_64/smp.rs @@ -1,6 +1,6 @@ use core::arch::global_asm; use core::mem::size_of; -use core::sync::atomic::{AtomicU32, AtomicU64, Ordering}; +use core::sync::atomic::{AtomicU64, Ordering}; use alloc::alloc::{alloc_zeroed, Layout}; @@ -9,7 +9,7 @@ use crate::arch::{apic, percpu, syscall}; use crate::clock; use crate::drivers::acpi::MadtInfo; use crate::smp_roster::Roster; -use crate::time::{Budget, Delay, Duration}; +use crate::time::{Delay, Duration, AP_START}; use crate::{log, process}; const TRAMPOLINE_PAGE: u64 = 0x8000; @@ -17,10 +17,6 @@ const TRAMPOLINE_VECTOR: u8 = 0x08; const AP_STACK_SIZE: usize = 64 * 1024; const DATA_OFFSET: usize = 0xF00; -/// The token the latest-launched AP echoed at `ap_entry`, not a flag, so a stale -/// AP cannot be read as this one; `0` means none has. -static AP_STARTED: AtomicU32 = AtomicU32::new(0); - static ROSTER: Roster = Roster::new(); /// The `rdtsc` the AP being started read on its first instruction in Rust. @@ -42,7 +38,7 @@ pub fn cpu_count() -> u32 { /// LAPIC id of `cpu_id`; panics if `cpu_id` is not online. pub fn apic_id_for(cpu_id: u32) -> u32 { assert!(cpu_id < cpu_count(), "apic_id_for: cpu {cpu_id} not online"); - ROSTER.apic_id(cpu_id) + ROSTER.hardware_id(cpu_id) } /// True once a shootdown must wait for siblings; the word the APs are released by. @@ -237,7 +233,6 @@ pub fn boot_aps(madt: &MadtInfo, boot_cr3: u64) { // SAFETY: `target` is reserved physical memory written only here; unaligned because `TrampolineData` is `repr(C, packed)`; this AP has not been sent its SIPI yet, and the loop reaches a second write only after the previous AP committed — which happens in `ap_entry` past every trampoline read — so no CPU is reading it; a failed AP breaks the loop below instead of reaching another write. unsafe { core::ptr::write_unaligned(target, data); } - AP_STARTED.store(0, Ordering::Release); AP_TSC.store(0, Ordering::Release); // The bracket's lower edge: taken after the previous AP committed and // before this one is sent anything, so nothing this AP does precedes it. @@ -256,24 +251,14 @@ pub fn boot_aps(madt: &MadtInfo, boot_cr3: u64) { apic::send_sipi(ap_id, TRAMPOLINE_VECTOR); delay(BETWEEN_SIPIS); - if AP_STARTED.load(Ordering::Acquire) != attempt.token() { + if !ROSTER.echoed(attempt) { apic::send_sipi(ap_id, TRAMPOLINE_VECTOR); } } - // Budget, not Tripwire or panic: a slow AP degrades the machine rather than crashing it, and a panic here would be a behaviour change C1's gate forbids. - const AP_START: Budget = Budget::of( - Duration::from_millis(100), - "the machine boots with the CPUs that came up before the first that did not", - ); let deadline = clock::nanos_since_boot() + AP_START.nanos(); - while AP_STARTED.load(Ordering::Acquire) != attempt.token() { - if clock::nanos_since_boot() >= deadline { break; } - core::hint::spin_loop(); - } - // Commit only on this attempt's own token, so `0..cpu_count()` stays dense. - if AP_STARTED.load(Ordering::Acquire) == attempt.token() { + if ROSTER.await_echo(attempt, || clock::nanos_since_boot() >= deadline) { let bracket_hi = cpu::rdtsc(); if tsc_inside(attempt.id(), bracket_lo, bracket_hi) { bracketed += 1; @@ -351,30 +336,8 @@ extern "C" fn ap_entry() -> ! { apic::init_ap(); // Echo this attempt's token, so the BSP counts this AP for its own attempt. - AP_STARTED.store(percpu::ap_token(), Ordering::Release); + ROSTER.echo(percpu::ap_token()); - while !ROSTER.released() { - core::hint::spin_loop(); - } - - // Only a committed CPU may join: an uncommitted AP has no scheduler slot and - // no shootdown targets it, so it parks. The acquire above makes the count visible. - let me = percpu::cpu_id(); - if me >= cpu_count() { - log!("CPU {me}: bring-up did not commit; parking"); - park_ap(); - } - - // A parked AP could not take a shootdown IPI, so this flush stands in for the ones missed while spinning. - // Must run before touching anything not self-mapped: the acquire on `released` makes the BSP's mappings visible, and this flush discards what the spin cached over them. - crate::arch::tlb::join(); - - // Once this CPU is committed and about to run something: the counter and - // the LVT are per logical CPU, so a CPU nobody arms here is one the - // hard-lockup bound does not cover. - crate::hardlockup::arm_this_cpu(); - - log!("CPU {me}: joining scheduler"); process::ap_idle(); } @@ -384,14 +347,6 @@ fn delay(span: Delay) { while clock::nanos_since_boot() - start < span.nanos() {} } -/// An uncommitted AP halts for the life of the machine. -fn park_ap() -> ! { - loop { - // SAFETY: `cli; hlt` on this CPU only; it takes no lock and touches no shared state. - unsafe { core::arch::asm!("cli; hlt", options(nomem, nostack)) }; - } -} - // Real mode → protected mode → long mode → Rust entry, copied to 0x8000 at runtime; addresses below are offsets into TrampolineData at 0x8F00. global_asm!( ".global _trampoline_start", diff --git a/kernel/src/log/mod.rs b/kernel/src/log/mod.rs index 07dceac9db2..1c4085059a0 100644 --- a/kernel/src/log/mod.rs +++ b/kernel/src/log/mod.rs @@ -33,12 +33,24 @@ pub static BOOT_SHARD: Shard = Shard::new(); // The ABI fixes how many shards a cursor can name; the kernel must not exceed it. const _: () = assert!(crate::sched::MAX_CPUS <= MAX_LOG_SHARDS); -/// Makes an AP's shard reachable to a reader. -/// # Safety -/// `shard` must be a live, initialised [`Shard`] that is never freed. -pub unsafe fn publish_ap_shard(cpu: u32, shard: *mut Shard) { - // SAFETY: the caller's contract is this one. - unsafe { registry::publish(registry::kernel_slots(), cpu, shard) }; +/// The shard `cpu` reserves in, published before that CPU runs an instruction, +/// since a reader finds it only through the registry: the boot shard for CPU +/// 0, and a fresh one, never freed, for every other. +pub fn shard_for(cpu: u32) -> &'static Shard { + if cpu == 0 { + return &BOOT_SHARD; + } + let layout = alloc::alloc::Layout::new::(); + // SAFETY: a `Shard` is not zero-sized; the block is never freed. + let ptr = unsafe { alloc::alloc::alloc_zeroed(layout) }.cast::(); + assert!(!ptr.is_null(), "log: no memory for cpu{cpu}'s shard"); + // SAFETY: fresh, zeroed and aligned for a `Shard`, and not yet published; + // then live and initialised for the machine's life, as `publish` needs. + unsafe { + Shard::initialize_zeroed(ptr); + registry::publish(registry::kernel_slots(), cpu, ptr); + &*ptr + } } /// The newest records the stop's account carries whole. The page's account diff --git a/kernel/src/process.rs b/kernel/src/process.rs index 6625dbbede2..cc6e9bdc2ef 100644 --- a/kernel/src/process.rs +++ b/kernel/src/process.rs @@ -1589,7 +1589,30 @@ pub fn handle_fault(error: crate::object::HandleError) -> ! { exit(HANDLE_FAULT_EXIT_CODE) } -/// AP entry into the scheduler. Called from smp::ap_entry once the machine is released. +/// An AP's way into the idle loop on either architecture, once it has answered +/// the BSP: it waits for the machine's release, then joins the scheduler. pub fn ap_idle() -> ! { + while !crate::arch::smp::is_ready() { + core::hint::spin_loop(); + } + + // Only a committed CPU may join: an uncommitted AP has no scheduler slot and + // no shootdown targets it, so it halts. The acquire above makes the count visible. + let me = percpu::cpu_id(); + if me >= crate::arch::smp::cpu_count() { + log!("CPU {me}: bring-up did not commit; halting"); + crate::arch::cpu::halt(); + } + + // Before touching anything not self-mapped: the acquire on the release + // makes the BSP's mappings visible, and this settles every TLB + // invalidation the wait could not take. + crate::arch::tlb::join(); + + // Once this CPU is committed and about to run something: the counter is + // per CPU, so a CPU nobody arms here is one the hard-lockup bound does not cover. + crate::hardlockup::arm_this_cpu(); + + log!("CPU {me}: joining scheduler"); scheduler::enter_idle_loop(); } diff --git a/kernel/src/smp_roster.rs b/kernel/src/smp_roster.rs index 2a010e1e15e..b85d444b5bb 100644 --- a/kernel/src/smp_roster.rs +++ b/kernel/src/smp_roster.rs @@ -1,7 +1,8 @@ //! The CPU roster and the one release/answer word, two invariants held as a type: //! an id commits only after its AP's handshake and `commit` publishes the slot //! before the count, so `0..count()` has no dead slot; and the word `release` sets -//! is the word `answering` reads. Compiled into `kernel-loom/`, so no `crate::`. +//! is the word `answering` reads. Both architectures bring their APs up through +//! it. Compiled into `kernel-loom/`, so no `crate::`. #[cfg(not(feature = "loom"))] use core::sync::atomic::{AtomicBool, AtomicU32, Ordering}; @@ -12,7 +13,7 @@ use loom::sync::atomic::{AtomicBool, AtomicU32, Ordering}; /// Matches `sched::MAX_CPUS`; the roster refuses an id at or above it. pub const MAX_CPUS: usize = 8; -const NO_LAPIC: u32 = u32::MAX; +const NO_HARDWARE_ID: u32 = u32::MAX; /// A CPU id and the token of the attempt bringing it up. #[derive(Clone, Copy)] @@ -34,14 +35,17 @@ impl Attempt { pub struct Roster { /// Committed CPUs; the BSP is 1 from the start, every other write a `commit`. count: AtomicU32, - /// `apic_ids[i]` is committed iff `i < count`; [`NO_LAPIC`] until then. - apic_ids: [AtomicU32; MAX_CPUS], + /// `hardware_ids[i]` is committed iff `i < count`; [`NO_HARDWARE_ID`] until then. + hardware_ids: [AtomicU32; MAX_CPUS], /// Released and answering: one fact, one store. ready: AtomicBool, #[cfg(feature = "smp-ready-split")] answer: AtomicBool, /// Source of per-attempt tokens; `0` is "no attempt". next_token: AtomicU32, + /// The token the latest-started AP echoed, not a flag, so a stale AP cannot + /// be read as this one; `0` means none has. + echoed: AtomicU32, } impl Roster { @@ -50,11 +54,12 @@ impl Roster { pub const fn new() -> Self { Self { count: AtomicU32::new(1), - apic_ids: [const { AtomicU32::new(NO_LAPIC) }; MAX_CPUS], + hardware_ids: [const { AtomicU32::new(NO_HARDWARE_ID) }; MAX_CPUS], ready: AtomicBool::new(false), #[cfg(feature = "smp-ready-split")] answer: AtomicBool::new(false), next_token: AtomicU32::new(1), + echoed: AtomicU32::new(0), } } @@ -63,17 +68,18 @@ impl Roster { pub fn new() -> Self { Self { count: AtomicU32::new(1), - apic_ids: core::array::from_fn(|_| AtomicU32::new(NO_LAPIC)), + hardware_ids: core::array::from_fn(|_| AtomicU32::new(NO_HARDWARE_ID)), ready: AtomicBool::new(false), #[cfg(feature = "smp-ready-split")] answer: AtomicBool::new(false), next_token: AtomicU32::new(1), + echoed: AtomicU32::new(0), } } - /// The BSP's own LAPIC id in slot 0; the count already covers it. - pub fn set_bsp(&self, lapic: u32) { - self.apic_ids[0].store(lapic, Ordering::Relaxed); + /// The BSP's own `arch::cpu::hardware_id` in slot 0; the count already covers it. + pub fn set_bsp(&self, hardware_id: u32) { + self.hardware_ids[0].store(hardware_id, Ordering::Relaxed); } pub fn count(&self) -> u32 { @@ -82,8 +88,8 @@ impl Roster { } /// Caller guarantees `id < count()`. - pub fn apic_id(&self, id: u32) -> u32 { - self.apic_ids[id as usize].load(Ordering::Relaxed) + pub fn hardware_id(&self, id: u32) -> u32 { + self.hardware_ids[id as usize].load(Ordering::Relaxed) } /// Reserve the next dense id and a token, committing nothing; `None` at MAX_CPUS. @@ -96,11 +102,36 @@ impl Roster { Some(Attempt { id, token }) } + /// The AP's half of the handshake, once it can take its first interrupt: + /// the token of the attempt that started it. + pub fn echo(&self, token: u32) { + self.echoed.store(token, Ordering::Release); + } + + /// Whether `at`'s AP has echoed; an acquire, so what the AP did before its + /// echo is visible after. + pub fn echoed(&self, at: Attempt) -> bool { + self.echoed.load(Ordering::Acquire) == at.token + } + + /// Whether `at`'s AP echoed before `spent` said the BSP's budget for it is gone. + pub fn await_echo(&self, at: Attempt, spent: impl Fn() -> bool) -> bool { + loop { + if self.echoed(at) { + return true; + } + if spent() { + return false; + } + core::hint::spin_loop(); + } + } + /// Fill a started AP's slot, then publish the count that covers it. Only the /// BSP calls this, one at a time, so `at.id` is the current count. - pub fn commit(&self, at: Attempt, lapic: u32) { + pub fn commit(&self, at: Attempt, hardware_id: u32) { debug_assert!(at.id == self.count.load(Ordering::Relaxed)); - self.apic_ids[at.id as usize].store(lapic, Ordering::Relaxed); + self.hardware_ids[at.id as usize].store(hardware_id, Ordering::Relaxed); // Release: the slot store above lands before the count exposes it; the // control drops it to relaxed and the model finds the unfilled slot. #[cfg(not(feature = "roster-commit-relaxed"))] diff --git a/kernel/src/time.rs b/kernel/src/time.rs index 852e95b6974..85d2d4ac9de 100644 --- a/kernel/src/time.rs +++ b/kernel/src/time.rs @@ -248,6 +248,12 @@ impl fmt::Display for Budget { } } +/// How long the boot CPU waits for an AP it started to echo its token. +pub const AP_START: Budget = Budget::of( + Duration::from_millis(100), + "the machine boots with the CPUs that came up before the first that did not", +); + /// A duration used as a bound on *another* duration, never as a wait. /// Nothing expires; there is no caller and no register. #[derive(Clone, Copy)] diff --git a/src/build.rs b/src/build.rs index 6f25adc1d9b..034ab1a4d22 100644 --- a/src/build.rs +++ b/src/build.rs @@ -3161,6 +3161,7 @@ mod tests { "tests/toolkitcase/system.toml", "tests/updatecase/system.toml", "tests/virtjobcase/system.toml", + "tests/virtsmpcase/system.toml", ]; fn load(cfg: &str) -> SystemConfig { diff --git a/tests/toyos.rs b/tests/toyos.rs index b86007fbf0f..7dc53060249 100644 --- a/tests/toyos.rs +++ b/tests/toyos.rs @@ -541,6 +541,7 @@ const SCREEN_TESTS: &[(&str, Sched, Tier)] = &[ ("virt_unmap_touch", Sched::Parallel, Tier::Local), ("virt_debug_refused", Sched::Parallel, Tier::Local), ("virt_readonly_copyout", Sched::Parallel, Tier::Local), + ("virt_smp", Sched::Parallel, Tier::Local), ]; /// What `screen_console_shell` types, and what it then looks for on its own. @@ -3570,9 +3571,11 @@ fn check_no_stale_cells(dump: &screen::Ppm, console: &str) -> Result<(), String> /// `test_rs_abuse_readonly_copyout`. const VIRT_COPYOUT: &str = "abuse_readonly_copyout"; -/// Boot `tests/virtjobcase` under the EL2 profile and judge its job `job`: -/// it ends with exit 0, having said `said`. The kernel carries `SYS_DEBUG` -/// for `debug_refused`, and every job runs in every boot of the case. +/// Boot `tests/virtjobcase` on one CPU under the EL2 profile and judge its job +/// `job`: it ends with exit 0, having said `said`. One CPU because `preempt` +/// and `fp_isolation` see a sibling run only when it took theirs. The kernel +/// carries `SYS_DEBUG` for `debug_refused`, and every job runs in every boot +/// of the case. fn virt_job(job: &str, said: &str) -> Result<(), String> { let config = compile::repo_root().join("tests/virtjobcase/system.toml"); let case = config.parent().expect("system.toml has a directory"); @@ -3581,18 +3584,25 @@ fn virt_job(job: &str, said: &str) -> Result<(), String> { let copyout = COPYOUT.get_or_init(|| { qemu::build_toyos_bin(profile.arch(), &compile::repo_root().join("tests/toyos-rust-tests"), VIRT_COPYOUT) }); - let mut qemu = QemuInstance::boot_with_options( + let qemu = QemuInstance::boot_with_options( case, &[], &[], BootOptions { profile, + smp: 1, kernel_features: toyos_build::build::TEST_KERNEL, ready_marker: "control registers: SCTLR_EL1=", extra_root_files: vec![(format!("bin/test_rs_{VIRT_COPYOUT}"), copyout.clone())], ..Default::default() }, ); + judge_virt_job(qemu, job, said).map(drop) +} + +/// Wait for `job`'s end on a guest booted with it, and judge it: it ends with +/// exit 0, having said `said`. Answers everything the PL011 carried. +fn judge_virt_job(mut qemu: QemuInstance, job: &str, said: &str) -> Result { let end = format!("===TEST_END {job} "); let mut rest = String::new(); let waited = await_marker(&mut qemu, &mut rest, &end, &format!("the job {job} to end")); @@ -3611,6 +3621,50 @@ fn virt_job(job: &str, said: &str) -> Result<(), String> { if !ended.contains(&format!("===TEST_END {job} exit=0===")) { return Err(format!("{ended}\nserial:\n{serial}")); } + Ok(serial) +} + +/// The CPUs `virt_smp` boots: as many as the roster holds. +const VIRT_CPUS: u32 = 8; + +/// Boot `tests/virtsmpcase` on [`VIRT_CPUS`] CPUs under the EL2 profile, where +/// PSCI is reached through `SMC`: each is started by `CPU_ON`, holds the +/// control-register declaration and joins the scheduler, and the case's job +/// `unmap_seen` ends with exit 0 once a page unmapped beside its reader has +/// ended the reader's process each time. +fn virt_smp() -> Result<(), String> { + let config = compile::repo_root().join("tests/virtsmpcase/system.toml"); + let case = config.parent().expect("system.toml has a directory"); + let qemu = QemuInstance::boot_with_options( + case, + &[], + &[], + BootOptions { + profile: qemu::Profile::VirtEl2, + smp: VIRT_CPUS, + ready_marker: "control registers: SCTLR_EL1=", + ..Default::default() + }, + ); + let serial = judge_virt_job(qemu, "unmap_seen", "unmap_seen: 4 reads on another thread of a page just unmapped")?; + let psci = serial.lines().find(|l| l.contains("PSCI: ")).unwrap_or_default(); + if !psci.contains(" through SMC") { + return Err(format!("PSCI is not said to be reached through SMC: {psci:?}\nserial:\n{serial}")); + } + let mut want = vec![ + format!("SMP: {VIRT_CPUS} of {VIRT_CPUS} MADT CPUs online"), + format!("control registers: {VIRT_CPUS} of {VIRT_CPUS} CPUs hold the declaration"), + ]; + for cpu in 1..VIRT_CPUS { + want.push(format!("SMP: cpu{cpu} mpidr={cpu:#x} online")); + want.push(format!("CPU {cpu}: joining scheduler")); + } + for want in want { + if !serial.contains(&want) { + return Err(format!("{want:?} not on the PL011\nserial:\n{serial}")); + } + } + eprintln!(" [virt] {VIRT_CPUS} CPUs online and scheduling"); Ok(()) } @@ -5157,6 +5211,7 @@ fn run_screen_test( virt_selftest(test_config, &["irq-storm"]) } "virt_timer_floor" => virt_selftest(test_config, &["timer-floor"]), + "virt_smp" => virt_smp(), "screen_late_panic" => { // The ordinary fatal panic, which no userland process can produce: // crash_report, capture, panic_flush, halt_all_cpus, render. The diff --git a/tests/virtsmpcase/system.toml b/tests/virtsmpcase/system.toml new file mode 100644 index 00000000000..c24d4a34efb --- /dev/null +++ b/tests/virtsmpcase/system.toml @@ -0,0 +1,19 @@ +# Every CPU the machine names, and the one job whose premise is more than one +# CPU: a page unmapped beside a thread that is still reading it. +[boot] +start = ["logd", "test-runner"] + +[programs.logd] +service = true +syscap = ["logread"] + +[programs.test-runner] +receives = ["power"] +# Four minutes from boot: every job here runs on emulated CPUs, and a job past +# the bound reboots the guest with the job named. +args = ["--bound-ms=240000", "unmap_seen"] + +[programs.toybox] + +[symlinks] +"bin/unmap_seen" = "/system/bin/toybox" diff --git a/toyos-acpi/src/fadt.rs b/toyos-acpi/src/fadt.rs index 604cde0e75a..db284cf0f3d 100644 --- a/toyos-acpi/src/fadt.rs +++ b/toyos-acpi/src/fadt.rs @@ -11,8 +11,14 @@ const FADT_FLAGS: usize = 112; /// A Generic Address Structure (§5.2.3.2): space at +0, bit width at +1, bit offset at +2, address at +4. const FADT_RESET_REG: usize = 116; const FADT_RESET_VALUE: usize = 128; +const FADT_ARM_BOOT_ARCH: usize = 129; +const FADT_MINOR_VERSION: usize = 131; pub const FADT_X_DSDT: usize = 140; +/// `ARM_BOOT_ARCH`'s two flags. +const PSCI_COMPLIANT: u16 = 1 << 0; +const PSCI_USE_HVC: u16 = 1 << 1; + /// The bytes [`reset_register`] reads to the end of, so a caller opening the /// table for that decode alone asks for what it needs and nothing after it. pub const FADT_FOR_RESET: usize = FADT_RESET_VALUE + size_of::(); @@ -118,6 +124,34 @@ pub fn reset_register(fadt: &Table

) -> Reset { } } +/// How the FADT says the Power State Coordination Interface is reached. +#[derive(Clone, Copy, PartialEq, Eq, Debug)] +pub enum Psci { + /// `PSCI_COMPLIANT`, through `SMC`. + Smc, + /// `PSCI_COMPLIANT` and `PSCI_USE_HVC`. + Hvc, + /// `PSCI_COMPLIANT` clear: firmware offers no PSCI. + Absent, + /// A FADT before ACPI 5.1, or one too short to reach the field: those + /// bytes are reserved, so nothing is said about PSCI at all. + Undefined { revision: u8, minor: u8 }, +} + +/// `ARM_BOOT_ARCH`, read only from a FADT whose version defines it: 5.1 on, +/// the version its header's major and `FADT Minor Version`'s low nibble spell. +pub fn psci(fadt: &Table

) -> Psci { + let revision = fadt.byte(SDT_REVISION).unwrap_or(0); + let minor = fadt.byte(FADT_MINOR_VERSION).map_or(0, |byte| byte & 0xF); + let defined = revision > 5 || (revision == 5 && minor >= 1); + match fadt.u16_at(FADT_ARM_BOOT_ARCH) { + Some(flags) if defined && flags & PSCI_COMPLIANT == 0 => Psci::Absent, + Some(flags) if defined && flags & PSCI_USE_HVC != 0 => Psci::Hvc, + Some(_) if defined => Psci::Smc, + _ => Psci::Undefined { revision, minor }, + } +} + pub fn dsdt_address(fadt: &Table

) -> u64 { let x_dsdt = match fadt.byte(SDT_REVISION) { Some(r) if r >= 2 => fadt.u64_at(FADT_X_DSDT).filter(|a| *a != 0), diff --git a/toyos-acpi/src/lib.rs b/toyos-acpi/src/lib.rs index 13df39bd4d7..9e237850066 100644 --- a/toyos-acpi/src/lib.rs +++ b/toyos-acpi/src/lib.rs @@ -23,8 +23,8 @@ mod spcr; use toyos_bootmap::DirectMapEnd; pub use fadt::{ - century_of, dsdt_address, iapc_boot_arch, reset_register, rtc_century, Century, Reset, - CMOS_RAM, FADT_FOR_RESET, FADT_PM1A_CNT_BLK, FADT_X_DSDT, + century_of, dsdt_address, iapc_boot_arch, psci, reset_register, rtc_century, Century, Psci, + Reset, CMOS_RAM, FADT_FOR_RESET, FADT_PM1A_CNT_BLK, FADT_X_DSDT, }; pub use madt::{ madt_entries, Gicc, IoApicEntry, MadtEntries, MadtEntry, MadtHalt, SourceOverride, MADT_ENTRIES, diff --git a/toyos-acpi/tests/corpus.rs b/toyos-acpi/tests/corpus.rs index 896789ac93a..0f9f73dc0c0 100644 --- a/toyos-acpi/tests/corpus.rs +++ b/toyos-acpi/tests/corpus.rs @@ -12,8 +12,8 @@ use common::{declare_len, entry, madt, rsdp, sdt, xsdt, Machine}; use toyos_abi::boot::RootBridgeWindow; use toyos_acpi::{ dsdt_address, ecam_base, find_table, hpet_base, iapc_boot_arch, madt_entries, memory_windows, - reset_register, rtc_century, Century, MadtEntry, MadtHalt, Phys, Reset, Table, TableError, - MADT_ENTRIES, MAX_TABLE_LEN, + psci, reset_register, rtc_century, Century, MadtEntry, MadtHalt, Phys, Psci, Reset, Table, + TableError, MADT_ENTRIES, MAX_TABLE_LEN, }; const RSDP_AT: u64 = 0x1_0000; @@ -395,6 +395,7 @@ fn no_single_byte_mutation_of_a_real_table_panics_or_runs_away() { let _ = iapc_boot_arch(m, rsdp_at); if let Ok(t) = find_table(m, rsdp_at, b"FACP", 36) { let _ = reset_register(&t); + let _ = psci(&t); } if let Ok(t) = find_table(m, rsdp_at, b"APIC", MADT_ENTRIES) { // Bounded by the table's own length, so a walk that has @@ -528,6 +529,47 @@ fn reset_fields_are_read_only_from_a_revision_that_defines_them() { assert_eq!(reset_of(&short), Reset::Absent); } +/// A revision-`rev` FADT carrying `ARM_BOOT_ARCH` `flags` at 129 and `FADT +/// Minor Version` `minor` at 131. +fn arm_facp(rev: u8, minor: u8, flags: u16) -> Vec { + let mut body = vec![0u8; 132 - 36]; + body[129 - 36..131 - 36].copy_from_slice(&flags.to_le_bytes()); + body[131 - 36] = minor; + sdt(b"FACP", rev, &body) +} + +fn psci_of(table: &[u8]) -> Psci { + let regions: &[(u64, &[u8])] = &[(TABLE_AT, table)]; + let fadt = Table::open(Machine { regions }, TABLE_AT, b"FACP", 36).expect("a sealed FADT"); + psci(&fadt) +} + +/// `ARM_BOOT_ARCH` on a FADT that defines it. QEMU 11.1.1's `virt` +/// publishes revision 6, minor 3 and `PSCI_COMPLIANT`, with `PSCI_USE_HVC` +/// where the guest has no EL2 (`hw/arm/virt-acpi-build.c:1135-1148`, +/// `hw/arm/virt.c:2956-2962`). +#[test] +fn the_psci_conduit_is_read_off_a_fadt_that_defines_it() { + assert_eq!(psci_of(&arm_facp(6, 3, 0b01)), Psci::Smc); + assert_eq!(psci_of(&arm_facp(6, 3, 0b11)), Psci::Hvc); + assert_eq!(psci_of(&arm_facp(6, 3, 0b10)), Psci::Absent, "USE_HVC without COMPLIANT is no PSCI"); + assert_eq!(psci_of(&arm_facp(6, 3, 0)), Psci::Absent); + assert_eq!(psci_of(&arm_facp(5, 1, 0b01)), Psci::Smc, "5.1 is the first version with the field"); +} + +/// Before ACPI 5.1 the three bytes are reserved, and flags written there are +/// not read; the minor version is the low nibble, the high one an errata letter. +#[test] +fn a_fadt_before_acpi_5_1_says_nothing_about_psci() { + assert_eq!(psci_of(&arm_facp(5, 0, 0b11)), Psci::Undefined { revision: 5, minor: 0 }); + assert_eq!(psci_of(&arm_facp(3, 0, 0b01)), Psci::Undefined { revision: 3, minor: 0 }); + assert_eq!(psci_of(&arm_facp(5, 0x10, 0b01)), Psci::Undefined { revision: 5, minor: 0 }); + + let mut short = arm_facp(6, 3, 0b01); + declare_len(&mut short, 130); + assert_eq!(psci_of(&short), Psci::Undefined { revision: 6, minor: 0 }); +} + /// Where the two firmwares' descriptor lists sit for the sweep below. const ROOT_BRIDGE_AT: u64 = 0x4_0000; const ROOT_BRIDGES: &[(&str, &[u8])] = &[ diff --git a/toyos-acpi/tests/fixtures.rs b/toyos-acpi/tests/fixtures.rs index 1e8edfecdf5..6ff04725bf7 100644 --- a/toyos-acpi/tests/fixtures.rs +++ b/toyos-acpi/tests/fixtures.rs @@ -7,8 +7,8 @@ use common::Machine; use toyos_abi::boot::RootBridgeWindow; use toyos_acpi::{ century_of, dsdt_address, ecam_base, find_table, hpet_base, iapc_boot_arch, madt_entries, - memory_windows, reset_register, rtc_century, Century, IoApicEntry, MadtEntry, Reset, - SourceOverride, TableError, FADT_PM1A_CNT_BLK, MADT_ENTRIES, + memory_windows, psci, reset_register, rtc_century, Century, IoApicEntry, MadtEntry, Psci, + Reset, SourceOverride, TableError, FADT_PM1A_CNT_BLK, MADT_ENTRIES, }; /// Where each table sat in that guest's physical memory. The XSDT's entries @@ -118,6 +118,14 @@ fn the_fadt_names_the_reset_register_qemu_acts_on() { assert_eq!(reset_register(&fadt), Reset::Port { port: 0xcf9, value: 0x0f }); } +/// Revision 3, so bytes 129-131 are the reserved ones QEMU 11.1.1 writes zero +/// (`hw/acpi/aml-build.c:2544-2550`): the table says nothing about PSCI. +#[test] +fn the_q35_fadt_predates_arm_boot_arch() { + let fadt = find_table(machine(), RSDP, b"FACP", 36).expect("FADT"); + assert_eq!(psci(&fadt), Psci::Undefined { revision: 3, minor: 0 }); +} + #[test] fn the_boot_architecture_flags_come_off_a_revision_that_defines_them() { let (revision, flags) = iapc_boot_arch(machine(), RSDP).expect("FADT"); diff --git a/toyos-gicv3/src/lib.rs b/toyos-gicv3/src/lib.rs index 40e73e90d20..43bec42c2d4 100644 --- a/toyos-gicv3/src/lib.rs +++ b/toyos-gicv3/src/lib.rs @@ -34,6 +34,13 @@ pub const fn sgi1r(intid: u32, target: u32) -> u64 { aff3 << 48 | (aff0 >> 4) << 44 | aff2 << 32 | (intid as u64) << 24 | aff1 << 16 | 1 << (aff0 & 0xF) } +/// `ICC_SGI1R_EL1` raising SGI `intid` on every CPU but the one writing it: +/// `IRM` set, under which the affinity fields and the target list are not read. +pub const fn sgi1r_others(intid: u32) -> u64 { + assert!(intid < 16, "an SGI's INTID is below 16"); + 1 << 40 | (intid as u64) << 24 +} + /// The offset, in a redistributor region `length` bytes long, of the /// redistributor whose affinity is `me`; `typer` reads `GICR_TYPER` of the /// redistributor at an offset. The walk steps by each one's own frames and diff --git a/toyos-gicv3/tests/gicv3.rs b/toyos-gicv3/tests/gicv3.rs index 73a55f8b754..1348e21b600 100644 --- a/toyos-gicv3/tests/gicv3.rs +++ b/toyos-gicv3/tests/gicv3.rs @@ -1,4 +1,4 @@ -use toyos_gicv3::{find_redistributor, packed_affinity, sgi1r, FRAME}; +use toyos_gicv3::{find_redistributor, packed_affinity, sgi1r, sgi1r_others, FRAME}; #[test] fn affinity_drops_mpidr_flags_between_aff2_and_aff3() { @@ -25,6 +25,14 @@ fn sgi_to_cpu_zero_is_bit_zero_of_range_zero() { assert_eq!(sgi1r(0, 0), 1); } +#[test] +fn sgi_to_every_other_cpu_is_irm_and_the_intid_alone() { + let value = sgi1r_others(15); + assert_eq!(value >> 40 & 1, 1, "IRM"); + assert_eq!(value >> 24 & 0xF, 15, "INTID"); + assert_eq!(value & !(1 << 40 | 0xF << 24), 0, "no affinity, target list or RES0 bit"); +} + /// A region of redistributors, each `(affinity, vlpis, last)`, as `GICR_TYPER` reads at each one's offset. fn region(cpus: &[(u32, bool, bool)]) -> (u64, impl Fn(u64) -> u64 + '_) { let stride = |vlpis: bool| if vlpis { 4 * FRAME } else { 2 * FRAME }; diff --git a/userland/toybox/src/main.rs b/userland/toybox/src/main.rs index 449ce6e89d4..a60c28b5e7d 100644 --- a/userland/toybox/src/main.rs +++ b/userland/toybox/src/main.rs @@ -21,6 +21,7 @@ mod shutdown; mod spin; mod stats; mod tone; +mod unmap_seen; mod unmap_touch; macro_rules! commands { @@ -36,7 +37,7 @@ macro_rules! commands { use arch::{first_entry, fp_isolation}; -commands!(cat, cp, debug_refused, echo, first_entry, fp_isolation, free, grep, hexdump, locale, ls, mkdir, mv, net, preempt, ps, pwd, reboot, rm, screen, shutdown, spin, stats, tone, unmap_touch); +commands!(cat, cp, debug_refused, echo, first_entry, fp_isolation, free, grep, hexdump, locale, ls, mkdir, mv, net, preempt, ps, pwd, reboot, rm, screen, shutdown, spin, stats, tone, unmap_seen, unmap_touch); fn main() { let args: Vec = std::env::args().collect(); diff --git a/userland/toybox/src/unmap_seen.rs b/userland/toybox/src/unmap_seen.rs new file mode 100644 index 00000000000..7fef30d0b97 --- /dev/null +++ b/userland/toybox/src/unmap_seen.rs @@ -0,0 +1,89 @@ +//! `unmap_seen`: whether an unmap reaches a translation another thread holds. +//! A child maps a page and a second thread reads it in a loop; once that +//! thread has read it, the first unmaps it and says so through memory, and the +//! reader's next read must end the child. On a machine of more CPUs than busy +//! threads the reader runs beside the unmap, on a CPU whose TLB only a +//! broadcast invalidation reaches. + +use std::process::Command; +use std::sync::atomic::{AtomicBool, AtomicU64, Ordering::{Acquire, Release}}; + +use toyos_abi::syscall::{mmap, munmap, MmapFlags, MmapProt}; + +const TRIALS: usize = 4; +const PAGE: usize = 2 * 1024 * 1024; +/// Reads the reader makes before the unmap, each through the translation its +/// CPU caches. +const WARM: u64 = 1_000; +/// The status `kill_process(-1)` gives a process whose fault nothing serves; +/// a refused unmap or a panic ends the child with another. +const FAULTED: i32 = -1; +/// The child's last line before the unmap. +const UNMAPPING: &str = "unmap_seen: the reader holds the page; unmapping"; +/// The reader's line if its read after the unmap came back. +const READ_BACK: &str = "unmap_seen: the reader read"; + +static READS: AtomicU64 = AtomicU64::new(0); +static UNMAPPED: AtomicBool = AtomicBool::new(false); + +pub fn main(args: Vec) { + match args.first().map(String::as_str) { + None => judge(), + Some("child") => child(), + Some(other) => panic!("unmap_seen: unknown mode {other:?}"), + } +} + +fn judge() { + for trial in 0..TRIALS { + let out = Command::new("/system/bin/unmap_seen") + .arg("child") + .output() + .unwrap_or_else(|e| panic!("unmap_seen: the child would not spawn: {e}")); + let said = String::from_utf8_lossy(&out.stdout); + assert!(said.contains(UNMAPPING), "trial {trial}: the child never reached the unmap: {said:?}"); + assert!( + out.status.code() == Some(FAULTED) && !said.contains(READ_BACK), + "trial {trial}: the child exited {:?}, not faulted ({FAULTED}), once its reader read a page \ + unmapped beside it: {said:?}", + out.status.code(), + ); + } + println!("unmap_seen: {TRIALS} reads on another thread of a page just unmapped each ended their process"); +} + +fn child() { + // SAFETY: a fresh anonymous mapping this function alone names. + let page = unsafe { mmap(core::ptr::null_mut(), PAGE, MmapProt::READ | MmapProt::WRITE, MmapFlags::ANONYMOUS | MmapFlags::PRIVATE) }; + assert!(!page.is_null(), "unmap_seen: mmap refused"); + // SAFETY: inside the mapping just made. + unsafe { page.write_volatile(0x5A) }; + let at = page as usize; + std::thread::spawn(move || read_until_unmapped(at as *const u8)); + while READS.load(Acquire) < WARM { + std::hint::spin_loop(); + } + println!("{UNMAPPING}"); + // SAFETY: the mapping made above, whole; its reader's next read is the probe. + unsafe { munmap(page, PAGE) }.expect("unmap_seen: munmap refused"); + UNMAPPED.store(true, Release); + // The reader's fault ends this process; this thread waits for it. + loop { + std::thread::park(); + } +} + +/// Read `page` until the unmap is said to have returned, then once more. +fn read_until_unmapped(page: *const u8) -> ! { + loop { + let unmapped = UNMAPPED.load(Acquire); + // SAFETY: mapped until `unmapped` reads true; the read after that is + // the probe, and a kernel that keeps its word ends the process there. + let byte = unsafe { page.read_volatile() }; + if unmapped { + println!("{READ_BACK} {byte:#x} after the unmap had returned"); + std::process::exit(0); + } + READS.fetch_add(1, Release); + } +} From 466d2e6979e59b2d8e5500022b2c04c3197d7bb3 Mon Sep 17 00:00:00 2001 From: japabu Date: Tue, 29 Sep 2026 23:01:06 +0200 Subject: [PATCH 2/6] virt_user_mode's comment no longer says one CPU The Virt profiles boot two CPUs by default and the kernel now starts the second, so "on one CPU" was false of virt_user_mode; the claim goes. Co-Authored-By: Claude Opus 5.5 --- tests/toyos.rs | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/toyos.rs b/tests/toyos.rs index 7dc53060249..9fca6c3a0de 100644 --- a/tests/toyos.rs +++ b/tests/toyos.rs @@ -5156,7 +5156,7 @@ fn run_screen_test( Ok(()) } "virt_user_mode" => { - // The port's stage 4 on one CPU, under the EL2 profile whose + // The port's stage 4, under the EL2 profile whose // entry also writes what the drop leaves EL2 holding: the kernel's // own tables, the GIC and the timer, and a process at EL0 — init, // whose every page arrives by a demand fault and whose spawn of From a72442ba47c83775c0416ec7e6c939a7aeb0da72 Mon Sep 17 00:00:00 2001 From: japabu Date: Wed, 30 Sep 2026 14:26:04 +0200 Subject: [PATCH 3/6] Answer the first review of #633: the EL1 entry, the halt and the failed AP are tested B2. `virt_el1_smp` boots eight CPUs on `virt` without EL2 under TCG, the new `Profile::VirtTcg`: firmware enters every CPU at EL1, as HVF does, and QEMU's FADT names PSCI's conduit `HVC`. It asserts `through HVC` and every CPU's `as declared; entered at EL1`, which `virt_smp` now asserts at EL2 for every CPU too. The EL1 arm of `apply_declaration` is the one that fetches at a physical address under the bring-up root, so the identity half's deletion (M4, green under VirtEl2) reds here, as does `hvc #0` made `smc #0`, which is undefined at EL1 on a machine without EL3. B3. `virt_fatal_halts_the_others_first` is x86-64's `panic_halts_the_others_first` on four VirtEl2 CPUs, from the new `tests/virtpaniccase`, whose one job is `test_rs_panic_halts_first`; its refused syscall gains an `svc` form. The judge after the boot is now one function, `the_others_halt_first`, over `qemu::stopped_cpus`, which takes the architecture. QEMU prints no halt state for an AArch64 vCPU and an idle one waits in `wfi` with `I` masked too, so a stopped CPU is told by where it waits: EL1, `PSTATE.I` set, `wfi` (d503207f) just behind the PC and the branch back to it (17ffffff) at it. Every inlined `cpu::halt` in the built kernel is that pair; `hw::halt`'s idle wait is `wfi` then `msr daifclr` (d50343ff), measured with llvm-objdump. QEMU's halted PC is the instruction after `wfi` (`gen_a64_update_pc(dc, 4)` before `gen_helper_wfi`). B4. AArch64's `start` honours `smp-skip-ap`, which now lives beside the roster as `smp::skip_startup` for both architectures, and `virt_failed_ap_leaves_no_hole` runs x86-64's shape on four VirtEl2 CPUs: cpu1 online, cpu2 never echoing, `2 of 4`, no CPU 2 or 3 joining, and the job ending 0. B5. `paging::check_joined` reads `TTBR1_EL1` and `TTBR0_EL1` back after every AP's `join` and refuses a CPU on any other roots, so `join`'s deletion panics the AP and `virt_smp` loses its eighth CPU. B6. One `ROSTER`, in the new generic `kernel/src/smp.rs`, with `cpu_count`, `hardware_id`, `answering`, `set_ready` and `is_ready`, the `MAX_CPUS` assert once, and one bring-up stack (`bringup_stack`, x86-64's 64 KiB, which the AP leaves for its idle stack either way). Both architectures' copies, x86-64's `apic_id_for` and AArch64's `hardware_id_of` among them, are deleted, and every caller names `crate::smp`. B7. `unmap_seen` is gone into `unmap_touch`: one judge runs four `touch` trials and four `seen` trials, the second a child mode, so the one-CPU `virt_unmap_touch` and the eight-CPU `virt_smp` run the same job. Notes: `shard_for` refuses a CPU whose shard is already published; `toyos-acpi`'s `psci` answers `Short` for a table that ends before `ARM_BOOT_ARCH` and the minor version, so the log no longer names a 6.3 table "before 5.1"; `psci.rs`'s two asm arms are one `smccc!`; `kick_all_but_self` stays one SGI per roster CPU, since `IRM`'s broadcast also reaches an AP that echoed too late and halted with its SGIs enabled, where a pending kick wakes the masked `wfi` for good. The missing barrier before an SGI or x2APIC ICR write is filed as `issues/kernel/an-ipi-can-overtake-the-stores-it-announces.md`. The track records the owner's word for running this stage, and owes the dump's `send_nmi` reach in place of the halt test that now exists. Removed as the review listed: the comments naming `alloc_log_shard` and `AP_STARTED`, the roster header's sentence, the track's narration of the diff, and the QEMU file:line provenance in `toyos-acpi`'s tests. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L --- ...pi-can-overtake-the-stores-it-announces.md | 31 ++++ issues/kernel/toyos-runs-on-arm64.md | 24 +-- kernel-loom/tests/log_publish.rs | 2 - kernel/src/actuator.rs | 2 +- kernel/src/arch/aarch64/irqchip.rs | 9 +- kernel/src/arch/aarch64/paging.rs | 21 +++ kernel/src/arch/aarch64/psci.rs | 35 ++--- kernel/src/arch/aarch64/smp.rs | 48 ++---- kernel/src/arch/x86_64/apic.rs | 8 +- kernel/src/arch/x86_64/nmi_gate.rs | 2 +- kernel/src/arch/x86_64/percpu.rs | 1 - kernel/src/arch/x86_64/smp.rs | 56 +------ kernel/src/arch/x86_64/tlb.rs | 3 +- kernel/src/drivers/xhci/mod.rs | 2 +- kernel/src/hardlockup/mod.rs | 3 +- kernel/src/hardlockup/probe.rs | 3 +- kernel/src/heartbeat.rs | 2 +- kernel/src/hw.rs | 2 +- kernel/src/irq_census.rs | 2 +- kernel/src/log/mod.rs | 4 + kernel/src/log/registry.rs | 1 - kernel/src/main.rs | 3 +- kernel/src/process.rs | 4 +- kernel/src/quiesce.rs | 2 +- kernel/src/sched/driver.rs | 6 +- kernel/src/sched/dump.rs | 5 +- kernel/src/smp.rs | 56 +++++++ kernel/src/smp_roster.rs | 3 +- kernel/src/syscall/dispatch.rs | 2 +- kernel/src/syscall/machine.rs | 2 +- src/build.rs | 1 + tests/common/qemu.rs | 61 ++++++-- .../src/bin/panic_halts_first.rs | 7 + tests/toyos.rs | 139 ++++++++++++++---- tests/virtpaniccase/system.toml | 13 ++ tests/virtsmpcase/system.toml | 4 +- toyos-acpi/src/fadt.rs | 24 +-- toyos-acpi/tests/corpus.rs | 16 +- toyos-acpi/tests/fixtures.rs | 4 +- userland/toybox/src/main.rs | 3 +- userland/toybox/src/unmap_seen.rs | 89 ----------- userland/toybox/src/unmap_touch.rs | 94 +++++++++--- 42 files changed, 479 insertions(+), 320 deletions(-) create mode 100644 issues/kernel/an-ipi-can-overtake-the-stores-it-announces.md create mode 100644 kernel/src/smp.rs create mode 100644 tests/virtpaniccase/system.toml delete mode 100644 userland/toybox/src/unmap_seen.rs diff --git a/issues/kernel/an-ipi-can-overtake-the-stores-it-announces.md b/issues/kernel/an-ipi-can-overtake-the-stores-it-announces.md new file mode 100644 index 00000000000..8a9703ba757 --- /dev/null +++ b/issues/kernel/an-ipi-can-overtake-the-stores-it-announces.md @@ -0,0 +1,31 @@ +--- +status: open +kind: defect +opened: 2026-09-30 +--- + +# An IPI can overtake the stores it announces + +Neither architecture orders a CPU's earlier stores before the interrupt it +raises to announce them, so on hardware the target can take the interrupt and +read the word it was sent for before that word has landed. + +- **AArch64**: `kernel/src/arch/aarch64/irqchip.rs`'s `raise` writes + `ICC_SGI1R_EL1` with no barrier before it. Only a `DSB` orders a system + register write after stores to Normal memory; Linux's `gic_ipi_send_mask` + (`drivers/irqchip/irq-gic-v3.c`) issues `dsb(ishst)` before the write for + exactly this. +- **x86-64**: `kernel/src/arch/x86_64/apic.rs`'s `Reg::write` of the ICR is a + `wrmsr`, which in x2APIC mode is not serializing (SDM Vol. 3A, "MSR Access + in x2APIC Mode"); Linux fences every x2APIC IPI with `weak_wrmsr_fence` + (`mfence; lfence`). + +One reader that depends on the order: `kernel/src/sched/dump.rs` stores +`OWES` for each sibling, then kicks it, and a sibling that takes its kick +before its `OWES` is visible answers nothing, which the dump then reports as +a silent CPU. + +No guest here reds on it, and none is expected to: under TCG the write is a +call into QEMU's interrupt controller model and under HVF a trap, and neither +is the hardware's ordering. The evidence is the architecture and Linux's code, +and the exit is the barrier on both, one instruction each. diff --git a/issues/kernel/toyos-runs-on-arm64.md b/issues/kernel/toyos-runs-on-arm64.md index 44531bf91a3..ac1a81c864c 100644 --- a/issues/kernel/toyos-runs-on-arm64.md +++ b/issues/kernel/toyos-runs-on-arm64.md @@ -28,6 +28,8 @@ Every x86 guest on this host runs under TCG emulation instead — there is no ## Owner rulings, 2026-09-26 +- **Running** (owner, 2026-09-29): "can we get the most important tracks + running in parallel? arm, network and llvm?" - **Hardware discovery is ACPI only** (edk2 MADT, GTDT and SPCR under QEMU); no devicetree. - **Start.** Stage 0 (shared groundwork on x86) and stages 1-3 (toolchain, loader, kernel reaching serial on QEMU `virt`) start now. Stage 0 waits for @@ -292,21 +294,19 @@ Each stage names its exit; "measured" means a number from a run. `dlopen` test; until it does the kernel refuses `R_AARCH64_TLSDESC` by name (`toyos_elf::rela::ExeRefusal::TlsDescriptor` for an executable, `toyos_elf::RelocError::TlsDescriptor` for a library). - **Every CPU starts, ahead of small-kernel stage 6 as stage 4 did:** it - ports no device interrupt, so nothing of the relay. The boot CPU starts - each GIC CPU interface the MADT enables with `CPU_ON`, through the conduit - the FADT's `ARM_BOOT_ARCH` names; each AP applies the declaration through - the boot CPU's own routine, installs its per-CPU block, redistributor and - timer, and echoes its attempt's token into the roster both architectures - share before it joins the scheduler. SGIs are the kick and the halt. - `virt_smp` judges eight CPUs, each holding the declaration, and a page - unmapped beside a thread still reading it. Owed before the exit holds: + **Every CPU starts, ahead of small-kernel stage 6 by the owner's word, as + stage 4 did:** it ports no device interrupt, so nothing of the relay. + Owed before the exit holds: `SYSTEM_RESET`, `SYSTEM_OFF` and `CPU_OFF` behind a reset and power-off seam that takes x86-64's reset register and PM1a out of `drivers/acpi.rs`, and the stop shown on eight CPUs ending in that - power-off; a test of the halt SGI; the TLS-descriptor resolver; - `issues/kernel/the-crash-evidence-records-x86-fault-registers.md`; and, - for the first HVF run, the clean of an AP's start block to the point of + power-off; the TLS-descriptor resolver; + `issues/kernel/the-crash-evidence-records-x86-fault-registers.md`; the + blocked-task dump's probe of a CPU that ignored its kick + (`sched/dump.rs`'s `probe_silent`), which reaches `irqchip::send_nmi`'s + `owed!` on a machine of more than one CPU, and which nothing but the + `dump-deaf-cpu` actuator asks for until AArch64 has a keyboard; and, for + the first HVF run, the clean of an AP's start block to the point of coherency, which TCG cannot fail on. 6. **Virtio on `virt`.** virtio-pci (ECAM from MCFG) for blk, net, gpu, diff --git a/kernel-loom/tests/log_publish.rs b/kernel-loom/tests/log_publish.rs index 31479208111..cdf8bf48139 100644 --- a/kernel-loom/tests/log_publish.rs +++ b/kernel-loom/tests/log_publish.rs @@ -40,8 +40,6 @@ fn a_reader_that_finds_a_shard_finds_it_built() { let publisher = registry.clone(); let w = loom::thread::spawn(move || { - // What `alloc_log_shard` does, in the order it does it: build the - // shard, then make it reachable. let shard = Box::into_raw(Box::new(Shard::new())); // SAFETY: leaked, so it outlives every reader; published once. unsafe { publish(&publisher[..], 1, shard) }; diff --git a/kernel/src/actuator.rs b/kernel/src/actuator.rs index 7e65a02bbf9..11b3182eb08 100644 --- a/kernel/src/actuator.rs +++ b/kernel/src/actuator.rs @@ -372,7 +372,7 @@ actuators! { /// Leave every AP holding the CR0/CR4 that INIT left it. no_ap_control_regs = "no-ap-control-regs"; - /// Skip the startup IPI for the AP that would be cpu2, so a non-last AP never starts. + /// Skip the startup for the AP that would be cpu2, so a non-last AP never starts. smp_skip_ap = "smp-skip-ap"; /// Time the same read loop on every CPU, either side of the `mov cr0` that enables caching. diff --git a/kernel/src/arch/aarch64/irqchip.rs b/kernel/src/arch/aarch64/irqchip.rs index a7731e738d3..df3f354febd 100644 --- a/kernel/src/arch/aarch64/irqchip.rs +++ b/kernel/src/arch/aarch64/irqchip.rs @@ -291,12 +291,15 @@ fn sgi(intid: u32, target: u32) { /// Wake `cpu` so it runs a scheduler pass. pub fn kick_cpu(cpu: u32) { - sgi(SGI_KICK, super::smp::hardware_id_of(cpu)); + sgi(SGI_KICK, crate::smp::hardware_id(cpu)); } +// Each by name, not with `IRM`'s broadcast: that reaches an AP which echoed too +// late and halted with its SGIs enabled, where a kick left pending wakes the +// masked `wfi` at once, for good. pub fn kick_all_but_self() { let me = percpu::cpu_id(); - for cpu in (0..super::smp::cpu_count()).filter(|&cpu| cpu != me) { + for cpu in (0..crate::smp::cpu_count()).filter(|&cpu| cpu != me) { kick_cpu(cpu); } } @@ -316,7 +319,7 @@ pub fn send_nmi(_cpu: u32) { /// that each is waiting for the release with interrupts masked, and the boot /// CPU's own interface may not be up to raise anything. pub fn stop_other_cpus() { - if super::smp::is_ready() { + if crate::smp::is_ready() { raise(toyos_gicv3::sgi1r_others(SGI_HALT)); } } diff --git a/kernel/src/arch/aarch64/paging.rs b/kernel/src/arch/aarch64/paging.rs index 06d954e9b75..4d40e40dc3b 100644 --- a/kernel/src/arch/aarch64/paging.rs +++ b/kernel/src/arch/aarch64/paging.rs @@ -697,6 +697,27 @@ pub(crate) fn join() { } } +/// What [`join`] wrote, read back: a CPU on any other roots is refused. +pub(crate) fn check_joined() { + let (ttbr1, ttbr0): (u64, u64); + // SAFETY: reads two translation registers; touches no memory. + unsafe { + core::arch::asm!( + "mrs {ttbr1}, ttbr1_el1", + "mrs {ttbr0}, ttbr0_el1", + ttbr1 = out(reg) ttbr1, + ttbr0 = out(reg) ttbr0, + options(nomem, nostack, preserves_flags), + ); + } + let root1 = KERNEL_TTBR1.load(core::sync::atomic::Ordering::Acquire); + let root0 = KERNEL_TTBR0.load(core::sync::atomic::Ordering::Acquire); + assert!( + (ttbr1, ttbr0) == (root1, root0), + "paging: this CPU holds TTBR1_EL1={ttbr1:#x} TTBR0_EL1={ttbr0:#x}, and the kernel's roots are {root1:#x} and {root0:#x}" + ); +} + /// A root for a CPU to turn its MMU on under, running at physical addresses: /// the kernel root's direct map, and the same entries again at identity below /// it, so one root serves both `TTBR`s as the loader's does. Its tables below diff --git a/kernel/src/arch/aarch64/psci.rs b/kernel/src/arch/aarch64/psci.rs index 4abbfd8dc21..e4df148df5e 100644 --- a/kernel/src/arch/aarch64/psci.rs +++ b/kernel/src/arch/aarch64/psci.rs @@ -62,29 +62,26 @@ impl Conduit { /// callee's to clobber under SMCCC 1.0, so they are declared so. fn call(self, function: u32, a1: u64, a2: u64, a3: u64) -> i32 { let answer: u64; - // SAFETY: an SMCCC call into firmware or a hypervisor, which returns - // to the next instruction; not `nomem`, since a callee may read what - // this CPU stored before it. - unsafe { - match self { - Self::Smc => core::arch::asm!( - "smc #0", + macro_rules! smccc { + ($instruction:literal) => { + core::arch::asm!( + $instruction, inout("x0") u64::from(function) => answer, inout("x1") a1 => _, inout("x2") a2 => _, inout("x3") a3 => _, out("x4") _, out("x5") _, out("x6") _, out("x7") _, out("x8") _, out("x9") _, out("x10") _, out("x11") _, out("x12") _, out("x13") _, out("x14") _, out("x15") _, out("x16") _, out("x17") _, options(nostack), - ), - Self::Hvc => core::arch::asm!( - "hvc #0", - inout("x0") u64::from(function) => answer, - inout("x1") a1 => _, inout("x2") a2 => _, inout("x3") a3 => _, - out("x4") _, out("x5") _, out("x6") _, out("x7") _, out("x8") _, out("x9") _, - out("x10") _, out("x11") _, out("x12") _, out("x13") _, out("x14") _, - out("x15") _, out("x16") _, out("x17") _, - options(nostack), - ), + ) + }; + } + // SAFETY: an SMCCC call into firmware or a hypervisor, which returns + // to the next instruction; not `nomem`, since a callee may read what + // this CPU stored before it. + unsafe { + match self { + Self::Smc => smccc!("smc #0"), + Self::Hvc => smccc!("hvc #0"), } } answer as i32 @@ -129,6 +126,10 @@ pub fn init(rsdp_addr: u64) -> Option { log!("PSCI: none: the FADT is version {revision}.{minor}, and ARM_BOOT_ARCH begins at 5.1"); return None; } + Psci::Short => { + log!("PSCI: none: the FADT ends before ARM_BOOT_ARCH and the minor version after it"); + return None; + } }; let version = conduit.call(PSCI_VERSION, 0, 0, 0); if version < 0 { diff --git a/kernel/src/arch/aarch64/smp.rs b/kernel/src/arch/aarch64/smp.rs index dff23623ca0..a2ef20417d0 100644 --- a/kernel/src/arch/aarch64/smp.rs +++ b/kernel/src/arch/aarch64/smp.rs @@ -10,14 +10,10 @@ use toyos_gicv3::packed_affinity; use super::percpu::{self, PerCpu}; use super::{cache, control_regs, cpu, irqchip, paging, psci}; use crate::mm::DirectMap; -use crate::smp_roster::Roster; +use crate::smp::{self, ROSTER}; use crate::time::AP_START; use crate::{clock, log}; -static ROSTER: Roster = Roster::new(); - -const _: () = assert!(crate::smp_roster::MAX_CPUS == crate::scheduler::MAX_CPUS); - /// What an AP's entry reads with its MMU off, at the physical address `CPU_ON` /// hands it in `x0`, and what its first Rust reads after. #[repr(C)] @@ -32,30 +28,6 @@ pub struct ApStart { token: u32, } -/// A stack for an AP's bring-up, never freed: the AP runs on it. -#[repr(C, align(16))] -struct Stack([u8; crate::process::KERNEL_STACK_SIZE]); - -pub fn cpu_count() -> u32 { - ROSTER.count() -} - -/// Release the APs into the scheduler. -pub fn set_ready() { - ROSTER.release(); -} - -/// Whether [`set_ready`] has run. -pub fn is_ready() -> bool { - ROSTER.released() -} - -/// `cpu`'s packed MPIDR affinity; panics if `cpu` is not online. -pub fn hardware_id_of(cpu: u32) -> u32 { - assert!(cpu < cpu_count(), "smp: cpu{cpu} is not online"); - ROSTER.hardware_id(cpu) -} - /// Start every other CPU `gic` names, in the MADT's order, until one does not /// echo within [`AP_START`] or the roster is full. pub fn start(gic: &irqchip::Gic, psci: Option) { @@ -64,20 +36,19 @@ pub fn start(gic: &irqchip::Gic, psci: Option) { let others = gic.cpus.iter().filter(|gicc| packed_affinity(gicc.mpidr) != me); let Some(psci) = psci else { log!("SMP: no PSCI to start the other {} CPUs with; the boot CPU runs alone", others.count()); - control_regs::report(cpu_count()); + control_regs::report(smp::cpu_count()); return; }; let root = paging::bringup_root(); let entry = DirectMap::phys_of(super::boot::ap_start as *const u8); for gicc in others { let Some(attempt) = ROSTER.begin_attempt() else { - log!("SMP: roster full at {} CPUs; ignoring further MADT entries", cpu_count()); + log!("SMP: roster full at {} CPUs; ignoring further MADT entries", smp::cpu_count()); break; }; - let stack = Box::leak(Box::::new_uninit()); let start: &'static ApStart = Box::leak(Box::new(ApStart { root, - stack_top: stack.as_ptr() as u64 + size_of::() as u64, + stack_top: smp::bringup_stack(), percpu: percpu::alloc(attempt.id()), redistributor: gic.redistributor(packed_affinity(gicc.mpidr)), token: attempt.token(), @@ -85,9 +56,11 @@ pub fn start(gic: &irqchip::Gic, psci: Option) { // Read with the MMU off, so from memory and never from this CPU's cache. cache::write_back(start as *const ApStart as u64, size_of::()); let context = DirectMap::phys_of(start as *const ApStart); - if let Err(refused) = psci.cpu_on(gicc.mpidr, entry, context) { - log!("SMP: CPU_ON refused cpu{} mpidr={:#x} ({refused:?}); the rest stay off", attempt.id(), gicc.mpidr); - break; + if !smp::skip_startup(attempt.id()) { + if let Err(refused) = psci.cpu_on(gicc.mpidr, entry, context) { + log!("SMP: CPU_ON refused cpu{} mpidr={:#x} ({refused:?}); the rest stay off", attempt.id(), gicc.mpidr); + break; + } } let deadline = clock::nanos_since_boot() + AP_START.nanos(); // Committed only on this attempt's own token, so `0..cpu_count()` stays dense. @@ -98,7 +71,7 @@ pub fn start(gic: &irqchip::Gic, psci: Option) { ROSTER.commit(attempt, packed_affinity(gicc.mpidr)); log!("SMP: cpu{} mpidr={:#x} online", attempt.id(), gicc.mpidr); } - let online = cpu_count(); + let online = smp::cpu_count(); log!("SMP: {online} of {} MADT CPUs online", gic.cpus.len()); control_regs::report(online); } @@ -109,6 +82,7 @@ pub(super) extern "C" fn ap_entry(start: &'static ApStart, el: u64) -> ! { // First: every log line, a fault's report among them, reads it. percpu::install(start.percpu); paging::join(); + paging::check_joined(); control_regs::check(el); irqchip::init_cpu(start.redistributor); ROSTER.echo(start.token); diff --git a/kernel/src/arch/x86_64/apic.rs b/kernel/src/arch/x86_64/apic.rs index 31b79029685..e8f54fd34c9 100644 --- a/kernel/src/arch/x86_64/apic.rs +++ b/kernel/src/arch/x86_64/apic.rs @@ -147,14 +147,14 @@ pub(super) fn tlb_ipi() { // Targeted, not broadcast: a broadcast kick would preempt every sibling per wake and cannot scale. pub fn kick_cpu(cpu_id: u32) { if !X2APIC_ENABLED.load(Ordering::Relaxed) { return; } - let apic_id = crate::arch::smp::apic_id_for(cpu_id); + let apic_id = crate::smp::hardware_id(cpu_id); Reg::Icr.write(((apic_id as u64) << 32) | 0x4000 | TIMER_VECTOR as u64); } // Kicked, and not left to arrive on their own: a CPU halted in the idle path has stopped its own timer, so nothing else brings it to the next scheduler pass. pub fn kick_all_but_self() { let me = percpu::cpu_id(); - for cpu in 0..crate::arch::smp::cpu_count() { + for cpu in 0..crate::smp::cpu_count() { if cpu != me { kick_cpu(cpu); } @@ -165,7 +165,7 @@ pub fn kick_all_but_self() { // Diagnostic only: an NMI can land inside any critical section, which this kernel cannot make NMI-safe. pub fn send_nmi(cpu_id: u32) { if !X2APIC_ENABLED.load(Ordering::Relaxed) { return; } - let apic_id = crate::arch::smp::apic_id_for(cpu_id); + let apic_id = crate::smp::hardware_id(cpu_id); Reg::Icr.write(((apic_id as u64) << 32) | 0x4400); } @@ -189,7 +189,7 @@ pub fn arm_perf_nmi() { /// that no sibling has been sent its `SIPI`, so there are only CPUs still waiting /// for one rather than CPUs that need halting. pub fn stop_other_cpus() { - if X2APIC_ENABLED.load(Ordering::Relaxed) && crate::arch::smp::is_ready() { + if X2APIC_ENABLED.load(Ordering::Relaxed) && crate::smp::is_ready() { Reg::Icr.write(0x000C_0000 | 0xFD); } } diff --git a/kernel/src/arch/x86_64/nmi_gate.rs b/kernel/src/arch/x86_64/nmi_gate.rs index 42b47045fde..d1d49b24753 100644 --- a/kernel/src/arch/x86_64/nmi_gate.rs +++ b/kernel/src/arch/x86_64/nmi_gate.rs @@ -291,7 +291,7 @@ fn hold_one(me: usize, cpus: usize) -> bool { pub fn storm() { static FIRED: AtomicBool = AtomicBool::new(false); - let cpus = (crate::arch::smp::cpu_count() as usize).min(MAX_CPUS); + let cpus = (crate::smp::cpu_count() as usize).min(MAX_CPUS); let me = percpu::cpu_id() as usize; if cpus < 2 { return; diff --git a/kernel/src/arch/x86_64/percpu.rs b/kernel/src/arch/x86_64/percpu.rs index 8e5df63dd82..f4ba2446624 100644 --- a/kernel/src/arch/x86_64/percpu.rs +++ b/kernel/src/arch/x86_64/percpu.rs @@ -102,7 +102,6 @@ pub struct PerCpu { log_shard: u64, /// Non-zero inside this CPU's NMI handler, written only by `arch::idt::nmi`'s entry; IST2 isn't re-entrant, so this proves no second NMI lands on it. nmi_active: u32, - /// The token of the attempt that booted this AP; the AP echoes it into `AP_STARTED` so a stale AP cannot answer for a later attempt. Zero on the BSP. ap_token: u32, /// `nmi_gate::hold`'s word: the storm asks in it from another CPU, and `arch::syscall`'s entry acknowledges and spins on it inside its window, through [`OFF_NMI_HOLD`]. #[cfg(feature = "boot-actuators")] diff --git a/kernel/src/arch/x86_64/smp.rs b/kernel/src/arch/x86_64/smp.rs index a38c792f584..987e8ff4719 100644 --- a/kernel/src/arch/x86_64/smp.rs +++ b/kernel/src/arch/x86_64/smp.rs @@ -2,23 +2,18 @@ use core::arch::global_asm; use core::mem::size_of; use core::sync::atomic::{AtomicU64, Ordering}; -use alloc::alloc::{alloc_zeroed, Layout}; - use crate::arch::cpu; use crate::arch::{apic, percpu, syscall}; use crate::clock; use crate::drivers::acpi::MadtInfo; -use crate::smp_roster::Roster; +use crate::smp::{self, ROSTER}; use crate::time::{Delay, Duration, AP_START}; use crate::{log, process}; const TRAMPOLINE_PAGE: u64 = 0x8000; const TRAMPOLINE_VECTOR: u8 = 0x08; -const AP_STACK_SIZE: usize = 64 * 1024; const DATA_OFFSET: usize = 0xF00; -static ROSTER: Roster = Roster::new(); - /// The `rdtsc` the AP being started read on its first instruction in Rust. /// /// **The whole of the TSC-trail measurement.** `clock::nanos_since_boot` @@ -29,36 +24,6 @@ static ROSTER: Roster = Roster::new(); /// and a skewed one is outside it by at least the skew. static AP_TSC: AtomicU64 = AtomicU64::new(0); -const _: () = assert!(crate::smp_roster::MAX_CPUS == crate::scheduler::MAX_CPUS); - -pub fn cpu_count() -> u32 { - ROSTER.count() -} - -/// LAPIC id of `cpu_id`; panics if `cpu_id` is not online. -pub fn apic_id_for(cpu_id: u32) -> u32 { - assert!(cpu_id < cpu_count(), "apic_id_for: cpu {cpu_id} not online"); - ROSTER.hardware_id(cpu_id) -} - -/// True once a shootdown must wait for siblings; the word the APs are released by. -pub(crate) fn answering() -> bool { - ROSTER.answering() -} - -/// Release the APs into the scheduler and, by the same store, start answering their shootdowns. -pub fn set_ready() { - ROSTER.release(); -} - -/// Whether [`set_ready`] has run: the machine's own word for "the scheduler is -/// what runs now". `kernel_main` calls it immediately before -/// `scheduler::enter_idle_loop`, so a `false` here means no task and no -/// kernel thread can make progress, and the panic path waits for none of them. -pub fn is_ready() -> bool { - ROSTER.released() -} - // Field offsets are hardcoded in the global_asm! trampoline below; the static assertion at the bottom checks the match. #[derive(Clone, Copy)] @@ -215,19 +180,13 @@ pub fn boot_aps(madt: &MadtInfo, boot_cr3: u64) { // Refused at MAX_CPUS, bounding a firmware that over-reports CPUs. let Some(attempt) = ROSTER.begin_attempt() else { - log!("SMP: roster full at {} CPUs; ignoring further MADT entries", cpu_count()); + log!("SMP: roster full at {} CPUs; ignoring further MADT entries", smp::cpu_count()); break; }; - let stack_layout = Layout::from_size_align(AP_STACK_SIZE, 4096).unwrap(); - // SAFETY: `AP_STACK_SIZE` is non-zero and 4096 is a power of two, `alloc_zeroed`'s whole contract. - // Never freed: the block becomes an AP's `rsp`, so no owning handle can hold it. - let stack_base = unsafe { alloc_zeroed(stack_layout) }; - assert!(!stack_base.is_null(), "SMP: failed to allocate AP stack"); - let ap_percpu = percpu::alloc_ap(attempt.id(), attempt.token()); - data.stack_top = stack_base as u64 + AP_STACK_SIZE as u64; + data.stack_top = smp::bringup_stack(); data.entry = ap_entry as *const () as u64; data.percpu_ptr = ap_percpu as u64; // SAFETY: `target` is reserved physical memory written only here; unaligned because `TrampolineData` is `repr(C, packed)`; this AP has not been sent its SIPI yet, and the loop reaches a second write only after the previous AP committed — which happens in `ap_entry` past every trampoline read — so no CPU is reading it; a failed AP breaks the loop below instead of reaching another write. @@ -244,7 +203,7 @@ pub fn boot_aps(madt: &MadtInfo, boot_cr3: u64) { /// SDM §8.4.4.1 asks 200us here; a millisecond gives room and is paid once per AP. const BETWEEN_SIPIS: Delay = Delay::from_spec(Duration::from_millis(1), "SDM §8.4.4.1, between the two SIPIs"); - if !skip_startup(attempt.id()) { + if !smp::skip_startup(attempt.id()) { apic::send_init(ap_id); delay(AFTER_INIT); @@ -271,7 +230,7 @@ pub fn boot_aps(madt: &MadtInfo, boot_cr3: u64) { break; } } - let online = cpu_count(); + let online = smp::cpu_count(); log!( "SMP: {online} of {} MADT cpus online, {bracketed} of {} APs inside the BSP's TSC bracket", madt.apic_ids.len(), @@ -310,11 +269,6 @@ fn tsc_inside(cpu_id: u32, lo: u64, hi: u64) -> bool { false } -/// The actuator staging a non-last AP that never starts; `false` without `boot-actuators`. -fn skip_startup(id: u32) -> bool { - crate::actuator::smp_skip_ap() && id == 2 -} - extern "C" fn ap_entry() -> ! { // First, so the reading is this CPU's TSC and not a measure of what runs // after it; `rdtsc` needs nothing this AP has not got out of the trampoline. diff --git a/kernel/src/arch/x86_64/tlb.rs b/kernel/src/arch/x86_64/tlb.rs index df1ecc0ddcd..a6565cd2f12 100644 --- a/kernel/src/arch/x86_64/tlb.rs +++ b/kernel/src/arch/x86_64/tlb.rs @@ -14,7 +14,8 @@ use core::sync::atomic::{AtomicU64, Ordering}; use crate::shootdown::{Generation, Shootdown}; -use super::{apic, percpu, smp}; +use super::{apic, percpu}; +use crate::smp; static SHOOTDOWN: Shootdown = Shootdown::new(); diff --git a/kernel/src/drivers/xhci/mod.rs b/kernel/src/drivers/xhci/mod.rs index 781923eef4d..479a5e6b146 100644 --- a/kernel/src/drivers/xhci/mod.rs +++ b/kernel/src/drivers/xhci/mod.rs @@ -680,7 +680,7 @@ pub(super) fn ports_wanted(kick: bool) { return; } let me = crate::arch::percpu::cpu_id(); - for cpu in 0..crate::arch::smp::cpu_count() { + for cpu in 0..crate::smp::cpu_count() { if cpu != me { crate::arch::irqchip::kick_cpu(cpu); } diff --git a/kernel/src/hardlockup/mod.rs b/kernel/src/hardlockup/mod.rs index b91a0022b23..e3546745db5 100644 --- a/kernel/src/hardlockup/mod.rs +++ b/kernel/src/hardlockup/mod.rs @@ -59,7 +59,8 @@ use core::fmt; use core::sync::atomic::{AtomicBool, AtomicU64, Ordering::{Acquire, Relaxed, Release}}; -use crate::arch::{cpu, percpu, pmu, smp, trap}; +use crate::arch::{cpu, percpu, pmu, trap}; +use crate::smp; use crate::sched::MAX_CPUS; /// The negative control, in a file of its own because it says what it staged diff --git a/kernel/src/hardlockup/probe.rs b/kernel/src/hardlockup/probe.rs index c68a6204df4..4cc41fab50c 100644 --- a/kernel/src/hardlockup/probe.rs +++ b/kernel/src/hardlockup/probe.rs @@ -25,7 +25,8 @@ use core::sync::atomic::{AtomicBool, AtomicU32, Ordering}; -use crate::arch::{irqchip, cpu, percpu, smp}; +use crate::arch::{irqchip, cpu, percpu}; +use crate::smp; /// What the staged CPU says before it stops answering, and the witness a sealed /// record carries in its tail: a `WEDGED` page whose text does not hold this diff --git a/kernel/src/heartbeat.rs b/kernel/src/heartbeat.rs index 30d21b1920f..d20758ad757 100644 --- a/kernel/src/heartbeat.rs +++ b/kernel/src/heartbeat.rs @@ -68,7 +68,7 @@ pub fn poll() { { return; } - let cpus = (crate::arch::smp::cpu_count() as usize).min(MAX_CPUS); + let cpus = (crate::smp::cpu_count() as usize).min(MAX_CPUS); let dispatched: u64 = (0..cpus).map(|c| DISPATCHED[c].load(Ordering::Relaxed)).sum(); // Saturating, not `-`: a diagnostic must not be the thing that panics. let ran = dispatched.saturating_sub(LAST_DISPATCHED.swap(dispatched, Ordering::Relaxed)); diff --git a/kernel/src/hw.rs b/kernel/src/hw.rs index c3268132ba5..041eaeb2fd4 100644 --- a/kernel/src/hw.rs +++ b/kernel/src/hw.rs @@ -120,7 +120,7 @@ pub(crate) fn note_running(ctx: *const KernelCtx) { /// could try to take. pub fn report_contexts(sp: u64, subject: Option) { let me = percpu::cpu_id() as usize; - let count = (crate::arch::smp::cpu_count() as usize).min(crate::sched::MAX_CPUS); + let count = (crate::smp::cpu_count() as usize).min(crate::sched::MAX_CPUS); let mine = RUNNING_CTX.get(me).map_or(0, |slot| slot.load(Relaxed)); let subject = subject.unwrap_or(mine); crate::log!(" Contexts: cpu{me} crashed at sp={sp:#018x}, asking about ctx {subject:#x}"); diff --git a/kernel/src/irq_census.rs b/kernel/src/irq_census.rs index 29782cdcfc0..6187df24d8f 100644 --- a/kernel/src/irq_census.rs +++ b/kernel/src/irq_census.rs @@ -135,7 +135,7 @@ pub fn taken_here() -> u64 { /// Logs one `irq: cpuN total=… =…` line per online CPU; counts are cumulative since boot. /// Allocates nothing, takes no lock, touches no device. pub fn log_census() { - for cpu in 0..crate::arch::smp::cpu_count() { + for cpu in 0..crate::smp::cpu_count() { let Some(counts) = read(cpu) else { continue }; crate::log!( "irq: cpu{cpu} total={}{}", diff --git a/kernel/src/log/mod.rs b/kernel/src/log/mod.rs index 1c4085059a0..1b12db52290 100644 --- a/kernel/src/log/mod.rs +++ b/kernel/src/log/mod.rs @@ -40,6 +40,10 @@ pub fn shard_for(cpu: u32) -> &'static Shard { if cpu == 0 { return &BOOT_SHARD; } + assert!( + registry::published(registry::kernel_slots(), cpu as usize - 1).is_none(), + "log: cpu{cpu} already has a shard, and a second would hide every record written to the first", + ); let layout = alloc::alloc::Layout::new::(); // SAFETY: a `Shard` is not zero-sized; the block is never freed. let ptr = unsafe { alloc::alloc::alloc_zeroed(layout) }.cast::(); diff --git a/kernel/src/log/registry.rs b/kernel/src/log/registry.rs index 2113cd2e031..fa595ae5f3c 100644 --- a/kernel/src/log/registry.rs +++ b/kernel/src/log/registry.rs @@ -37,7 +37,6 @@ pub fn slots() -> Slots { /// Publishes `shard` as the shard for `cpu`. /// # Safety /// `shard` must be a live, initialised [`Shard`] that is never freed. -// cpu0 never calls this: `alloc_log_shard` returns before reaching it, so a zero `cpu` is a caller bug, not a valid case. pub unsafe fn publish(slots: &[AtomicPtr], cpu: u32, shard: *mut Shard) { let slot = (cpu as usize) .checked_sub(1) diff --git a/kernel/src/main.rs b/kernel/src/main.rs index e77c5dd8281..cfdf2b03b66 100644 --- a/kernel/src/main.rs +++ b/kernel/src/main.rs @@ -15,6 +15,7 @@ pub use mm::{UserAddr, DirectMap, PHYS_OFFSET}; mod invalidation; mod shootdown; mod sleeplock; +mod smp; mod smp_roster; mod sync; mod hasher; @@ -115,7 +116,7 @@ mod late_panic { use crate::mm::policy::MmioPolicy; use alloc::boxed::Box; use alloc::sync::Arc; -use arch::{cpu, percpu, smp}; +use arch::{cpu, percpu}; use drivers::{acpi, gop, nvme, pci, serial, virtio_console, virtio_gpu, virtio_sound, xhci}; use toyos_abi::boot::{KernelArgs, MemoryMapEntry}; use toyos_rootimage::handoff::{held, Descriptor}; diff --git a/kernel/src/process.rs b/kernel/src/process.rs index cc6e9bdc2ef..c7ebfbdc078 100644 --- a/kernel/src/process.rs +++ b/kernel/src/process.rs @@ -1592,14 +1592,14 @@ pub fn handle_fault(error: crate::object::HandleError) -> ! { /// An AP's way into the idle loop on either architecture, once it has answered /// the BSP: it waits for the machine's release, then joins the scheduler. pub fn ap_idle() -> ! { - while !crate::arch::smp::is_ready() { + while !crate::smp::is_ready() { core::hint::spin_loop(); } // Only a committed CPU may join: an uncommitted AP has no scheduler slot and // no shootdown targets it, so it halts. The acquire above makes the count visible. let me = percpu::cpu_id(); - if me >= crate::arch::smp::cpu_count() { + if me >= crate::smp::cpu_count() { log!("CPU {me}: bring-up did not commit; halting"); crate::arch::cpu::halt(); } diff --git a/kernel/src/quiesce.rs b/kernel/src/quiesce.rs index b343638287e..9f0b463c14c 100644 --- a/kernel/src/quiesce.rs +++ b/kernel/src/quiesce.rs @@ -173,7 +173,7 @@ pub fn stop() -> Record { .expect("quiesce::stop: the caller holds no task to park"); // The kick is the timer vector, whose return to Ring 3 is the gate. crate::arch::irqchip::kick_all_but_self(); - let cpus = crate::arch::smp::cpu_count(); + let cpus = crate::smp::cpu_count(); let began = crate::clock::now(); let deadline = Deadline::at(began + PARK.duration()); diff --git a/kernel/src/sched/driver.rs b/kernel/src/sched/driver.rs index 7b9797324a3..d917a8c5ff9 100644 --- a/kernel/src/sched/driver.rs +++ b/kernel/src/sched/driver.rs @@ -169,7 +169,7 @@ fn next_key() -> TaskKey { } pub fn total_cpu_ns() -> u64 { - (0..crate::arch::smp::cpu_count() as usize) + (0..crate::smp::cpu_count() as usize) .map(|i| CPU_TIME_NS[i].0.load(Ordering::Relaxed)) .sum() } @@ -352,7 +352,7 @@ fn try_with_cpu(f: impl FnOnce(&CpuSched) -> R) -> Option { /// Build every CPU's mailbox and handle, and the BSP's `CpuSched`. Called once, before any task exists. pub fn init() { - let count = crate::arch::smp::cpu_count() as usize; + let count = crate::smp::cpu_count() as usize; assert!(count <= MAX_CPUS, "cpu count {count} exceeds MAX_CPUS"); let mut handles = Vec::with_capacity(count); // A CPU number, not a walk of `SCHEDS`: `SCHEDS` is `MAX_CPUS` long whatever `count` is. @@ -389,7 +389,7 @@ fn idle_ctx() -> KernelCtx { /// boot since every init program is spawned before any CPU has published a load. fn placement(now: Nanos) -> CpuId { static ROTATE: AtomicU64 = AtomicU64::new(0); - let count = crate::arch::smp::cpu_count() as u64; + let count = crate::smp::cpu_count() as u64; let start = CpuId((ROTATE.fetch_add(1, Ordering::Relaxed) % count) as u32); cpus().place(start, now) } diff --git a/kernel/src/sched/dump.rs b/kernel/src/sched/dump.rs index 3a5a8ccd2fb..a8ff114a85f 100644 --- a/kernel/src/sched/dump.rs +++ b/kernel/src/sched/dump.rs @@ -18,7 +18,8 @@ use core::sync::atomic::{AtomicBool, AtomicU64, Ordering}; pub use super::dump_request::Entered; use super::dump_request::{DumpRequest, Left}; -use crate::arch::{irqchip, percpu, smp}; +use crate::arch::{irqchip, percpu}; +use crate::smp; use crate::sched::payload::{SCHED_BLOCKED, SCHED_READY, SCHED_RUNNING}; use crate::time::{Budget, Duration, Floor}; @@ -450,7 +451,7 @@ pub mod staged { }; if !ARMED.load(Ordering::Acquire) { // The release: every CPU has joined, so the count above is the machine's. - if crate::arch::smp::is_ready() && !ARMED.swap(true, Ordering::AcqRel) { + if crate::smp::is_ready() && !ARMED.swap(true, Ordering::AcqRel) { log!("dump-in-blocking-pass: armed"); } return; diff --git a/kernel/src/smp.rs b/kernel/src/smp.rs new file mode 100644 index 00000000000..6c9c667a625 --- /dev/null +++ b/kernel/src/smp.rs @@ -0,0 +1,56 @@ +//! The machine's CPUs: the one [`ROSTER`] each architecture's bring-up +//! (`arch::smp`) commits its APs to, what the rest of the kernel asks of it, +//! and what starting an AP takes on either architecture. + +use core::mem::size_of; + +use alloc::boxed::Box; + +use crate::smp_roster::{Roster, MAX_CPUS}; + +pub static ROSTER: Roster = Roster::new(); + +const _: () = assert!(MAX_CPUS == crate::scheduler::MAX_CPUS); + +pub fn cpu_count() -> u32 { + ROSTER.count() +} + +/// `cpu`'s hardware id — its LAPIC id on x86-64, its packed MPIDR affinity on +/// AArch64; panics if `cpu` is not online. +pub fn hardware_id(cpu: u32) -> u32 { + assert!(cpu < cpu_count(), "smp: cpu{cpu} is not online"); + ROSTER.hardware_id(cpu) +} + +/// True once a shootdown must wait for siblings; the word the APs are released by. +pub fn answering() -> bool { + ROSTER.answering() +} + +/// Release the APs into the scheduler and, by the same store, start answering their shootdowns. +pub fn set_ready() { + ROSTER.release(); +} + +/// Whether [`set_ready`] has run: the machine's own word for "the scheduler is +/// what runs now". `kernel_main` calls it immediately before +/// `scheduler::enter_idle_loop`, so a `false` here means no task and no +/// kernel thread can make progress, and the panic path waits for none of them. +pub fn is_ready() -> bool { + ROSTER.released() +} + +/// The actuator staging a non-last AP that never starts; `false` without `boot-actuators`. +pub fn skip_startup(id: u32) -> bool { + crate::actuator::smp_skip_ap() && id == 2 +} + +/// The top of a stack an AP runs on from its entry until its idle loop, which +/// has a stack of its own. Never freed: nothing says when the AP has left it. +pub fn bringup_stack() -> u64 { + #[repr(C, align(16))] + struct Stack([u8; 64 * 1024]); + let stack = Box::leak(Box::::new_uninit()); + stack.as_ptr() as u64 + size_of::() as u64 +} diff --git a/kernel/src/smp_roster.rs b/kernel/src/smp_roster.rs index b85d444b5bb..c881952f1d6 100644 --- a/kernel/src/smp_roster.rs +++ b/kernel/src/smp_roster.rs @@ -1,8 +1,7 @@ //! The CPU roster and the one release/answer word, two invariants held as a type: //! an id commits only after its AP's handshake and `commit` publishes the slot //! before the count, so `0..count()` has no dead slot; and the word `release` sets -//! is the word `answering` reads. Both architectures bring their APs up through -//! it. Compiled into `kernel-loom/`, so no `crate::`. +//! is the word `answering` reads. Compiled into `kernel-loom/`, so no `crate::`. #[cfg(not(feature = "loom"))] use core::sync::atomic::{AtomicBool, AtomicU32, Ordering}; diff --git a/kernel/src/syscall/dispatch.rs b/kernel/src/syscall/dispatch.rs index 07110371438..cfe4349ac6a 100644 --- a/kernel/src/syscall/dispatch.rs +++ b/kernel/src/syscall/dispatch.rs @@ -349,7 +349,7 @@ pub(crate) fn syscall_dispatch(num: u64, a1: u64, a2: u64, a3: u64, a4: u64) -> Err(e) => e.to_u64(), } } - SYS_CPU_COUNT => crate::arch::smp::cpu_count() as u64, + SYS_CPU_COUNT => crate::smp::cpu_count() as u64, SYS_FUTEX_WAIT => match UserAddr::checked(a1) { Some(addr) => process::futex_wait(addr, a2 as u32, a3), None => bad_addr, diff --git a/kernel/src/syscall/machine.rs b/kernel/src/syscall/machine.rs index 5cbb532931d..f813cb84f5f 100644 --- a/kernel/src/syscall/machine.rs +++ b/kernel/src/syscall/machine.rs @@ -236,7 +236,7 @@ pub(super) fn sys_sysinfo(syscap: RawHandle, out: &mut UserBytesMut) -> u64 { } let (total_mem, used_mem) = crate::mm::pmm::stats(); - let cpu_count = crate::arch::smp::cpu_count(); + let cpu_count = crate::smp::cpu_count(); let uptime = crate::clock::nanos_since_boot(); let total_cpu_ns = crate::scheduler::total_cpu_ns(); let total_available_ns = uptime * cpu_count as u64; diff --git a/src/build.rs b/src/build.rs index 034ab1a4d22..0f2178d7958 100644 --- a/src/build.rs +++ b/src/build.rs @@ -3161,6 +3161,7 @@ mod tests { "tests/toolkitcase/system.toml", "tests/updatecase/system.toml", "tests/virtjobcase/system.toml", + "tests/virtpaniccase/system.toml", "tests/virtsmpcase/system.toml", ]; diff --git a/tests/common/qemu.rs b/tests/common/qemu.rs index 5149182ea10..87f62eace2b 100644 --- a/tests/common/qemu.rs +++ b/tests/common/qemu.rs @@ -703,23 +703,54 @@ pub fn await_marker_new( await_guest(qemu, log, doing, |log| log[from.min(log.len())..].contains(marker)) } -/// Per vCPU in `info registers -a`, whether it is halted with interrupts off: -/// the stop's `cli; hlt`, which no interrupt ends. An idle CPU halts with `IF` -/// set, and a running one is not halted, so neither is this. -pub fn stopped_cpus(registers: &str) -> Vec { +/// `wfi`, and the branch back to the instruction before it: the loop every +/// copy of the kernel's `arch::cpu::halt` compiles to on AArch64. +const WFI: u32 = 0xd503_207f; +const BACK_TO_WFI: u32 = 0x17ff_ffff; + +/// Per vCPU in `info registers -a`, whether it is halted with interrupts off, +/// which no interrupt ends. +/// +/// x86-64: `HLT=1` with `IF` clear, the stop's `cli; hlt`. An idle CPU halts +/// with `IF` set, and a running one is not halted, so neither is this. +/// +/// AArch64: QEMU prints no halt, and an idle CPU waits in `wfi` with `I` set +/// too (`arch::hw::halt`), so the wait is told apart by where it is: at EL1 +/// with `PSTATE.I` set, a `wfi` just behind the PC (QEMU's halted PC is the +/// instruction after it) and [`BACK_TO_WFI`] at it, where an idle CPU's next +/// instruction unmasks instead. +pub fn stopped_cpus(monitor: &mut QmpMonitor, arch: Arch) -> Vec { + let registers = monitor.human("info registers -a"); registers .split("CPU#") .skip(1) .map(|cpu| { - let field = |name: &str| -> Option { - Some(cpu.split(name).nth(1)?.chars().take_while(char::is_ascii_hexdigit).collect()) + let field = |name: &str| -> Option { + let digits: String = cpu.split(name).nth(1)?.chars().take_while(char::is_ascii_hexdigit).collect(); + u64::from_str_radix(&digits, 16).ok() }; - let flags = field("RFL=").and_then(|f| u64::from_str_radix(&f, 16).ok()); - field("HLT=").as_deref() == Some("1") && flags.is_some_and(|f| f & (1 << 9) == 0) + match arch { + Arch::X86_64 => field("HLT=") == Some(1) && field("RFL=").is_some_and(|f| f & 1 << 9 == 0), + Arch::Aarch64 => { + let (Some(pc), Some(pstate)) = (field(" PC="), field("PSTATE=")) else { + return false; + }; + let (el, masked) = (pstate >> 2 & 3, pstate & 1 << 7 != 0); + el == 1 && masked && pc.checked_sub(4).is_some_and(|at| words(monitor, at) == [WFI, BACK_TO_WFI]) + } + } }) .collect() } +/// The two 32-bit words at guest virtual address `at`, as the monitor's CPU +/// translates it; none where it cannot. +fn words(monitor: &mut QmpMonitor, at: u64) -> Vec { + let dump = monitor.human(&format!("x/2wx {at:#x}")); + let Some((_, words)) = dump.split_once(':') else { return Vec::new() }; + words.split_whitespace().filter_map(|word| u32::from_str_radix(word.strip_prefix("0x")?, 16).ok()).collect() +} + /// The fatal path's last line, which `panic_reboot::reboot_now` writes to the /// 16550 raw just before it resets the machine. pub const PANIC_REBOOTING: &str = "panic: no key inside the bound, so nobody is here"; @@ -1078,13 +1109,18 @@ pub enum Profile { /// firmware then hands the loader the CPU at EL2, and the kernel's entry /// has to drop from it. HVF gives a guest EL1 only. VirtEl2, + /// [`Profile::Virt`] emulated on `-cpu max` whatever the host: firmware + /// hands the loader the CPU at EL1, as HVF does, and QEMU's FADT names + /// PSCI's conduit `HVC`; unlike HVF's, the CPU has RNDR for the kernel's + /// hash seed. + VirtTcg, } impl Profile { /// The architecture this machine is. pub fn arch(self) -> Arch { match self { - Self::Virt | Self::VirtEl2 => Arch::Aarch64, + Self::Virt | Self::VirtEl2 | Self::VirtTcg => Arch::Aarch64, Self::Headless | Self::HeadlessNoIommu | Self::VirtioNetNoMsix @@ -1125,11 +1161,10 @@ impl Profile { } /// How this host provides the machine: [`Profile::VirtEl2`] emulated, - /// since only emulation gives a guest EL2; every other as its - /// architecture's own. + /// since only emulation gives a guest EL2. pub fn accel(self) -> Accel { match self { - Self::VirtEl2 => Accel::Tcg, + Self::VirtEl2 | Self::VirtTcg => Accel::Tcg, _ => self.arch().accel(), } } @@ -1452,7 +1487,7 @@ pub const NVME_T14_BLOCKS: u64 = NVME_T14_BYTES / 4096; impl Profile { fn shape(self) -> Shape { match self { - Self::VirtEl2 => Self::Virt.shape(), + Self::VirtEl2 | Self::VirtTcg => Self::Virt.shape(), Self::Virt => Shape { vga: "std", panel: None, diff --git a/tests/toyos-rust-tests/src/bin/panic_halts_first.rs b/tests/toyos-rust-tests/src/bin/panic_halts_first.rs index 0dd2fd7d685..aa2873e50a2 100644 --- a/tests/toyos-rust-tests/src/bin/panic_halts_first.rs +++ b/tests/toyos-rust-tests/src/bin/panic_halts_first.rs @@ -15,11 +15,18 @@ const SIBLINGS: usize = 3; fn retired() { let ret: u64; + #[cfg(target_arch = "x86_64")] // SAFETY: a register-only `syscall` whose number the kernel refuses // without reading any argument; nothing in this process is touched. unsafe { core::arch::asm!("syscall", in("rdi") RETIRED, lateout("rax") ret, out("rcx") _, out("r11") _); } + #[cfg(target_arch = "aarch64")] + // SAFETY: the same call through `svc`, which answers in `x0` and keeps + // every other register. + unsafe { + core::arch::asm!("svc #0", inlateout("x0") RETIRED => ret); + } assert_ne!(ret, 0, "syscall {RETIRED} answered as if it were live"); } diff --git a/tests/toyos.rs b/tests/toyos.rs index 3da5855b7e2..661c32ea54e 100644 --- a/tests/toyos.rs +++ b/tests/toyos.rs @@ -542,6 +542,9 @@ const SCREEN_TESTS: &[(&str, Sched, Tier)] = &[ ("virt_debug_refused", Sched::Parallel, Tier::Local), ("virt_readonly_copyout", Sched::Parallel, Tier::Local), ("virt_smp", Sched::Parallel, Tier::Local), + ("virt_el1_smp", Sched::Parallel, Tier::Local), + ("virt_failed_ap_leaves_no_hole", Sched::Parallel, Tier::Local), + ("virt_fatal_halts_the_others_first", Sched::Parallel, Tier::Local), ]; /// What `screen_console_shell` types, and what it then looks for on its own. @@ -3630,32 +3633,42 @@ fn judge_virt_job(mut qemu: QemuInstance, job: &str, said: &str) -> Result Result<(), String> { +/// Boot `tests/virtsmpcase` under `profile` on `cpus` CPUs, with `params` armed. +fn boot_virt_smp(profile: qemu::Profile, cpus: u32, params: &'static [&'static str]) -> QemuInstance { let config = compile::repo_root().join("tests/virtsmpcase/system.toml"); let case = config.parent().expect("system.toml has a directory"); - let qemu = QemuInstance::boot_with_options( + QemuInstance::boot_with_options( case, &[], &[], BootOptions { - profile: qemu::Profile::VirtEl2, - smp: VIRT_CPUS, + profile, + smp: cpus, + kernel_params: params, ready_marker: "control registers: SCTLR_EL1=", ..Default::default() }, - ); - let serial = judge_virt_job(qemu, "unmap_seen", "unmap_seen: 4 reads on another thread of a page just unmapped")?; + ) +} + +/// Boot `tests/virtsmpcase` on [`VIRT_CPUS`] CPUs under `profile`, whose +/// firmware enters every CPU at EL`el` and whose FADT names PSCI's `conduit`: +/// each CPU is started by `CPU_ON`, holds the control-register declaration as +/// entered there and joins the scheduler, and the case's job `unmap_touch` +/// ends with exit 0. +fn virt_smp(profile: qemu::Profile, conduit: &str, el: u32) -> Result<(), String> { + let serial = judge_virt_job(boot_virt_smp(profile, VIRT_CPUS, &[]), "unmap_touch", UNMAP_TOUCH_SAID)?; let psci = serial.lines().find(|l| l.contains("PSCI: ")).unwrap_or_default(); - if !psci.contains(" through SMC") { - return Err(format!("PSCI is not said to be reached through SMC: {psci:?}\nserial:\n{serial}")); + if !psci.contains(&format!(" through {conduit}")) { + return Err(format!("PSCI is not said to be reached through {conduit}: {psci:?}\nserial:\n{serial}")); } let mut want = vec![ format!("SMP: {VIRT_CPUS} of {VIRT_CPUS} MADT CPUs online"), @@ -3670,10 +3683,70 @@ fn virt_smp() -> Result<(), String> { return Err(format!("{want:?} not on the PL011\nserial:\n{serial}")); } } - eprintln!(" [virt] {VIRT_CPUS} CPUs online and scheduling"); + let entered = format!("as declared; entered at EL{el}"); + for cpu in 0..VIRT_CPUS { + if !serial.lines().any(|l| record_cpu(l) == Some(cpu) && l.contains(&entered)) { + return Err(format!("cpu{cpu} never said its registers are {entered:?}\nserial:\n{serial}")); + } + } + eprintln!(" [virt] {VIRT_CPUS} CPUs entered at EL{el}, started through {conduit}, and scheduling"); + Ok(()) +} + +/// `smp_failed_ap_leaves_no_hole` on AArch64: `smp-skip-ap` keeps `CPU_ON` +/// from the CPU that would be cpu2 of four, and the bring-up stops there, so +/// cpu2's id goes to no CPU behind it. The case's job, which the scheduler +/// places across the CPUs that came up, ends with exit 0. +fn virt_failed_ap_leaves_no_hole() -> Result<(), String> { + const CPUS: u32 = 4; + let qemu = boot_virt_smp(qemu::Profile::VirtEl2, CPUS, &["smp-skip-ap"]); + let serial = judge_virt_job(qemu, "unmap_touch", UNMAP_TOUCH_SAID)?; + // The premise, not just a small machine: cpu1 came up and cpu2 did not. + for premise in ["SMP: cpu1 mpidr=0x1 online", "SMP: cpu2 mpidr=0x2 did not echo within"] { + if !serial.contains(premise) { + return Err(format!("{premise:?} not on the PL011, so no non-last AP failed\nserial:\n{serial}")); + } + } + let online = format!("SMP: 2 of {CPUS} MADT CPUs online"); + if !serial.contains(&online) { + return Err(format!("{online:?} not on the PL011\nserial:\n{serial}")); + } + for phantom in ["CPU 2: joining scheduler", "CPU 3: joining scheduler"] { + if serial.contains(phantom) { + return Err(format!( + "a CPU past the failed AP joined, so `0..cpu_count()` is not the online set: {phantom:?}\nserial:\n{serial}" + )); + } + } + eprintln!(" [virt] a non-last AP never started and the dense machine ran its job"); Ok(()) } +/// [`panic_halts_the_others_first`] on AArch64: `tests/virtpaniccase` runs +/// `test_rs_panic_halts_first` as its one job on [`STOP_CPUS`] CPUs under the +/// EL2 profile, and the halt SGI stops every CPU but the one going fatal. +fn virt_fatal_halts_the_others_first() -> Result<(), String> { + let config = compile::repo_root().join("tests/virtpaniccase/system.toml"); + let case = config.parent().expect("system.toml has a directory"); + let profile = qemu::Profile::VirtEl2; + let fatal = qemu::build_toyos_bin(profile.arch(), &compile::repo_root().join("tests/toyos-rust-tests"), "panic_halts_first"); + let qemu = QemuInstance::boot_with_options( + case, + &[], + &[], + BootOptions { + profile, + smp: STOP_CPUS, + kernel_features: ACTUATOR_KERNEL, + qmp: true, + ready_marker: "control registers: SCTLR_EL1=", + extra_root_files: vec![("bin/test_rs_panic_halts_first".to_string(), fatal)], + ..Default::default() + }, + ); + the_others_halt_first(qemu, profile.arch()) +} + /// Boot `test_config` under the EL2 profile with the kernel selftest `armed` /// names, and judge its one line: `: PASS`. fn virt_selftest(test_config: &Path, armed: &'static [&'static str; 1]) -> Result<(), String> { @@ -5201,7 +5274,7 @@ fn run_screen_test( } "virt_fp_isolation" => virt_job("fp_isolation", "fp_isolation: v0-v31, FPCR and FPSR survived"), "virt_first_entry" => virt_job("first_entry", "first_entry: x1-x30 were zero"), - "virt_unmap_touch" => virt_job("unmap_touch", "unmap_touch: 4 reads of a page just unmapped"), + "virt_unmap_touch" => virt_job("unmap_touch", UNMAP_TOUCH_SAID), "virt_debug_refused" => virt_job( "debug_refused", "debug_refused: SYS_DEBUG's double fault and TLB acknowledgement delay were refused", @@ -5217,7 +5290,12 @@ fn run_screen_test( virt_selftest(test_config, &["irq-storm"]) } "virt_timer_floor" => virt_selftest(test_config, &["timer-floor"]), - "virt_smp" => virt_smp(), + "virt_smp" => virt_smp(qemu::Profile::VirtEl2, "SMC", 2), + // The EL1 entry's own arm, which fetches at a physical address under + // the bring-up root, and PSCI through `HVC`: the path HVF takes. + "virt_el1_smp" => virt_smp(qemu::Profile::VirtTcg, "HVC", 1), + "virt_failed_ap_leaves_no_hole" => virt_failed_ap_leaves_no_hole(), + "virt_fatal_halts_the_others_first" => virt_fatal_halts_the_others_first(), "screen_late_panic" => { // The ordinary fatal panic, which no userland process can produce: // crash_report, capture, panic_flush, halt_all_cpus, render. The @@ -16318,23 +16396,17 @@ fn hda_two_live_refused( /// path that waited first — for a log, a drain, anything — would leave every /// other CPU running userland under it. `test_rs_panic_halts_first` keeps three /// siblings making kernel records while its main thread goes fatal. -/// -/// **QEMU is the judge, and no clock is in it.** Once the fatal path has said -/// its line past the stop, `panic_reboot::arm`'s, every vCPU but the one that -/// went fatal must show [`qemu::stopped_cpus`]' `cli; hlt`. The fatal one holds -/// its panel under the shipped minute, so the machine is still there to ask. fn panic_halts_the_others_first( test_config: &Path, c_bins: &[(String, Vec)], rust_bins: &[(String, Vec)], ) -> Result<(), String> { - const SMP: usize = 4; let mut qemu = QemuInstance::boot_with_options( test_config, c_bins, rust_bins, BootOptions { - smp: SMP as u32, + smp: STOP_CPUS, kernel_features: ACTUATOR_KERNEL, qmp: true, ..Default::default() @@ -16342,6 +16414,19 @@ fn panic_halts_the_others_first( ); writeln!(qemu.stdin_mut(), "run test_rs_panic_halts_first").map_err(|e| format!("stdin: {e}"))?; qemu.flush_stdin(); + the_others_halt_first(qemu, qemu::SUITE_ARCH) +} + +/// The CPUs `test_rs_panic_halts_first` runs on: one for it and one per sibling. +const STOP_CPUS: u32 = 4; + +/// **QEMU is the judge, and no clock is in it.** Once the fatal path of the +/// `test_rs_panic_halts_first` that `qemu` runs has said its line past the +/// stop, `panic_reboot::arm`'s, every vCPU but the one that went fatal must be +/// one [`qemu::stopped_cpus`] calls halted. The fatal one holds its panel, so +/// the machine is still there to ask. +fn the_others_halt_first(mut qemu: QemuInstance, arch: toyos_build::arch::Arch) -> Result<(), String> { + let cpus = STOP_CPUS as usize; let mut console = String::new(); await_guest(&mut qemu, &mut console, "the fatal path's line past the stop", |c| { fatal_past_the_stop(c).is_some() @@ -16360,19 +16445,19 @@ fn panic_halts_the_others_first( // taken the stop, and this is how long it is waited for. let give_up = Instant::now() + qemu::GUEST_QUIET; loop { - let stopped = qemu::stopped_cpus(&monitor.human("info registers -a")); - if stopped.len() == SMP && stopped.iter().filter(|&&cpu| cpu).count() >= SMP - 1 { + let stopped = qemu::stopped_cpus(&mut monitor, arch); + if stopped.len() == cpus && stopped.iter().filter(|&&cpu| cpu).count() >= cpus - 1 { break; } if Instant::now() >= give_up { return Err(format!( "{STALLED} waiting for the other CPUs to halt after the fatal path on cpu{fatal} \ - stopped them — QEMU shows each vCPU in `cli; hlt` as {stopped:?}\n{console}" + stopped them — QEMU shows each vCPU halted with interrupts masked as {stopped:?}\n{console}" )); } console.push_str(&qemu.drain_serial(Duration::from_millis(200))); } - eprintln!(" [panic] the fatal path on cpu{fatal} left every other CPU in `cli; hlt`"); + eprintln!(" [panic] the fatal path on cpu{fatal} left every other CPU halted with interrupts masked"); Ok(()) } diff --git a/tests/virtpaniccase/system.toml b/tests/virtpaniccase/system.toml new file mode 100644 index 00000000000..0797700ae27 --- /dev/null +++ b/tests/virtpaniccase/system.toml @@ -0,0 +1,13 @@ +# The fatal path on a machine of more than one CPU: the one job goes fatal +# while its siblings make kernel records on the other CPUs, and takes the +# machine with it, so nothing runs after it. +[boot] +start = ["logd", "test-runner"] + +[programs.logd] +service = true +syscap = ["logread"] + +[programs.test-runner] +receives = ["power"] +args = ["--bound-ms=240000", "test_rs_panic_halts_first"] diff --git a/tests/virtsmpcase/system.toml b/tests/virtsmpcase/system.toml index c24d4a34efb..e7dfb0bc46d 100644 --- a/tests/virtsmpcase/system.toml +++ b/tests/virtsmpcase/system.toml @@ -11,9 +11,9 @@ syscap = ["logread"] receives = ["power"] # Four minutes from boot: every job here runs on emulated CPUs, and a job past # the bound reboots the guest with the job named. -args = ["--bound-ms=240000", "unmap_seen"] +args = ["--bound-ms=240000", "unmap_touch"] [programs.toybox] [symlinks] -"bin/unmap_seen" = "/system/bin/toybox" +"bin/unmap_touch" = "/system/bin/toybox" diff --git a/toyos-acpi/src/fadt.rs b/toyos-acpi/src/fadt.rs index db284cf0f3d..6080dff8e10 100644 --- a/toyos-acpi/src/fadt.rs +++ b/toyos-acpi/src/fadt.rs @@ -133,22 +133,28 @@ pub enum Psci { Hvc, /// `PSCI_COMPLIANT` clear: firmware offers no PSCI. Absent, - /// A FADT before ACPI 5.1, or one too short to reach the field: those - /// bytes are reserved, so nothing is said about PSCI at all. + /// A FADT before ACPI 5.1: those bytes are reserved, so nothing is said + /// about PSCI at all. Undefined { revision: u8, minor: u8 }, + /// A FADT that ends before `ARM_BOOT_ARCH` and the minor version after it. + Short, } /// `ARM_BOOT_ARCH`, read only from a FADT whose version defines it: 5.1 on, /// the version its header's major and `FADT Minor Version`'s low nibble spell. pub fn psci(fadt: &Table

) -> Psci { let revision = fadt.byte(SDT_REVISION).unwrap_or(0); - let minor = fadt.byte(FADT_MINOR_VERSION).map_or(0, |byte| byte & 0xF); - let defined = revision > 5 || (revision == 5 && minor >= 1); - match fadt.u16_at(FADT_ARM_BOOT_ARCH) { - Some(flags) if defined && flags & PSCI_COMPLIANT == 0 => Psci::Absent, - Some(flags) if defined && flags & PSCI_USE_HVC != 0 => Psci::Hvc, - Some(_) if defined => Psci::Smc, - _ => Psci::Undefined { revision, minor }, + let (Some(flags), Some(minor)) = (fadt.u16_at(FADT_ARM_BOOT_ARCH), fadt.byte(FADT_MINOR_VERSION)) else { + return Psci::Short; + }; + let minor = minor & 0xF; + if revision < 5 || (revision == 5 && minor < 1) { + return Psci::Undefined { revision, minor }; + } + match flags { + flags if flags & PSCI_COMPLIANT == 0 => Psci::Absent, + flags if flags & PSCI_USE_HVC != 0 => Psci::Hvc, + _ => Psci::Smc, } } diff --git a/toyos-acpi/tests/corpus.rs b/toyos-acpi/tests/corpus.rs index 0f9f73dc0c0..ee953dcccec 100644 --- a/toyos-acpi/tests/corpus.rs +++ b/toyos-acpi/tests/corpus.rs @@ -546,8 +546,7 @@ fn psci_of(table: &[u8]) -> Psci { /// `ARM_BOOT_ARCH` on a FADT that defines it. QEMU 11.1.1's `virt` /// publishes revision 6, minor 3 and `PSCI_COMPLIANT`, with `PSCI_USE_HVC` -/// where the guest has no EL2 (`hw/arm/virt-acpi-build.c:1135-1148`, -/// `hw/arm/virt.c:2956-2962`). +/// where the guest has no EL2. #[test] fn the_psci_conduit_is_read_off_a_fadt_that_defines_it() { assert_eq!(psci_of(&arm_facp(6, 3, 0b01)), Psci::Smc); @@ -564,10 +563,17 @@ fn a_fadt_before_acpi_5_1_says_nothing_about_psci() { assert_eq!(psci_of(&arm_facp(5, 0, 0b11)), Psci::Undefined { revision: 5, minor: 0 }); assert_eq!(psci_of(&arm_facp(3, 0, 0b01)), Psci::Undefined { revision: 3, minor: 0 }); assert_eq!(psci_of(&arm_facp(5, 0x10, 0b01)), Psci::Undefined { revision: 5, minor: 0 }); +} - let mut short = arm_facp(6, 3, 0b01); - declare_len(&mut short, 130); - assert_eq!(psci_of(&short), Psci::Undefined { revision: 6, minor: 0 }); +/// A table that ends inside `ARM_BOOT_ARCH`, or before the minor version that +/// says whether it has one, is short whatever its revision says. +#[test] +fn a_fadt_that_ends_before_arm_boot_arch_is_short() { + for len in [130, 131] { + let mut short = arm_facp(6, 3, 0b01); + declare_len(&mut short, len); + assert_eq!(psci_of(&short), Psci::Short, "a FADT of {len} bytes"); + } } /// Where the two firmwares' descriptor lists sit for the sweep below. diff --git a/toyos-acpi/tests/fixtures.rs b/toyos-acpi/tests/fixtures.rs index 6ff04725bf7..b4a769e455a 100644 --- a/toyos-acpi/tests/fixtures.rs +++ b/toyos-acpi/tests/fixtures.rs @@ -118,8 +118,8 @@ fn the_fadt_names_the_reset_register_qemu_acts_on() { assert_eq!(reset_register(&fadt), Reset::Port { port: 0xcf9, value: 0x0f }); } -/// Revision 3, so bytes 129-131 are the reserved ones QEMU 11.1.1 writes zero -/// (`hw/acpi/aml-build.c:2544-2550`): the table says nothing about PSCI. +/// Revision 3, so bytes 129-131 are the reserved ones QEMU 11.1.1 writes zero: +/// the table says nothing about PSCI. #[test] fn the_q35_fadt_predates_arm_boot_arch() { let fadt = find_table(machine(), RSDP, b"FACP", 36).expect("FADT"); diff --git a/userland/toybox/src/main.rs b/userland/toybox/src/main.rs index a60c28b5e7d..449ce6e89d4 100644 --- a/userland/toybox/src/main.rs +++ b/userland/toybox/src/main.rs @@ -21,7 +21,6 @@ mod shutdown; mod spin; mod stats; mod tone; -mod unmap_seen; mod unmap_touch; macro_rules! commands { @@ -37,7 +36,7 @@ macro_rules! commands { use arch::{first_entry, fp_isolation}; -commands!(cat, cp, debug_refused, echo, first_entry, fp_isolation, free, grep, hexdump, locale, ls, mkdir, mv, net, preempt, ps, pwd, reboot, rm, screen, shutdown, spin, stats, tone, unmap_seen, unmap_touch); +commands!(cat, cp, debug_refused, echo, first_entry, fp_isolation, free, grep, hexdump, locale, ls, mkdir, mv, net, preempt, ps, pwd, reboot, rm, screen, shutdown, spin, stats, tone, unmap_touch); fn main() { let args: Vec = std::env::args().collect(); diff --git a/userland/toybox/src/unmap_seen.rs b/userland/toybox/src/unmap_seen.rs deleted file mode 100644 index 7fef30d0b97..00000000000 --- a/userland/toybox/src/unmap_seen.rs +++ /dev/null @@ -1,89 +0,0 @@ -//! `unmap_seen`: whether an unmap reaches a translation another thread holds. -//! A child maps a page and a second thread reads it in a loop; once that -//! thread has read it, the first unmaps it and says so through memory, and the -//! reader's next read must end the child. On a machine of more CPUs than busy -//! threads the reader runs beside the unmap, on a CPU whose TLB only a -//! broadcast invalidation reaches. - -use std::process::Command; -use std::sync::atomic::{AtomicBool, AtomicU64, Ordering::{Acquire, Release}}; - -use toyos_abi::syscall::{mmap, munmap, MmapFlags, MmapProt}; - -const TRIALS: usize = 4; -const PAGE: usize = 2 * 1024 * 1024; -/// Reads the reader makes before the unmap, each through the translation its -/// CPU caches. -const WARM: u64 = 1_000; -/// The status `kill_process(-1)` gives a process whose fault nothing serves; -/// a refused unmap or a panic ends the child with another. -const FAULTED: i32 = -1; -/// The child's last line before the unmap. -const UNMAPPING: &str = "unmap_seen: the reader holds the page; unmapping"; -/// The reader's line if its read after the unmap came back. -const READ_BACK: &str = "unmap_seen: the reader read"; - -static READS: AtomicU64 = AtomicU64::new(0); -static UNMAPPED: AtomicBool = AtomicBool::new(false); - -pub fn main(args: Vec) { - match args.first().map(String::as_str) { - None => judge(), - Some("child") => child(), - Some(other) => panic!("unmap_seen: unknown mode {other:?}"), - } -} - -fn judge() { - for trial in 0..TRIALS { - let out = Command::new("/system/bin/unmap_seen") - .arg("child") - .output() - .unwrap_or_else(|e| panic!("unmap_seen: the child would not spawn: {e}")); - let said = String::from_utf8_lossy(&out.stdout); - assert!(said.contains(UNMAPPING), "trial {trial}: the child never reached the unmap: {said:?}"); - assert!( - out.status.code() == Some(FAULTED) && !said.contains(READ_BACK), - "trial {trial}: the child exited {:?}, not faulted ({FAULTED}), once its reader read a page \ - unmapped beside it: {said:?}", - out.status.code(), - ); - } - println!("unmap_seen: {TRIALS} reads on another thread of a page just unmapped each ended their process"); -} - -fn child() { - // SAFETY: a fresh anonymous mapping this function alone names. - let page = unsafe { mmap(core::ptr::null_mut(), PAGE, MmapProt::READ | MmapProt::WRITE, MmapFlags::ANONYMOUS | MmapFlags::PRIVATE) }; - assert!(!page.is_null(), "unmap_seen: mmap refused"); - // SAFETY: inside the mapping just made. - unsafe { page.write_volatile(0x5A) }; - let at = page as usize; - std::thread::spawn(move || read_until_unmapped(at as *const u8)); - while READS.load(Acquire) < WARM { - std::hint::spin_loop(); - } - println!("{UNMAPPING}"); - // SAFETY: the mapping made above, whole; its reader's next read is the probe. - unsafe { munmap(page, PAGE) }.expect("unmap_seen: munmap refused"); - UNMAPPED.store(true, Release); - // The reader's fault ends this process; this thread waits for it. - loop { - std::thread::park(); - } -} - -/// Read `page` until the unmap is said to have returned, then once more. -fn read_until_unmapped(page: *const u8) -> ! { - loop { - let unmapped = UNMAPPED.load(Acquire); - // SAFETY: mapped until `unmapped` reads true; the read after that is - // the probe, and a kernel that keeps its word ends the process there. - let byte = unsafe { page.read_volatile() }; - if unmapped { - println!("{READ_BACK} {byte:#x} after the unmap had returned"); - std::process::exit(0); - } - READS.fetch_add(1, Release); - } -} diff --git a/userland/toybox/src/unmap_touch.rs b/userland/toybox/src/unmap_touch.rs index 7acc4b9697a..ef4892b2599 100644 --- a/userland/toybox/src/unmap_touch.rs +++ b/userland/toybox/src/unmap_touch.rs @@ -1,52 +1,72 @@ -//! `unmap_touch`: whether a page's unmapping reaches the TLB. A child writes -//! a page, unmaps it and reads it at once; the read must end the child, -//! because a translation left cached past the unmap still reaches the frame. +//! `unmap_touch`: whether a page's unmapping reaches every TLB that may hold +//! it. Each trial is a child that writes a page, unmaps it and reads it again; +//! that read must end the child, because a translation left cached past the +//! unmap still reaches the frame. In `touch` the unmapping thread reads it; in +//! `seen` a second thread that has been reading it all along, which on a +//! machine of more CPUs than busy threads runs beside the unmap, on a CPU whose +//! TLB only a broadcast invalidation reaches. use std::process::Command; +use std::sync::atomic::{AtomicBool, AtomicU64, Ordering::{Acquire, Release}}; -use toyos_abi::syscall::{mmap, MmapFlags, MmapProt}; +use toyos_abi::syscall::{mmap, munmap, MmapFlags, MmapProt}; const TRIALS: usize = 4; const PAGE: usize = 2 * 1024 * 1024; /// The status `kill_process(-1)` gives a process whose fault nothing serves; /// a refused unmap or a panic ends the child with another. const FAULTED: i32 = -1; -/// The child's last line before the unmap and the read. -const UNMAPPING: &str = "unmap_touch: written; unmapping and reading"; -/// The child's line if the read came back. +/// A child's last line before the unmap. +const UNMAPPING: &str = "unmap_touch: written; unmapping"; +/// A child's line if its read after the unmap came back. const READ_BACK: &str = "unmap_touch: read"; +/// Reads `seen`'s reader makes before the unmap, each through the translation +/// its CPU caches. +const WARM: u64 = 1_000; + +static READS: AtomicU64 = AtomicU64::new(0); +static UNMAPPED: AtomicBool = AtomicBool::new(false); pub fn main(args: Vec) { match args.first().map(String::as_str) { None => judge(), Some("touch") => touch(), + Some("seen") => seen(), Some(other) => panic!("unmap_touch: unknown mode {other:?}"), } } fn judge() { - for trial in 0..TRIALS { - let out = Command::new("/system/bin/unmap_touch") - .arg("touch") - .output() - .unwrap_or_else(|e| panic!("unmap_touch: the child would not spawn: {e}")); - let said = String::from_utf8_lossy(&out.stdout); - assert!(said.contains(UNMAPPING), "trial {trial}: the child never reached the unmap: {said:?}"); - assert!( - out.status.code() == Some(FAULTED) && !said.contains(READ_BACK), - "trial {trial}: the child exited {:?}, not faulted ({FAULTED}), on the read of a page it had just unmapped: {said:?}", - out.status.code(), - ); + for mode in ["touch", "seen"] { + for trial in 0..TRIALS { + let out = Command::new("/system/bin/unmap_touch") + .arg(mode) + .output() + .unwrap_or_else(|e| panic!("unmap_touch: the child would not spawn: {e}")); + let said = String::from_utf8_lossy(&out.stdout); + assert!(said.contains(UNMAPPING), "{mode} trial {trial}: the child never reached the unmap: {said:?}"); + assert!( + out.status.code() == Some(FAULTED) && !said.contains(READ_BACK), + "{mode} trial {trial}: the child exited {:?}, not faulted ({FAULTED}), on a read of a page it had just unmapped: {said:?}", + out.status.code(), + ); + } } - println!("unmap_touch: {TRIALS} reads of a page just unmapped each ended their process"); + println!("unmap_touch: {TRIALS} reads of a page just unmapped on the unmapping thread, and {TRIALS} on another, each ended their process"); } -fn touch() { +/// A fresh page, written. +fn written() -> *mut u8 { // SAFETY: a fresh anonymous mapping this function alone names. let page = unsafe { mmap(core::ptr::null_mut(), PAGE, MmapProt::READ | MmapProt::WRITE, MmapFlags::ANONYMOUS | MmapFlags::PRIVATE) }; assert!(!page.is_null(), "unmap_touch: mmap refused"); // SAFETY: inside the mapping just made. unsafe { page.write_volatile(0x5A) }; + page +} + +fn touch() { + let page = written(); println!("{UNMAPPING}"); // After the last print: a print can switch the CPU, and a switch can drop // the translation the first read caches. @@ -55,3 +75,35 @@ fn touch() { let (first, answer, second) = unsafe { crate::arch::read_unmap_read(page.cast(), PAGE) }; println!("{READ_BACK} {second:#x} after the unmap answered {answer:#x}, {first:#x} before it"); } + +fn seen() { + let page = written(); + let at = page as usize; + std::thread::spawn(move || read_until_unmapped(at as *const u8)); + while READS.load(Acquire) < WARM { + std::hint::spin_loop(); + } + println!("{UNMAPPING}"); + // SAFETY: the mapping made above, whole; its reader's next read is the probe. + unsafe { munmap(page, PAGE) }.expect("unmap_touch: munmap refused"); + UNMAPPED.store(true, Release); + // The reader's fault ends this process; this thread waits for it. + loop { + std::thread::park(); + } +} + +/// Read `page` until the unmap is said to have returned, then once more. +fn read_until_unmapped(page: *const u8) -> ! { + loop { + let unmapped = UNMAPPED.load(Acquire); + // SAFETY: mapped until `unmapped` reads true; the read after that is + // the probe, and a kernel that keeps its word ends the process there. + let byte = unsafe { page.read_volatile() }; + if unmapped { + println!("{READ_BACK} {byte:#x} after the unmap had returned"); + std::process::exit(0); + } + READS.fetch_add(1, Release); + } +} From 3f062d4073f3081ac7a9ffa417e14d64298fb4fb Mon Sep 17 00:00:00 2001 From: japabu Date: Wed, 30 Sep 2026 16:21:21 +0200 Subject: [PATCH 4/6] Answer the second review of #633: M2 and M3 red on virt_smp, the halt judge names the fatal CPU B8: the orchestrator ran r2-m2-unmap-invalidates-locally and r2-m3-kick-goes-to-self on virt_smp at 1fa75ba26. M2 EXIT=1 at `seen trial 0: the child exited Some(0), not faulted (-1)`, the child saying `read 0x5a after the unmap had returned`; M3 EXIT=1, STALLED before userland with every AP's scheduler empty. unmap_touch.rs is unchanged: the `seen` reader lands off the unmapper's CPU in the first trial as the applet stands. AArch64's control_regs CHECKED and report are deleted with virt_smp's want line for them. They restated the roster: ap_entry echoes only after check, virt_smp already judges every CPU's own `entered at ELn` line, and an AP that echoed too late still counted, so the line could name more CPUs than the roster holds. the_others_halt_first requires every vCPU but the fatal one halted, by index, rather than any cpus - 1 of them. QEMU's CPU#n is the kernel's cpun: QEMU builds its MADT in CPU# order with the boot CPU first, and both architectures number APs in MADT order. panic_halts_first's syscall asm moves into tests/toyos-rust-tests/src/arch/, one module per architecture behind one selector, as toybox's src/arch/ is. Deleted as the review listed: R1 (irqchip's "the caller ends every one it is given", trap's "acknowledged, handled, ended"), R2 (stopped_cpus's "which no interrupt ends"), R3 (accel's doc naming VirtEl2 alone), R4 (corpus.rs's QEMU 11.1.1 sentence), R5 (VIRT_CPUS's "as many as the roster holds"). Not done: a test of psci.rs's 0.2 floor. The review's home, toyos-acpi, is ACPI table decoding by its description, and a PSCI_VERSION answer is an SMCCC register, not a table; no other host-tested crate owns PSCI. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L --- kernel/src/arch/aarch64/control_regs.rs | 12 +----------- kernel/src/arch/aarch64/irqchip.rs | 2 +- kernel/src/arch/aarch64/smp.rs | 5 +---- kernel/src/arch/aarch64/trap.rs | 2 +- tests/common/qemu.rs | 6 ++---- tests/toyos-rust-tests/src/arch/aarch64.rs | 16 ++++++++++++++++ tests/toyos-rust-tests/src/arch/mod.rs | 11 +++++++++++ tests/toyos-rust-tests/src/arch/x86_64.rs | 16 ++++++++++++++++ .../src/bin/panic_halts_first.rs | 19 ++++++------------- tests/toyos.rs | 11 +++++------ toyos-acpi/tests/corpus.rs | 4 +--- 11 files changed, 61 insertions(+), 43 deletions(-) create mode 100644 tests/toyos-rust-tests/src/arch/aarch64.rs create mode 100644 tests/toyos-rust-tests/src/arch/mod.rs create mode 100644 tests/toyos-rust-tests/src/arch/x86_64.rs diff --git a/kernel/src/arch/aarch64/control_regs.rs b/kernel/src/arch/aarch64/control_regs.rs index 2940a56d268..0cad5e024e2 100644 --- a/kernel/src/arch/aarch64/control_regs.rs +++ b/kernel/src/arch/aarch64/control_regs.rs @@ -9,7 +9,7 @@ //! and the GIC architecture specification (IHI 0069H), chapter 12, for the //! `ICC_` registers. -use core::sync::atomic::{AtomicU32, AtomicU64, Ordering}; +use core::sync::atomic::AtomicU64; use crate::log; @@ -102,9 +102,6 @@ pub fn tcr() -> u64 { TCR | (read!("id_aa64mmfr0_el1") & 0xF) << TCR_IPS_SHIFT } -/// CPUs whose registers [`check`] found as declared. -static CHECKED: AtomicU32 = AtomicU32::new(0); - /// Read every EL1 register the declaration names back, and refuse a CPU that /// holds anything else; then say what it holds, and that it was entered at `el`. pub fn check(el: u64) { @@ -125,7 +122,6 @@ pub fn check(el: u64) { // What the drop from EL2 left, or the EL1 entry kept: EL1, on `SP_EL1`. let (now, spsel) = (read!("CurrentEL") >> 2 & 3, read!("SPSel") & 1); assert_eq!((now, spsel), (1, 1), "control registers: running at EL{now} on SP_EL{spsel}, not EL1 on SP_EL1"); - CHECKED.fetch_add(1, Ordering::Relaxed); log!( "control registers: SCTLR_EL1={SCTLR:#x} TCR_EL1={:#x} MAIR_EL1={:#x} CPACR_EL1={CPACR:#x} \ CNTKCTL_EL1={CNTKCTL:#x}, as declared; entered at EL{el}{}", @@ -138,9 +134,3 @@ pub fn check(el: u64) { }, ); } - -/// How many CPUs hold the declaration, said once after the last was started: a -/// CPU that never reached [`check`] shows as a count below the roster's. -pub fn report(cpus: u32) { - log!("control registers: {} of {cpus} CPUs hold the declaration", CHECKED.load(Ordering::Relaxed)); -} diff --git a/kernel/src/arch/aarch64/irqchip.rs b/kernel/src/arch/aarch64/irqchip.rs index df3f354febd..eb7f49b1468 100644 --- a/kernel/src/arch/aarch64/irqchip.rs +++ b/kernel/src/arch/aarch64/irqchip.rs @@ -260,7 +260,7 @@ pub fn init_cpu(frame: u64) { pub(super) fn acknowledge() -> Option { let intid: u64; // SAFETY: reads `ICC_IAR1_EL1`, which acknowledges the highest pending - // Group 1 interrupt; the caller ends every one it is given. + // Group 1 interrupt. unsafe { core::arch::asm!("mrs {}, S3_0_C12_C12_0", out(reg) intid, options(nomem, nostack, preserves_flags)) }; let intid = (intid & 0xFF_FFFF) as u32; (intid != SPURIOUS).then_some(intid) diff --git a/kernel/src/arch/aarch64/smp.rs b/kernel/src/arch/aarch64/smp.rs index a2ef20417d0..0f294097f43 100644 --- a/kernel/src/arch/aarch64/smp.rs +++ b/kernel/src/arch/aarch64/smp.rs @@ -36,7 +36,6 @@ pub fn start(gic: &irqchip::Gic, psci: Option) { let others = gic.cpus.iter().filter(|gicc| packed_affinity(gicc.mpidr) != me); let Some(psci) = psci else { log!("SMP: no PSCI to start the other {} CPUs with; the boot CPU runs alone", others.count()); - control_regs::report(smp::cpu_count()); return; }; let root = paging::bringup_root(); @@ -71,9 +70,7 @@ pub fn start(gic: &irqchip::Gic, psci: Option) { ROSTER.commit(attempt, packed_affinity(gicc.mpidr)); log!("SMP: cpu{} mpidr={:#x} online", attempt.id(), gicc.mpidr); } - let online = smp::cpu_count(); - log!("SMP: {online} of {} MADT CPUs online", gic.cpus.len()); - control_regs::report(online); + log!("SMP: {} of {} MADT CPUs online", smp::cpu_count(), gic.cpus.len()); } /// An AP's first Rust: at EL1, on its bring-up stack with the vectors diff --git a/kernel/src/arch/aarch64/trap.rs b/kernel/src/arch/aarch64/trap.rs index 8b510ae42ee..d85e2ca15a8 100644 --- a/kernel/src/arch/aarch64/trap.rs +++ b/kernel/src/arch/aarch64/trap.rs @@ -129,7 +129,7 @@ fn exception(frame: &Frame, entry: u64) -> ! { panic!("{entry}: {} at {:#x}", class_name(frame.esr), frame.elr); } -/// One interrupt: acknowledged, handled, ended. From EL0 a tick or a kick +/// One interrupt. From EL0 a tick or a kick /// preempts here, where the interrupted context holds nothing; from EL1 it /// only asks for the pass the context will run when it may. fn irq(from_el0: bool) { diff --git a/tests/common/qemu.rs b/tests/common/qemu.rs index 87f62eace2b..33ad24f0f65 100644 --- a/tests/common/qemu.rs +++ b/tests/common/qemu.rs @@ -708,8 +708,7 @@ pub fn await_marker_new( const WFI: u32 = 0xd503_207f; const BACK_TO_WFI: u32 = 0x17ff_ffff; -/// Per vCPU in `info registers -a`, whether it is halted with interrupts off, -/// which no interrupt ends. +/// Per vCPU in `info registers -a`, whether it is halted with interrupts off. /// /// x86-64: `HLT=1` with `IF` clear, the stop's `cli; hlt`. An idle CPU halts /// with `IF` set, and a running one is not halted, so neither is this. @@ -1160,8 +1159,7 @@ impl Profile { } } - /// How this host provides the machine: [`Profile::VirtEl2`] emulated, - /// since only emulation gives a guest EL2. + /// How this host provides the machine. pub fn accel(self) -> Accel { match self { Self::VirtEl2 | Self::VirtTcg => Accel::Tcg, diff --git a/tests/toyos-rust-tests/src/arch/aarch64.rs b/tests/toyos-rust-tests/src/arch/aarch64.rs new file mode 100644 index 00000000000..89bf330a22a --- /dev/null +++ b/tests/toyos-rust-tests/src/arch/aarch64.rs @@ -0,0 +1,16 @@ +//! AArch64's half of the test binaries' assembly. + +/// Syscall `number` with no argument set, past every `toyos_abi` wrapper; the +/// kernel's answer. +/// +/// # Safety +/// The kernel refuses `number` without reading any argument register. +pub unsafe fn bare_syscall(number: u64) -> u64 { + let answer: u64; + // SAFETY: the caller's; the `svc` is the ABI's (`toyos_abi::syscall`), + // which answers in x0 and keeps every other register. + unsafe { + core::arch::asm!("svc #0", inlateout("x0") number => answer); + } + answer +} diff --git a/tests/toyos-rust-tests/src/arch/mod.rs b/tests/toyos-rust-tests/src/arch/mod.rs new file mode 100644 index 00000000000..665e1444620 --- /dev/null +++ b/tests/toyos-rust-tests/src/arch/mod.rs @@ -0,0 +1,11 @@ +//! What the test binaries must say in assembly, one module per architecture. + +#[cfg(target_arch = "aarch64")] +mod aarch64; +#[cfg(target_arch = "aarch64")] +pub use aarch64::*; + +#[cfg(target_arch = "x86_64")] +mod x86_64; +#[cfg(target_arch = "x86_64")] +pub use x86_64::*; diff --git a/tests/toyos-rust-tests/src/arch/x86_64.rs b/tests/toyos-rust-tests/src/arch/x86_64.rs new file mode 100644 index 00000000000..a5a6ec42719 --- /dev/null +++ b/tests/toyos-rust-tests/src/arch/x86_64.rs @@ -0,0 +1,16 @@ +//! x86-64's half of the test binaries' assembly. + +/// Syscall `number` with no argument set, past every `toyos_abi` wrapper; the +/// kernel's answer. +/// +/// # Safety +/// The kernel refuses `number` without reading any argument register. +pub unsafe fn bare_syscall(number: u64) -> u64 { + let answer: u64; + // SAFETY: the caller's; the `syscall` is the ABI's (`toyos_abi::syscall`), + // which clobbers rax, rcx and r11. + unsafe { + core::arch::asm!("syscall", in("rdi") number, lateout("rax") answer, out("rcx") _, out("r11") _); + } + answer +} diff --git a/tests/toyos-rust-tests/src/bin/panic_halts_first.rs b/tests/toyos-rust-tests/src/bin/panic_halts_first.rs index aa2873e50a2..b97eb5403c6 100644 --- a/tests/toyos-rust-tests/src/bin/panic_halts_first.rs +++ b/tests/toyos-rust-tests/src/bin/panic_halts_first.rs @@ -7,6 +7,9 @@ use std::sync::atomic::{AtomicUsize, Ordering}; use std::sync::Arc; +#[path = "../arch/mod.rs"] +mod arch; + /// A retired syscall's number: each call is refused and is one kernel record /// naming it. const RETIRED: u64 = 26; @@ -14,19 +17,9 @@ const RETIRED: u64 = 26; const SIBLINGS: usize = 3; fn retired() { - let ret: u64; - #[cfg(target_arch = "x86_64")] - // SAFETY: a register-only `syscall` whose number the kernel refuses - // without reading any argument; nothing in this process is touched. - unsafe { - core::arch::asm!("syscall", in("rdi") RETIRED, lateout("rax") ret, out("rcx") _, out("r11") _); - } - #[cfg(target_arch = "aarch64")] - // SAFETY: the same call through `svc`, which answers in `x0` and keeps - // every other register. - unsafe { - core::arch::asm!("svc #0", inlateout("x0") RETIRED => ret); - } + // SAFETY: a retired number, which the kernel refuses without reading any + // argument; nothing in this process is touched. + let ret = unsafe { arch::bare_syscall(RETIRED) }; assert_ne!(ret, 0, "syscall {RETIRED} answered as if it were live"); } diff --git a/tests/toyos.rs b/tests/toyos.rs index 55835984fe3..b2b0a9a8863 100644 --- a/tests/toyos.rs +++ b/tests/toyos.rs @@ -3636,7 +3636,7 @@ fn judge_virt_job(mut qemu: QemuInstance, job: &str, said: &str) -> Result Result<(), String if !psci.contains(&format!(" through {conduit}")) { return Err(format!("PSCI is not said to be reached through {conduit}: {psci:?}\nserial:\n{serial}")); } - let mut want = vec![ - format!("SMP: {VIRT_CPUS} of {VIRT_CPUS} MADT CPUs online"), - format!("control registers: {VIRT_CPUS} of {VIRT_CPUS} CPUs hold the declaration"), - ]; + let mut want = vec![format!("SMP: {VIRT_CPUS} of {VIRT_CPUS} MADT CPUs online")]; for cpu in 1..VIRT_CPUS { want.push(format!("SMP: cpu{cpu} mpidr={cpu:#x} online")); want.push(format!("CPU {cpu}: joining scheduler")); @@ -16445,7 +16442,9 @@ fn the_others_halt_first(mut qemu: QemuInstance, arch: toyos_build::arch::Arch) let give_up = Instant::now() + qemu::GUEST_QUIET; loop { let stopped = qemu::stopped_cpus(&mut monitor, arch); - if stopped.len() == cpus && stopped.iter().filter(|&&cpu| cpu).count() >= cpus - 1 { + // QEMU's CPU#n is the kernel's cpun: its MADT lists them in that order, + // the boot CPU first, and the kernel numbers them as the MADT lists them. + if stopped.len() == cpus && stopped.iter().enumerate().all(|(cpu, &halted)| halted || cpu == fatal as usize) { break; } if Instant::now() >= give_up { diff --git a/toyos-acpi/tests/corpus.rs b/toyos-acpi/tests/corpus.rs index ee953dcccec..231856db8be 100644 --- a/toyos-acpi/tests/corpus.rs +++ b/toyos-acpi/tests/corpus.rs @@ -544,9 +544,7 @@ fn psci_of(table: &[u8]) -> Psci { psci(&fadt) } -/// `ARM_BOOT_ARCH` on a FADT that defines it. QEMU 11.1.1's `virt` -/// publishes revision 6, minor 3 and `PSCI_COMPLIANT`, with `PSCI_USE_HVC` -/// where the guest has no EL2. +/// `ARM_BOOT_ARCH` on a FADT that defines it. #[test] fn the_psci_conduit_is_read_off_a_fadt_that_defines_it() { assert_eq!(psci_of(&arm_facp(6, 3, 0b01)), Psci::Smc); From bacab58d0628e911dffc2934b640e18e87f3bf14 Mon Sep 17 00:00:00 2001 From: japabu Date: Wed, 30 Sep 2026 16:54:19 +0200 Subject: [PATCH 5/6] issues: the ARM track lists a kernel NVMe driver #536 moved out Found reading the track through the merge of #536: stage 4 still names the kernel's NVMe driver among those `arch::msi_message` refuses, and `kernel/src/drivers/nvme.rs` is gone; `drivers/pci.rs` is the one caller left. The lines are stage 4's, not this branch's, so they are filed rather than cut here. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L --- ...rnel-nvme-driver-that-storage-moved-out.md | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) create mode 100644 issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md diff --git a/issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md b/issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md new file mode 100644 index 00000000000..56b9004d5b7 --- /dev/null +++ b/issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md @@ -0,0 +1,19 @@ +--- +status: open +kind: tooling +opened: 2026-09-30 +--- + +# The ARM track lists a kernel NVMe driver that storage moved out + +`issues/kernel/toyos-runs-on-arm64.md`'s stage 4 says `arch::msi_message` +refuses "so the kernel's xHCI (`virt`'s boot stick), NVMe, HDA, virtio-sound, +virtio-console and virtio-gpu drivers each refuse their function by name". +#536 deleted `kernel/src/drivers/nvme.rs`: an NVMe controller is +`/system/bin/blockd`'s, a claimed function, and `kernel/src/drivers/pci.rs` is +`msi_message`'s one caller. The track's "NVMe, xHCI and virtio have no ISA +dependence" in its assessment names the same driver. + +The exit is the track's lines saying what the tree holds: NVMe out of the +kernel's list, and the claimed function's MSI named where it is owed +(`issues/kernel/nothing-reaches-the-msi-arm-of-a-claimed-function.md`). From cd125a18fbed055c12d9f024929ca4279dbeb048 Mon Sep 17 00:00:00 2001 From: japabu Date: Wed, 30 Sep 2026 18:02:27 +0200 Subject: [PATCH 6/6] Answer the final review of #633: the ARM track's false NVMe word goes, and its issue with it The track's stage 4 named the kernel's NVMe driver among the drivers `arch::msi_message` refuses on AArch64; #536 deleted that driver. The word is deleted in place, and the issue that tracked deleting it goes: stale prose is deleted, not tracked. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L --- ...rnel-nvme-driver-that-storage-moved-out.md | 19 ------------------- issues/kernel/toyos-runs-on-arm64.md | 2 +- 2 files changed, 1 insertion(+), 20 deletions(-) delete mode 100644 issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md diff --git a/issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md b/issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md deleted file mode 100644 index 56b9004d5b7..00000000000 --- a/issues/kernel/the-arm-track-lists-a-kernel-nvme-driver-that-storage-moved-out.md +++ /dev/null @@ -1,19 +0,0 @@ ---- -status: open -kind: tooling -opened: 2026-09-30 ---- - -# The ARM track lists a kernel NVMe driver that storage moved out - -`issues/kernel/toyos-runs-on-arm64.md`'s stage 4 says `arch::msi_message` -refuses "so the kernel's xHCI (`virt`'s boot stick), NVMe, HDA, virtio-sound, -virtio-console and virtio-gpu drivers each refuse their function by name". -#536 deleted `kernel/src/drivers/nvme.rs`: an NVMe controller is -`/system/bin/blockd`'s, a claimed function, and `kernel/src/drivers/pci.rs` is -`msi_message`'s one caller. The track's "NVMe, xHCI and virtio have no ISA -dependence" in its assessment names the same driver. - -The exit is the track's lines saying what the tree holds: NVMe out of the -kernel's list, and the claimed function's MSI named where it is owed -(`issues/kernel/nothing-reaches-the-msi-arm-of-a-claimed-function.md`). diff --git a/issues/kernel/toyos-runs-on-arm64.md b/issues/kernel/toyos-runs-on-arm64.md index cd418fcceb9..b93f81a9881 100644 --- a/issues/kernel/toyos-runs-on-arm64.md +++ b/issues/kernel/toyos-runs-on-arm64.md @@ -277,7 +277,7 @@ Each stage names its exit; "measured" means a number from a run. stage's SMMUv3 first. Stubbed on AArch64, each owned by the small-kernel track, which moves the driver out of the kernel: - `arch::msi_message` refuses, so the kernel's xHCI (`virt`'s boot stick), - NVMe, HDA, virtio-sound, virtio-console and virtio-gpu drivers each + HDA, virtio-sound, virtio-console and virtio-gpu drivers each refuse their function by name. - `drivers::gop` refuses a scanout that is not whole 2 MiB pages of its own, which a `ramfb` scanout carved out of RAM need not be.