Skip to content
Merged
31 changes: 31 additions & 0 deletions issues/kernel/an-ipi-can-overtake-the-stores-it-announces.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
---
status: open
kind: defect
opened: 2026-09-30
---

# An IPI can overtake the stores it announces

Neither architecture orders a CPU's earlier stores before the interrupt it
raises to announce them, so on hardware the target can take the interrupt and
read the word it was sent for before that word has landed.

- **AArch64**: `kernel/src/arch/aarch64/irqchip.rs`'s `raise` writes
`ICC_SGI1R_EL1` with no barrier before it. Only a `DSB` orders a system
register write after stores to Normal memory; Linux's `gic_ipi_send_mask`
(`drivers/irqchip/irq-gic-v3.c`) issues `dsb(ishst)` before the write for
exactly this.
- **x86-64**: `kernel/src/arch/x86_64/apic.rs`'s `Reg::write` of the ICR is a
`wrmsr`, which in x2APIC mode is not serializing (SDM Vol. 3A, "MSR Access
in x2APIC Mode"); Linux fences every x2APIC IPI with `weak_wrmsr_fence`
(`mfence; lfence`).

One reader that depends on the order: `kernel/src/sched/dump.rs` stores
`OWES` for each sibling, then kicks it, and a sibling that takes its kick
before its `OWES` is visible answers nothing, which the dump then reports as
a silent CPU.

No guest here reds on it, and none is expected to: under TCG the write is a
call into QEMU's interrupt controller model and under HVF a trap, and neither
is the hardware's ordering. The evidence is the architecture and Linux's code,
and the exit is the barrier on both, one instruction each.
26 changes: 20 additions & 6 deletions issues/kernel/toyos-runs-on-arm64.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,9 +28,8 @@ Every x86 guest on this host runs under TCG emulation instead — there is no

## Owner rulings, 2026-09-26

- **Parked** (owner ruling, 2026-09-29). No later stage starts now. The Exit
(LLVM under emulation) needs stages 4-7; until they are
picked up it is unmet.
- **Running** (owner, 2026-09-29): "can we get the most important tracks
running in parallel? arm, network and llvm?"
- **Hardware discovery is ACPI only** (edk2 MADT, GTDT and SPCR under QEMU); no devicetree.
- **Start.** Stage 0 (shared groundwork on x86) and stages 1-3 (toolchain,
loader, kernel reaching serial on QEMU `virt`) start now. Stage 0 waits for
Expand Down Expand Up @@ -278,7 +277,7 @@ Each stage names its exit; "measured" means a number from a run.
stage's SMMUv3 first. Stubbed on AArch64, each owned by the small-kernel
track, which moves the driver out of the kernel:
- `arch::msi_message` refuses, so the kernel's xHCI (`virt`'s boot stick),
NVMe, HDA, virtio-sound, virtio-console and virtio-gpu drivers each
HDA, virtio-sound, virtio-console and virtio-gpu drivers each
refuse their function by name.
- `drivers::gop` refuses a scanout that is not whole 2 MiB pages of its
own, which a `ramfb` scanout carved out of RAM need not be.
Expand All @@ -292,6 +291,20 @@ Each stage names its exit; "measured" means a number from a run.
`dlopen` test; until it does the kernel refuses `R_AARCH64_TLSDESC` by name
(`toyos_elf::rela::ExeRefusal::TlsDescriptor` for an executable,
`toyos_elf::RelocError::TlsDescriptor` for a library).
**Every CPU starts, ahead of small-kernel stage 6 by the owner's word, as
stage 4 did:** it ports no device interrupt, so nothing of the relay.
Owed before the exit holds:
`SYSTEM_RESET`, `SYSTEM_OFF` and `CPU_OFF` behind a reset and power-off
seam that takes x86-64's reset register and PM1a out of
`drivers/acpi.rs`, and the stop shown on eight CPUs ending in that
power-off; the TLS-descriptor resolver;
`issues/kernel/the-crash-evidence-records-x86-fault-registers.md`; the
blocked-task dump's probe of a CPU that ignored its kick
(`sched/dump.rs`'s `probe_silent`), which reaches `irqchip::send_nmi`'s
`owed!` on a machine of more than one CPU, and which nothing but the
`dump-deaf-cpu` actuator asks for until AArch64 has a keyboard; and, for
the first HVF run, the clean of an AP's start block to the point of
coherency, which TCG cannot fail on.

6. **Virtio on `virt`.** virtio-pci (ECAM from MCFG) for blk, net, gpu,
sound, input and rng. virtio-input replaces the i8042 as the
Expand Down Expand Up @@ -323,8 +336,9 @@ ARM is done when LLVM compiles under emulation on the development Mac (owner
ruling, 2026-09-29). One recipe is timed three ways: macOS natively, a Linux
arm64 guest under QEMU with HVF, and ToyOS arm64 under QEMU with HVF. All
three times are recorded, and it passes when ToyOS is at least as fast as the
Linux guest. ToyOS compiling LLVM is `issues/build/toyos-builds-itself.md`'s
work on AArch64.
Linux guest. The timing is one measurement, taken by hand and not automated
(owner ruling, 2026-09-29). ToyOS compiling LLVM is
`issues/build/toyos-builds-itself.md`'s work on AArch64.

## Interactions with other tracks

Expand Down
2 changes: 0 additions & 2 deletions kernel-loom/tests/log_publish.rs
Original file line number Diff line number Diff line change
Expand Up @@ -40,8 +40,6 @@ fn a_reader_that_finds_a_shard_finds_it_built() {
let publisher = registry.clone();

let w = loom::thread::spawn(move || {
// What `alloc_log_shard` does, in the order it does it: build the
// shard, then make it reachable.
let shard = Box::into_raw(Box::new(Shard::new()));
// SAFETY: leaked, so it outlives every reader; published once.
unsafe { publish(&publisher[..], 1, shard) };
Expand Down
12 changes: 6 additions & 6 deletions kernel-loom/tests/smp_bringup.rs
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@ use std::sync::atomic::{AtomicBool, Ordering::SeqCst};
use kernel_loom::smp_roster::Roster;
use loom::sync::Arc;

/// A LAPIC id no uncommitted slot holds (the roster fills a slot with `u32::MAX`).
const LAPIC: u32 = 7;
/// A hardware id no uncommitted slot holds (the roster fills a slot with `u32::MAX`).
const HARDWARE_ID: u32 = 7;

/// Bounded for loom (an unbounded spin never finishes); each turn a scheduling point.
const POLLS: usize = 4;
Expand All @@ -32,16 +32,16 @@ fn a_committed_count_never_outruns_its_slot() {

let committer = {
let r = r.clone();
loom::thread::spawn(move || r.commit(attempt, LAPIC))
loom::thread::spawn(move || r.commit(attempt, HARDWARE_ID))
};

let reader = loom::thread::spawn(move || {
if r.count() >= 2 {
let got = r.apic_id(1);
let got = r.hardware_id(1);
SAW.store(true, SeqCst);
assert_eq!(
got, LAPIC,
"a reader saw cpu1 counted while its LAPIC slot was still unfilled",
got, HARDWARE_ID,
"a reader saw cpu1 counted while its hardware-id slot was still unfilled",
);
}
});
Expand Down
2 changes: 1 addition & 1 deletion kernel/src/actuator.rs
Original file line number Diff line number Diff line change
Expand Up @@ -339,7 +339,7 @@ actuators! {
/// Leave every AP holding the CR0/CR4 that INIT left it.
no_ap_control_regs = "no-ap-control-regs";

/// Skip the startup IPI for the AP that would be cpu2, so a non-last AP never starts.
/// Skip the startup for the AP that would be cpu2, so a non-last AP never starts.
smp_skip_ap = "smp-skip-ap";

/// Time the same read loop on every CPU, either side of the `mov cr0` that enables caching.
Expand Down
149 changes: 100 additions & 49 deletions kernel/src/arch/aarch64/boot.rs
Original file line number Diff line number Diff line change
@@ -1,25 +1,29 @@
//! The AArch64 steps of the boot: the entry the loader jumps to, and what
//! `kernel_main` asks of this architecture at the points where one differs
//! from another.
//! The AArch64 steps of the boot: the entries firmware and the loader jump
//! to, and what `kernel_main` asks of this architecture at the points where
//! one differs from another.
//!
//! **The entry.** The loader leaves the CPU as firmware ran it — at EL2 or
//! EL1, on firmware's identity tables — cleans the kernel image and its own
//! **The entries.** The loader leaves the boot CPU as firmware ran it — at EL2
//! or EL1, on firmware's identity tables — cleans the kernel image and its own
//! tables to the point of coherency, and jumps to [`_start`]'s physical
//! address with `x0 = &KernelArgs`. `_start` writes the
//! address with `x0 = &KernelArgs`. PSCI starts every other CPU at
//! [`ap_start`]'s physical address, MMU off, at the level firmware gives an
//! operating system, with `x0` its [`ApStart`]'s physical address. Each names
//! a root table and branches to [`apply_declaration`], which writes the
//! [`control_regs`](super::control_regs) declaration whole: at EL2 it writes
//! `HCR_EL2` first, halts in a named refusal unless it reads back as declared,
//! then programs EL1's registers with the MMU already on and drops with `ERET`; at EL1 it
//! turns the MMU off first, so no translation register changes under a live
//! walk. Either way it arrives at the kernel's link address in the view at
//! `PHYS_OFFSET`, on the kernel's own stack, with the vectors installed, and
//! calls `kernel_main`.
//! then programs EL1's registers with the MMU already on and drops with
//! `ERET`; at EL1 it turns the MMU off first, so no translation register
//! changes under a live walk. Either way the entry resumes at its link address
//! in the view at `PHYS_OFFSET`, installs the vectors on its own stack, and
//! calls `kernel_main` or `smp::ap_entry`.

use core::mem::offset_of;

use toyos_abi::boot::{KernelArgs, MemoryMapEntry};
use toyos_acpi::MadtEntry;

use super::control_regs as regs;
use super::smp::ApStart;
use crate::drivers::acpi::direct_phys;
use crate::log;
use crate::mm::Region;
Expand All @@ -39,12 +43,80 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! {
"add x20, x20, x1",
"ldr x1, [x19, #{stack_size}]",
"add x20, x20, x1",
// x2 = TCR_EL1 whole, x3 = the loader's L0 table, x4 = MAIR_EL1.
"ldr x3, [x19, #{root}]",
"ldr x22, =1f",
"b {declare}",
"1:",
"ldr x1, ={phys_offset}",
"add x20, x20, x1",
"mov sp, x20",
"add x19, x19, x1",
"adrp x1, {entry_el}",
"str x21, [x1, :lo12:{entry_el}]",
"bl {install}",
"mov x0, x19",
"mov x29, xzr",
"mov x30, xzr",
"bl {kernel_main}",
kernel_memory = const offset_of!(KernelArgs, kernel_memory_addr),
stack_offset = const offset_of!(KernelArgs, kernel_stack_addr),
stack_size = const offset_of!(KernelArgs, kernel_stack_size),
root = const offset_of!(KernelArgs, boot_pml4_addr),
declare = sym apply_declaration,
phys_offset = const crate::PHYS_OFFSET,
entry_el = sym regs::ENTRY_EL,
install = sym super::trap::install,
kernel_main = sym crate::kernel_main,
);
}

/// Where PSCI `CPU_ON` starts every other CPU, at its physical address with
/// its MMU off and `x0` the physical address of its [`ApStart`].
/// # Safety
/// Only firmware may enter this, as `super::smp::start` asked it to.
#[unsafe(naked)]
pub(super) unsafe extern "C" fn ap_start() -> ! {
core::arch::naked_asm!(
"mov x19, x0",
"ldr x20, [x19, #{stack_top}]",
"ldr x3, [x19, #{root}]",
"ldr x22, =1f",
"b {declare}",
"1:",
"mov sp, x20",
"ldr x1, ={phys_offset}",
"add x19, x19, x1",
"bl {install}",
"mov x0, x19",
"mov x1, x21",
"mov x29, xzr",
"mov x30, xzr",
"bl {ap_entry}",
stack_top = const offset_of!(ApStart, stack_top),
root = const offset_of!(ApStart, root),
declare = sym apply_declaration,
phys_offset = const crate::PHYS_OFFSET,
install = sym super::trap::install,
ap_entry = sym super::smp::ap_entry,
);
}

/// The declaration written on this CPU and its MMU turned on under the root
/// in `x3`, whichever level it was entered at; then a branch to the link
/// address in `x22` at EL1 on `SP_EL1`, with `x21` the level entered at.
/// Keeps `x19`, `x20` and `x22`. Entered by a branch from an entry running at
/// its physical address, and never returns.
/// # Safety
/// Only [`_start`] and [`ap_start`] branch here, with `x3` a root that maps the
/// kernel image at both its physical and its link address.
#[unsafe(naked)]
unsafe extern "C" fn apply_declaration() -> ! {
core::arch::naked_asm!(
// x2 = TCR_EL1 whole, x4 = MAIR_EL1, x5 = SCTLR_EL1.
"ldr x2, ={tcr}",
"mrs x1, id_aa64mmfr0_el1",
"and x1, x1, #0xf",
"orr x2, x2, x1, lsl #{ips}",
"ldr x3, [x19, #{root}]",
"ldr x4, ={mair}",
"ldr x5, ={sctlr}",
"mrs x21, CurrentEL",
Expand All @@ -70,8 +142,7 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! {
"isb",
"msr sctlr_el1, x5",
"isb",
"ldr x1, =3f",
"br x1",
"br x22",
// EL2: `HCR_EL2` first, since with `E2H` set every `_el1` name
// below is an EL2 register; refused unless it reads back as declared.
// Then EL1's registers with the MMU on, EL2's others, and the drop.
Expand All @@ -83,7 +154,7 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! {
"cmp x6, x1",
"b.ne {refuse_hcr}",
// `ICC_SRE_EL2`, whose `Enable` lets EL1 write its own `ICC_SRE_EL1`
// (`super::irqchip::init`). Skipped where `ID_AA64PFR0_EL1.GIC` names
// (`super::irqchip::init_cpu`). Skipped where `ID_AA64PFR0_EL1.GIC` names
// no system-register interface: there the write is an undefined
// instruction under firmware's vectors and ends the boot silently,
// while EL1's own access fails under this kernel's, which report it.
Expand Down Expand Up @@ -112,31 +183,13 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! {
"dsb nsh",
"mov x1, #{spsr}",
"msr spsr_el2, x1",
"ldr x1, =3f",
"msr elr_el2, x1",
"msr elr_el2, x22",
"isb",
"eret",
// At the link address, EL1, MMU on.
"3:",
"ldr x1, ={phys_offset}",
"add x20, x20, x1",
"mov sp, x20",
"add x19, x19, x1",
"adrp x1, {entry_el}",
"str x21, [x1, :lo12:{entry_el}]",
"bl {install}",
"mov x0, x19",
"mov x29, xzr",
"mov x30, xzr",
"bl {kernel_main}",
// EL3, or anything else: nothing here may run there.
"9:",
"wfe",
"b 9b",
kernel_memory = const offset_of!(KernelArgs, kernel_memory_addr),
stack_offset = const offset_of!(KernelArgs, kernel_stack_addr),
stack_size = const offset_of!(KernelArgs, kernel_stack_size),
root = const offset_of!(KernelArgs, boot_pml4_addr),
tcr = const regs::TCR,
ips = const regs::TCR_IPS_SHIFT,
mair = const regs::MAIR,
Expand All @@ -150,10 +203,6 @@ pub unsafe extern "C" fn _start(_kernel_args: &KernelArgs) -> ! {
cntkctl = const regs::CNTKCTL,
icc_sre_el2 = const regs::ICC_SRE_EL2,
spsr = const regs::SPSR_EL2_TO_EL1,
phys_offset = const crate::PHYS_OFFSET,
entry_el = sym regs::ENTRY_EL,
install = sym super::trap::install,
kernel_main = sym crate::kernel_main,
);
}

Expand Down Expand Up @@ -183,7 +232,7 @@ pub fn before_panel() {}
/// and what this boot found — the memory map and the tables the rest of the
/// port reads.
pub fn after_console(args: &KernelArgs, maps: &[MemoryMapEntry]) {
regs::check();
regs::check(regs::ENTRY_EL.load(core::sync::atomic::Ordering::Relaxed));
for entry in maps {
log!("memory: {:#014x}..{:#014x} uefi type {}", entry.start, entry.end, entry.uefi_type);
}
Expand Down Expand Up @@ -259,17 +308,20 @@ pub fn reserved() -> Region {
Region { start: 0, end: 0 }
}

/// What the boot learns bringing interrupts up and hands later steps: nothing
/// yet, since the other CPUs the MADT names are the port's stage 5's to read.
pub struct Platform;
/// What the boot learns bringing interrupts up and hands later steps: the
/// other CPUs the MADT names, and how to start them.
pub struct Platform {
gic: super::irqchip::Gic,
psci: Option<super::psci::Conduit>,
}

/// Interrupt delivery: this CPU's per-CPU block, the GIC and the timer's
/// interrupt, and interrupts unmasked. The syscall gate is the vectors' own.
pub fn interrupts(rsdp_addr: u64) -> Platform {
super::percpu::init_bsp();
super::irqchip::init(rsdp_addr);
let gic = super::irqchip::init(rsdp_addr);
super::cpu::enable_interrupts();
Platform
Platform { gic, psci: super::psci::init(rsdp_addr) }
}

/// The clock: the generic timer's count, at the rate firmware states in
Expand All @@ -295,10 +347,9 @@ pub fn timer() {
/// drives on an ACPI Arm machine.
pub fn platform_devices(_rsdp_addr: u64) {}

/// Every other CPU, running: the port's stage 5, which starts each with PSCI
/// `CPU_ON`. Until then the boot CPU runs alone.
pub fn start_other_cpus(_platform: &Platform, _args: &KernelArgs) {
log!("smp: the boot CPU runs alone; the other CPUs are the port's stage 5 (PSCI CPU_ON)");
/// Every other CPU, running.
pub fn start_other_cpus(platform: &Platform, _args: &KernelArgs) {
super::smp::start(&platform.gic, platform.psci);
}

/// The interrupt-controller selftests an actuator asks for.
Expand Down
Loading
Loading