Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 10 additions & 3 deletions issues/a-counters-read-under-host-load-can-go-silent-for-15-s.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,9 +17,9 @@ compiler; the branch `wt/toyos-counterstall` at `bc5f36c7b`, whose harness
keeps every line the guest said. `test_rs_counters_read` started at 4.031,
spawned at 4.044, two of its threads exited (4.174 and 4.576), and then the
console carried nothing until the harness gave up 15 s later. No CPU printed
`sched: cpu=` again, though each last printed one at 2.17-2.22 s and prints
again on its first idle trip 10 s on (`scheduler::log_health`): no CPU went
idle, or the console stopped. No register capture of that guest exists. One
`sched: cpu=` again, though each last printed one at 2.17-2.22 s and that
kernel printed again on a CPU's first idle trip 10 s on: no CPU went idle, or
the console stopped. No register capture of that guest exists. One
in 30 guests of that loop; none in the 1724 guests that followed on the same
host at load 15-60.

Expand All @@ -39,6 +39,13 @@ once per answering CPU (`762a524f0`, reverted in the next commit) changed
neither the kicks taken during the read (median 116/122/118, fix/base/fix)
nor its span, so the waiter-list contention is not shown to be the cause.

**The idle report is gone.** The owner ruled on 2026-10-04, choosing "Remove
it entirely": "Delete the periodic report and its counters; hang triage uses
the trace diary and panic records instead." A silent guest no longer says
whether its CPUs went idle by a `sched:` line's absence; its diary
(`kernel/src/trace.rs`, read by `/system/bin/trace`) and its panic records
(`kernel/src/panic.rs`, `kernel/src/blackbox.rs`) do.

Exit: the cause of the silence is named from a capture of a silent guest
(registers over QMP before anything else touches it), and fixed with the
evidence, or shown to be the host stopping the guest.
12 changes: 10 additions & 2 deletions issues/a-shared-boot-stopped-answering-and-no-capture-says-why.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,8 +50,16 @@ would leave a whole boot idle straight after a sibling thread's clean exit.
path's two posts.

**Exit condition.** A capture taken from a boot that has actually stopped, which
names the first waiter and the subject it waits on — the blocked-task dump, or a
guest whose own last line is not the periodic reporter.
names the first waiter and the subject it waits on — the blocked-task dump, or
that boot's trace diary and panic records.

**The periodic reporter is gone.** The owner ruled on 2026-10-04, choosing
"Remove it entirely": "Delete the periodic report and its counters; hang
triage uses the trace diary and panic records instead." The `sched:` and
`PMM:` lines the sightings below quote are no longer printed, so a stopped
boot is told from a healthy idle one by its diary (`kernel/src/trace.rs`, read
by `/system/bin/trace`) and its panic records (`kernel/src/panic.rs`,
`kernel/src/blackbox.rs`).

**Sighting, 2026-09-25.** `cargo test` (the full fast tier) on the dev host, 12 wide, TCG, on
`wt/toyos-inspect` at `9ff0d254`. That branch touches no file under `kernel/`,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,10 @@ A thousand such processes is the whole 16 GB machine. Packing a process's
small regions into one shared page lowers the floor to about 4 MB, since every
stack still needs its own page with an unmapped neighbour as its guard; that
moves the ceiling to a few thousand and not past it. x86-64 offers no page
size between 4 KiB and 2 MiB.
size between 4 KiB and 2 MiB. That `PMM:` record and its rows went with the
kernel's idle report (the owner's ruling of 2026-10-04: "Delete the periodic
report and its counters"); stage 3's floor is read from the used memory
`SYS_SYSINFO` reports.

**What this removes as a side effect:** a device window mapped at 4 KiB no
longer shares a 2 MiB page with a neighbour's registers, so the relocation
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,9 @@ running through the silence; the capture cannot tell a dropped marker from a `lo
forwarding. The tree moved `rust` from `aca5f527f` to `9151571ca`, which changes std's exported C
`malloc`, `free` and `realloc` in every Rust guest program, so reading it as this loss rests on
that change being off its path. Owner: the orchestrator.
That kernel's ten-second report is gone: the owner ruled on 2026-10-04, choosing "Remove it
entirely": "Delete the periodic report and its counters; hang triage uses the trace diary and
panic records instead." A later sighting tells a live kernel from a stopped one by those.

## Exit condition

Expand Down

This file was deleted.

2 changes: 1 addition & 1 deletion issues/the-global-pipe-lock-spans-a-user-copy.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ The bulk copy is inside that closure, not outside it.

The size is bounded only by the ring: `PIPE_SIZE = PAGE_2M` (`pipe.rs:104`), `PAGE_2M = 2 * 1024 * 1024` (`toyos-userbound/src/span.rs:27`), and `capacity = total_size - size_of::<RingHeader>()` (`ring.rs:60`) with `RingHeader` `#[repr(C, align(64))]` holding one `AtomicU32` (`ring.rs:24-27`) — so **2,097,088 bytes** is the largest single copy under the lock. Nothing above caps it: `SYS_READ`/`SYS_WRITE` pass the userland length straight through (`kernel/src/syscall/dispatch.rs:110-116`), `object::ops::try_read` hands the full window to `pipe::try_read` (`kernel/src/object/ops.rs:339-340`), and the only other bound, `user_ptr::window` (`user_ptr.rs:269`), requires physical contiguity — which a demand-paged 2 MiB frame satisfies exactly.

A pipe's **first** write is worse, because the page is allocated lazily under the same lock. `try_write` calls `pipe.back()` (`pipe.rs:250`), which calls `pmm::alloc_page(pmm::Category::Pipe)` (`pipe.rs:148`). That takes `BITMAP` nested inside `PIPES` (`kernel/src/mm/pmm.rs:221`), linearly scans up to the whole physical bitmap for a free frame (`pmm.rs:224-241`), and then `write_bytes(..., 0, PAGE_2M)` — a 2 MiB zeroing (`pmm.rs:233-237`) — before `Ring::new` and the user copy that follows it.
A pipe's **first** write is worse, because the page is allocated lazily under the same lock. `try_write` calls `pipe.back()` (`pipe.rs:250`), which calls `pmm::alloc_page()` (`pipe.rs:148`). That takes `BITMAP` nested inside `PIPES` (`kernel/src/mm/pmm.rs:221`), linearly scans up to the whole physical bitmap for a free frame (`pmm.rs:224-241`), and then `write_bytes(..., 0, PAGE_2M)` — a 2 MiB zeroing (`pmm.rs:233-237`) — before `Ring::new` and the user copy that follows it.

## What queues behind it

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,14 @@ lines are quoted on #681 (comment 5962619169; the boot itself in comment
9.925 s all eight CPUs carry one at once. That the CPU stops is shown by two
CPUs waiting on a lock across a step, which went 4,503,694 and 4,554,675 ns
between two turns of their own spin.
- **MPERF reads the stop as about 4.55 ms of C0 on every CPU.** One T14 boot
of a scout image without ACPI mode read twelve back-to-back idle seconds
between counters rounds; `35cd63142`, which added this bullet, carries its
image hash and per-second lines. In the five whose SMI count moved by one,
cpu2, cpu4, cpu5 and cpu6, which ran none of the log's work, read 4535 to
4571 ppm busy; in the seven where it did not, 6 to 285 ppm. So the
`counters` row's idle second reads one of two floors, about 0.45% or 0.03%
and less: a 1 s second catches a 2.2 s-period SMI or does not.
- **Linux, in one condition, read none.** On the same machine under Ubuntu's
`6.8.0-142-generic`, `perf stat -a -A -e msr/smi/` read 0 on every CPU over
120.378 s beside `rtla timerlat top -q -d 2m --dma-latency 0`, which holds
Expand Down
2 changes: 1 addition & 1 deletion issues/toyos-explains-itself.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ opened: 2026-10-04

ToyOS answers what it is doing, what it did and what it is made of from
inside itself, with programs it ships. Today the answers are scattered: the
log carries numbers in prose (`irq:`, `tlb:`, `PMM:`, `sched:`, `syscalls:`),
log carries numbers in prose (`irq:`, `tlb:`, `syscalls:`),
the diary computes no lateness, nothing reads RAPL or C-state
residency, and a process's memory is a byte sum that reads 0 under
contention.
Expand Down
49 changes: 20 additions & 29 deletions kernel/pure/sched/cpu.rs
Original file line number Diff line number Diff line change
Expand Up @@ -398,18 +398,10 @@ impl<X: SchedPayload> CpuSched<X> {
self.dying.iter().map(|corpse| &corpse.task)
}

pub fn dying_len(&self) -> usize {
self.dying.len()
}

pub fn stopped(&self) -> impl Iterator<Item = &ReadyTask<X>> + '_ {
self.stopped.iter()
}

pub fn stopped_len(&self) -> usize {
self.stopped.len()
}

pub fn zombie_key(&self) -> Option<TaskKey> {
self.zombie.as_ref().map(|z| z.key())
}
Expand Down Expand Up @@ -1687,9 +1679,8 @@ impl<H: Hw, P: PreemptGuard> SchedPass<'_, '_, H, P, Disposed> {
// wants everything a new task would queue behind, and a corpse
// mid-unwind is exactly that: it is dispatched ahead of the fair
// band, so counting `rq` alone makes a CPU holding two teardowns
// look as empty as an idle one — the same blindness `dying_len`
// closes in the dump. The steal probe wants what this CPU could
// hand over, which is the fair band and only the fair band;
// look as empty as an idle one. The steal probe wants what this
// CPU could hand over, which is the fair band and only the fair band;
// publishing the first number to the second reader sends thieves to
// CPUs with nothing to give.
//
Expand Down Expand Up @@ -2727,7 +2718,7 @@ mod tests {
w.cpus[0].parked_task(key).is_some(),
"the retire lost the claim: the entry stays for the wake to find",
);
assert_eq!(w.cpus[0].dying_len(), 0, "the retire placed nothing itself");
assert_eq!(w.cpus[0].dying().count(), 0, "the retire placed nothing itself");
assert!(w.cpus[0].rq.is_empty(), "and queued nothing either");

// Now the wake it lost to lands, and *it* places the task — in the
Expand Down Expand Up @@ -2877,14 +2868,14 @@ mod tests {
}

assert!(stopped_shared.stop_pending(), "the safe point takes the mark");
assert_eq!(w.cpus[0].stopped_len(), 1);
assert_eq!(w.cpus[0].stopped().count(), 1);
assert_eq!(w.cpus[0].stopped[0].key(), stopped);
assert_eq!(
w.cpus[0].running().map(|t| t.key()),
Some(other),
"the CPU keeps working; only the stopped task is out",
);
assert_eq!(w.cpus[0].dying_len(), 0, "stopping is not dying");
assert_eq!(w.cpus[0].dying().count(), 0, "stopping is not dying");

// Every later pass, including ones where the CPU has nothing else.
w.run_a_pass_at(C0, Nanos(NOW.0 + QUANTUM_NS + 1));
Expand All @@ -2894,7 +2885,7 @@ mod tests {
Some(stopped),
"no pick serves the band",
);
assert_eq!(w.cpus[0].stopped_len(), 1);
assert_eq!(w.cpus[0].stopped().count(), 1);
w.abandon();
}

Expand Down Expand Up @@ -2928,7 +2919,7 @@ mod tests {
w.post_claimed_wake(C0, &parked_shared, WakeReason::Woken);
w.run_a_pass(C0);

assert_eq!(w.cpus[0].stopped_len(), 1, "the wake reached the band");
assert_eq!(w.cpus[0].stopped().count(), 1, "the wake reached the band");
assert_eq!(w.cpus[0].stopped[0].key(), parked);
assert!(w.cpus[0].rq.is_empty(), "and never the run queue");
assert!(
Expand Down Expand Up @@ -2970,9 +2961,9 @@ mod tests {
w.post_claimed_wake(C0, &shared, WakeReason::Woken);
w.run_a_pass(C0);

assert_eq!(w.cpus[0].stopped_len(), 1);
assert_eq!(w.cpus[0].stopped().count(), 1);
assert_eq!(w.cpus[0].stopped[0].key(), key);
assert_eq!(w.cpus[0].dying_len(), 0, "never dispatched to unwind");
assert_eq!(w.cpus[0].dying().count(), 0, "never dispatched to unwind");
assert!(w.cpus[0].running().is_none());
w.abandon();
}
Expand All @@ -2994,7 +2985,7 @@ mod tests {
let pass = SchedPass::begin(&mut cpus[0], env, NOW);
let _ = pass.dispose_stop().finish();
}
assert_eq!(w.cpus[0].stopped_len(), 1);
assert_eq!(w.cpus[0].stopped().count(), 1);
assert_eq!(
w.handles.get(C0).load(),
0,
Expand Down Expand Up @@ -3044,7 +3035,7 @@ mod tests {
"the corpse that has been waiting longest unwinds next",
);
assert_eq!(
w.cpus[0].dying_len(),
w.cpus[0].dying().count(),
1,
"and the one whose quantum expired went back to the dying list",
);
Expand Down Expand Up @@ -3087,7 +3078,7 @@ mod tests {
Some(queued),
"the waiting corpse runs; the killed one did not keep the CPU",
);
assert_eq!(w.cpus[0].dying_len(), 1);
assert_eq!(w.cpus[0].dying().count(), 1);
assert_eq!(w.cpus[0].dying[0].task.key(), expiring);
assert!(w.cpus[0].rq.is_empty(), "never through the fair queue");
w.abandon();
Expand Down Expand Up @@ -3131,7 +3122,7 @@ mod tests {
"the yield hands the CPU to the corpse that was waiting",
);
assert_eq!(
w.cpus[0].dying_len(),
w.cpus[0].dying().count(),
1,
"and the yielder went back to the dying list",
);
Expand Down Expand Up @@ -3193,7 +3184,7 @@ mod tests {
cpus[1].drain(env, NOW);
}

assert_eq!(w.cpus[1].dying_len(), 1, "the arriving corpse is placed to unwind");
assert_eq!(w.cpus[1].dying().count(), 1, "the arriving corpse is placed to unwind");
assert_eq!(w.cpus[1].dying[0].task.key(), key);
assert!(
w.cpus[1].rq.is_empty(),
Expand Down Expand Up @@ -3262,7 +3253,7 @@ mod tests {
Some(rt),
"the RT task got the CPU on the first pass after it became ready",
);
assert_eq!(w.cpus[0].dying_len(), 1, "the corpse is queued, not running");
assert_eq!(w.cpus[0].dying().count(), 1, "the corpse is queued, not running");
assert!(w.released().is_empty(), "and nothing was discarded");
w.abandon();
}
Expand All @@ -3284,7 +3275,7 @@ mod tests {
Some(rt),
"the expiring quantum is not a fresh one for the corpse",
);
assert_eq!(w.cpus[0].dying_len(), 1);
assert_eq!(w.cpus[0].dying().count(), 1);
w.abandon();
}

Expand Down Expand Up @@ -3388,7 +3379,7 @@ mod tests {
Some(rt),
"the RT task still takes the CPU on the pass that makes it ready",
);
assert_eq!(w.cpus[0].dying_len(), 1, "and the corpse is queued");
assert_eq!(w.cpus[0].dying().count(), 1, "and the corpse is queued");

// Follow the armed timer, which is the only thing that takes the CPU
// away from a task nothing preempts — a real machine does exactly this.
Expand Down Expand Up @@ -3537,7 +3528,7 @@ mod tests {
Some(rt),
"the grant ends on its own boundary and not a nanosecond later",
);
assert_eq!(w.cpus[0].dying_len(), 1, "the corpse is queued again");
assert_eq!(w.cpus[0].dying().count(), 1, "the corpse is queued again");
w.abandon();
}

Expand Down Expand Up @@ -3577,7 +3568,7 @@ mod tests {
"the corpse is unwinding, not doing real-time work, so the sibling \
that is doing real-time work gets the CPU at the next pass",
);
assert_eq!(w.cpus[0].dying_len(), 1, "and the corpse waits its age out");
assert_eq!(w.cpus[0].dying().count(), 1, "and the corpse waits its age out");
assert!(
w.cpus[0].dying[0].task.is_rt(),
"with its right intact — this is about the band it competes in, not \
Expand Down Expand Up @@ -3646,7 +3637,7 @@ mod tests {
"and the task came out of cpu2's surplus",
);
assert_eq!(
w.cpus[1].dying_len(),
w.cpus[1].dying().count(),
3,
"while cpu1's corpses stayed exactly where they were",
);
Expand Down
4 changes: 2 additions & 2 deletions kernel/src/arch/x86_64/vtd/table.rs
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@
use alloc::vec::Vec;

use crate::iommu::{AddressWidth, IommuError, Iova, StreamId};
use crate::mm::pmm::{self, Category, PhysPage};
use crate::mm::pmm::{self, PhysPage};
use crate::mm::{DirectMap, Mmio, PAGE_2M};

/// 4 KiB per table: 256 16-byte entries (root/context) or 512 8-byte entries (second-level).
Expand Down Expand Up @@ -49,7 +49,7 @@ impl Tables {
/// Returns one zeroed 4 KiB table, usable as a root, context, second-level, or invalidation-queue table.
pub fn alloc(&mut self) -> Table {
if self.used + TABLE_BYTES > PAGE_2M as usize {
let page = pmm::alloc_page(Category::Dma)
let page = pmm::alloc_page()
.expect("iommu: no physical memory for a remapping table");
self.pages.push(page);
self.used = 0;
Expand Down
2 changes: 1 addition & 1 deletion kernel/src/clock.rs
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ fn publish_page(counter_at_boot: u64, period_fs: u64) {
use toyos_abi::clock::{ClockPage, CLOCK_MAGIC};
let bytes = crate::mm::PAGE_2M as usize;
// Held for the machine's life: every process maps it.
let frame = crate::process::PageAlloc::new(bytes, crate::mm::pmm::Category::SharedMemory)
let frame = crate::process::PageAlloc::new(bytes)
.expect("clock: no 2 MiB frame for the clock page");
// SAFETY: a fresh allocation this function owns, `bytes` long, that no
// address space maps yet; zeroed whole because all of it is mapped, and
Expand Down
2 changes: 1 addition & 1 deletion kernel/src/drivers/gop.rs
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ pub fn init(
);
log!("GOP: scanout memory type {memory_type}");

let cursor_pages = crate::mm::pmm::alloc_contiguous(1, crate::mm::pmm::Category::Framebuffer).expect("GOP: cursor alloc failed");
let cursor_pages = crate::mm::pmm::alloc_contiguous(1).expect("GOP: cursor alloc failed");
let cursor_phys = cursor_pages[0].direct_map().phys();
// Cursor buffer is plain system RAM, not scanout, so it keeps the default write-back type.
let cursor = Region {
Expand Down
6 changes: 3 additions & 3 deletions kernel/src/drivers/virtio_gpu.rs
Original file line number Diff line number Diff line change
Expand Up @@ -415,8 +415,8 @@ impl GpuController {
let fb_pages = fb_size.div_ceil(PAGE_2M as usize);
let fb_aligned = (fb_pages * PAGE_2M as usize) as u64;
let all_pages =
[crate::mm::pmm::alloc_contiguous(fb_pages, crate::mm::pmm::Category::Framebuffer)?,
crate::mm::pmm::alloc_contiguous(fb_pages, crate::mm::pmm::Category::Framebuffer)?];
[crate::mm::pmm::alloc_contiguous(fb_pages)?,
crate::mm::pmm::alloc_contiguous(fb_pages)?];
let regions = all_pages.map(|pages| {
let phys = pages[0].direct_map().phys();
Region {
Expand Down Expand Up @@ -617,7 +617,7 @@ pub fn init(devices: &[PciDevice]) -> Option<(Box<dyn Gpu>, GpuInfo)> {
gpu.set_scanout(0, gpu.resource, rect);

let cursor_bytes = (CURSOR_SIZE * CURSOR_SIZE * 4) as usize;
let cursor_pages = crate::mm::pmm::alloc_contiguous(1, crate::mm::pmm::Category::Framebuffer).expect("VirtIO GPU: cursor alloc failed");
let cursor_pages = crate::mm::pmm::alloc_contiguous(1).expect("VirtIO GPU: cursor alloc failed");
let cursor_ptr = cursor_pages[0].direct_map().as_mut_ptr::<u8>();
let cursor_phys = cursor_pages[0].direct_map().phys();
gpu.cursor = Region {
Expand Down
Loading
Loading