Skip to content

A process writes two records, its spawn's and its exit's, and the machine's census is taken once, where the machine ends - #776

Merged
Japabu merged 4 commits into
mainfrom
wt/toyos-exitlog
Oct 8, 2026
Merged

Japabu merged 4 commits into
mainfrom
wt/toyos-exitlog

Conversation

@Japabu

@Japabu Japabu commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

The owner's rulings this lands under: "Fix the kernel's exit logging first", and "Every production log line but earn its keep and the logs mist follow a strategy."

Head ffe11fedc. The virt_smp measurement and the first T14 section below are at c6269f885; this head's gates and its own T14 run are under "Gates at ffe11fedc" and "The T14, at ffe11fedc", and what the head changes since is under "The review of 31958d800".

What was wrong

Every process exit wrote a census of the whole machine into the log: one irq: cpuN line per CPU, tlb:, the unclaimed vectors and, after an fsync, the flush census, beside syscalls:, memory: and exit: records of its own and three records at spawn. On the T14's eight CPUs that was 14 records and 1,989 bytes per process, 1,288 of them the eight irq: lines (the testcases readback at 809c33c0c). counters_metal's loaded phase spawns about 22,000 children in 20 s, so the boot outlogged logkeeper's sixteen one-megabyte files and its own middle was deleted, with the rows whose lines sat there.

The strategy, as the module headers state it

kernel/src/process.rs:

  • A process writes two records: one where its spawn lands, and one where it ends, exit: <name> pid=N code=N cpu=Nms …, the verdict first and then what that process itself consumed.
  • Each is written every time and charged to the process it names. Nothing there is rate-limited, and a thread's end writes nothing.
  • A reading of the whole machine is crate::census's, taken once, where the machine ends, by its stop or by its death. No process's start or end repeats one, so the log's volume is a function of what ran, never of how many CPUs watched it.

kernel/src/census.rs: one shape on every boot (irq: per CPU, tlb:, the unclaimed line, the panel); every reading a relaxed load of an atomic, with no lock, no allocation and no device, because a death seals from an NMI or an interrupt entry. Each source hands its lines to a sink: the stop logs them, a death writes them into the record it seals.

kernel/src/loader/mod.rs: a spawn that lands writes spawn: <path> pid=N unresolved=N (…ms); one that is refused writes one record naming why; nothing is said on the way, by the loader or by crate::elf under it.

The review of 31958d800, finding by finding

  • BLOCKER, the panic seal's census had no reader. screen_fatal_halt_composited requires, on the PANIC page it already reads out of the halted guest, a line that begins irq: cpu0 and one that begins tlb: shootdowns=. The mutation that deletes the census write from record_panic is the patch in this pull request's comment; its red and the greens either side are under the gates. It stays a guest test because no metal row seals a panic (both staged deaths are bound-ended) and the host cannot run record_panic, which seals from the kernel's panic path into the machine's page.
  • The loader's header was false at three sites, not two. All three are deleted, none rides the record:
    • dlopen: prescan … not caching (elf/cache.rs): the library still loads, the function's other uncached arm (no memory for its window) never said so, and nothing read the line. A count on spawn: was the alternative and would have been a field that reads 0 in 22,000 records a boot.
    • ELF: .symtab … no symbol map (loader/symbols.rs): its consequence is already in the one record. An executable whose table cannot be read exports nothing, and what a library wanted of it is counted in unresolved=.
    • ELF: {counts} refused (elf/index.rs), which the review did not name: a second record on a refused spawn, whose one record already carries the reason.
  • process.rs's "two records". Named, not folded. The isa: line is the device function's record, the pair of the one its claim wrote at the bind, bounded by the ISA table and absent for a process that held no ports: as a field it would be on every exit: for a fact about a device. A fault's report is several records (registers, the fault trace, frames) and cannot ride one line. The header says both are another owner's and charged to it.
  • The dropping arm of the stop's seal. The accounting is toyos_blackbox::Whole, beside Report::tail; the kernel keeps the sink. Three host tests: the exact fit and one byte less, a line that does not fit taking a later one that would have, and every room over two kilobytes holding the kept lines and the count's line.
  • The census's bound in a death's record. The census takes a budget, toyos_blackbox::CENSUS_BYTES, a quarter of the box as the recovery section's share is, through the same Whole: a line that does not fit is dropped whole with every later one and census: lines dropped to fit this record: N says so. Not a const assertion: the worst case is the sum of four lines in two architectures' formats, one of which can name 256 vectors, and a hand-kept sum is what the assertion would check.
  • snapshot_committed's range. The stop's records alone are meant: seal_tail takes the newest record's stamp as after and reads from the nanosecond past it. The tail on the T14 loses its fifteenth record, the spawn of /system/bin/reboot.
  • The fixtures. tests/checks.rs and src/metaldevices.rs open the tail with the head the kernel writes.
  • The issues. See "Records".

The review of f2b337afd, finding by finding

  • --ci host and the guest suite at no head. Still not run; see "Not run". This finding is open.
  • The contract was false for dynamically linked programs. The three dynamic: lines are gone, and so are the lines crate::elf wrote on the same path: dlopen: cache hit, dlopen: cached, dlopen: base=, three dlopen: applied … relocs, and one record per unresolved symbol in four places. An unresolved symbol does not refuse the spawn: elf/reloc.rs's stated policy is that it is untrusted input, never fatal, and faults only if used. Each function answers how many it left and the spawn's record carries the sum as unresolved=N. dlopen shares those functions, so it writes one record where it lands, dlopen: <path> pid=N unresolved=N.
  • The thread record's rate limit. The record, log_limited!, Limited, LIMIT_BURST, LIMIT_WINDOW_NS and the kernel's dependency on toyos-elide are deleted. No threads= is added: the count at teardown is of threads still in the table, not of threads the process ran, and nothing reads it.
  • The sixteen-record tail. The stop seals its own records: every record from a stamp taken before the census, newest first, bounded by toyos_blackbox::REPORT_BYTES, saying how many older ones it dropped in the words a death's tail uses. The constant and the ordering around it are gone. The unclaimed line is written at zero. The flush census is deleted with the counters only it read and its two feeders in usb_storage.rs; nothing read it.
  • A machine that dies took no census. The panic seal, the hard-lockup seal and the deadline's seal write crate::census into their record. They do not log it: a seal can have interrupted the log's own commit. deadline_wedge_chain, usb_load_chain and hard_lockup_chain now require it on the page, and the host test that feeds them a wedge page refuses a page without it and one missing only tlb:.
  • Why virt_mask_windows stays a guest test. What it judges is that an AArch64 mask-windows kernel on eight CPUs writes a report naming every CPU at each process's end and at its stop: a kernel's runtime output, which no type holds; a host test can only feed the judge text, which mask_windows_verdict does; and the one metal machine is x86-64.
  • NOTE, the spawn: record's fields: tid=, dst=, base=, entry= and root= are gone, none having a reader. The timings stay: issues/a-t14-wedge-ran-the-deadline-out-and-sealed-nothing.md reads a slow spawn by total=.
  • NOTE, mask-windows at the stop: it reports there too.
  • NOTE, the issues: see "Records".

Readers moved or deleted

reader what happened
irq_census_conservation (T14, testcases) Reads the black-box page the loader pass after the reset prints. The judge's logic is unchanged; irq_census_verdict (host) holds it.
mask_windows (T14) and virt_mask_windows (QEMU) A mask-windows kernel reports on its own (windows::report) at each process's end and at the stop. The census-pairing check had nothing left to pair and is replaced by: every CPU is in every report. mask_windows_verdict (host) moved with it.
the three death judges Now require the census on the sealed page.
the suite's irq summary (irqcensus::observe, summary) Deleted with its issue: the harness kills its guests, so none reaches a stop. issues/every-interrupt-lands-on-the-boot-cpu.md says the track has no instrument for the loaded suites today.
syscalls:, memory:, the flush census, the thread record, the dynamic:/dlopen: lines No judge read them.

The measurement

virt_smp under QEMU, eight CPUs; the window is its job unmap_touch, start marker to end marker, 9 processes ended in it; only the records a process's start and end write and the readings that rode its end are counted. Base is the whole change reverted (the negative control).

records per process bytes per process
base 809c33c0c 14.9 1,657.0
branch c6269f885 2.0 285.1

The base's 134 records: 72 irq:, 9 each of ELF:, spawn: TLS, spawn:, syscalls:, memory:, exit:, and 8 thread exits. The branch's 18: 9 spawn: (1,101 bytes) and 9 exit: (1,465 bytes).

The T14, at c6269f885

Five boots (testcases, testcases-watchdog, windowscase, deadlinewedge, hardlockup), each returned 0; the judge exited 0 with 24 passed and 0 failed. Read off the readbacks:

  • testcases' kernel.log is 7,562,577 bytes against 16,645,533 at the base, whole, with no was deleted line: 22,178 spawn: records of 163 bytes and 22,168 exit: … pid= records of 173, 336 bytes a process against 1,989. No irq: cpu, thread-exit, dynamic:, dlopen: or suppression line.
  • Its sealed page carries fifteen records, none dropped: the stop's fourteen, eight irq: cpuN, tlb:, the unclaimed line and the panel among them, and the record before the stop, which this head no longer seals.
  • The pages of deadlinewedge and hardlockup each carry eight irq: cpuN, one tlb:, the unclaimed line and the panel, written by the seal, and each says how many older ring records it dropped (166 on deadlinewedge, 180 on hardlockup).
  • windowscase: 24 windows: cpuN lines in the file, three reports of eight, and the stop's eight on the page.
  • At f2b337afd, before the fix round, the same boot wrote 22,120 exit: lines, 22,081 of them process ends and 39 thread ends (the orchestrator's reading of that run, in his comment here).

Gates at c6269f885

build-request-r4.sh: before virt_smp 0, measure-before 0, cargo metadata --locked 0, cargo test --test toyos-checks 0, cargo test --lib -- bootlog metal 0, cargo run -- --build-only 0, after virt_smp 0, measure-after 0, guests virt_mask_windows, machine_shutdown, screen_panic_muted, screen_fatal, nested_nmi_is_loud, virt_early_panic each 0, stop-census order 0 on the AArch64 and the x86-64 console, tree clean, staging done. The same request at 58f95ccce was red: one compile error and one host fixture, both fixed in c6269f885.

Gates at ffe11fedc

build-request-r5.sh, each step's own exit code:

  • cargo metadata --locked 0; cargo test --test toyos-checks 0 (running 37 tests); cargo test -p toyos-blackbox 0 (running 36 tests, and 0 doc tests); cargo test --lib -- bootlog metal 0 (running 96 tests); cargo run -- --build-only 0.
  • Guests by filter: screen_fatal 0 (screen_fatal_behind_a_painter and screen_fatal_halt_composited, the second with "sealed in the black box (14238 bytes)"), screen_panic_muted 0, nested_nmi_is_loud 0, virt_early_panic 0, virt_smp 0, virt_mask_windows 0, machine_shutdown 0.
  • The stop's census order (irq: lines, tlb:, the unclaimed line, the panel, the last word, no flush census): 0 on virt_smp's console and on machine_shutdown's.
  • The mutation, record_panic writing no census (the patch is in this pull request's comment): applied 0, screen_fatal_halt_composited exit 1 with FAIL screen_fatal_halt_composited: the panic's sealed record has no line that begins "irq: cpu0 ", restored 0, tree clean, and the same test on the restored tree 0.
  • git status --porcelain empty before the staging and after it.

The same request at 07ebd1e3a was red at --build-only on one unused import, removed in this head.

The T14, at ffe11fedc

Staged because this head changes what the stop's seal writes and the code every death's census goes through. Four boots (testcases, testcases-watchdog, deadlinewedge, hardlockup), each returned 0; the judge exited 0 with 23 passed and 0 failed. Read off the readbacks:

  • testcases' kernel.log is 7,606,002 bytes, whole, with no was deleted line: 22,306 spawn: records and 22,296 exit: … pid= records, and no exit: … tid=, ELF:, dlopen:, dynamic:, suppression or irq: cpu line.
  • Its sealed page carries the head and fourteen log-tail: records, none dropped: eight irq: cpuN, tlb: shootdowns=, irq: unclaimed vectors no-isr=0 and panel: paints= among them. The oldest is irq: cpu0, the stop's first; no spawn: /system/bin/reboot is on it.
  • The pages of deadlinewedge and hardlockup each carry eight irq: cpuN, one tlb: shootdowns=, the unclaimed line and the panel as the seal's own lines, in that order, above usb-recovery: and the ring's tail, with 166 and 180 older ring records dropped. No page carries census: lines dropped to fit this record.

High-risk checks (the kernel)

  • Negative control: the base arm of the measurement; for the panic seal's census, the mutation above.
  • Independent oracle: the T14 runs above, at c6269f885 and at this head.

Not run

  • cargo run -- --ci host and the whole guest suite, at any head. The pull request is a draft, so CI's host is skipped.
  • Whole's dropping arm in the kernel: host-tested in toyos-blackbox, reached by no boot, since no stop and no death here fills its room.
  • usbload was not staged; its seal is deadlinewedge's.
  • The dlopen: record was seen on the T14 at ffe11fedc, on the shared boot (the only tier that runs dlopen_dedup: it is not a test of CI's guest suite): 32 records of the shape dlopen: <path> pid=N unresolved=1, and refusals dlopen: ELF: …; test_rs_dlopen_dedup ended exit=0, the boot's judge 62 passed. A non-zero unresolved= is seen there too. No judge reads either record.

Records

  • issues/a-t14-boot-that-outlogs-its-retention-loses-its-middle-and-the-rows-whose-lines-sat-there.md takes the T14 reading and stays open: 7.56 MB is 45% of what is kept, so about 2.2 times the children outlogs it again, and the harness is as silent about a hole as it was.
  • issues/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md: its two landmark records are gone; the window runs from the job's exit: to the one spawn: record and takes in the VFS-lock sites it had eliminated for run 19; the WEDGED record it waits for carries the census.
  • issues/a-120000-ms-boot-deadline-fired-132859-ms-late-on-the-t14.md and issues/an-xhci-storm-starves-the-cpu-that-takes-it.md say where their census comes from now; the second says a storm the machine survives leaves no irq: line until the stop.
  • issues/a-thread-ended-after-the-boots-last-word-and-no-record-can-say-so-now.md, renamed from a-jobs-exit-record-landed-under-the-boots-last-word.md: the record that showed it is deleted, and the file says what is unanswered, its owner and an exit a test can fail.
  • issues/a-refused-syscall-writes-a-log-record-per-call-and-nothing-bounds-the-caller.md, filed: dlopen's and spawn's refusals write a record per call. Older than this branch, not fixed here.
  • issues/toyos-elides-limit-has-one-user-and-lives-in-a-shared-crate.md, filed: to be moved into logkeeper once The T14 rows that need the host to reach the machine under ToyOS are deleted, with the swap chain and the machine's end of the stream: 27 boots to 26 #773 has landed.
  • issues/a-processs-syscall-profile-is-one-threads.md names the exit record's syscalls=; its defect stands.

Growth

45 files, +776 −775 (git diff --shortstat origin/main...ffe11fedc). Kernel, toyos-blackbox, Cargo.lock and toyos-elide's header +402 −490; tests, toyos-blackbox's 66 lines of host tests among them, +151 −196; issues +223 −89. This round alone: 16 files, +323 −115.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A

…hine's census is the stop's

Every process exit wrote the machine's whole census into the log: one
`irq: cpuN` line per CPU, `tlb:`, the unclaimed vectors and, after an fsync,
the flush census, beside `syscalls:`, `memory:` and `exit:` records of its
own. On the T14's eight CPUs that is 14 records and 1,989 bytes per process
counting the three records its spawn writes, 1,288 of them the eight `irq:`
lines (the `testcases` readback at 809c33c, 8,304 children of
`counters_metal`'s `loaded` phase). That phase spawns about 21,900 children in
20 s, about 43 MB of log against sixteen one-megabyte files, and the boot's
own middle was deleted with the rows whose lines sat there. On a quiet boot
the census was still half the log: 2,520 `irq:` lines in `shared`'s 837,108
bytes.

What a process's end says now is one record: `exit: <name> pid=N code=N
cpu=Nms peak=NMB allocs=N frees=N syscalls=N syscall_wall=Nms <number>=<count>
...`. The verdict leads, so a profile that fills the record cuts only itself.
Nothing is dropped: the fields are what `syscalls:` and `memory:` carried.

A reading of the whole machine belongs to no process. The stop already took
the interrupt census (`syscall::machine`'s `quiesce`), and it now takes the
shootdowns' issuer census, the unclaimed vectors and the flush census there
too, once a boot. The `tlb:` line is said at zero as well: an eight-CPU
AArch64 boot that issued no invalidation stopped without one, and a census
whose shape depends on what the boot happened to do is not one a reader can
hold to. Their once-per-batch statics go: one caller, once. The
flush census leads because the sealed tail keeps the newest sixteen records
and it is the one no judge reads.

Who read what, and where each went:

- `irq_census_conservation` (the T14's `testcases`) read the census out of
  the log file. The stop writes after `logkeeper` has stopped, so the judge
  now reads the black-box page the next loader pass prints, as the panel
  census is read. The saved readback's page carries all eight `irq:` lines.
- `common::irqcensus::observe` and `summary`, the suite's per-run table of
  where interrupts landed, read every guest's exit census. The harness kills
  its guests, so none reaches a stop, and the table would read almost
  nothing: it is deleted, with
  `issues/the-irq-census-summary-takes-a-cpus-last-stamped-line-as-its-newest-read.md`,
  whose subject it was. `issues/every-interrupt-lands-on-the-boot-cpu.md` says
  what reads the census now and that its baseline was taken with the table.
- A `mask-windows` kernel printed each CPU's `windows:` line under its census
  line, and both metal and QEMU judges read a report per exit. That kernel
  now reports at each process's end on its own (`windows::report`), and no
  census rides with it. The judge's pairing of census and windows lines is
  replaced by what is left to hold: every CPU is in every report.
- `syscalls:`, `memory:` and the flush census had no judge.

A spawn writes one record too. `ELF: ... relocations indexed` and `spawn:
TLS ...` are deleted: no judge read them, and a line a spawn writes on its
way is not what a spawn says. They were the landmark of
`issues/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md`, which
now says its window opens at the job's exit record and closes at the one
`spawn:` record. The rule is the contract in `process.rs`'s header: two
records per process, each charged to it, and no reading of the whole machine
at either.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

T14 run at f2b337afd (orchestrator): testcases, testcases-watchdog, windowscase, each staged at this head and booted once.

boot toyos-metal exit image sha256
testcases 0 7d204d96…0f71278c
testcases-watchdog 0 8a4626eb…5ee34b9
windowscase 0 ba05dcbf…50a06d01

Judge (--metal --metal-readback … boot:testcases boot:windowscase): EXIT=0, 22 passed, 0 failed, 3 boots; irq_census_conservation and mask_windows among the PASS lines.

The testcases log on the stick, read from the readback's kernel.log:

809c33c0c f2b337afd
bytes that came back 16,645,533 (of about 43 MB written; parts 2 to 26 deleted) 8,730,435 (all of it)
logkeeper was deleted lines 15 0
irq: cpu, syscalls: pid=, memory: pid=, ELF:, spawn: TLS records present 0 of each
spawn: / exit: records — 22,091 / 22,120

The boot's log is whole and under the sixteen-part retention; the phase spawned about as many children as before (highest pid 22,091 against 21,897).

@Japabu
Japabu marked this pull request as ready for review October 8, 2026 18:21
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review of wt/toyos-exitlog at f2b337afd against origin/main 809c33c0c, round 1.

Growth: 19 files, +216 −337. Kernel (production) +97 −89; tests +44 −192; issues +75 −56.

BLOCKER

  • Evidence — cargo run -- --ci host is green at no head, and the guest suite has run at none: the body's "Not run" says both. CI's host and guest / suite on run 37823615634 were pending when this was written (toolchain / build passed). The change is the exit path of every process, so every guest test is one it reaches. What closes this is under "What the pending checks must show".
  • kernel/src/loader/mod.rs:10 and kernel/src/process.rs:10 — the contract is false of the tree for every dynamically linked program — the header says a spawn that lands writes one record and "Nothing is said on the way", and the same file writes dynamic: {} exe symbols available to libraries at :475 on every spawn of an executable with a DT_NEEDED, dynamic: loaded {} base=… ({} syms, {}ms) at :786 on each library's first load, and dynamic: unresolved exe symbol: {} at :487 once per unresolved GLOB_DAT, a count the file chooses and nothing bounds. No measurement saw them: virt_smp's children and all three T14 boots are statically linked (0 dynamic: lines in each log). The first is a per-spawn record no judge reads and goes, or its count rides the spawn: record; the second is a library's record, once per cached image, and the contract either names it or loses it; the third is a per-symbol record a program drives at any length and is the shape this branch exists to end.
  • kernel/src/process.rs:1326 with the header at :17 — the thread record's rate limit is a suppression, and the header writes it into the strategy — log_limited!'s LIMIT is one static for the site, so it is machine-wide: sixteen thread ends a second across every process, and process B's record is dropped because process A ended threads. That contradicts the sentence two lines above it ("A record here is charged to the process it names") and the body's "No rate limiter, no drop-under-load knob", which is false of the tree. It fired on the T14 at this head: exit: <name> tid=1 code=0 cpu=22ms (after 49 like it suppressed). It is also inconsistent with the rest of the contract: the same boot wrote 22,081 process-exit records at about 1,100 a second with no limit, because their volume is "a function of what ran"; a thread end is no different. No judge reads the record (src/bootlog.rs:736 uses one as a sample line for a stamp parser). Either it earns its keep and is written every time, or it does not and goes: the second deletes the record, log_limited!, Limited, LIMIT_BURST and LIMIT_WINDOW_NS (kernel/src/log/mod.rs:266-311, this site their only kernel user), with a threads=N on the exit record if the count is wanted. A limiter is neither.
  • kernel/src/syscall/machine.rs:95-102 against kernel/src/log/mod.rs:63 — the machine's only census has one carrier, a tail of sixteen records, and the stop writes fifteen on the T14 — read off all three readbacks at this head: two flush-census:, eight irq: cpuN, tlb:, panel:, stop:, usb-quiesce:, Rebooting., and the sixteenth is whatever came before the stop. The count is cpu_count + 7 plus one for each further flushed device, each further USB disk, and an unclaimed vector; the tail is a constant. Ten CPUs cut a flush-census: line on every boot, twelve cut cpu0's irq: line, and on the T14 one unclaimed vector and two more lines do the same and red irq_census_conservation with "no cpu0 in the census", which names neither the tail nor the cut. The body's "Unsure of" says this and nothing records it; the comment at :96 ("Oldest first is first cut from the sealed tail, so the flushes lead") orders the kernel's census around the constant. That is a design sized to the test machine and a compromise the branch found, neither removed nor filed. The page is not the bound: a death's seal carries the ring's tail by what fits, 190 records in the wedge issue's run 21. Seal the stop's own records, every one from a stamp taken before the census, bounded by the page and saying what it dropped, as the death's seal does. The same hunk is not consistent with its own reason for the tlb: fix: "a census whose shape depends on what the boot happened to do cannot be held to anything", while the flush-census: lines (kernel/src/block.rs:573, :584) and the unclaimed line stay omitted at zero, which is what makes the count move (the AArch64 virt_smp stop wrote no flush-census: line, machine_shutdown wrote two).
  • kernel/src/syscall/machine.rs:95 and kernel/src/process.rs:13-16 — a machine that dies takes no census, and nothing says so — quiesce is the only caller of the tlb, flush and unclaimed censuses and, with the blocked-task dump, of irq_census::log_census. A panic, the hard-lockup seal and the deadline's WEDGED seal take none, so a boot that does not reach its stop leaves no interrupt, shootdown or flush reading in any file or on the page. On main the ring's tail and the file both carried the last process end's, and two open T14 defects were read from exactly that: issues/a-120000-ms-boot-deadline-fired-132859-ms-late-on-the-t14.md ("The tail's irq: census, every line as sealed — printed twice, by two dying processes") and issues/an-xhci-storm-starves-the-cpu-that-takes-it.md (the irq: block out of the log file at 34.6 s). The wedge issue's next occurrence ends in a WEDGED seal and would carry none either. The contract's "taken once, where the machine stops" is the place the reading is not wanted most. Two legal outcomes: the paths that end the machine without a stop take the same census before they seal (irq_census::log_census documents itself as allocating nothing, taking no lock and touching no device; the other three read atomics), one function with quiesce as its other caller; or the loss is filed with an owner and an exit, and the three issues above say their evidence channel is gone. The first is the one the contract's own words ask for.
  • Guest tests — virt_mask_windows's judged behaviour changes (tests/common/irqcensus.rs:154-172, tests/toyos.rs:1922) and the body does not say why a type, a host test and a metal row cannot reach it — the body says "No guest test is added or cut" and stops. The reason is one sentence the body owes; whether it stands is judged when it is given.

NOTE

  • kernel/src/loader/mod.rs:687 — the spawn: record is 55% of the T14 log (4,801,663 of 8,730,435 bytes, 217 a record against the exit's 172) and carries tid=0 base=0x10000000000 on all 22,091 of its records in that boot, and a root= physical address no judge or tool reads — two fields with one value and one with no reader, about 27 + 16 bytes a spawn. Under "every production log line must earn its keep" they go, or the body names who reads each.
  • kernel/src/block.rs:566 — the flush census is written where no file carries it and is placed first so that it is the first cut; the body says no judge reads it — a record whose position is chosen for being expendable has not been shown to earn its keep. Name its reader (the loader's page after a reset, for whom) or delete it with the counters only it reads.
  • kernel/src/windows.rs:167 — a mask-windows kernel no longer reports at the stop, so every window closed between the last process's end and the reset is reported by nothing, and the module header's "a window is reported by the report after it closes" is false of them — no judge reads a duration there, so nothing reds.
  • tests/common/irqcensus.rs:168 — the replaced check is not weaker for any defect the kernel side can have: a CPU missing from every report still reds in mask_windows (longest.len() != cpus), one missing from some reds on the count, windows_under still refuses a report out of CPU order; neither the old nor the new check sees a whole report lost. mask_windows_verdict moved with it and still holds all three refusals.
  • issues/a-t14-boot-that-outlogs-its-retention-loses-its-middle-and-the-rows-whose-lines-sat-there.md:34 — "Owed: testcases' log bytes and parts on the T14 with that kernel" is false of the record: the reading exists at this head (8,730,435 bytes, 0 was deleted lines, 22,091 spawn: and 22,081 exit: … pid= records, 390 bytes a child against 1,989). The issue takes it in this diff.
  • The same issue does not close, and the branch is right not to close it. Its exit has three legs and one boot's reading meets none: "no boot of the metal profile" was read on one boot of 27; the alternative leg, a readback with a deleted part redding by name, does not exist, so the harness is as silent about a hole as it was; and acpi_server_events still rides testcases-hold. The margin is a factor, not a bound: 8.73 MB of 16 MiB is 52%, 98.7% of it still one job's spawn: and exit: records, and the phase spawns a child per CPU for twenty seconds, so about 1.9 times the children (a sixteen-CPU machine, or a faster one) outlogs the retention again. The issue should say that in those words.
  • issues/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md — still workable, and it says its loss plainly: the window is wider and the VFS-lock elimination no longer holds for a new occurrence. Its exit can be read. It does not say that the WEDGED tail it waits for now carries no census (the BLOCKER above).
  • issues/every-interrupt-lands-on-the-boot-cpu.md — deleting observe and summary is right: they read lines no killed guest prints any more, and keeping them would be dead code. The track says its present state plainly ("no instrument for the loaded suites today"). The deleted issue's subject went with the code and no citation of its slug remains at this head.
  • Pull request body — "Independent oracle: the T14. Owed" and "Not run: Anything on the T14" are false of the record: the reading is in the orchestrator's comment at this head (judge EXIT=0, 22 passed, 3 boots). The body carries it, with the comment's exit: count read as 22,120 lines of which 22,081 are process ends and 39 thread ends.
  • Pull request body — the measurement table is at 552cde186, one commit before the head; it stands on the T14 reading at the head, not on its own.

What the pending checks must show

At f2b337afd, on the run the ready pull request started: host concluded success with its cargo run -- --ci host step exit 0 (not skipped: the earlier run's host was SKIPPED on the draft and is no evidence); toolchain / build success (it is); guest / suite success on every shard, none skipped, with no test deleted or filtered to get there. A red in either is a defect of this branch until shown otherwise. They close the first BLOCKER only; the others need a new head, and that head needs these checks and the three T14 boots again, since every other BLOCKER changes what the kernel writes at a spawn, an exit or the stop.

SEND BACK

…es takes the census too

The review of f2b337a found the header's contract false in five places.
Each is closed by making the kernel do what the header says.

A spawn writes one record for a dynamically linked program too. The loader
and `crate::elf` under it wrote `dynamic: N exe symbols available`, `dynamic:
loaded <lib>` and, on every spawn and every library, `dlopen: cache hit`,
`dlopen: cached`, `dlopen: base=`, `dlopen: applied N ... relocs`, and one
record per unresolved symbol (`dynamic: unresolved exe symbol`, `dynamic: lib
unresolved symbol`, `dtpmod:`/`tls: unresolved TLS symbol`), a count the file
chooses and nothing bounded. All go. `elf::reloc`'s stated policy is that an
unresolved symbol is untrusted input, never fatal, and faults only if used,
so a spawn is not refused for one: each function answers how many it left and
the spawn's record carries the sum as `unresolved=N`. `dlopen` shares those
functions, so it says one record where it lands, `dlopen: <path> pid=N
unresolved=N`, in place of the lines they wrote for it.

The `spawn:` record loses `tid=`, `dst=`, `base=`, `entry=` and `root=`: no
judge or tool reads them, and two were one value on every record. Its
timings stay; a slow spawn is read by them
(`issues/a-t14-wedge-ran-the-deadline-out-and-sealed-nothing.md`).

A thread's end writes nothing. Its record went through `log_limited!`, one
static per site, so one process's thread ends suppressed another's: a
limiter is not a strategy. The record, `log_limited!`, `Limited`,
`LIMIT_BURST` and `LIMIT_WINDOW_NS` are deleted, and the kernel's dependency
on `toyos-elide` with them. No judge read the record; no `threads=` is added,
since the count at teardown is of threads still in the table, not of threads
the process ran, and nothing reads it.

The census is one module, `kernel/src/census.rs`, with one shape on every
boot: `irq:` per CPU, `tlb:`, the unclaimed line (said at zero now), the
panel. Each source hands its lines to a sink. The stop logs them; a panic,
the hard-lockup seal and the deadline's seal write them into the record they
seal, because a death seals from an NMI or an interrupt entry and may not log
there. Every reading is a relaxed load of an atomic: no lock, no allocation,
no device. The flush census is deleted with the counters only it read:
nothing read it, and it had been placed to be the first thing cut.

The stop seals its own records: every record from a stamp taken before the
census, newest first, bounded by what a report may spend of the page, and
saying how many older ones it dropped in the words a death's tail uses. The
sixteen-record constant is gone, and with it the ordering of the census
around it.

A `mask-windows` kernel reports at the stop as well as at each process's end.

The three death judges (`deadline_wedge_chain`, `usb_load_chain`,
`hard_lockup_chain`) now require the census on the sealed page.

Issues kept true: the retention issue takes the T14 reading at f2b337a and
says its margin is a factor; the wedge issue, the late-deadline issue and the
xHCI storm issue say where their census comes from now; the issue about a
thread's exit record after the last word says the record is gone and its
question is not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

T14 run at c6269f885 (orchestrator): five boots, each staged at this head, hash-checked against the request and booted once; two of them staged deaths.

boot toyos-metal exit image sha256
testcases 0 535d8026…6eaa1235
testcases-watchdog 0 d6979907…0bbae681
windowscase 0 13245941…744a2cdd
deadlinewedge 0 b2046e56…1b0d10bd
hardlockup 0 d59f6882…75ca6b0f

Judge (--metal --metal-readback … boot:testcases boot:windowscase boot:deadlinewedge boot:hardlockup): EXIT=0, 24 passed, 0 failed, 5 boots.

Read from the readbacks:

809c33c0c f2b337afd c6269f885
testcases log that came back, bytes 16,645,533 of about 43 MB written 8,730,435, whole 7,562,577, whole
logkeeper was deleted lines 15 0 0
spawn: / exit: records — 22,091 / 22,120 22,178 / 22,168
dynamic:, dlopen:, suppressed, irq: cpu lines in kernel.log present 0 dynamic: seen, suppressed present 0 of each

The sealed page (loader.log) of each of deadlinewedge, hardlockup and testcases carries eight irq: cpuN lines and one tlb: line; the two deaths' pages each say what older records they dropped, the stop's drops none.

Read off the readback of `testcases`: 7,562,577 bytes whole, 22,178 spawn and
22,168 exit records, 336 bytes a process, and the margin as the factor it is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu Japabu changed the title A process writes two records, its spawn's and its exit's, and the machine's census is taken once, at the stop A process writes two records, its spawn's and its exit's, and the machine's census is taken once, where the machine ends Oct 8, 2026
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review of wt/toyos-exitlog at 31958d800 against origin/main 26f5ef205, round 2. Reviewed as new: git diff f2b337afd 31958d800.

Growth: 38 files, +500 −706. Production (kernel, Cargo.lock, toyos-elide) +322 −456; tests +72 −194; issues +106 −56. git diff c6269f885 31958d800 is one issue file, +12 −8, so the measurements at c6269f885 are of this head's code.

Merge: git merge-tree --write-tree origin/main 31958d800 exits 0; so does this head onto origin/main merged with #773's head 19b660d76. In that tree toyos-elide keeps two users, toyos-symbols and logkeeper, and nothing names log_limited, LIMIT_BURST, the flush census or the deleted functions.

Round 1's BLOCKERs

  1. Evidence, --ci host and the guest suite at no head: OPEN. gh pr checks 776 at this head: host, toolchain and guest all skipping (run 37828827702, the draft). The body's "Not run" says the same.
  2. The contract false for dynamically linked programs: CLOSED for what it named. By the diff: the three dynamic: sites, dlopen: cache hit, dlopen: cached, dlopen: base=, the three dlopen: applied, dlopen: resolved and the five per-symbol sites are gone, and each of the six counting sites adds on its None arm. The values written are the ones written before (unwrap_or(0) for both TPOFF widths and DTPOFF64, unwrap_or(module_id) for DTPMOD64), and sys_dlopen still skips DTPOFF64 after a refused TPOFF. No measurement reached a non-zero count or a library; see the notes. Two sites the new header's wording is still false of are a NOTE below.
  3. The thread record's limiter: CLOSED. Deleted with log_limited!, Limited, both constants and the kernel's dependency. The T14's testcases log at c6269f885: 22,178 spawn: and 22,168 exit: … pid= records, 0 lines matching exit: … tid= or suppressed.
  4. The sixteen-record tail: CLOSED. The T14's testcases page: the head with no count, 15 log-tail: records newest first, irq: cpu7 down to irq: cpu0, tlb:, the unclaimed line at zero and the panel among them, no dropped line; windowscase's carries 23, the stop's eight windows: cpuN among them. The arithmetic, read: left = REPORT_BYTES − head − 1 − DROPPED_LINE_BYTES against a writer reopened at 0 with TEXT_BYTES as its limit, so head, records and the dropped line together are at most REPORT_BYTES and the reset's account keeps its 2,048; a line is measured by the Display that writes it; once one is dropped every older one is. The end dropped is the oldest, which at the stop is the census from cpu0 up; the count is said. The arm that drops ran nowhere (NOTE).
  5. A machine that dies takes no census: OPEN in part. Closed for the deadline's seal and the hard-lockup's: the T14 pages of deadlinewedge and hardlockup each carry irq: cpu0 to irq: cpu7, tlb: shootdowns=, irq: unclaimed vectors no-isr=0 and panel: paints= as the seal's own unstamped lines above usb-recovery: and the ring's tail, with 166 and 180 older records dropped; boot_deadline_ends_a_wedge and hard_lockup_ends_a_deaf_cpu PASS. Open for the panic's seal: the BLOCKER below.
  6. virt_mask_windows' reason: CLOSED. The body gives it and it stands: the AArch64 instrument's feed points are reached only by an AArch64 kernel running, the host half is mask_windows_verdict, and the T14 is x86-64.

Round 1's NOTEs: the spawn: fields, the flush census, mask-windows at the stop, the owed T14 reading in the retention issue and the wedge issue's census sentence are all taken.

BLOCKER

  • Evidence — cargo run -- --ci host green at no head, and the guest suite run at none — as round 1. What closes it is under "What the checks must show".
  • kernel/src/blackbox.rs:101 — the panic seal's census is on no page anyone has read, and one existing test reads that page — the mutation: delete the write_fmt(&mut report, format_args!("{}", crate::census::Sealed)) statement from record_panic. No test at any tier turns red: death_took_the_census is called by the three bound-ended judges only, and the fixture feeds a WEDGED page. The body's reason, "No QEMU test reads a death's sealed page back", is false of the tree: screen_fatal (tests/toyos.rs:2754-2776) reads the PANIC page out of the halted guest's memory with toyos_blackbox::recover and asserts PANIC: and its marker in it. It ran green at c6269f885 (r4-guest-screen_fatal.log: "sealed in the black box (14239 bytes)") and was not asked. A panic is the death a kernel bug takes, the first of the three round 1 named, and the measurement costs one assertion in a boot already paid for. The assertion: the sealed text carries a line that begins irq: cpu0 and one that begins tlb: shootdowns=, line-anchored because that boot's ring tail carries a blocked-task dump's stamped irq: cpu0 records, which a bare contains would take for the seal's. The implementer runs the mutation and shows screen_fatal red on it and green without. The body then says in one sentence why this stays a guest test (no metal row seals a panic; the host cannot run record_panic).

NOTE

  • kernel/src/loader/mod.rs:10-13 — "Nothing is said on the way, by this file or by crate::elf under it" is false at two sites on a load that lands — kernel/src/elf/cache.rs:49 (dlopen: prescan … will not fit one allocation, not caching, once per load of such a library, by every spawn that names it, since it is never cached) and kernel/src/loader/symbols.rs:68 (ELF: .symtab … exceed one kernel allocation, no symbol map, once per spawn of such an executable). Each is bounded per load and neither is the per-symbol shape, so this does not reopen BLOCKER 2; each goes, rides the one record, or is named by the header. kernel/src/process.rs's "two records" likewise does not cover isa: …'s ports went back with pid N at a device owner's end, or a faulting process's report. A header-only change leaves the T14 reading standing.
  • kernel/src/syscall/vm.rs:410 — the dlopen: record, as asked: it is not a record per call. A repeat of a held name returns at :241 or :372 before it, nothing unloads, and a load past MAX_LIBRARIES is refused at :384, so a process writes at most MAX_LIBRARIES of them, one per library it maps. That is right. The refusals are the unbounded ones: :256, :263, :274, :307 and :349 each write a record per call, and dlopen of a missing path in a loop is a log storm at syscall rate from any program; the same holds for a refused spawn. Older than this branch and unchanged by it, but the header now states the refusal record as policy and the tree's only limiter is gone from the kernel: filed in issues/ with an owner and an exit, not fixed here.
  • kernel/src/log/mod.rs:101-104 — the arm that drops ran on no machine and under no test — the T14's stops wrote 15 and 23 records, about 2.5 KB of 14,296, and with MAX_CPUS = 8 only a stop that fell short of its budget can fill the page. The accounting (fits whole; the first that does not fit drops itself and every older one; the count's line fits its reserve) is pure and belongs beside Report::tail in toyos-blackbox, which holds the same rule for bytes and has host tests; the kernel keeps the sink.
  • kernel/src/census.rs:34-45 with kernel/src/blackbox.rs:101 and kernel/src/drivers/panic_console/mod.rs:698 — what bounds the census in a death's record is MAX_CPUS = 8 and nothing says so — a line is at most about 337 bytes (twelve u64 counts), so eight are under 2.8 KB of REPORT_BYTES and it cannot fill the page today. It is written through Report::write, which cuts mid-line and says nothing, ahead of the recovery section and the ring's tail: at a MAX_CPUS near forty the worst case takes the page and the tail with it, silently, where the stop's seal counts what it drops. A const assertion against REPORT_BYTES beside kernel/src/smp.rs:13's, or the census taking a budget, before MAX_CPUS moves.
  • The census in a seal, as asked, by reading: every reading is a load of an atomic (BLOCKS and the per-CPU words, ISSUED, WAIT_NS, MAX_NS, TAKEN, NO_ISR, UNCLAIMED, the panel's four, the roster's count, the clock's period); no lock, no allocation, no device, no log!; the formatting is integer Display into a writer that returns Ok; before the clock is calibrated the panel's spans read zero and an unbuilt CPU is skipped. The REPORTED statics are deleted, so the census holds no state and a second run, by a panic inside a seal or an NMI over the stop's census::log, reads the same words and harms nothing. nanos_of_ticks' 128-bit divide was already in the wedge seal through the panel's line.
  • tests/common/power.rs:312-317 — the three death judges can fail on the claim: none of those boots reaches the stop's census::log (usbload's sweep starts above it in quiesce and never returns), so tlb: shootdowns=, the unclaimed line and the panel's can only be the seal's. irq: cpu0 alone could be a blocked-task dump's in the tail.
  • usbload unstaged — not owed before landing — its seal is deadline::expire's, the one deadlinewedge read, and what its judge needs from the tail is usb-load: sweeping disk 0, which in the last usbload readback (an earlier head) is the page's newest record, so the census's 1.4 KB comes off the other end. The first full profile reads it; a red there is this branch's.
  • unresolved= and dlopen: — the body's "has not been seen" names no tier: the guest suite's dlopen_dedup and abuse_dlopen_ledger reach sys_dlopen's record (and the T14's process_bound_libraries), so the suite owed above is its first sighting and the body should say what those logs carry. A non-zero unresolved= is reached by no test at any tier and no judge reads the field; not owed a guest test.
  • kernel/src/log/mod.rs:116 — snapshot_committed is from..=to, so the tail's oldest record is the one that was newest before the stop, not the stop's — on the T14 the fifteenth is spawn: /system/bin/reboot pid=N unresolved=0 (…). The body and the retention issue say "the stop's fifteen records"; fourteen are.
  • tests/checks.rs:970, src/metaldevices.rs:106 — both fixtures still open the tail with … newest first (16), a head no kernel writes.
  • issues/a-jobs-exit-record-landed-under-the-boots-last-word.md — its exit can no longer be read — the first leg is met by the record's deletion and the second names a deleted test; the file says the question survives and gives it no exit. Rewritten to the surviving question, or folded into the issue that records the test's restoring commit.
  • toyos-elide/src/limit.rs — one user is left, userland/logkeeper/src/origin.rs, in a crate two programs share for Elided — what one program alone uses belongs in its package. Filed, or moved once The T14 rows that need the host to reach the machine under ToyOS are deleted, with the swap chain and the machine's end of the stream: 27 boots to 26 #773 has landed.

What the checks must show

At the head that lands, on the run the ready pull request starts and again on the merge queue's: host concluded success with its cargo run -- --ci host step exit 0, not skipped; toolchain / build success, cold or not; guest / suite success on every shard, none skipped, with no test deleted or filtered to get there, screen_fatal with the page assertion, dlopen_dedup, abuse_dlopen_ledger, virt_mask_windows and machine_shutdown among them. A red in any is this branch's until shown otherwise.

The T14 reading at c6269f885 stands for a later head whose git diff c6269f885 <head> -- kernel toyos-blackbox bootloader is comments alone. A head that moves the tail's accounting or deletes a log site is read on the guest suite, and on the T14 only if it changes what a seal writes.

The checks do not land this by themselves: the second BLOCKER needs a new head with a changed guest test and its mutation run, and that diff is read in round 3 with the checks.

SEND BACK

…, and the load says nothing on the way

Review round 2 of #776.

The panic's census had no reader. `screen_fatal_halt_composited` already
reads the PANIC page out of the halted guest; it now requires a line that
begins `irq: cpu0 ` and one that begins `tlb: shootdowns=`, anchored at the
line's start because the ring's tail under the census can carry a
blocked-task dump's stamped `irq: cpu0`.

The stop's tail kept its accounting in the kernel, where the arm that drops
ran on no machine. It is `toyos_blackbox::Whole` now, beside `Report::tail`:
a line goes in whole or not at all, the first that does not fit drops itself
and every later one, and the count's line comes off the room first. Three
host tests reach the dropping arm, the exact fit and the reserve. The kernel
keeps the sink.

A death's census was bounded by `MAX_CPUS = 8` and nothing said so; written
through `Report::write` it would have been cut mid-line, silently, ahead of
the ring's tail. It goes through the same `Whole` within
`toyos_blackbox::CENSUS_BYTES`, a quarter of the box as the recovery
section's share is, and says what it dropped. A `const` assertion was the
other choice and was not taken: the worst case of four lines in two
architectures' formats (the unclaimed line alone can name 256 vectors) is a
hand-kept sum, where the budget is one accounting already under test.

The tail's range was `from..=to` with `from` the newest record before the
stop, so the tail opened with a record that was not the stop's (on the T14,
the spawn of `/system/bin/reboot`). `seal_tail` takes that stamp as `after`
and reads from the nanosecond past it.

Three records were written on the way of a load, against the loader's
header:
- `dlopen: prescan … not caching` is deleted. The library still loads; the
  sibling arm that leaves one uncached, no memory for its window, never said
  so; nothing read the line.
- `ELF: .symtab … no symbol map` is deleted. Its consequence is already in
  the spawn's one record: an executable that exports nothing leaves what a
  library wanted of it in `unresolved=`.
- `ELF: {counts} refused` is deleted: it was a second record on a refused
  spawn, whose one record already names the reason.

`kernel/src/process.rs`'s header names the two records written at a
process's end that are not the process's: a device function's ports going
back (`crate::isa`, the pair of its bind record, bounded by the ISA table
and absent for a process that held none, so not a field of every `exit:`),
and a fault's report, which is several records and cannot ride one.

Records: the two fixtures drop the `(16)` no kernel writes; the retention
issue says fourteen of the fifteen records were the stop's; the exit-record
issue is renamed to what survives of it, with an owner and an exit; refusal
records per call and `toyos-elide`'s one-user `limit` are filed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

T14 run at ffe11fedc (orchestrator): four boots staged at this head, hash-checked against the request, booted once each.

boot toyos-metal exit image sha256
testcases 0 e5559d93…acb8c7e8
testcases-watchdog 0 01d975fb…56aa99a13
deadlinewedge 0 4971bcdd…27b6f16f
hardlockup 0 54f3fff0…044de0c1

Judge (--metal --metal-readback … boot:testcases boot:deadlinewedge boot:hardlockup): EXIT=0, 23 passed, 0 failed, 4 boots.

Read from the readbacks: the testcases log is 7,606,002 bytes, whole, 0 was deleted lines; 22,306 spawn: and 22,296 exit: records; 0 dynamic:, dlopen:, suppressed or irq: cpu lines in kernel.log. The sealed page of each of the three boots carries eight irq: cpuN lines and one tlb: line; the stop's page carries fourteen log-tail records and drops none; neither death's page nor the stop's has a census: lines dropped line; the two deaths' pages each say what older ring records they dropped.

@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

The mutation round 2's BLOCKER names, as run at ffe11fedc by build-request-r5.sh: record_panic writes no census. Applied with git apply after --check, cargo test --test toyos-build -- screen_fatal_halt_composited --nocapture exited 1 with FAIL screen_fatal_halt_composited: the panic's sealed record has no line that begins "irq: cpu0 "; restored with git apply -R, tree clean, the same command exited 0.

diff --git a/kernel/src/blackbox.rs b/kernel/src/blackbox.rs
index 6f3913e06..40192418b 100644
--- a/kernel/src/blackbox.rs
+++ b/kernel/src/blackbox.rs
@@ -98,7 +98,6 @@ pub fn record_panic(records: &[u8]) {
         crate::panic::first_words(&mut report);
         // The machine's census, which a death takes as a stop does: atomics
         // alone, so it is inside this region's rules (`crate::census`).
-        let _ = core::fmt::Write::write_fmt(&mut report, format_args!("{}", crate::census::Sealed));
         crate::log::recovery::seal_into(&mut report);
         report.tail(records, toyos_blackbox::RECORD_OPENS_WITH);
         report.seal(State::Panic, stamp, identity);

@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review of wt/toyos-exitlog at ffe11fedc against origin/main 3c49b847d, round 3. Reviewed as new kernel code: git diff 31958d800 ffe11fedc.

Growth: 45 files, +776 −775 (git diff --shortstat origin/main...ffe11fedc). Production (kernel, toyos-blackbox less its tests, Cargo.lock, toyos-elide's header) +402 −490; tests (toyos-blackbox's 66 lines among them) +151 −196; issues +223 −89. This round: 16 files, +323 −116. Whole is 76 lines of production for 52 the kernel lost, with two callers and three host tests; accepted, round 2 asked for it there.

Round 2's BLOCKERs

  1. Evidence, --ci host and the guest suite at no head: OPEN. gh pr checks 776 at ffe11fedc: host, toolchain and guest all skipping (run 37832885535, the draft). Nothing else is open; what closes it is under "What the checks must show".
  2. The panic seal's census read by no test: CLOSED. exitlog-r5b.log at ffe11fedc: guest screen_fatal EXIT=0 with PASS screen_fatal_halt_composited and "sealed in the black box (14238 bytes)"; the mutation applied, screen_fatal_halt_composited EXIT=1, and r5-mutation-screen_fatal.log shows the guest booted and took the probe's panic (PANIC: panicked at kernel/src/sched/driver.rs:705:13) before FAIL screen_fatal_halt_composited: the panic's sealed record has no line that begins "irq: cpu0 ", so the red is the page assertion and not a boot; restored, tree clean, EXIT=0 again. The patch in the comment is the statement round 2 named. The assertion is line-anchored as asked, and the body gives the reason it stays a guest test.

Round 2's NOTEs, each taken: the three log sites are gone and process.rs names the two records that are another owner's (crate::isa::process_ends and dump_crash_diagnostics both exist); the tail's accounting is in toyos-blackbox with host tests; the census has a budget; the tail's range excludes the record before the stop; both fixtures open with the head the kernel writes; the exit-record issue is renamed with no citation of the old slug left at this head or on origin/main; the two issues are filed with status, kind, opened, an owner and an exit.

As asked

  • Whole, the accounting. within takes the count's widest line off the room first; put measures the line with its newline, drops it when it does not fit or when one was dropped before, and writes it otherwise; close says the count only where one was dropped. Each of the three tests fails on the defect it names, by reading: without the reserve the exact fit less one byte keeps all three lines and the 2,048-room sweep runs a room over; without the sticky dropped > 0 the second test keeps short; the sweep holds kept + counted = offered and kept as a prefix. r5-blackbox.log: 36 passed, the three among them.
  • The right end. The stop: snapshot_committed walks newest first, so what is dropped is the oldest and the count is said last in the tail's words. A death: each hands the lines oldest reading first, so what is dropped is the highest CPUs, then tlb:, the unclaimed line and the panel, and census: lines dropped to fit this record: N says so; the three death judges then red by name on a missing tlb:.
  • A death's census through it. No allocation, no lock, no index, no unwrap: saturating_sub/saturating_add, one subtraction guarded by the comparison above it, results of write_fmt discarded into a Report whose write cannot fail. Sealed::fmt answers Ok whatever was dropped. Safe from the panic path, the NMI and the deadline's entry as the round-2 reading of the sources was.
  • The budget. CENSUS_BYTES is 4,096 of REPORT_BYTES' 14,296; the recovery section is capped at 4,096 and its head; so a full census and a full section still leave the crash's first words and about 6 KB of ring tail, and ACCOUNT_BYTES is untouched. On the T14 the census is eleven lines and neither death's page carries the dropped line.
  • One thing Whole assumes. put renders a line twice, once to measure and once to write. Three sources load their words before format_args!; the panel's (kernel/src/drivers/panic_console/mod.rs:1074) loads four atomics inside Display, so a paint on another CPU between the two renders can write a digit more than was measured. No effect at this head: the panel's line is the census's last, the count's 63-byte reserve is unspent whenever that line is kept, and Report::write bounds the page whatever arrives. Not owed before landing; a source added after the panel's loads before it formats, as the other three do.
  • seal_tail(after). from is after + 1 ns, saturating. A stop's record cannot share the stamp of the newest record before it: quiesce::stop() and the console drain run between them. The T14's page: the head, fourteen log-tail: records from Rebooting. down to irq: cpu0, no spawn: /system/bin/reboot, no dropped line.
  • The deleted ELF: … refused. A refused spawn still writes one record with the path and the reason, spawn: <path>: ELF: a relocation group does not fit one allocation or the TLS-descriptor reason (kernel/src/loader/mod.rs:294-297), and the caller is told ResourceExhausted or InvalidArgument (Refused::error). What went is the seven counts, which the file the record names answers. parse_rela_entries has that one caller. The symtab arm changes no behaviour: it returned None before, and None falls to dynamic_map over an empty .dynsym, so the consequence is unresolved=. The prescan arm still returns None and the library still loads uncached. The one log! left in crate::elf is the cache budget's refusal, which is the one record of that refused load. No test, tool or issue at this head reads any of the three deleted lines.
  • What the body says is not measured. Accurate, with one addition. dlopen_dedup is not in MACHINE_TESTS or SCREEN_TESTS, so a filter naming it exits 1 at tests/toyos.rs:5108; abuse_dlopen_ledger is in RUST_SKIP (tests/toyos.rs:154) and runs only in the T14's testcases-bounds boot, which this round did not stage (0 dlopen and 0 dlopen_dedup in the testcases readback). CI's guest suite reaches the dlopen: record first, by dlopen_dedup on the shared boot of the ready pull request's run; the T14 reaches it only on a full profile. No judge on either reads the record, so the sighting is a person reading that boot's console. A non-zero unresolved= is reached at no tier and read by no judge; not owed a guest test. usbload stays as round 2 left it: its seal is the Sealed deadlinewedge read at this head.
  • The T14 at ffe11fedc. Read from the readbacks: four boots returned 0, judge EXIT=0, 23 passed, 0 failed. testcases' kernel.log is 7,606,002 bytes with 22,306 spawn: and 22,296 exit: … pid= records and 0 of exit: … tid=, ELF:, dlopen, dynamic:, was deleted, suppressed, irq: cpu. deadlinewedge and hardlockup: irq: cpu0 to irq: cpu7, tlb: shootdowns=, irq: unclaimed vectors no-isr=0, panel: paints= as unstamped lines in that order above usb-recovery:, then older records dropped to fit this page: 166 and 180. No page carries census: lines dropped.
  • The merge. git merge-tree --write-tree origin/main ffe11fedc exits 0 at 3c49b847d; so does this head onto origin/main merged with The host's toolchains live in one store outside every checkout, and every compiler is keyed #769's head 9a7d15900, and The host's toolchains live in one store outside every checkout, and every compiler is keyed #769 shares no file with this branch. Both sides touched Cargo.lock, tests/checks.rs and tests/toyos.rs and each merges without a hand: the merged lock differs from main's by the kernel's toyos-elide line alone. No file needs a resolution, so none would be new code. In the merged tree kernel, toyos-blackbox, bootloader, toyos-abi, toyos-elf and toyos-elide are byte-identical to ffe11fedc, and nothing names log_limited, LIMIT_BURST, the flush census or the deleted functions.

BLOCKER

  • Evidence — cargo run -- --ci host green at no head, and the guest suite run at none — as rounds 1 and 2; the draft skips all three checks.

NOTE

What the checks must show

The head that lands is ffe11fedc merged with origin/main by git alone, or ffe11fedc itself taken by the merge queue: no hand resolution, no commit of the branch's own after ffe11fedc, and git diff ffe11fedc <head> -- kernel toyos-blackbox bootloader toyos-abi toyos-elf toyos-elide empty. At that head, on the run the ready pull request starts and again on the merge queue's:

  • host: concluded success, its cargo run -- --ci host step exit 0, not skipped.
  • toolchain / build: success, cold or not.
  • guest / suite: success on every shard, none skipped, no test deleted or filtered to get there, and among the passes by name screen_fatal_halt_composited, screen_fatal_behind_a_painter, screen_panic_muted, nested_nmi_is_loud, virt_early_panic, virt_smp, virt_mask_windows, machine_shutdown, machine_shutdown_short_stop and dlopen_dedup.

With those three read, the evidence BLOCKER is closed and the orchestrator lands without a fourth round, correcting the body's numbers himself. A red in any of them is this branch's until shown otherwise; a red answered by a commit, a merge that needs a hand, or any change under the six paths above is a new head and is read in round 4.

The T14 reading at ffe11fedc stands for that merged head: what was read is what the kernel and the black box write, and both are byte-identical there. What the merge brings from #773 is logkeeper, sshserver and supervisor, read on the T14 under that pull request. If the shared boot's console is in the suite's log, the body takes one dlopen: /system/lib/libtls_dlopen_lib.so pid=N unresolved=0 line from it; if not, "seen at no tier" stays true and stays said.

SEND BACK

@Japabu
Japabu marked this pull request as ready for review October 8, 2026 19:41
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

CI and the last T14 reading at ffe11fedc, against round 3's list (orchestrator).

CI, run 37833784047, event pull_request, conclusion success, each job by its own conclusion:

  • host: success; the step Run cargo run -- --ci host concluded success.
  • toolchain / build: success.
  • guest / suite: success, not skipped; [ci] the suite: test result: ok. 36 passed, 36 total; [ci] Guest: 5 step(s), all green. Among the PASS lines by name: screen_fatal_halt_composited, screen_fatal_behind_a_painter, screen_panic_muted, nested_nmi_is_loud, virt_early_panic, virt_smp, virt_mask_windows, machine_shutdown, machine_shutdown_short_stop.

Round 3's list also named dlopen_dedup. It is not a test of CI's guest suite: it is a member of the T14's shared boot and runs at no other tier. So that boot was staged and booted on the T14 at this head:

  • shared: toyos-metal exit 0, image sha256 228367bb…df42b43c; judge (--metal --metal-readback … boot:shared) EXIT=0, 62 passed, 0 failed; TEST_END test_rs_dlopen_dedup exit=0.
  • The dlopen: record, seen for the first time: 32 records in that boot's log, of the shape dlopen: /system/lib/libtls_dlopen_lib.so pid=N unresolved=1, and refusals of the shape dlopen: ELF: .dynstr lies on a page of the module's writable window. A non-zero unresolved= is thereby seen as well. No judge reads either; this is a reading of the log.
  • The log is 299,037 bytes and whole.

With this the evidence BLOCKER is closed; the body's "seen at no tier" is no longer true of the dlopen: record and is corrected.

@Japabu
Japabu added this pull request to the merge queue Oct 8, 2026
Merged via the queue into main with commit 5335317 Oct 8, 2026
6 checks passed
@Japabu
Japabu deleted the wt/toyos-exitlog branch October 8, 2026 20:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant