diff --git a/.claude/agents/implementer.md b/.claude/agents/implementer.md index 76e4b54a29f..18c161f3403 100644 --- a/.claude/agents/implementer.md +++ b/.claude/agents/implementer.md @@ -44,6 +44,10 @@ guess. Then build, then test on root `CLAUDE.md`'s tiers before anyone reviews: ## A fork +A fork is for a crate, `rust/` or LLVM. A standalone C program has none: its ToyOS changes are patch +files in a recipe in this repository, against an upstream archive pinned by hash, and only a delta +too large to read as patches moves to a fork repository pinned by commit. + To edit a fork, clone it beside the monorepo and list it in `.cargo/config.toml`. Fork clones are shared by every worktree: explicit paths, never `stash`, never switch a branch in one. A fork keeps one branch per upstream base; a fix is a commit appended to it, never a new branch. A fork depends diff --git a/issues/build/a-toyos-hosted-llvm-is-configured-as-running-on-the-build-machine.md b/issues/build/a-toyos-hosted-llvm-is-configured-as-running-on-the-build-machine.md new file mode 100644 index 00000000000..2c8706929fc --- /dev/null +++ b/issues/build/a-toyos-hosted-llvm-is-configured-as-running-on-the-build-machine.md @@ -0,0 +1,33 @@ +--- +status: open +kind: defect +opened: 2026-10-03 +--- + +# A ToyOS-hosted LLVM is configured as running on the build machine + +LLVM takes `LLVM_HOST_TRIPLE` from `config.guess`, run by the CMake that +configures it (`llvm/cmake/modules/GetHostTriple.cmake`, +`llvm/cmake/config-ix.cmake`), unless the configuration names one, and +bootstrap names none: the `rust/` fork's +`src/bootstrap/src/core/build_steps/llvm.rs` defines +`LLVM_DEFAULT_TARGET_TRIPLE` and no `LLVM_HOST_TRIPLE`. So the LLVM that M2 +cross-builds to run on ToyOS (`issues/build/toyos-builds-itself.md`) records +the machine that built it as its host. The configure of that build on the +development Mac, the fork at the commit `main` pins (`95960d6c214`), logged +(#682, comment 5968053849) + +``` +-- LLVM host triple: arm64-apple-darwin27.0.0 +-- LLVM default target triple: x86_64-unknown-toyos +``` + +`config.guess` knows no ToyOS, and `get_host_triple` ends the configure where +it exits non-zero, so a configure run inside ToyOS stops there. That half is +read from `GetHostTriple.cmake`, not run. + +Owner: `issues/build/toyos-builds-itself.md`: M2 for the triple the host +build records, M5 for the configure inside ToyOS. + +Exit condition: the configure of a ToyOS-hosted LLVM logs `LLVM host triple: +x86_64-unknown-toyos`, and a configure inside ToyOS reaches its end. diff --git a/issues/build/code-used-by-one-program-lives-in-that-program.md b/issues/build/code-used-by-one-program-lives-in-that-program.md index 8735d4afb0b..1a3d63d3f8a 100644 --- a/issues/build/code-used-by-one-program-lives-in-that-program.md +++ b/issues/build/code-used-by-one-program-lives-in-that-program.md @@ -22,7 +22,9 @@ program named is neither. No step here merges one; each stays a crate under the no-panic track. Every step lands green on `--ci host` and `--build-only`. Test counts are what -`cargo test -p -- --list` lists today. +`cargo test -p -- --list` lists today. Step 2 lands right after #592 +and #631, of which #631 has landed, before the latency work +(`issues/kernel/toyos-beats-linuxs-latency-on-the-t14.md`; owner, 2026-10-03). 1. **Delete `toyos-userpin`.** It models the pin invariant and names nothing the kernel defines; `munmap_reissues_read_window` holds the kernel to it. diff --git a/issues/build/mio-deregister-fd-leaves-a-pending-poll-live.md b/issues/build/mio-deregister-fd-leaves-a-pending-poll-live.md index 75d98169392..3d87f16d7d6 100644 --- a/issues/build/mio-deregister-fd-leaves-a-pending-poll-live.md +++ b/issues/build/mio-deregister-fd-leaves-a-pending-poll-live.md @@ -55,22 +55,18 @@ first (grepped the estate for a submitter — found none, which is true and remains true) rather than asking whether `deregister_fd`'s behavior was itself correct. -## Why it matters for pipeline 2 +## What a fix has to promise -The completion track makes every kernel wait a cancellable completion, answered -by `Cancelled` rather than by discarding the stack -(`issues/kernel/every-wait-in-this-kernel-is-a-spin.md`). The 2026-08-15 -mechanism-consolidation audit called the userland-facing half of that "pipeline -2" and named its cancellation rewrite the highest-risk piece of the whole -track. mio's selector is exactly -the kind of consumer that rewrite has to get right: a `deregister` that cannot -promise "no event after this point" is the userland mirror of the kernel-side -bug pipeline 2 exists to remove. Re-adding a cancel op (or otherwise making -`deregister_fd` synchronous with the kernel's pending state — draining the SQ/CQ -for that fd, or having the kernel answer a targeted query) is in scope for -whoever designs that half; retiring `IORING_OP_POLL_REMOVE` a second time -without addressing this would repeat the same reasoning gap. +A kernel wait a kill ends is answered `Cancelled` (`kernel/src/watch.rs`). +mio's selector is the userland mirror of that, and a `deregister` that cannot +promise "no event after this point" is what it lacks: no ring op withdraws a +watch (`toyos-abi/src/inbox.rs` has `OP_NOP`, `OP_WATCH` and `OP_ACCEPT`). +Re-adding a cancel op (or otherwise making `deregister_fd` synchronous with +the kernel's pending state — draining the SQ/CQ for that fd, or having the +kernel answer a targeted query) is in scope for whoever fixes this; retiring +`IORING_OP_POLL_REMOVE` a second time without addressing this would repeat the +same reasoning gap. Filed rather than fixed: this is a fork (mio, not this repository) and a -design question (what the cancel primitive should look like once pipeline 2 -exists), not a local bug fix. +design question (what the cancel primitive should look like), not a local bug +fix. diff --git a/issues/build/n2-does-not-compile-for-toyos.md b/issues/build/n2-does-not-compile-for-toyos.md new file mode 100644 index 00000000000..5998969a89f --- /dev/null +++ b/issues/build/n2-does-not-compile-for-toyos.md @@ -0,0 +1,26 @@ +--- +status: open +kind: defect +opened: 2026-10-03 +--- + +# n2 does not compile for ToyOS + +n2 is the only Ninja the build runs (`src/n2.rs`), pinned at `b1fead52ccda`. +At that commit its `process::run_command` exists under `cfg(unix)`, +`cfg(windows)` and for `wasm32`, and its `terminal::use_fancy` and `get_cols` +under the same three; its `task.rs`, `run.rs` and `progress_fancy.rs` call +them under no `cfg`. `x86_64-unknown-toyos` is none of the three: its target +names no family (`compiler/rustc_target/src/spec/base/toyos.rs` in the `rust/` +fork). Read from the source, not built. + +Its unix arm is `posix_spawn` of `/bin/sh -c ` with `/dev/null` as +stdin, a pipe and `waitpid`, so a ToyOS arm waits on what +`issues/build/ninja-runs-every-command-through-a-bin-sh-toyos-does-not-have.md` +and `issues/filesystem/there-is-no-dev-null.md` wait on. + +Owner: `issues/build/toyos-builds-itself.md`, whose M5 builds LLVM in the +guest. + +Exit condition: n2 builds for `x86_64-unknown-toyos`, and inside ToyOS it +builds a file of two edges, one depending on the other. diff --git a/issues/build/nothing-launches-a-script-by-its-interpreter-line.md b/issues/build/nothing-launches-a-script-by-its-interpreter-line.md new file mode 100644 index 00000000000..41aa5f986dd --- /dev/null +++ b/issues/build/nothing-launches-a-script-by-its-interpreter-line.md @@ -0,0 +1,21 @@ +--- +status: open +kind: defect +opened: 2026-10-03 +--- + +# Nothing launches a script by its interpreter line + +A file that begins `#!` is not a program on ToyOS. A spawn of one is refused +`InvalidArgument` with every other file that is not an ELF +(`kernel/src/loader/mod.rs`, where `elf::parse_layout` fails), and nothing +above the kernel reads the line: `git grep -n -e '"#!"' -e ENOEXEC -- kernel +userland toyos toyos-abi` finds only `ENOEXEC`'s definition in +`userland/libc/include/errno.h`. + +Owner: `issues/build/toyos-builds-itself.md`, whose tools (a POSIX `sh`, Perl, +make, CMake) are the first programs here that start a script by its path. + +Exit condition: inside ToyOS, a spawn by path of an executable file whose +first line names an interpreter runs that interpreter on the file, shown by a +test that runs such a script and reads its exit code. diff --git a/issues/build/one-failed-request-for-the-crates-io-token-reds-a-publish-run.md b/issues/build/one-failed-request-for-the-crates-io-token-reds-a-publish-run.md new file mode 100644 index 00000000000..12411f54e06 --- /dev/null +++ b/issues/build/one-failed-request-for-the-crates-io-token-reds-a-publish-run.md @@ -0,0 +1,26 @@ +--- +status: open +kind: tooling +opened: 2026-10-03 +--- + +# One failed request for the crates.io token reds a publish run + +`publish.yml`'s `rust-lang/crates-io-auth-action` step trades the job's OIDC +token for a crates.io token in one request, and the job ends where that +request fails. Run 36998451606, `main` at `7c4a648b3`: the step logged +`Requesting token from: https://crates.io/api/v1/trusted_publishing/tokens` +and, 89 ms later, `##[error]fetch failed`; `cargo run -- --ci publish` was +skipped, and the run was over 12 s after it was created. Nothing ran it +again. The next landing's run, 37000495781 at `c1c504835`, was green, and +36998451606 is the one red among the forty `publish` runs on `main` from +36772021165 to 37110586853. + +Until the next landing, what that tip changed in the SDK crates is not on +crates.io, and the nightly's `release` refuses such a tip +(`issues/build/a-release-that-decides-before-its-tips-crates-are-up-reds-the-nightly.md`). + +Owner: the orchestrator. + +**Exit**: a publish run whose first request for the token fails asks again, +within a bound, before it reds. diff --git a/issues/build/parallel-tests-red-under-other-suites.md b/issues/build/parallel-tests-red-under-other-suites.md index 246e92ddaa4..bdb5567824f 100644 --- a/issues/build/parallel-tests-red-under-other-suites.md +++ b/issues/build/parallel-tests-red-under-other-suites.md @@ -218,7 +218,7 @@ changes. runs were 270/270 on a tree differing from the red one by two doc comments and one removed `#[track_caller]`; the branch's kernel delta touches no TLB, no shootdown and no `munmap` path. `ALONE … GREEN` in 145 ms. `screen_early_panic` failed in the same - run and is already this file's and the redlist's, `ALONE … GREEN` there too. + run and is already this file's, `ALONE … GREEN` there too. The assertion that went red is the test's own disarmed control: `munmap still took 11740090ns with the delay disarmed, so the numbers above measured @@ -245,8 +245,7 @@ changes. never ran. Exactly this file's typed-on-a-wall-clock class, and the class fix had already landed when the row was read back: `shell_type_line` (7a033450, 2026-08-26) bounds each burst by the queue and takes the guest's - own echo of the whole line as the verdict, with three tries. The redlist row - is retired against it. + own echo of the whole line as the verdict, with three tries. - **`syscall_window_nmi`** — added 2026-08-27, one sighting, dev host, a 288-name `cargo test` run at 92 guests with a second worktree's suite on the @@ -452,20 +451,6 @@ mechanism for it. seconds. The test asserts an ordering of two lines, and the second line never came. - **It contradicts a retirement rather than joining a class.** All three of the - name's earlier rows in `src/redlist.rs` are retired: the two dev-host - `ALONE: GREEN` rows by the single-word tally in - `kernel/src/arch/x86_64/i8042/tally.rs` (2026-08-17), which made `N interrupts - and 0 bytes` unprintable, and the CI row by "the verdict revises itself once" - (2026-08-28) — a mute line said while a decoder still holds the run is - `HEALTH_MUTE_BLIND`, the first blamed byte moves it to `HEALTH_MUTE_SAID` - with the line that names the bytes, and `i8042-split-burst` stages that - interleaving on every run. Four bytes and not zero puts this sighting on the - CI row's producer — the test's own Pause, reported after the first interrupt - delivered four of its six bytes — which is exactly what the 2026-08-28 clause - says is no longer waited for. Under a loaded host the second line did not - arrive, so the revision that retirement rests on is not unconditional. - Second sighting, the logd branch's fast tier at `eef19bd1`, on a dev host running two other worktrees' suites: the same words at `[kernel 5.865 cpu0] ... 2 interrupts and 4 bytes, nothing decoded`, the one red of 357, and @@ -474,8 +459,7 @@ mechanism for it. Owner: the i8042 tally. **Exit condition**: the verdict revising itself in a parallel run — a loaded full suite in which `i8042_undecoded_bytes`' - first mute line names nothing and its second names the sequence, or the - retirement's clause narrowed to the conditions under which it holds. + first mute line names nothing and its second names the sequence. ## Deleted as flaky tests diff --git a/issues/build/the-guest-check-gates-nothing.md b/issues/build/the-guest-check-gates-nothing.md deleted file mode 100644 index b88174c5277..00000000000 --- a/issues/build/the-guest-check-gates-nothing.md +++ /dev/null @@ -1,27 +0,0 @@ ---- -status: assigned -kind: tooling -opened: 2026-10-02 ---- - -# The guest check gates nothing - -`ci.yml`'s `guest` runs the guest suite as `guest / suite` on every non-draft -pull request and every merge group, and main's ruleset does not name it: a -pull request or a merge group whose `guest / suite` is red still lands. -`gh api repos/ToyOSOrg/ToyOS/rulesets/20589156` at 13:59:20Z on 2026-10-02 -reads `check_response_timeout_minutes` 240 and one required check, `host`. -Pull request #671, which made the check, says so of itself: "Until step 4 the -guest check gates nothing." - -The naming is a ruleset edit, with no commit, and it waits on main. Main's -cache scope holds no toolchain layer (`actions/caches` for `refs/heads/main` -at 14:00:46Z on 2026-10-02: two `host-sealed-` entries and nothing else), so -until `publish.yml`'s `toolchain` has saved the four there, every merge group -builds all four, as pull request run 36934214557 did in 2:11:46, and a check -named before that makes every landing wait on it. - -Owner: the orchestrator. - -**Exit**: `gh api repos/ToyOSOrg/ToyOS/rulesets/20589156` lists `guest / suite` -among the required checks. diff --git a/issues/build/the-guest-job-builds-every-image-on-every-run.md b/issues/build/the-guest-job-builds-every-image-on-every-run.md new file mode 100644 index 00000000000..c11698030a3 --- /dev/null +++ b/issues/build/the-guest-job-builds-every-image-on-every-run.md @@ -0,0 +1,27 @@ +--- +status: open +kind: tooling +opened: 2026-10-03 +--- + +# The guest job builds every image on every run + +`guest.yml`'s `suite` restores one cache entry, the sysroot, and builds the +rest in the job: the build driver, the harness, and each kernel, loader, ROOT +and test binary a guest boots. Job 111020980149 of run 37061831501 (#649's +head `b6635359a`, a pull request run with the sysroot restored from cache) +ran 18 min 11 s (its log's lines: #682, comment 5968053849): + +- 2 min 46 s in `deps`, before the checkout; +- 2 min 15 s compiling the driver and the harness (`Finished` in 1m 26s and + in 47.62s); +- 767.5 s in the suite by its own line, of which its 26 `PASS` lines account + for 208 s. The other 559 s is the log's silence, up to 209 s at a time, + between a `cleaning` line or a verdict and the next guest's first line: + three kernel builds by the suite's summary, and the loaders, ROOTs and test + binaries beside them. + +Owner: the orchestrator. + +**Exit**: what of those builds an entry restored from `main`'s cache scope +would save is measured on a pull request run. diff --git a/issues/build/the-hosted-rustcs-stage2-carries-seventeen-proc-macro-libraries-no-image-needs.md b/issues/build/the-hosted-rustcs-stage2-carries-seventeen-proc-macro-libraries-no-image-needs.md new file mode 100644 index 00000000000..3b9157cd488 --- /dev/null +++ b/issues/build/the-hosted-rustcs-stage2-carries-seventeen-proc-macro-libraries-no-image-needs.md @@ -0,0 +1,27 @@ +--- +status: open +kind: defect +opened: 2026-10-03 +--- + +# The hosted rustc's stage2 carries seventeen proc-macro libraries no image needs + +The ToyOS-hosted stage2's `lib/` in release +`toolchain-linux-x86_64-48dd24f826263d6c` holds eighteen shared objects: +`librustc_driver`, 486,137,984 bytes, and seventeen proc-macro libraries the +compiler's own build loaded, 179,575,736 bytes together (the listing: #682, +comment 5968053849). The Linux-hosted +stage2 beside it holds `librustc_driver` alone, 594,483,472 bytes, which +`issues/build/toyos-builds-itself.md` measures at 184.7 MB stripped. + +`collect_hosted_rustc` (`src/build.rs`) puts every `lib/*.so` of that +directory, and `bin/rustc`, into ROOT as built. `hosted-rustc` is `false` in +`system.toml` and a config that sets it is refused until the compiler's +licences are read (`build::shipped`), so no image carries them today. #629 +deletes the function, the key and the refusal, and this paragraph goes in its +merge. + +Owner: `issues/build/toyos-builds-itself.md`, M3. + +Exit condition: an image that ships the hosted rustc carries `librustc_driver` +stripped and no proc-macro library. diff --git a/issues/build/toyos-builds-itself.md b/issues/build/toyos-builds-itself.md index d4dd3114d8a..2f0e7633e11 100644 --- a/issues/build/toyos-builds-itself.md +++ b/issues/build/toyos-builds-itself.md @@ -139,3 +139,14 @@ M3's exit then waits on a linker in the guest (`issues/build/the-hosted-rustc-names-a-linker-toyos-does-not-have.md`), which toyos-ld is not: it refuses every executable with thread-local storage, so every std program. + +**Also owed, by milestone.** +- M2: the host triple a ToyOS-hosted LLVM records + (`issues/build/a-toyos-hosted-llvm-is-configured-as-running-on-the-build-machine.md`, + whose other half, the configure inside ToyOS, is M5's). +- M3: `issues/build/the-hosted-rustcs-stage2-carries-seventeen-proc-macro-libraries-no-image-needs.md`. +- M4: cargo's file locks, + `issues/design-debt/std-file-lock-answers-ok-and-locks-nothing.md`. +- M5: `issues/build/n2-does-not-compile-for-toyos.md`. +- The tools of "To build", the first programs that start a script by its + path: `issues/build/nothing-launches-a-script-by-its-interpreter-line.md`. diff --git a/issues/design-debt/redesign-the-log-subsystem.md b/issues/design-debt/redesign-the-log-subsystem.md index bc4264bb0b7..c3d4a99a434 100644 --- a/issues/design-debt/redesign-the-log-subsystem.md +++ b/issues/design-debt/redesign-the-log-subsystem.md @@ -7,12 +7,15 @@ opened: 2026-08-08 # Redesign the log subsystem, and re-shape `kernel/src` **The owner decided on 2026-08-19: go.** Both halves are approved as planned -work. Sequencing, set by the orchestrator with the ruling: the log core waits -until pipeline 2's second pull request (the lock conversions behind -`issues/kernel/every-wait-in-this-kernel-is-a-spin.md`) has landed — the two -touch the same scheduler-adjacent paths and land one at a time. The directory -re-shape is a scheduling matter exactly as priced below: one clean pass in a -window with few worktrees in flight, and never interleaved with a code change. +work. With the ruling the orchestrator sequenced the log core behind the +kernel's lock conversions (`vfs::VFS`, the FAT volumes' lock, `xhci::XHCI` and +`ProcessData` becoming sleep locks), which touch the same scheduler-adjacent +paths. Nothing plans those conversions any more: +`issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md` moves +storage and USB out of the kernel instead (owner, 2026-09-25), so that wait +names nothing that will land. The directory re-shape is a scheduling matter +exactly as priced below: one clean pass in a window with few worktrees in +flight, and never interleaved with a code change. The target shape stands as reviewed: a log core (ring + context stamping, once) with serial, file and screen as independent sinks carrying explicit diff --git a/issues/design-debt/std-file-lock-answers-ok-and-locks-nothing.md b/issues/design-debt/std-file-lock-answers-ok-and-locks-nothing.md new file mode 100644 index 00000000000..af236182e12 --- /dev/null +++ b/issues/design-debt/std-file-lock-answers-ok-and-locks-nothing.md @@ -0,0 +1,25 @@ +--- +status: open +kind: defect +opened: 2026-10-03 +--- + +# std's `File::lock` answers `Ok` and locks nothing + +The ToyOS file pal in the std fork (`library/std/src/sys/fs/toyos.rs`) answers +`Ok(())` from `File::lock`, `lock_shared`, `try_lock`, `try_lock_shared` and +`unlock` and takes no lock. Two processes that each ask for the exclusive lock +on one file are both told they hold it, and nothing in the kernel ABI or the +file servers' protocol could tell them otherwise: neither has a lock call. + +cargo is one such program. Every file lock it takes is one of these calls +(its `src/util/flock.rs` at `7c83d4cc0953`, the cargo the fork pins), so two +cargo processes in one ToyOS are kept apart by nothing. + +Owner: `issues/build/toyos-builds-itself.md`, whose M4 runs cargo in the guest; +the locks libc owes for the same milestone are +`issues/build/libc-refuses-what-toyos-cannot-yet-answer.md`'s. + +Exit condition: a second process's `try_lock` on a file another holds locked +answers `WouldBlock`, and its `lock` waits for the holder's `unlock` or its +end, shown by a test with two processes. diff --git a/issues/design-debt/the-internet-clients-work-unchanged.md b/issues/design-debt/the-internet-clients-work-unchanged.md index 6aa2b22830b..f766ad27769 100644 --- a/issues/design-debt/the-internet-clients-work-unchanged.md +++ b/issues/design-debt/the-internet-clients-work-unchanged.md @@ -33,8 +33,25 @@ netd, `toyos::net` and std's ToyOS networking in the `rust/` fork. program on `rustls`; the program and its judge's server installed `rustls-rustcrypto` 0.0.2-alpha, and #660 cut both. The row comes back on `ring` and waits for it; `rustls-rustcrypto` does not come back (owner, - 2026-10-02). Whether `ring` builds for the ToyOS target is open and - untried: no build for that target compiles it. `rustls-rustcrypto` is + 2026-10-02). Published `ring` 0.17.14 does not link for + `x86_64-unknown-toyos`: its `build.rs` picks its assembly by an OS list + that lacks `toyos`, and `src/rand.rs` has an OS list of its own. With + `"toyos"` added to `build.rs`'s `LINUX_ABI` and `target_os = "toyos"` to + `src/rand.rs` it builds for both ToyOS targets. In an x86-64 QEMU guest + it passes known answers (SHA-256, + SHA-512, HMAC, X25519, ChaCha20-Poly1305, AES-GCM, Ed25519, P-256, + `SystemRandom`), and `ureq` 3.4.2 on `rustls` 0.23.45 fetches 320000 + bytes over TLS 1.3 from a server on the host and refuses a wrong name and + an untrusted root (#682, comment 5968053596). What stands before it + lands: a git fork's + `build.rs` runs `perl`, an arrival + `issues/build/the-build-runs-host-tools-outside-rust-and-qemu.md` does not + declare, and leaves C asserts on, so `__assert_fail` is undefined unless + ring builds with `debug = false` or `toyos_c` is linked; `src/build.rs` + gives the C compiler's environment (`cc_env`) to userland builds alone, so + a test crate's C is compiled by the host's `cc`; and not run: AArch64, the + T14, a `git =` dependency, the licence gate over ring in an image. + `rustls-rustcrypto` is still named by doom's build script, which installs it on the host, and by a line the cut left in `tests/toyos-rust-tests` (`issues/build/the-guest-test-crate-depends-on-three-crates-no-test-uses.md`). diff --git a/issues/design-debt/the-shipped-image-carries-three-test-programs.md b/issues/design-debt/the-shipped-image-carries-three-test-programs.md index f43b8c0cc7d..4b3f4ce2ffb 100644 --- a/issues/design-debt/the-shipped-image-carries-three-test-programs.md +++ b/issues/design-debt/the-shipped-image-carries-three-test-programs.md @@ -13,7 +13,8 @@ of the shipped `system.toml` in three rows: nothing in the tree runs `/system/bin/proctest` but `proctest` itself. - `input-test = {}`: a test utility by its README, and nothing runs `/system/bin/input-test`; `toyos-symbols/tests/real.rs` reads a frozen copy - under `toyos-symbols/tests/fixtures/`, not the shipped binary. + under `toyos-symbols/tests/fixtures/`, not the shipped binary. It leaves the + shipped image (owner, 2026-10-03). - `"bin/spin" = "/system/bin/toybox"`: `userland/toybox/src/spin.rs` loops forever, and nothing in the tree runs it; its one named use is a test, the metal `sshd_exec` row diff --git a/issues/diagnostics/blocked-dump-cannot-fire-on-a-total-freeze.md b/issues/diagnostics/blocked-dump-cannot-fire-on-a-total-freeze.md index c84dae1b3c5..da08a94831d 100644 --- a/issues/diagnostics/blocked-dump-cannot-fire-on-a-total-freeze.md +++ b/issues/diagnostics/blocked-dump-cannot-fire-on-a-total-freeze.md @@ -126,7 +126,7 @@ for six consecutive green hosted runs and has **three**: nightly `ci` runs screen_blocked_dump (4s)`. The loaded-dev-host half asked for three loaded observations and has them — green 3 of 3 beside a full `cargo test` fast tier in one worktree, twelve guests up in each, painting `== VERDICT:` every time — -so its redlist row is retired and this shape's remaining count is the CI one. +so this shape's remaining count is the CI one. Three more green nightlies close it; one red under this name restarts it. The recorded repaint mechanism does not cover this shape — that one names a diff --git a/issues/diagnostics/no-cyclictest.md b/issues/diagnostics/no-cyclictest.md deleted file mode 100644 index cf63c8cda6d..00000000000 --- a/issues/diagnostics/no-cyclictest.md +++ /dev/null @@ -1,34 +0,0 @@ ---- -status: open -kind: track -opened: 2026-08-08 ---- - -# There is no cyclictest, so nobody can ask this machine what its wake latency is - -`git grep -in cyclictest` over the tree finds it only in this file. Until a -cyclictest-equivalent exists, **no honest claim can be made about this machine's -wake latency in either direction** — which makes it the instrument that turns -the first metal boot into a measurement rather than an impression. - -**What to build.** A userland program that enters the RT band, arms an absolute -timer, sleeps, and histograms `actual − programmed` at 1 µs resolution over -enough samples to have percentiles rather than a maximum. - -**What it needs, and it is not blocked on any of it.** `SYS_RT_ENTER` (112) is -the privilege: a `SysCap` carrying `Rights::RT`, granted by a `system.toml` row -(`[programs.soundd] syscap = ["rt"]` is the shape), gated at the dispatch site -and not in the scheduler. The gate this file was opened against is gone — -`SYS_SET_RT_PRIORITY` (96) demanded a `VirtioSound` or `HdaAudio` device claim, -so a latency tool could only reach the band by taking the sound card away from -soundd and measuring a different machine. - -**What exists is not a substitute, and each instrument fails differently.** -soundd's `max_wake_lat_ns` (`toyos-mixer/src/stats.rs`, printed by -`userland/soundd/src/mix.rs`) is a **max over a ~2 s window, not a distribution** — no -percentiles, no sample count; it measures against a DLL's *prediction of a DMA -completion* rather than against a programmed timer, so it folds in the device -model; and it needs soundd plus a sound card to exist at all, which is exactly -what the T14 has not got. `toyos-sched`'s invariant I4 -(`toyos-sched/sim/src/invariants.rs`) bounds the same quantity but runs in the -simulator, so it can never see TCG distortion, real IPI delivery, or metal. diff --git a/issues/diagnostics/nothing-in-the-machine-can-read-the-trace-ring.md b/issues/diagnostics/nothing-in-the-machine-can-read-the-trace-ring.md index 39e1cf7c6c8..b60c3078669 100644 --- a/issues/diagnostics/nothing-in-the-machine-can-read-the-trace-ring.md +++ b/issues/diagnostics/nothing-in-the-machine-can-read-the-trace-ring.md @@ -6,40 +6,46 @@ opened: 2026-07-30 # Nothing inside the machine can read the trace ring, and nothing samples RIP -Rewritten 2026-08-24 from three sentences on unbuilt profiling layers 2 and 3 -that pointed at a "diagnostics roadmap" in CLAUDE.md which no longer exists, -in a layer numbering unreadable without that document. -Read against the tree instead, one of the two things it said was missing is -built and the other's stated blocker is gone. +`kernel/src/trace.rs` is a per-CPU ring of `RING_CAPACITY` 24-byte `repr(C)` +records, enabled from `main.rs` at boot and written by the timer, the +scheduler, `irq_ring` and every `toyos_sched::hw::TraceEvent` the core emits. +Its one reader is LLDB: `p &TRACE_RINGS` and then `memory read`, which is why +the discriminants are fixed by hand and held there by const assertions. A +booted machine can ask itself nothing: no syscall, no tool, no gate. -**What answers "where did this process's time go".** `SYS_PROCESS_STATS` (94) -takes a `Process` handle and fills `ProcessStats`: wall and CPU, syscall count -and total, demand and zero faults with their time, read ops and bytes, blocked -time split five ways (io, futex, pipe, ipc, other), runqueue wait, peak memory, -allocation count. `/system/bin/stats ` spawns, waits and prints it. Per-syscall -counts exist too — `ProcessData::syscall_counts`, 128 bins — and nothing reads -them out. What is still owed there is who may ask: -`issues/diagnostics/the-kernel-keeps-nothing-it-enumerates.md`, a policy gap now rather -than an ABI one, because no diagnostic tool holds a handle to a daemon. +**Ruled** (owner, 2026-10-03): the ring is finished in three steps, and the +work then stops and is judged before anything more is built on it. -**Event tracing is built, and only a debugger can read it.** -`kernel/src/trace.rs` is a per-CPU ring of `RING_CAPACITY` 24-byte `repr(C)` -records, enabled from `main.rs` at boot, written by the APIC timer, the IDT, the -scheduler, `irq_ring` and every `toyos_sched::hw::TraceEvent` the core emits — -pick, idle, preempt, block, wake, timer arm/stop/fire, IRQ entry, and IRQ drain -carrying its measured service latency. Its one reader is LLDB: `p &TRACE_RINGS` -and then `memory read`, which is why the discriminants are fixed by hand and -held there by const assertions. So a wedged kernel under a debugger can be asked -what the schedule was, and a booted machine can ask itself nothing at all — no -syscall, no tool, no gate. **That is what is to be built**, and the ring is the -expensive half of it, already paid for. +1. **The ring is readable**: the log's slot protocol, raw counter stamps, + `SYS_TRACE_READ` in the shape of the log's read, the `trace` right it + needs, and a decoder crate. **Exit**: host tests on the decoder; a guest + test in which a spawned child's wake precedes its pick, a second cursor is + undisturbed, loss is counted after a flood, and a process without the right + is ended; and on the T14, the ticks a record costs, measured by a boot that + writes a million records. +2. **Timer and thread lateness**, computed by a reader from the arm, fire and + pick events, with no tracer in the kernel. **Exit**, on the T14: the 50 ms + for which `hold_once` keeps both windows open once a boot + (`kernel/src/windows.rs`) reads back as at least that much lateness. +3. **Windows over a threshold, and shootdown records.** **Exit**: `SYS_DEBUG`'s + shootdown-acknowledgement delay names the delayed CPU and its time in the + trace. A window's record names its opener by address, which + `issues/kernel/the-cpu-that-spawns-a-toybox-applet-reads-1-4-ms-of-interrupts-and-preemption-off-on-the-t14.md` + and + `issues/kernel/inits-claim-of-a-pci-function-the-t14-lacks-holds-interrupts-off-for-3-8-ms.md` + need. + +Behind that judgement, and not before it: interrupt enter and exit records +with a noise reader, the profiling sample, and the censuses moved onto the +ring. -**RIP sampling is not built, and the blocker `trace.rs` names for it is gone.** -That header says layer 3 "builds on this ring once in-kernel call-stack -unwinding is available". It is available: `arch/idt/exceptions.rs` walks the -`rbp` chain for kernel and for user frames, with a fault-tolerant variant for -the double-fault path, and `symbols.rs` resolves each frame against the loaded -symbol table with the elision `toyos-elide` decides. +The lines he drew: -**Residual, one line**: `kernel/src/trace.rs:13` still cites "the diagnostics -roadmap (see CLAUDE.md)", which points at a deleted document. +- The ring stays always written in the shipping kernel. +- Reading it needs a new `trace` right. +- The reader tool is not in the shipped image. +- The LLDB reading path and its pinned numbers go. +- A T14 judge may read a binary trace file, the decoder printing the lines a + pull request quotes. +- The profiling sample shares the ring. +- 2 MiB of rings at eight CPUs is accepted. diff --git a/issues/filesystem/a-package-is-a-directory-under-apps-and-the-installer-is-a-program.md b/issues/filesystem/a-package-is-a-directory-under-apps-and-the-installer-is-a-program.md index 4ed14ef83d9..52abd05b77e 100644 --- a/issues/filesystem/a-package-is-a-directory-under-apps-and-the-installer-is-a-program.md +++ b/issues/filesystem/a-package-is-a-directory-under-apps-and-the-installer-is-a-program.md @@ -32,6 +32,21 @@ the binary in a ToyOS guest is this track's harness's job, not gbae's. desktop's launcher lists `/apps` and reads each manifest; init grants an app its namespace from that row the same way `system.toml` grants a system program's. Removal is deleting the directory. +- **A package is a program, whole** (owner, 2026-10-02): self-contained and + fixed at build time. No library packages, no dependency solver, no install + scripts; a fix in what a package carries is a rebuild of the package. +- Not his ruling, a refinement put to him with it that he did not object to: a + library can be a build input named exactly by hash, never an installed + thing; shared objects stay, only library packages go. +- **`/apps/` is immutable by stage-then-commit** (owner, 2026-10-02). + `pkg` unpacks into a private staging directory and commits it in one step, + one operation the file server gains; nothing writes a committed package. +- **A standalone C program's ToyOS changes are patch files in a recipe** + (owner, 2026-10-02), against an upstream archive pinned by hash; a delta too + large to read as patches moves to a fork repository pinned by commit. Forks + stay for crates, `rust/` and LLVM. Recipes and patches live in this + repository, where `NOTICE` and the licence gate carry them, not in a ports + repository. - **The installer is an ordinary program**, `/system/bin/pkg`, with no authority a shell does not have: it fetches, verifies, unpacks and writes under `/apps` because `/apps` is writable to it, and it asks the user @@ -57,17 +72,19 @@ the binary in a ToyOS guest is this track's harness's job, not gbae's. The storage track's users and mount-protocol stages do not block this one. -1. Landed: `pkg install `, `pkg remove`, `pkg list`, judged by - `pkg_install_gbae` against gbae's release archive committed under - `tests/fixtures`. **What a package holds is the image's `[apps]` row, never +1. Landed: `pkg install `, `pkg remove`, `pkg list`. Its judge, + `pkg_install_gbae`, is a metal row + `issues/build/the-guest-suite-runs-only-what-no-cheaper-tier-reaches.md` + owes. **What a package holds is the image's `[apps]` row, never its own `manifest.toml`** — that file names which of the directory's binaries a launch starts and nothing else, so a writable `/apps` cannot name - a device or a right. + a device or a right. The stage-then-commit above amends it: `pkg` writes + `/apps//` in place today, its `manifest.toml` last + (`userland/pkg/src/main.rs`). 2. The HTTPS fetch: TLS client under `pkg`, the GitHub redirect, the sums file from the same release. This is the internet-client track's last stage (`issues/design-debt/the-internet-clients-work-unchanged.md`). Judged in QEMU against a server the harness - runs on the host in Rust, serving the bytes already committed under - `tests/fixtures`; then once against GitHub itself, by hand, with the owner + runs on the host in Rust; then once against GitHub itself, by hand, with the owner watching. No registered test fetches anything. 3. Updates: `pkg install` of a newer version replaces the directory whole after the new archive verified; the old one is gone only after the new @@ -76,8 +93,9 @@ The storage track's users and mount-protocol stages do not block this one. project. 5. The users track's per-user `/home` (`issues/filesystem/a-user-is-a-home-tree-and-a-login-row.md`) decides - where a package's own data goes; until then a package writes under - `/apps//` only. + where a package's own data goes. Until then nothing says where: a + committed `/apps/` is written by nothing, and that directory is where + a package wrote before the stage-then-commit ruling. 6. **An app's rights are its request ∩ the user's grant ∩ the image's ceiling** (owner ruling, 2026-09-24; the ceiling's shape, 2026-09-26). The package's manifest *requests* rights; the user *grants* them per user @@ -103,3 +121,10 @@ The storage track's users and mount-protocol stages do not block this one. `assets/soundfont.sf2`, which doom alone opens, leave the image as one installable package. Until they do, `src/licence.rs` names each as an exception pending this stage. + +## Open with the owner + +- The libraries the machine provides: GPU, video. +- A C program's optional requests. +- ABI stability. +- Scale. diff --git a/issues/hardware/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md b/issues/hardware/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md index a5c00238902..420da8563d0 100644 --- a/issues/hardware/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md +++ b/issues/hardware/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md @@ -63,14 +63,11 @@ not what these boots died in — what is left is what happens when the job list ran. **The VFS lock is not in the window.** Run 20 also shows `vfs::lock()` held -across a 32 s stick write with other CPUs at 200M spins -(`issues/kernel/the-vfs-lock-is-held-across-a-usb-write-for-thirty-seconds.md`), -which is within 2x of `Lock::lock`'s deadlock panic, and a spawn does take that -lock. But it takes it at `loader/mod.rs:370` and in `load_needed_libs`, both +across a 32 s stick write with other CPUs at 200M spins, and a spawn does take +that lock. But it takes it at `loader/mod.rs:370` and in `load_needed_libs`, both *before* the `ELF: … relocations indexed` and `spawn: TLS … modules` records that every hung boot wrote; everything after them reads the already-open backing -and takes no VFS lock. So it is a real hazard one step from a panic and it is -not this. +and takes no VFS lock. So it is not this. ## What now exists to answer it diff --git a/issues/hardware/eleven-names-red-on-ci.md b/issues/hardware/eleven-names-red-on-ci.md deleted file mode 100644 index b3ee56ca5e8..00000000000 --- a/issues/hardware/eleven-names-red-on-ci.md +++ /dev/null @@ -1,61 +0,0 @@ ---- -status: open -kind: tooling -opened: 2026-08-08 ---- - -# Eleven names are red on CI, at a rate that is now measured - -Supersedes *a runner reds a rotating handful every run, and the rate is -unmeasured*, whose whole ask was this run. `probe-rate.yml`, run `31258202923`, -tree `f8f73e1`: **five reps of the exact twelve-shard configuration `ci.yml` -runs** — same image, same accelerator, same `--jobs 1` — sixty jobs, **all sixty -finished**, 292 tests each, 1460 outcomes. **281 of the 292 names were green in -all five.** - -| test | red | shard | `Sched` | what it says | -|---|---|---|---|---| -| `std_unwind` | **5/5** | 10 | shared block | `exit code Some(-1)` — the #MF a waiting FP save raised from a pending x87 exception, whose fix landed after this probe and has not been re-measured on CI | -| `std_unwind_so` | **5/5** | 10 | shared block | the same | -| `metal_sim_null_audio` | **5/5** | 11 | Serial | soundd did not present a null sink on a device-less machine — **closed**: it always did, and the test read the line through a span of host wall clock | -| `hda_tone` | **4/5** | 4 | Serial | 1 mid-tone silence in the capture (`issues/audio/`) | -| `late_storage_connect` | 2/5 | 7 | Serial | the boot scan bound a disk, so the port was not held empty | -| `hda_two_live_refused` | 2/5 | 2 | Parallel | `"presenting a null sink" never reached the boot console` — **closed** with it | -| `blocked_dump` | 2/5 | 3 | Parallel | two *different* reasons — the census half, and /system/bin/terminal racing the compositor | -| `dump_nmi_probe` | 1/5 | 2 | Serial | the rip resolved to `u128_div_rem`, not to the spin | -| `kernel_heartbeat` | 1/5 | 5 | Serial | 2 of 12 heartbeats dropped a healthy CPU from the mask | -| `usb_disk_index_stable` | 1/5 | 2 | Parallel | nothing enumerated on the first controller | - -A twelfth name has been seen since and is not in the table because the probe did -not see it: `xhci_slow_connect`, red alone in run `31261669826`. Its margin was -*inside the guest's boot* — the controller started at 0.296–0.311 s against a -300 ms held-empty window — which is why running it alone moved it by -milliseconds and not by a verdict. Re-measured 2026-08-24 on the dev host, that -margin is gone: four boots put the controller at 0.109, 0.117, 0.122 and -0.227 s, i.e. 73–191 ms of clearance, on hosts independently measured at -2.31x–4.45x width. - -A thirteenth, with one sample each way on **one tree**: `desktop_audio_client` -stalled wide *and* alone in run `31264914759` and passed in run `31266194663`, -same commit, half an hour apart. It is 0 of 5 in the table's own probe, so this -is a rate and not a reproduction — but the capture is worth the note, because it -is #172's signature away from the T14: two clients connect, both tones say -`done`, and only one `client N removed` ever follows. The wait it blew is -`both clients to leave the mixer`. - -**The top five reproduce, so they are defects and not a rate.** The bottom six -fire one or two runs in five, which is 20–40% and is not "noise" either: the bar -this was measured against tolerates one in fifty *with the failure named*, and -none of these six has been looked at. - -**`metal_sim_null_audio` and `hda_two_live_refused` are the first two off this -table**, closed when soundd stopped racing to present its null sink. -**Re-taken 2026-09-04 against the last three nightly `ci` runs on `main`** -(`33485669019`, `33603832656`, `33728852421`): of the names above, only -`usb_disk_index_stable` reds in any of them. - -**Six of the eleven are `Sched::Serial`, and until 2026-08-08 the harness re-ran -none of them**: the retry loop was written for the parallel phase and branched on -the *run's* width. Half the list had no second sample at all, which is most of -why the earlier lists looked like they rotated. - diff --git a/issues/hardware/the-t14-answers-only-through-a-usb-stick.md b/issues/hardware/the-t14-answers-only-through-a-usb-stick.md index 58099445fbc..3bd30760768 100644 --- a/issues/hardware/the-t14-answers-only-through-a-usb-stick.md +++ b/issues/hardware/the-t14-answers-only-through-a-usb-stick.md @@ -6,19 +6,18 @@ opened: 2026-09-07 # The T14 answers only through a USB stick -Every result from the bench rides a stick and a reboot into Ubuntu. The laptop -is on a cable on the same LAN as the development Mac and its NIC is the onboard -Intel I219 at `00:1f.6`, `8086:15fc`, which the kernel enumerates and nothing -claims. The track is to make that cable the answer path. +Every boot of the bench is flashed to a stick under Ubuntu and judged by what +is read off it after a reboot into Ubuntu. The laptop is on a cable on the same +LAN as the development Mac and its NIC is the onboard Intel I219 at `00:1f.6`, +`8086:15fc`, which netd claims on the two LAN boots (`tests/lanleasecase`, +`tests/lantalkcase`); on the second the Mac reads the boot's log, runs a +command and hands the machine back over the cable. The track is to make that +cable the answer path. -The substrate a process needs to drive a PCI function itself is built -(`kernel/src/pcidev/mod.rs`, `userland/netd/src/virtio_net.rs`). What is left is -the I219 driver in netd, with DHCP under the hostname `toyos-t14` and a first -ping and ssh from the Mac; the log a boot serves, read from the Mac while it -is booting; command execution, file -transfer both ways and key auth in sshd, with the harness running userland tests -over ssh through a russh client; and a netboot spike in which the firmware -fetches the loader over HTTP so the stick leaves the boot path. +What is left is the harness running userland tests over ssh through a russh +client, which `issues/hardware/the-t14-reboots-through-ubuntu-for-every-test.md` +stages, and a netboot spike in which the firmware fetches the loader over HTTP +so the stick leaves the boot path. Constraints a reader would otherwise pay to re-derive: @@ -34,8 +33,8 @@ Constraints a reader would otherwise pay to re-derive: - **ssh is the bench's transport and a real feature**: sshd is built on russh and the harness's client is russh too. No host ssh binary, no fork. - **Addressing is DHCP with a hostname**, resolved through the router's DNS. The - T14's MAC is the same under ToyOS and Ubuntu, so the lease is the one `t14` - already resolves to. Wi-Fi is out — the AX210 needs a firmware image. + T14's MAC is the same under ToyOS and Ubuntu. Wi-Fi is out — the AX210 needs + a firmware image. - **QEMU's `virtio-net-pci-non-transitional` on `q35` advertises no PCIe function-level reset** — measured, not assumed: `pcidev`'s refusal on that ground reddened every netd registration at once. So a re-claim is made safe by diff --git a/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md b/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md index 3ebab4d84c2..77aedba3f46 100644 --- a/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md +++ b/issues/hardware/usb-disk-index-stable-nothing-enumerates-on-the-first-controller.md @@ -13,7 +13,6 @@ controller`. Re-taken 2026-09-04 against the last three nightly `ci` runs on `main` (`33485669019`, `33603832656`, `33728852421`): of the eleven names that probe found, only this one still reds. -Split out of `issues/hardware/eleven-names-red-on-ci.md`, which covers eleven -names and has no exit condition for this one in particular. +`usb_disk_index_stable` is deleted; `4505f872d`'s parent restores it. **Exit condition.** The cause of the empty first controller is fixed. Owner: orchestrator. diff --git a/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md b/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md index 55f1128816c..a342ceb7b2d 100644 --- a/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md +++ b/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md @@ -6,14 +6,16 @@ opened: 2026-09-28 # The supervisor is host-tested and owns the machine's stop -Held by the orchestrator. Every stage waits on PR #536 (`wt/toyos-fsd`). +Held by the orchestrator. Stage 1 lands right after #592 and #631, of which +#631 has landed, before the latency work +(`issues/kernel/toyos-beats-linuxs-latency-on-the-t14.md`; owner, 2026-10-03). ## Stages Every guest test this track names is registered. A deleted test covers nothing. -1. **The rename**, one mechanical PR, first after #536. Before stage 1 is +1. **The rename**, one mechanical PR. Before stage 1 is briefed, the exit's search below runs once over the `rust/` fork's delta as well as the superproject, so its hits are known going in. It touches `toyos/src`, `toyos-abi/src`, `userland/libc/src` and the `rust/` fork's delta, and its diff --git a/issues/kernel/deferred-release-outlives-its-syscall.md b/issues/kernel/deferred-release-outlives-its-syscall.md index 943fe2c616b..fd487e787f9 100644 --- a/issues/kernel/deferred-release-outlives-its-syscall.md +++ b/issues/kernel/deferred-release-outlives-its-syscall.md @@ -185,13 +185,14 @@ nothing about a batch another CPU is holding. The two honest shapes are makes "a killed process holds nothing" a fact rather than a race. The second is right and it is **not free-standing work**: it is the object -layer's release protocol, and the track that owns it is -`issues/kernel/every-wait-in-this-kernel-is-a-spin.md` — its one park site, its -cancellable kill and its sleep lock between them decide what a hook released -from this queue is allowed to do. The constraint that track's own reasoning -derived, and which anything touching this queue must not lose: **none of the -three drain sites can park, so no `on_zero_handles` hook may take a sleep lock -at all.** It belongs with that track, not beside it. +layer's release protocol, and no track owns it. What a hook released from this +queue is allowed to do is decided by three things the tree has: the proof a +park needs (`scheduler::Parkable`), a kill answered `Cancelled` at the park +(`kernel/src/watch.rs`) and the sleep lock (`kernel/src/sleeplock.rs`). The +constraint they leave, which `ZeroHandles::on_zero_handles`'s doc holds +(`kernel/src/object/mod.rs`) and which anything touching this queue must not +lose: **none of the three drain sites can park, so no `on_zero_handles` hook +may take a sleep lock at all.** ### What "give the batch an owner" costs, worked out 2026-08-20 @@ -207,8 +208,8 @@ for the design — *"this is the one statement that makes 'a hook cannot run und a lock' structural"*. A `close_all` that hands its objects back for the caller to run re-opens, for that one call site, precisely the guard-outlives-the-drop trap the queue exists to make unwritable. Any owner-shaped fix has to pay for -that property somewhere else rather than spend it, which is why this is the -track's work and not a fix at the site. +that property somewhere else rather than spend it, which is why this is a +redesign of the release protocol and not a fix at the site. `kernel-loom` is not the instrument for it, and that is worth writing down so nobody re-derives it: the models compile the real kernel files with @@ -216,38 +217,54 @@ nobody re-derives it: the models compile the real kernel files with `kobject!` set and every subsystem those hooks reach. A transliteration is what that crate's header exists to refuse. -That is the cost on the `deferred` rows. The `immediate` rows have their own, -and the section "What the sleep lock decided, 2026-08-20" is it. - -## What the sleep lock decided, 2026-08-20 - -That track's lock-conversion pass reached this and could not answer it, and why -is worth recording here, because it changes what "give the batch an owner" has -to cover. - -**The kind that most needs a release site is `File`, and `File` is not on this -queue.** `object/mod.rs` makes it an `immediate` row deliberately — *"A file's -flush and its cache reference ride the last `Arc`"* — so its release is -`OpenFileState::drop` (`kernel/src/object/file.rs:27-39`), which takes -`vfs::lock()` and flushes. Once `vfs::VFS` is a sleep lock that `Drop` has no -legal way to acquire it: a `Drop` impl cannot be handed a `Parkable`, and the -two contexts it actually runs in — `ops::close` inside -`process::with_process_data` (`arch/syscall.rs:1250`) and `ops::close_all` -inside `teardown_resources`'s Phase 2 (`process.rs:1126`–`1149`) — hold a -`Lock` guard, so the baseline assertion refuses one level before -the discipline rule does. `try_lock` is not an answer either: the holder it -would lose to is a thread inside a device round trip, so the failure is routine, -and what it loses is a modified file's write-back. - -So the two constraints meet. **The hook queue may not park, and the row that -would want to is not on the hook queue but in a `Drop` that also may not.** -Moving `File` to `deferred` swaps one illegal site for another. The second shape -above is still right for the `deferred` rows, and by itself it reaches neither -`File` nor `close_all`. - -The track carries this as **wall 4**, with the three shapes the owner has to -choose between. Nothing here should be built before that choice, because all -three of them move this queue. +That is the cost on the `deferred` rows. The one `immediate` row that had a +cost of its own was `File`, and the owner ruled on it. + +## What the owner ruled on `File`'s release, 2026-08-23 + +`File` is not on this queue. `object/mod.rs` makes it an `immediate` row — +*"A file's flush and cache reference ride the last `Arc`"* — so its release is +`OpenFileState::drop` (`kernel/src/object/file.rs`), which runs under +`Lock`. On 2026-08-20 that `Drop` took `vfs::lock()` and flushed +a modified file to the device, and the lock conversions then planned made +`vfs::VFS` a sleep lock, which no `Drop` could take: it cannot be handed a +`Parkable`, and moving `File` to `deferred` swapped one site that may not park +for another. The question went to the owner as "may a `Drop` impl park", in +three shapes: + +1. **The write-back queue first.** A file's dirty pages outlive the handle + that dirtied them until write-back reports complete, and `Drop` releases a + cache reference and nothing else. +2. **Let this one `Drop` park, and say so.** `ops::close` and + `teardown_resources` take the entry out and drop it once the `ProcessData` + guard is gone, and "a handle to a `File` may only be dropped from a context + that may park" becomes a rule nothing enforces. +3. **Give the batch an owner**, the second shape under "What to do", **and + make `File` deferred**, which needs `drain_zero_handles`'s two scheduler + sites to stop running hooks that can park: a redesign of this queue rather + than a use of it. + +The ruling, in the words it was recorded in, where "wall 4" is this question: + +> **Owner ruling on wall 4, 2026-08-23: shape 1 — the write-back queue chunk +> lands first.** `vfs::VFS`'s conversion is sequenced behind the +> write-back-queue chunk […]. That chunk's invariant — a file's dirty pages +> outlive the handle that dirtied them until write-back reports complete — +> leaves `OpenFileState::drop` releasing only a cache reference, which a +> `Drop` may do, so no new rule ("a `Drop` may take a sleep lock") is created +> and `scheduler::Parkable`'s header sentence stays true as written. Shapes 2 +> and 3 are declined precisely because they would create that rule or redesign +> the zero-handle drain; shape 1 is the one that needs neither. + +The queue landed with #257 (`7ab9367b0`) and left the kernel in `80a1f1ceb`, +when nothing the kernel flushed reached a device any more; the conversion it +was sequenced before is planned by nothing. Today `OpenFileState::drop` calls +`file_cache::release` and takes "the file cache's lock and no other", so +`File` needs no release site that parks. + +Open with the owner: he declined shape 3 because it would "redesign the +zero-handle drain", and whether that decline reaches "give the batch an owner" +for the `deferred` rows is not ruled. ## A fourth witness: a device claim, and init waiting it out diff --git a/issues/kernel/every-interrupt-lands-on-the-boot-cpu.md b/issues/kernel/every-interrupt-lands-on-the-boot-cpu.md index 36b3b4c72d6..adccaf9303e 100644 --- a/issues/kernel/every-interrupt-lands-on-the-boot-cpu.md +++ b/issues/kernel/every-interrupt-lands-on-the-boot-cpu.md @@ -33,8 +33,12 @@ the machine's, and every device shares it. and either kept (by keeping that device's delivery pinned) or rebuilt for its new producer story — the `i8042` module header lists its own; the audit finds the rest. This is the dangerous half, and the reason this - track sequences AFTER pipeline 2's lock conversions: the drain and wait - machinery under the ISRs must be settled ground first. + track was sequenced after the kernel's lock conversions, which nothing + plans any more. The reason stands: the drain and wait machinery under the + ISRs must be settled ground first, and what moves that machinery today is + stage 6 of + `issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md`, + where `irq_ring` and `drain_irqs` go. 4. The instrument before the change: measure interrupt distribution and the boot CPU's share under the loaded suites, so the improvement is a number against a number. **Done — see below.** diff --git a/issues/kernel/every-wait-in-this-kernel-is-a-spin.md b/issues/kernel/every-wait-in-this-kernel-is-a-spin.md deleted file mode 100644 index 9be6894a017..00000000000 --- a/issues/kernel/every-wait-in-this-kernel-is-a-spin.md +++ /dev/null @@ -1,544 +0,0 @@ ---- -status: assigned -kind: track -opened: 2026-08-12 ---- - -# Every wait in this kernel is a spin, and a killed task dies by having its stack discarded - -**The heading and the paragraph under it are the state at opening, not the -state of the tree.** What exists now: the completion core, where every wait in -the kernel rechecks one predicate and a waiter lends a watch to the object it -waits on (`aaddf38a^:kernel/src/completion/mod.rs:1-8`, since folded into `kernel/src/watch.rs`); typed durations; and a sleep -lock a contender parks on — written, loom-driven, and still -`#![allow(dead_code)]` "until a kernel static converts to it" -(`kernel/src/sleeplock.rs:1-9`), because none has. What has not moved is the -disk path: BOT runs one command in flight per controller and `with_disk` -(`kernel/src/drivers/xhci/mod.rs:1339`) holds the controller lock for the whole -of it (`kernel/src/drivers/xhci/wait/msc.rs:1-4`), which is why -`kernel/CLAUDE.md:12` still reads "No disk wait in this kernel can park". The -four-site count below has not been re-measured since it was written: it is a -baseline to re-take before the next chunk, not a number to quote. - -The state at opening, kept because every chunk below is written against it: -four wait sites still spin — two in xHCI, one in NVMe, one in virtio — so a CPU -waits for a device. A kill is answered by throwing the task's kernel stack away -at five separate reap-in-place arms, and the parked arm is the one that strands -a held lock. Every duration in the kernel is a bare number. - -**The first five chunks landed 2026-08-19 as #91** (`a5ccf14`): duration kinds, -the completion core, the one park site with the cancellable kill, the sleep -lock, and `usbd`/`iod`. **Owner ruling 2026-08-16: this work lands as two pull -requests, split after the kernel-thread chunk**, so #91 was the first and -everything from xHCI async onwards is the second. The split is at the last -boundary where no lock has converted yet; taking it anywhere later means landing -a tree in which one of the four locks is a sleep lock and the other three are -not, which the baseline assertion refuses by design. - -**The second pull request is #148, and it carries the owed tripwire rather than -the chunk.** `block::OPERATION` is landed — a two-second budget over one whole -block-device operation, minted by `UsbBlockDevice` and honoured between commands -in `XhciController::scsi` — which discharges the open decision recorded below -and is what has to exist before a transfer bound can go. The conversion itself -did not land; what stopped it is three walls, established rather than guessed, -recorded under "What xHCI async costs, measured against the tree" so nobody -derives them a third time. - -**The third pull request is #152, and it carries the third wall's answer.** -Owner ruling 1B is landed as `scheduler::Operation`: `Parkable::of_current` is -deleted, the type has no public constructor, and the two doors that replace it -each refuse the other's context. The first two walls did **not** land, and the -reason is one sentence: **the park cannot exist before all four locks convert, -and no partial conversion is legal.** At `wait_transfer` on the filesystem path -the preempt count is `vfs::VFS` + `fat32_adapter::VOLUMES` + `xhci::XHCI` above -its baseline, so `Parkable`'s assertion fails by construction until the three go -together — which is the same argument §21.1 makes and which the borrow-chain -rewrite (wall 1) and the per-disk claim (wall 2) both exist only to *serve*. -Landing either of them alone is a rewrite with no park behind it and a claim -with nothing to exclude, which is the technical debt this tree deletes. What -#152 adds instead is the design each of them now has, below, and two findings -the previous pass did not have. - -**The fourth pull request is the `Await::Transfer` half of wall 1, alone, and -the reason it is alone is two more walls.** The transfer arm now carries the TRB -address, which closes the recorded defect below without needing a park behind -it. The conversion did not land, and what stopped it this time is not walls 1–3: -**`OpenFileState::drop` takes the VFS**, and **the demand-paging path holds -`Lock` across a device round trip and can be a nested trap**. Both -are recorded as walls 4 and 5 under "What xHCI async costs", each with the -owner's options, because each is a decision at the level of ruling 1B and -neither is derivable from the chunk list. The lock count is corrected there too: -the fourth lock is `process::ProcessData` and `process::PROCESS_TABLE` does not -convert at all. - -**The commitment: one completion primitive, one inbox, one park site, one -recheck predicate, and a kill answered by `Cancelled` at the park rather than by -discarding the stack.** This kernel does not unwind, so a killed task holding a -live kernel stack must be schedulable at every safe point and must die by -returning through that stack. Cancellable waits come before sleep locks, and -both before any lock conversion; the order is forced, not preferred. - -## The chunks, and the invariant each must preserve - -- **Duration kinds.** Every duration carries a kind whose constructor demands - what justifies it, and a number nobody can cite is a tripwire or does not - exist. -- **The completion core behind the existing wait queue.** A poster stores its - record under the subject's own lock before it claims the waiter, so a parker - that publishes its intent and then reads its inbox can never miss a post. -- **The one park site, and the cancellable kill.** Inseparable. A killed task - holding a live kernel stack is schedulable at every safe point and dies by - returning through that stack, so no guard is ever abandoned. -- **The sleep lock.** A sleep-lock holder stays preemptible and raises no - preempt count, so the baseline assertion keeps meaning exactly "a spinlock is - held". -- **xHCI async, and the four lock conversions.** Inseparable. A CPU never waits - for a device: the lock is dropped before the park, and a completion is matched - to its asker by identity, never by arrival order. -- **The idle loop's declared end state.** A CPU halts only when nothing is - runnable and no deadline is armed, and the halt condition is a declared set a - test enumerates. -- **A poll kind for registers with no interrupt behind them.** Such a register - is re-read at a declared cadence inside a declared bound, written once. -- **Blocking syscalls, absolute deadlines, the ring as an inbox.** A deadline is - absolute and total over its whole range, so no value means "no timeout" and no - site can silently turn block-forever into return-immediately. -- **The write-back queue.** A file's dirty pages outlive the handle that dirtied - them until write-back reports complete, and a re-open before that sees the - pages and not the device. -- **The deletion commit and its source gate.** A spin exists only at a named - site with a stated reason, and shrinking that list is the only way to claim a - spin was removed. -- **The interleaved A/B.** The wake number is believable only beside a positive - assertion that the log still got written — a log that stopped improves it - identically. -- **Widening the pass-cost window.** It starts where the pass starts, and it is - turned on only against its own baseline. - -## Decisions already made, so they are not re-argued - -- A watch is a node the waiter lends to the object, and the subject is a - borrowed reference, never an id. **Rejected:** a global registry, a slot - arena, two park channels, multishot polls, - userspace-only blocking wrappers, a sleep lock that spins where it cannot - park, poisoning, and shootdown-as-completion. A freed object cannot be named. -- The park token proves the *context* may park and never encodes which locks are - held. **A `&mut` token is rejected**: it forbids the held-across-a-park shape - the whole refactor exists to create, and two stacked sleep locks need the - borrow to stack. -- A spinlock held across a park stays a *runtime* named panic. The type system - cannot see it, and raising a baseline to clear a red converts a boot failure - into a field investigation. -- Readiness is level or edge and the class belongs to the subject. **The - machine's log is edge by necessity**: no reader cursor exists in the kernel, - and one locked read-modify-write per log line costs **350 ms of boot under - TCG** (497–504 ms became 812–839 ms), which forbids moving the post to the - producer. -- The cancel is one-shot, consumed by the wait that reports it; the sticky kill - bit is what terminates, and a second cancel to one thread panics at the call - site. **Owed before implementation:** a killed thread cannot park at all - today, so teardown needs a non-cancellable park, a scoped clear, or a commit - that distinguishes the two. -- Four locks convert and there is no fifth. Blast radius: 29 filesystem sites, - 68 process-data sites, 259 lock calls over 45 statics. **The token cannot be - threaded through the block-device trait**, which lives in a pure host-tested - crate. - - **`process::PROCESS_TABLE` is not one of the four, and the count that said it - was counted the wrong static.** The list has always named - `process::ProcessData` — the per-process `Lock` — and the "26 on - `PROCESS_TABLE`" beside it is a different lock. Checked 2026-08-20 over all 26 - of its acquisitions: `PROCESS_TABLE` is never held while `vfs::lock()` is - taken, no VFS guard is ever live while it is taken, and it is never held - across a device round trip. `loader::spawn` is the one function that does both - and does them in sequence — `vfs::lock()` at `loader/mod.rs:437` and inside - `load_needed_libs` are statement-scoped and long dropped before - `PROCESS_TABLE.lock()` at `loader/mod.rs:694`. `release_process` and - `kill_process` bracket `teardown_resources` between two table acquisitions and - hold neither across it. So `PROCESS_TABLE` does not have to convert for the - park to be legal. - What *does* have to convert is `Lock`, and wall 5 is why. - - **Six, and the count above is the xHCI chunk's rather than the machine's.** - The USB path is `vfs::VFS` → `fat32_adapter::VOLUMES` → the disk's - `block::Handle` → `xhci::XHCI` → `process::ProcessData`; the NVMe path is a - `page_cache::Cached`'s own lock → the same kind of handle lock → - `NvmeBlockDevice::read_blocks`. A handle lock is held across the whole device - round trip, and the order is stated at `block.rs`, `page_cache.rs` and - `fat32_adapter.rs` alike. So whoever converts either wait inherits the cache - lock, which guards a `HashMap` and a slot vector every btree walk touches, and - the handle lock, which guards the `Box` whose trait provably - cannot carry a park token. Nobody may read "there is no fifth" as a property - of the kernel. -- **A bulk transfer has zero real cancellers once the transfer bound is - deleted** — the reset-recovery path's only trigger *is* that bound. So the - largest open decision is a tripwire on the transfer against a budget at the - filesystem layer, and deleting the bound with nothing in its place makes a - shipped daemon's give-up policy silently unreachable. **Decided and landed in - #148**, and the shape it took corrects the premise in one place: the budget - belongs to the *block-device operation* and not to a filesystem operation, - because that is the layer at which one call is one operation and the layer - above it cannot reach the driver at all (see the third wall below). - `block::OPERATION` bounds the composition `USB_TIMEOUT_NS` cannot see — the - batching, the retries and the recoveries one `read_blocks` is made of — and - the daemon it decides is `/system/bin/logd`, whose `LOG_WRITE_BUDGET` is measured in - userland around a syscall and so is reachable only if the syscall returns. Its - doc named `USB_TIMEOUT_NS` as what made that so, which was true of a dead - device and never of a slow one; both bounds are now named there. **What is - still owed at the conversion**: the park in `wait_transfer` takes its - `Deadline` from `min(the transfer bound, until)`, so the operation's budget - becomes the canceller rather than merely the refuser of the *next* command. -- Owner ruling on order: endowment, then the log, then completions. - -## What xHCI async costs, measured against the tree - -Read off the tree on 2026-08-20 while #148 was being written, so that the next -attempt starts from these rather than from the chunk list's one sentence. None -of the three is an argument against the chunk; each is work the chunk has to -contain. - -- **The borrow chain is the refactor.** "The lock is dropped before the park" - means the `&mut XhciController` the guard hands out may not be live across the - park — and at the moment a transfer is waited for it is live eight frames - below the acquire: `with_disk` → `msc_read` → `with_storage` → - `transfer_blocks` → `scsi` → `bot` → `framed_phase` → `bulk` → - `wait_transfer`. Every one of those takes `&mut self` (or hands one on). - Making the park legal means they stop taking it and take a handle - that can re-derive the controller after a re-acquire, which is a rewrite of - `wait/msc.rs`, `wait/mod.rs` and the control-transfer half of `device.rs` - rather than a conversion of `XHCI`'s declaration. - - **The shape it takes, read off the tree on 2026-08-20.** A session value - naming `(controller index, pool block, disk number)`, with a `with(|ctrl, dev| - …)` that re-acquires and re-derives for one short non-blocking critical - section, and a `wait` that holds none of it. The `MscDevice` copy stays on the - caller's stack for the whole operation rather than being written back per - step, which is what wall 2's claim makes safe. Every device-touching step — - the TRB enqueue, the doorbell, the DMA copies, the recovery commands — goes - inside a `with`; every wait goes outside one. - - **Two things the rewrite gets for free and one it does not.** Free: the - outstanding-operation slot already exists and is already identity-matched and - deadline-carrying — `toyos_xhci::job::Outstanding`, pure and host-tested, - is exactly "what ends the wait, when the wait stops being worth having, and - the answer once one arrives", and the MSC path needs one per pool block beside - the controller-wide one the port machine owns. Also free: something already - drains the event ring while a waiter is parked — `poll_if_pending` runs from - `drain_irqs` at the top of every pass on every CPU, on a `try_lock`, and the - xHCI ISR's `need_resched` is what makes a pass happen. Not free: **`Await::Transfer` - was `{ slot, dci }` and matched by endpoint, not by TRB.** Its own sibling arm's - doc says why that is wrong — "matching on anything coarser hands a command that - ran out its deadline and answered afterwards to whatever asked next" — and the - transfer arm did the coarser thing, with `Stages::DataThenStatus` - standing in for the ambiguity a control transfer's two completions on one - endpoint create. - - **Landed 2026-08-20, and it is the one part of this chunk that did not need a - park behind it.** The arm carries the TRB address, which a Transfer Event - names in its first two dwords exactly as a Command Completion Event names its - Command TRB (§6.4.2.1 with ED clear, which every TRB this driver enqueues - has). `Stages` stopped being a *count* in the same change and became the - second `Await`: a count can only be spent by whatever arrives next, and the - two stages are two TRBs at two addresses. `enqueue_control` answers both - addresses instead of a `bool`, and `wait_transfer` takes an `Await` rather - than `(slot, dci)`. -- **There is no per-disk exclusion to drop the lock under.** - `XhciController::with_storage` takes the `MscDevice` out of the pool block by - `Copy`, works on the copy and writes it back — deliberately, so a command can - borrow the controller and the device's rings at once. Today the `XHCI` guard - is what makes that safe. Drop it mid-transfer and a second CPU entering - `with_disk` for the same index reads the *stale* copy and enqueues its own TRBs - on the same endpoint ring. A per-block claim is new design the chunk owes, and - it is not implied by "convert the lock". - - **And the claim has a second job the first pass did not name: it has to stop - the teardown.** Today `XHCI` held for the whole operation is also what stops - `teardown_port` → `release_blocks` reclaiming a pool block, and - `slot_gone`/`let_go` clearing `msc[at].disk`, underneath a transfer. Drop the - lock and an unplug on another CPU can do both while the session's stack copy - of the `MscDevice` still names the block's rings and its DMA window. So a - claim that only excludes a second *reader* is not enough: the unplug path has - to see the claim and either defer or mark, and the session has to notice on - its next re-derivation. Pulling the boot stick mid-write is the test that - exists for it (`usb_boot_stick_pulled`). -- **Neither the token nor the deadline can be threaded to the leaf, and the - answer has to be one answer.** `BlockAccess::read_at` (`toyos-fat32`, a pure - host-tested crate) is the frame that takes `VOLUMES`, and `BlockDevice` is a - kernel trait but its implementors are reached from a `&mut self` that knows - nothing of the caller. So a leaf acquire cannot receive a `&Parkable` by - argument. It can still *mint* one — after the conversion a `SleepGuard` raises - no preempt count, so `Parkable::of_current` at that depth asserts exactly the - right thing and succeeds — but `scheduler::Parkable`'s own header claims the - opposite ("a function with no `Parkable` in scope cannot park… none of them - can make one"), which is a discipline rule and not a compile-time one, since - `of_current` is `pub`. **The owner's call, and it is one decision covering - both values**: either the leaves mint and that header is corrected to say what - it really buys, or the leaves recover both the token and the operation's - deadline from the running task, which needs a word on `TaskHandle` and makes - the ambient recovery explicit. #148 sidestepped it by minting the deadline one - frame *above* the trait, in `UsbBlockDevice`, which works for a value and - cannot work for the park token. - - **Owner ruling 1B, 2026-08-20: the leaves receive, and a leaf never mints.** - `Parkable::of_current` at a leaf is forbidden, and `scheduler::Parkable`'s - header claim becomes enforced truth rather than discipline: the frame that - owns the operation establishes parkability once, the depth reads the token and - the deadline off the running task, and a depth that asks without an - establishment above it is a loud refusal. **Landed in #152 as - `scheduler::Operation`, and the wall is closed.** `of_current` is deleted and - `Parkable` has no public constructor, so the forbidden call does not compile; - `Parkable::at_entry` refuses inside an establishment and `Operation::parkable` - refuses outside one. Three things the implementation settled that the ruling - left open: - - * **The word has two homes.** A task's is on its `TaskHandle`, so an operation - survives the migration a sleep-lock holder can take mid-transfer; a context - with no task — boot, and an idle CPU's pass — gets one slot per CPU, because - it has no handle and cannot be moved off its CPU. `sleeplock`'s `NOT_A_TASK` - is the same split. - * **Establishments nest and an inner one may only narrow**, which refusing to - nest would have forbidden. `fat32_adapter::VOLUMES` is taken *above* - `BlockDevice`, so the frame that must establish park permission on the - filesystem path sits above the frame that owns the block-device deadline: - the two are nested by construction. What nesting may not do is widen — an - inner establishment takes the earlier of the two deadlines — so nobody buys - device time by starting a second operation inside the first. - * **The deadline is the live half and the token is not.** `block::OPERATION` - is established by `UsbBlockDevice`'s three trait methods and recovered by - `msc_read`/`msc_write`/`msc_flush`, so `storage_read`, `storage_write` and - `storage_flush` lose an argument; `scsi` keeps its `until` parameter, which - is what leaves it usable by `bring_up` — an enumeration is not a - block-device operation, has no establishment, and passes `Deadline::never` - by name. `Operation::parkable` carries a named `allow(dead_code)` until the - conversion, in the shape `sleeplock.rs` already uses for the same reason. - -**Two more, and they are not xHCI's.** The three above are what the *driver* -costs, and the 2026-08-20 pass found them all payable. What it could not pay is -what the *locks* cost, and neither of the two below is visible from the driver -at all: one is in the object layer's file release and one is in the demand-paging -path. They are numbered on with the others because they are walls in the same -sense — established rather than guessed, and each an owner's decision rather -than an implementation. - -### Wall 4: `OpenFileState::drop` takes the VFS, and no `Drop` may take a sleep lock - -Read off the tree 2026-08-20. `object::file::OpenFileState::drop` -(`kernel/src/object/file.rs:27-39`) takes `vfs::lock()` and, for a modified -file, calls `Vfs::flush_file` — which for a FAT mount is a device round trip. -Its own doc has said so since it was written: *"**This takes the VFS lock** — -flushing needs it."* It is load-bearing rather than incidental, and the evidence -is that an unrelated file already has to reason about it: `loader::spawn` scopes -its own `vfs::lock()` to a single statement and says why — *"every `return` past -this point drops `pending`, and releasing a file object takes the VFS lock -(`object::file::OpenFileState::drop`)"* (`loader/mod.rs:442-445`). Three facts -make it fatal to `vfs::VFS`'s conversion, and none of them is one of the first -three walls. - -* **A `Drop` impl cannot receive a `Parkable`**, and the two contexts this drop - actually runs in cannot let it make one either. `ops::close` runs inside - `process::with_process_data` (`arch/syscall.rs:1250`) and `ops::close_all` - runs inside `teardown_resources`'s Phase 2 (`process.rs:1126` guard live - through `process.rs:1149`), so a `Lock` guard is live at both and - `Parkable::mint`'s baseline assertion refuses one level before the discipline - rule does. The third route, `sys_dup2`'s `drop(displaced)` - (`arch/syscall.rs:2645`), is deliberately outside the closure and is the only - one that is at baseline. -* **`try_lock` is not the answer, and the flush is why.** The holder this would - lose the race to is a thread inside a device round trip, so the failure is - routine rather than rare, and what is lost is a modified file's write-back - with a log line — the silent-degradation class the root `CLAUDE.md` forbids. -* **The queue it could move to is closed to it by this track's own derived - constraint.** None of `drain_zero_handles`'s three sites can park, so no - `on_zero_handles` hook may take a sleep lock; and `File` is an `immediate` row - *precisely so that its flush rides the last `Arc` instead* — - `object/mod.rs:296-298` says exactly that. Making `File` `deferred` therefore - moves the flush from a place it may not run to another place it may not run. - -**The decision is the owner's, and it is "may a `Drop` impl park".** -`scheduler::Parkable`'s header states the opposite as a property of the tree — -*"that is why `sched::dump`, `panic_console`, every ISR and every `Drop` impl -are structurally unable to block"* — so the answer changes what that sentence -means everywhere, not only here. Three shapes: - -1. **The write-back queue chunk first.** It is already on the chunk list above, - and its invariant is the one this needs: a file's dirty pages outlive the - handle that dirtied them until write-back reports complete, and a re-open - before that sees the pages and not the device. `Drop` then releases a cache - reference and nothing else, which is a thing a `Drop` may do. This makes the - conversion's prerequisite an existing chunk rather than new design, and it is - the only shape that needs no new rule. -2. **Let this one `Drop` park, and say so.** `ops::close` takes the entry out of - the table and drops it outside `with_process_data` — which `sys_dup2` already - does and documents — and `teardown_resources` drains the handles, drops the - `ProcessData` guard, and only then drops them. Every `File` handle then drops - at baseline and `Drop` mints `Parkable::at_entry`. What it costs is the - sentence above: "a handle to a `File` may only be dropped from a context that - may park" becomes a rule nothing enforces, on a type `Arc` will happily drop - anywhere. -3. **Give the batch an owner** — `deferred-release-outlives-its-syscall`'s own - second shape — and make `File` deferred. That buys a parkable release site on - the syscall path, and it needs `drain_zero_handles`'s two scheduler sites to stop - running hooks that can park, which is a redesign of that queue rather than a - use of it. - -**Owner ruling on wall 4, 2026-08-23: shape 1 — the write-back queue chunk -lands first.** `vfs::VFS`'s conversion is sequenced behind the write-back-queue -chunk, which is already on the chunk list above. That chunk's invariant — a -file's dirty pages outlive the handle that dirtied them until write-back -reports complete — leaves `OpenFileState::drop` releasing only a cache -reference, which a `Drop` may do, so no new rule ("a `Drop` may take a sleep -lock") is created and `scheduler::Parkable`'s header sentence stays true as -written. Shapes 2 and 3 are declined precisely because they would create that -rule or redesign the zero-handle drain; shape 1 is the one that needs neither. - -**Landed 2026-08-23 (#257): the write-back queue chunk.** `crate::writeback` is the -queue; the last close pins the file (`file_cache`'s `teardown_owed`) and -`writeback::enqueue`s it, and `OpenFileState::drop` now calls only -`file_cache::release_to_writeback` + `enqueue` — no VFS lock, no device, exactly -the cache-reference release wall 4 needs. `iod` (`iod.rs`) and `SYS_SHUTDOWN` -both drain it through `writeback::drain_all`, which pops each entry under the -VFS lock and runs the teardown the `Drop` used to (flush the dirty pages, drop -the file, `close_file`). The per-handle `OpenFileState.modified` flag is gone: -the file owns its dirty state (`file_cache`'s `dirty_meta`), so a reader closing -last still flushes what a writer dirtied, and `fsync` reads that same bit. **No -lock converted** — `vfs::VFS` stays a spinlock, `iod` spins in the driver, and -the drain pops under the VFS lock so `sync_all` cannot commit a device ahead of -a file's flush. This is the state that unblocks `vfs::VFS`'s conversion (the -next chunk): a `Drop` reaching this release site now touches neither a sleep -lock nor a device. - -### Wall 5: demand paging holds `ProcessData` across the device, and a nested trap is a level above the baseline - -Read off the tree 2026-08-20. `process::handle_page_fault` takes -`data_arc.lock()` at `process.rs:1621` — a `sync::Lock` — and holds -it across `backing.read_page(...)` at `process.rs:1658`, which for a `/boot` or -`/log` mapping is `FatBacking::read_page` (`fat32_adapter.rs:428`), acquires -`VOLUMES`, and goes to the device. Two things follow and they are separable. - -* **This is what makes `process::ProcessData` the fourth lock**, and it is a - much narrower claim than "68 process-data sites": the fault path is the one - place a `ProcessData` guard is live across the acquire that becomes the park. - The guard is taken for `elf.reloc_index` and `elf.elf_base` and then held for - the whole 2 MiB fill, beside an address-space guard the same function already - snapshots and drops before the reads. So there is a second answer that is not - a conversion: snapshot those two the same way and drop `ProcessData` before - the fill. Whether the lock converts or the frame is restructured is a - judgement about the other 67 sites, not about this one. -* **What no restructuring answers is the baseline.** `blocking_baseline()` - answers `BASELINE_TRAP` = 1 for a user thread, and a fault taken *inside* a - syscall runs at two. `scheduler.rs`'s own `BASELINE_TRAP` doc names this exact - path and calls it routine rather than hypothetical: *"The first demand-paging - path that parks instead of spinning … breaks that and trips this on a nested - trap holding no lock at all."* A park anywhere under `handle_page_fault` - therefore panics on a stack holding no lock, and it panics by construction - rather than under contention — so it is not something a green suite would - hide, and not something a larger constant fixes. - - The shape of an answer is that the entitlement stops being a constant: each - entry knows the level it raised, and `blocking_baseline` reads that rather - than assuming one. That turns the tripwire into what it already claims to be — - *"a spinlock is held"* — for nested traps as well as unnested ones. It is a - change to the assertion the whole design rests on, which is why it is recorded - here rather than taken. - -**Owner ruling on wall 5, 2026-08-23: both.** The demand-paging fault path is -restructured so it does not hold `Lock` across `read_page` and the -device round trip: what the fill needs — the two `elf` fields — is snapshotted -and the guard dropped before the fill, the way the function already snapshots -and drops its address-space guard. **And** `blocking_baseline` stops being a -constant: each trap entry records the level it raised and `blocking_baseline` -reads that, so the tripwire means *"a spinlock is held"* for a nested trap as -well as an unnested one. - -## What the conversion still owes, in the order it has to be done - -Read off the tree on 2026-08-20 with wall 3 landed and the TRB half of wall 1 -landed. Items 1–4 are one pull request, because the baseline assertion refuses -any split; **items 0a and 0b are owner decisions that come before all of it**, -and neither is derivable from the rest of this file. - -0a. **Wall 4's decision: may a `Drop` impl park?** Until it is answered - `vfs::VFS` cannot convert, so nothing below can either. **Ruled 2026-08-23: - shape 1** — the write-back-queue chunk, already on the chunk list, is the - prerequisite, and `vfs::VFS`'s conversion is sequenced behind it; the ruling - is recorded at wall 4. **Landed 2026-08-23 (#257)** (`crate::writeback`): - `OpenFileState::drop` now releases only a cache reference, so this - prerequisite is met and `vfs::VFS` is free to convert. See the "Landed" - note at wall 4. -0b. **Wall 5's decision: does the baseline stop being a constant, or does the - fault path stop holding `ProcessData` across the device — or both?** Until it - is answered `fat32_adapter::VOLUMES` cannot be parked under from the - demand-paging path. **Ruled 2026-08-23: both** — the fault path drops the - `ProcessData` guard before the fill, and the baseline is read off what each - trap entry recorded; the ruling is recorded at wall 5. Once done, this makes - `fat32_adapter::VOLUMES` and a `block::Handle`'s device lock — the NVMe pair, - behind the same wall by the same route — parkable under from the - demand-paging path. -1. **Wall 2's claim on the pool block**, including the teardown half above. -2. **Wall 1's session.** Its other half — `Await::Transfer` carrying the TRB - address — is landed; what is left is the borrow chain. -3. **The locks together** — `vfs::VFS`, `fat32_adapter::VOLUMES`, `xhci::XHCI`, - `process::ProcessData`. Counted 2026-08-20: 28 `vfs::lock()` sites (13 in - `arch/syscall.rs`, 9 in `main.rs` which are boot and become `try_lock`, 3 in - `loader`, 3 in `object/`), 8 `VOLUMES` acquisitions in `fat32_adapter.rs` (3 - in `BlockAccess for FatVolume`, 1 in `FileBacking for FatBacking`, 4 in - `mount`, which is boot), 4 on `XHCI`. `PROCESS_TABLE` is **not** among them — - see the correction under "Four locks convert and there is no fifth". - Everything but the two trait impls can take a `&Parkable` by argument; those - are the ambient recovery wall 3 exists for, and `UsbBlockDevice`'s - `block::begin_operation` is already the establishment `XHCI`'s acquire - recovers from, so only `VOLUMES` needs one placed above it. -4. **The park itself**, and only then the track's owed `min(the transfer bound, - until)`. **Not before**, and the reason is in `XhciController::scsi`'s own - doc: ending a *spin* at the caller's deadline abandons a transfer the device - is still going to answer, and the recovery then reads the wreckage as a - device that is not recovering — a slow disk marked permanently offline for - having been slow. At a park, waking on the operation's deadline is not - abandoning anything: the waiter wakes and *then* decides. So the owed item is - a property of the park and lands with it. - -Both findings from the 2026-08-20 pass are closed. Kernel clippy runs with the -actuator features on. The other — the two page-cache locks the count above did -not name — is now the paragraph under "Four locks convert and there is no -fifth", and the deadline half of it landed the same day: -`nvme::Queue::wait_completion` spun with nothing bounding it at all, and now -takes `drivers/nvme.rs`'s `COMMAND` inside a command and `block::OPERATION` -between two, exactly as the USB path does. The **conversion** of the cache lock -and the device handle's did not land and is not in the list above; it is the -NVMe chunk's, and it is owed on top of the four. - -**And it is behind wall 5 by the same route the USB path is**: -`file_backing::NvmeBacking::read_page` reaches `Cached::raw_read` -(`file_backing.rs:123`), which takes the handle's device lock — and `read_page` -is called from `process::handle_page_fault` with `Lock` held. So -the pair is not a second, independent piece of work that could go first: -whatever answers wall 5 for `VOLUMES` answers it for the device lock, and -whatever does not, does not. - -`drain_zero_handles`'s derived constraint — none of its three drain sites can -park, so no `on_zero_handles` hook may take a sleep lock — is **untouched by -#148 and by #152**, because neither converts a lock. Its ground was checked -rather than assumed: the hook that would want the VFS is a file's, and `File` is -an `immediate` row whose hook is empty; the flush rides `OpenFileState::drop` on -the last `Arc` instead. - -**And that last clause is wall 4, which the sentence it used to end — "the -constraint still costs nothing to keep" — got exactly backwards.** The flush -riding `OpenFileState::drop` is not a way of avoiding the constraint; it is the -same problem one row over. A `Drop` cannot take a sleep lock either, so the -constraint and the `immediate` row together leave the file flush with *no* legal -site once `vfs::VFS` converts. What it costs is the conversion, and -`deferred-release-outlives-its-syscall` — which points here for its own answer — -is the same decision seen from the object layer. - -Measured, and worth carrying: a 2 ms-per-transfer stick takes the worst wake -from **7,117 µs to 165,948 µs** at smp=1, and 6,174 µs to 250,912 µs under load -at smp=8. The audio period is 2.902 ms against a 23.219 ms pipeline. The -scheduler migration cost about **70 defects** in code whose own suites were -green, which is the calibration for how this one is landed. - -Six entries under `issues/design-debt/` recorded that the deleted document's own -citations had rotted — five against the tree, one against a log plan deleted -before it. All six closed with it. - -`usb_boot_stick_pulled` is deleted. diff --git a/issues/kernel/placement-is-blind-to-caches-and-topology.md b/issues/kernel/placement-is-blind-to-caches-and-topology.md index 1d28f1b9105..ac3b962d19d 100644 --- a/issues/kernel/placement-is-blind-to-caches-and-topology.md +++ b/issues/kernel/placement-is-blind-to-caches-and-topology.md @@ -26,4 +26,4 @@ placement story has three parts none of which we have: Composes with the reservation model (placement decides *where*, reservations decide *how much*) and with the toyos-sched sim, which would need a topology -model to gate any of this. Independent of pipeline 2's remaining chunks. +model to gate any of this. diff --git a/issues/kernel/syscall-preemption-is-incidental.md b/issues/kernel/syscall-preemption-is-incidental.md index e1a7fa60e56..728d8567aac 100644 --- a/issues/kernel/syscall-preemption-is-incidental.md +++ b/issues/kernel/syscall-preemption-is-incidental.md @@ -37,3 +37,15 @@ switch (the scheduler's own baselines needed it): before that the count drifted, so a lock drop inside a syscall reached zero at random and preempted at random. The behaviour is now deterministic, and deterministically weaker than the model assumes. + +Owner: `issues/kernel/toyos-beats-linuxs-latency-on-the-t14.md`, whose second +step is syscalls running with interrupts on (owner, 2026-10-03). + +**Exit**, on the T14: the longest interrupts-off window the `mask_windows` row +reads under its load (`herd irqs_off_ns=`, printed by `windows_on_metal` in +`tests/toyos.rs`) is no longer than the longest lateness of the timer's +interrupt in Linux's reading of this machine, which that track keeps: a masked +window makes a timer's interrupt late by at most its own length. No figure is +set here. Which of Linux's figures is the bar is that track's open question, +and what the row reads today, and which syscalls those windows are, is in +`issues/hardware/xhci-waits-are-spins.md`, "On the T14". diff --git a/issues/kernel/the-global-pipe-lock-spans-a-user-copy.md b/issues/kernel/the-global-pipe-lock-spans-a-user-copy.md index ae961f8815f..237e07851bc 100644 --- a/issues/kernel/the-global-pipe-lock-spans-a-user-copy.md +++ b/issues/kernel/the-global-pipe-lock-spans-a-user-copy.md @@ -26,7 +26,7 @@ Everything pipe-shaped in the kernel goes through the same static: `create` (`pi ## Impact, stated at its real size -This is **not** the device-I/O class. The copy is memory-speed and takes no device round trip, so it does not approach `sync::Lock`'s 500,000,000-spin `DEADLOCK` panic (`sync.rs:72-75`) the way a `block::Handle`'s device lock and `vfs::VFS` do — that class is tracked in `issues/kernel/every-wait-in-this-kernel-is-a-spin.md` and is not what this file is about. +This is **not** the device-I/O class. The copy is memory-speed and takes no device round trip, so it does not approach `sync::Lock`'s 500,000,000-spin `DEADLOCK` panic (`sync.rs:72-75`) the way a `block::Handle`'s device lock does — that class is tracked in `issues/hardware/xhci-waits-are-spins.md` and is not what this file is about. What it is: a global lock held with preemption disabled for a stretch that scales with a userland-chosen length up to 2 MiB, on the path that carries every IPC message in the system. The scheduler's own bar for an uninterruptible stretch is `MAX_PASS_NS = 200_000` ns (`toyos-sched/src/cpu.rs:1235`); a 2 MiB copy plus, on first write, a 2 MiB zeroing and a bitmap scan is on the wrong side of that at any plausible memory bandwidth. **The hold time is not measured and that is the first thing owed.** diff --git a/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md b/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md index 7751ab6b52c..303fbbf64df 100644 --- a/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md +++ b/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md @@ -8,7 +8,6 @@ opened: 2026-09-25 Owner ruling, 2026-09-25: the kernel is to be super small, super performant and safe. This track holds that, and supersedes the ordering of -`issues/kernel/every-wait-in-this-kernel-is-a-spin.md`, `issues/kernel/every-driver-is-still-in-the-kernel.md` and `issues/hardware/a-disk-plugged-in-after-boot-is-bound-inside-a-scheduling-pass.md`, which stay as the evidence each stage closes. diff --git a/issues/kernel/the-vfs-lock-is-held-across-a-usb-write-for-thirty-seconds.md b/issues/kernel/the-vfs-lock-is-held-across-a-usb-write-for-thirty-seconds.md deleted file mode 100644 index 959f5de76b1..00000000000 --- a/issues/kernel/the-vfs-lock-is-held-across-a-usb-write-for-thirty-seconds.md +++ /dev/null @@ -1,58 +0,0 @@ ---- -status: open -kind: defect -opened: 2026-09-07 ---- - -# One VFS lock is held across a USB write, and 32 s of it comes within 2x of the deadlock panic - -T14 run 20, `tests/metaldevicecase`, `metal-suite` f46f91eb. A 2.4 MiB write to -the boot stick took 32 s, and for its whole length other CPUs sat on -`vfs::lock()`: - -``` -[08:51:38 44.923 cpu4] LOCK CONTENTION: 50M spins at src/vfs.rs:32:18, ticket=57 now=56 -[08:51:42 48.634 cpu4] LOCK CONTENTION: 100M spins at src/vfs.rs:32:18, ticket=57 now=56 -[08:51:50 56.046 cpu4] LOCK CONTENTION: 200M spins at src/vfs.rs:32:18, ticket=57 now=56 -``` - -`VFS` is one `Lock>` for the whole machine (`kernel/src/vfs.rs:14`), -and `ops::fsync` (`kernel/src/object/ops.rs`) holds it across `flush_file` **and** -`sync_for_path` together — deliberately, so a volume cannot be unmounted between -them — which means across every device write and the SYNCHRONIZE CACHE under -them. `logd` fsyncs after every batch, on the same stick. - -**The number that matters is how close this is to a panic.** From the log's own -timestamps, 50M spins is 3.72 s (15.024 → 18.745 s), so `Lock::lock`'s -`DEADLOCK` threshold of 500M spins (`kernel/src/sync.rs`) is about **37 s of -continuous contention**. This boot reached 200M. A write 1.8x slower — a bigger -file, a slower stick, a device that retries — panics the machine inside a lock -acquisition, and the report goes to a `logd` that is on the far side of the lock -it is panicking about. That is a machine that cannot say why it died. - -**It is not the wedge in -`issues/hardware/a-t14-boot-wedges-after-a-jobs-exit-and-nothing-said-why.md`, -and that was checked rather than assumed.** A spawn does take `vfs::lock()` — -`open_backing` at `kernel/src/loader/mod.rs:370` and twice in `load_needed_libs` -— but all three are *before* the `ELF: … relocations indexed` and -`spawn: TLS … modules` records, which every hung boot wrote. What runs after -them reads the already-open backing directly (`read_file_range`, -`elf::read_backing_into`) and takes no VFS lock at all. So the contention is -upstream of the window the hangs stop in. - -**What rules the device out, and what it costs beside the panic distance.** The -payload is 2,457,600 B — `userland/metalprobe/src/usb.rs`'s `BYTES`, which is -`SLOWEST_KIB_S 512 x JOB_BOUND_MS 60_000 x SHARE_PERCENT 8 / 100 / 1_000 x 1024` -— and `exit: usbwrite pid=7 code=22091910 cpu=32100ms` makes that 108.6 KiB/s, -a fifth of the 512 KiB/s floor the same file calls "an order of magnitude under -any USB 2.0 flash device". The stick is not what is slow: `flush-census: dev=16 -flushes=9 ... p50<=512us p99<=1024us max=630us of 9 flushes`, and -`syscalls: pid=7 total=52 syscall_wall=31957ms` puts the whole cost inside 52 -syscalls. Two rows pay for it: `usbwrite.metaldevicecase.span_us` reads -22,091,910 against its committed ceiling of 4,800,000, and one bound covers the -whole job list — which is why run 20's list never reached its `reboot` job and -was ended by the runner's deadline instead. - -**Exit condition**: the sync half of `fsync` off the global VFS lock, or a bound -on how long that lock may be held that is under `Lock::lock`'s tripwire rather -than within 2x of it. diff --git a/issues/kernel/toyos-beats-linuxs-latency-on-the-t14.md b/issues/kernel/toyos-beats-linuxs-latency-on-the-t14.md new file mode 100644 index 00000000000..4805b66b62e --- /dev/null +++ b/issues/kernel/toyos-beats-linuxs-latency-on-the-t14.md @@ -0,0 +1,45 @@ +--- +status: open +kind: track +opened: 2026-10-03 +--- + +# ToyOS beats Linux's latency on the T14 + +Ruled (owner, 2026-10-03): the bar of the latency work is Linux's reading on +the same T14, and its order is the three steps below, with stage 4 of +`issues/design-debt/toyos-has-its-own-network-stack.md` running beside them. + +1. **The firmware's interrupt**: stage 1 of + `issues/kernel/toyos-runs-the-machine-in-acpi-mode-and-interprets-its-aml.md`. +2. **Syscalls run with interrupts on**: + `issues/kernel/syscall-preemption-is-incidental.md`. +3. **USB and storage leave the kernel**: usbd, step 10 of + `issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md`. + +**Linux's reading, the reference.** Ubuntu's `6.8.0-142-generic` on the same +machine, `rtla timerlat top -q -d 1m --dma-latency 0`, a 1 kHz timer on each +CPU (#649, comment 5961690204; its log whole in #682's comment 5968053374). +It samples: a window under 1 ms is seen only when a tick falls in it. + +- The timer interrupt's longest lateness on a CPU: 3 to 99 µs idle and 16 to + 131 µs with every CPU spawning `/bin/true`, on seven CPUs. cpu4 read 2952 + and 2880 µs, one event a run, its cause not identified; a two-minute idle + run after them read 4 to 53 µs on all eight (#681, comment 5962619169). +- A woken thread's longest lateness: 9 to 571 µs idle and 38 to 503 µs + loaded on those seven; cpu4 2964 and 2917. +- No firmware interrupt on any CPU in 120.378 s, with `SCI_EN` set (#681, + comments 5962619169 and 5962768253). + +Open with the owner: which figure of Linux's is the bar, cpu4's 2.9 ms or +the other CPUs' longest. + +**Exit**: each step's exit is met, in the file the step names, and on the T14 +the longest lateness of a 1 kHz timer's interrupt and of the thread it wakes, +on each CPU, reads under Linux's. Nothing takes that reading of ToyOS today: +`latency_wake` reads one thread's p99 and `mask_windows` each CPU's longest +masked windows. Two files owe it: step 2 of +`issues/diagnostics/nothing-in-the-machine-can-read-the-trace-ring.md`, a +reader that computes timer and thread lateness from the ring, and +`issues/hardware/no-program-measures-toyos-against-linux-on-one-machine.md`, +one program that reads lateness past a 1 ms timer under both systems. diff --git a/issues/kernel/toyos-runs-the-machine-in-acpi-mode-and-interprets-its-aml.md b/issues/kernel/toyos-runs-the-machine-in-acpi-mode-and-interprets-its-aml.md index 75a57722416..3cb371c922e 100644 --- a/issues/kernel/toyos-runs-the-machine-in-acpi-mode-and-interprets-its-aml.md +++ b/issues/kernel/toyos-runs-the-machine-in-acpi-mode-and-interprets-its-aml.md @@ -24,10 +24,26 @@ legacy mode today is lost; not the switch alone with the events masked, and not after the AML work. The AML interpreter follows as later stages: ToyOS writes its own, and the battery comes first (his direction of `0ee814f5a`). +**Ruled** (owner, 2026-10-03), on stage 1: + +- **Each embedded-controller event is taken off the controller.** The log gets + a line the first time a query number appears and a count at intervals, never + a line per event: under Linux the T14's EC raises about 2.6 a second (#682, + comment 5968053374). +- **The server holds the EC's ports unfiltered.** The kernel filters no EC + command; what the holder can do with those ports is recorded as a weakness + in an issue of its own, as #592 records the i8042 holder's reset line. +- **Stage 1 is built on #592's `isa` claim**, as one row the FADT and the ECDT + name, after #592 lands. +- **The attended press waits for the owner.** Everything else in stage 1 is + built and reviewed first; he then presses the button once, briefly, on a + boot held open for it. + **Stage 1: ACPI mode, its SCI served in userland.** The switch to ACPI mode, -and a userland server that claims the SCI and handles the power button, a -fixed event that needs no AML. **Exit**: on the T14, `MSR_SMI_COUNT` stays +and a userland server that claims the SCI, handles the power button, a +fixed event that needs no AML, and takes the EC's events. **Exit**: on the +T14, `MSR_SMI_COUNT` stays flat on every CPU over the interval the firmware issue's exit defines, and a press of the power button stops the machine cleanly, through ToyOS's own power-off path (`SYS_SHUTDOWN`), and the boot's log records the press and that -stop: a T14 row reads both there. +stop, and each EC query number once with its count: a T14 row reads them there.