diff --git a/.claude/agents/implementer.md b/.claude/agents/implementer.md index 32e4e7d8592..98db494669c 100644 --- a/.claude/agents/implementer.md +++ b/.claude/agents/implementer.md @@ -14,6 +14,10 @@ One brief, one worktree, one branch. What the brief does not name you do not tou find off your path is filed in `issues/`, never fixed. If something blocks you, stop and say so in one clause; do not work around it. +Before adding kernel behaviour, ask whether userland can own it. When the clean design changes the +ABI, change the ABI; never pick a lesser design to avoid that. A clean design that reaches past +your fence blocks you. + ## Measure, build, test Where hardware or anything uncertain is involved, take the cheap measurement before you build on a @@ -26,6 +30,9 @@ guess. Then build, then test before anyone reviews: - Long commands run in the background with output to a file under the job scratchpad the brief names. Stay inside one turn while anything runs: sleep at most two minutes, print a line, check again. Ten minutes of silence kills you, and ending a turn to announce a wait strands the work. +- Nothing a pull request's evidence rests on, mutation patches and run logs included, lives only in + a temporary directory: `/tmp` is wiped when the CLI restarts. Post mutation patches to the pull + request as a comment. - A mutation is a measurement only once the mutated tree is shown to build. Apply it as a checked patch, restore it in the same script, and leave the tree clean. - Never a flat wait, in code or in a test: wait on the event, bounded by a timeout that fails @@ -56,9 +63,9 @@ Fork sources live outside this repository: a search for callers must also cover `git commit -F `, never `-m`. No `--amend`, no rebase, no force: merge `origin/main`, never rebase onto it. Never run `git submodule` in a linked worktree: it writes `core.worktree` into the -fork's shared config and breaks git in the primary checkout's `rust/`. Never touch `toyos-abi/src`, `toyos/src` or `userland/libc/src` unless the brief is -an ABI brief. A new dependency is taken where it is the cleanest path: a general, widely used -crate (root `CLAUDE.md`, "Dependencies"), and the pull request says why. +fork's shared config and breaks git in the primary checkout's `rust/`. A new dependency is taken +where it is the cleanest path: a general, widely used crate (root `CLAUDE.md`, "Dependencies"), and +the pull request says why. Push from your branch, never `main`, with `git status --porcelain` empty: `git push -u origin `, and `gh pr create --draft` at the first push. The pull request body is the handoff the reviewer reads, diff --git a/.claude/agents/reviewer.md b/.claude/agents/reviewer.md index 5708adb3aa8..19330b044ab 100644 --- a/.claude/agents/reviewer.md +++ b/.claude/agents/reviewer.md @@ -42,7 +42,8 @@ above; otherwise it is a NOTE. - **Fit.** Does the tree already do this? Is each new thing where it belongs: a pure decision in a pure crate, the user/kernel boundary in `toyos-userbound`, a device claim in a userland server? One declaration read by every reader, refusal by name, authority moved in by the parent. Zero - legacy: no shim, no workaround, no silent default. A new dependency only where it is the + legacy: no shim, no workaround, no silent default. A BLOCKER each: a kernel addition that + userland could own; a design made worse to spare the ABI. A new dependency only where it is the cleanest path, a general and widely used crate the pull request says why it takes; no new fetch. Nothing outside the brief's fence. Assembly, a naked function and a `core::arch` or `std::arch` path live only in an @@ -75,12 +76,9 @@ above; otherwise it is a NOTE. A file added to or deleted from `tests/testcases/tinycc/` moves the count `tests/testcases/LICENSE` states in the same diff, and `46_grep.c` never comes back. Nothing else is tracked under `tests/testcases/` but that `LICENSE` and `system.toml`. -- **What no gate reads.** A BLOCKER each: a diff that declares a retired ABI name or reuses a - retired syscall, `SYS_DEBUG` action or inbox op number (the retired numbers are - `kernel/src/syscall/dispatch.rs`'s `retired_syscalls!` and the "formerly …" and "retired and - unused" entries in `toyos-abi/src/syscall.rs` and `toyos-abi/src/inbox.rs`; the retired names - include `SharedToken` and `services::connect`); a workspace member's `Cargo.toml` declaring `[profile]` or `[patch]`, which - cargo ignores with only a warning; a new package without a `description` saying what it is. +- **What no gate reads.** A BLOCKER each: a workspace member's `Cargo.toml` declaring `[profile]` + or `[patch]`, which cargo ignores with only a warning; a new package without a `description` + saying what it is. A new cargo feature or `cfg` arm of one, and every arm a changed `src/clippy.rs` shape stops building, is shown linted in the pull request body: a `mem::forget` planted in that arm turns `cargo run -- --clippy` red. - **Growth.** Every line is a responsibility, not an asset. State the branch's net lines (`git diff --shortstat origin/main...HEAD`), production and tests apart. Production code that grows diff --git a/CLAUDE.md b/CLAUDE.md index 0b910305b39..3c5f2ec6c0e 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,6 +1,6 @@ # ToyOS -An operating system built from scratch in Rust, held to a production-grade engineering bar — the bar is the changes, not yet the product. Modern x86-64 hardware (2020+), UEFI only; ARM64 planned — keep the architecture portable. The quality bar is shipping software: correct, efficient, minimal, zero silent debt. A tracked weakness is still a weakness: the honest answer about current state is "known, tracked, still true" — never "we have an issue for that." +A general-purpose operating system built from scratch in Rust, held to a production-grade engineering bar — the bar is the changes, not yet the product. Modern x86-64 hardware (2020+), UEFI only; ARM64 planned — keep the architecture portable. Its test machines decide its feature set, never its design. The quality bar is shipping software: correct, efficient, minimal, zero silent debt. A tracked weakness is still a weakness: the honest answer about current state is "known, tracked, still true" — never "we have an issue for that." ## Where the rest of this lives @@ -38,13 +38,13 @@ A subdirectory `CLAUDE.md` loads when a file in that subtree is `Read`, and not > A snapshot, deliberately shallow — always read the code. -**Kernel** — minimal; new additions are discussed and justified. Resource management, scheduling, process lifecycle, filesystem, device arbitration. 2 MB pages, demand paging, PIE binaries, full SMP. +**Kernel** — takes on only what userland cannot. 2 MB pages, demand paging, PIE binaries, full SMP. **Userspace daemons** — compositor, netd, soundd, sshd, logd. Each claims a device or capability from the kernel and serves its function; crash one and the kernel is fine. **The log is a userland file.** `/system/bin/logd` reads records on a cursor and owns `/log`; the kernel keeps the record ring, the console and the panel, and writes no file. `SYS_FSYNC` reaches the device's cache flush because logd's durability claim rests on it. -**Syscall ABI** — `toyos-abi/`: struct layouts, syscall numbers, typed wrappers; completely unstable, read the code. Never add or change a syscall without discussion; a deleted syscall's number is retired, never reused. `toyos/` builds on it with typed handles, IPC framing, ports, namespaces and `surface` — userland uses `toyos`, the kernel uses `toyos-abi` only. +**Syscall ABI** — `toyos-abi/`: struct layouts, syscall numbers, typed wrappers; completely unstable. The cleanest, most sustainable ABI beats convenience; a removed number is free. `toyos/` builds on it with typed handles, IPC framing, ports, namespaces and `surface` — userland uses `toyos`, the kernel uses `toyos-abi` only. **Capabilities** — a process holds exactly what its parent moved into it, and among kernel objects there is nothing it can name to get more. No registry, no connect-by-name, no pid-as-authority: `/system/bin/init` builds every program's namespace and device claims from `system.toml` before spawning it, and a handle a process does not hold is a bug in that process — the kernel ends it rather than answering a word it can ignore. **Isolation is non-negotiable, and the filesystem is inside it**: a process names only the paths in the view its parent built for it, the unit of isolation is the program, and a user is the part of the tree a session was handed. Not yet true of files: the kernel still resolves every path against one machine-wide tree until the storage track's per-program views land. @@ -60,7 +60,7 @@ A subdirectory `CLAUDE.md` loads when a file in that subtree is `Read`, and not **Rust** and **QEMU** for development, on any host OS and architecture — the development machine is nothing special. Beside them, where no Rust tool does the job, only C or C++ tools ToyOS can one day build and run (Python, Perl, CMake, make), each declared. No binary for one host OS alone: a macOS binary is a hard no, and "only for tests" does not soften it. ToyOS's own code is Rust; it writes no Python, Perl or shell of its own. Only general and widely used crates — one that does *our* job we write ourselves, and a driver crate never; third-party crates are used as published, and a fork carries a change written to upstream quality and goes when upstream has it. No upstream pull requests are sent for now: ToyOS needs more attention and more contributors before upstream projects take it seriously, and upstreams tend to refuse AI-first projects and their contributions. A third-party source ToyOS cannot build without changing it is carried as an unmodified-source packaging mirror with a byte-identity gate, not as a fork. The north star is **self-hosting**: nothing — build, test, or verification — rests on a host binary. Ask of anything new: could this ever run inside ToyOS? Self-hosting means ToyOS rebuilds itself on ToyOS and reproduces the host's bytes; a bootstrap from source with no binary seed is out of scope. -Vendor firmware a device verifies by its maker's signature may be shipped: pinned by version and hash, redistributable unmodified, recorded in `NOTICE`, and loaded only by that device's own driver through its IOMMU domain; it never executes on the CPU. +Vendor firmware a device or CPU verifies by its maker's signature may be shipped: pinned by version and hash, redistributable unmodified, recorded in `NOTICE`. A device's is loaded only by its own driver through its IOMMU domain and never executes on the CPU; CPU microcode is loaded by the kernel. The bar is not yet the tree: `.claude/agents/reviewer.md`, "Arrivals", says where every host tool and every standing failure is declared. `NOTICE` names every committed third-party file with its hash, upstream and licence; an image carrying `DOOM1.WAD` may not be sold. @@ -74,7 +74,7 @@ The testing rules live where they are enforced: the PR gate and the nightly in ` - `cargo run` builds everything (toolchain, kernel, bootloader, userland, image) and launches QEMU; `--build-only` skips the launch. `cargo test` runs the QEMU harness; `cargo run -- --ci host` runs every host suite, as the PR gate's required `host` check does. - **Agents never run QEMU.** An agent verifies with host tests and builds the image at most; the orchestrator runs every guest test, one suite at a time. - **Both produce large output**: run them in the background and read the output file — `[N characters truncated]` means data was lost. A full boot is under a second; incremental builds finish in seconds. -- **Leave the machine as you found it.** The development machine is shared: every agent stops what it started, removes the worktrees and scratch build output it no longer needs, and never leaves an emulator, a build or a watcher running. +- **Leave the machine as you found it.** The development machine is shared: every agent stops what it started, killing only by PID and waiting out a build that holds the global lock, and removes the worktrees and scratch build output it no longer needs. ## Repository layout diff --git a/Cargo.toml b/Cargo.toml index f7859848977..318e42c520e 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -198,8 +198,7 @@ opt-level = 2 # `opt-level = 0`. Nothing fail-fast is traded for it — `opt-level` does not # touch `debug-assertions`, which cargo still passes as `on` (checked with # `cargo build -p toyos-fat32 -v`), and rustc derives `overflow-checks` from -# that. The pure crates' hostile-input tests keep both knobs that *found* the -# two crafted-ELF kernel panics in `issues/`. +# that. # # **One member depends on this line for correctness rather than for speed, and # its own manifest says so at length — this is the other half of that @@ -218,10 +217,8 @@ opt-level = 3 # The one profile every guest binary is built with. Optimised, because an # unoptimised guest mismeasures everything under TCG; -# debug-assertions and overflow-checks on, because fail-fast beats speed here -# and both crafted-ELF kernel panics in `issues/` were *found* by -# an overflow check. `--release` would turn the last two off silently, so this -# build system no longer has the flag. +# debug-assertions and overflow-checks on, because fail-fast beats speed here. +# `--release` would turn the last two off silently. # # It is declared here rather than in `toyos-ld/Cargo.toml` because cargo # ignores `[profile]` in a workspace member and that is one now. It is the only diff --git a/bootloader/Cargo.toml b/bootloader/Cargo.toml index 7e57aadd824..089250dbe72 100644 --- a/bootloader/Cargo.toml +++ b/bootloader/Cargo.toml @@ -41,10 +41,8 @@ uefi-services = { version = "0.23.0", features = ["panic_handler", "logger"] } # The one profile every guest binary is built with. Optimised, because an # unoptimised guest mismeasures everything under TCG; -# debug-assertions and overflow-checks on, because fail-fast beats speed here -# and both crafted-ELF kernel panics in `issues/` were *found* by -# an overflow check. `--release` would turn the last two off silently, so this -# build system no longer has the flag. +# debug-assertions and overflow-checks on, because fail-fast beats speed here. +# `--release` would turn the last two off silently. [profile.toyos] inherits = "dev" opt-level = 2 diff --git a/issues/build/code-used-by-one-program-lives-in-that-program.md b/issues/build/code-used-by-one-program-lives-in-that-program.md new file mode 100644 index 00000000000..8735d4afb0b --- /dev/null +++ b/issues/build/code-used-by-one-program-lives-in-that-program.md @@ -0,0 +1,67 @@ +--- +status: open +kind: track +opened: 2026-09-29 +--- + +# Code used by one program lives in that program + +A crate exists because two programs share it. What only the kernel uses is the +kernel package's, what only one userland program uses is that program's, and +shared crates with one subject are one crate (owner, 2026-09-29). No gate holds +the layout: step 2 writes it into `.claude/agents/reviewer.md`'s Fit line, +which until then puts a pure decision in a pure crate. + +An input boundary is a crate of its own, whoever uses it: the no-panic +track (`issues/kernel/a-panic-is-never-an-accident.md`) forbids its tier 1 per +crate, and a crate that holds a tier-2 stop cannot forbid the set. A crate is +one when its own source decodes a word from outside its trust, or bounds it by +its form: hardware registers, firmware tables, disk bytes, network bytes, or +what another program sent, a syscall's arguments included. A lookup of a key a +program named is neither. No step here merges one; each stays a crate under +the no-panic track. + +Every step lands green on `--ci host` and `--build-only`. Test counts are what +`cargo test -p -- --list` lists today. + +1. **Delete `toyos-userpin`.** It models the pin invariant and names nothing + the kernel defines; `munmap_reissues_read_window` holds the kernel to it. + Check: `git grep toyos-userpin -- ':!issues/'` is empty. +2. **The kernel's library.** `kernel/pure/` is the `kernel` package's lib, and + its bin is `test = false`. `toyos-pcid`, `toyos-proclife` and `toyos-sched` + move in. `kernel-loom` and `toyos-sched/loom` become `kernel/loom/`, and + `toyos-sched/sim` `kernel/sim/`. The harness dev-depends on the kernel, and + the build system does not depend on it. The library has no `tests/`, since + an integration test builds the binary for the host. `--ci host` tests it + with `sched-check`, the feature scheduler tests need. The Fit line + states this track's rule. + Closes `issues/build/the-pcid-negative-control-runs-nowhere.md`. + Check: the library lists at least 135 tests, `kernel/loom` 79 and + `kernel/sim` 53, and `--clippy` lints the library's tests on the host. + Every moved control, and pcid's `counting-allocator`, reds with its verdict, + and `declared_model_controls` reads the kernel's manifest and every one in + the host workspace, not a list. An `unsafe {}` planted in a module that was + `forbid(unsafe_code)` does not compile. `cargo tree -e normal -p + toyos-build` names no `kernel`. +3. **libc's host test moves to the tests.** `toyos-libc-copies` moves to + `tests/libc-arch/`. + Check: `--ci host` runs there every test the package lists today, and + `--clippy` lints them. +4. **A crate one package uses goes under it, a crate of its own.** A crate of + this tree with exactly one consumer moves under it. Its consumers are the + packages that name it as a dependency of any kind, under any `cfg`, as + `cargo metadata --no-deps` reads every manifest `git ls-files '*Cargo.toml'` + lists, excluded packages and `tests/` included; a crate the images ship as a + program of its own counts as its own consumer. A move under a userland + program lands with `src/userlandhost.rs`'s survey gating a nested crate's + tests, which it lists as escapes today. + Check: `--ci host` runs every test each package lists today, and `--clippy` + lints them. +5. **`toyos-fat32-check` goes under `toyos-fat32/`.** The build and + `toyos-fat32`'s tests both use it, and it moves to `toyos-fat32/check/`, + beside the one subject it judges. + Check: `--ci host` runs every test the package lists today, and `--clippy` + lints them. + +**Exit:** no directory this file names as moved or merged still exists, and +step 4's count finds no crate outside its one consumer. diff --git a/issues/build/doom-jpg-shows-ids-art-under-no-recorded-terms.md b/issues/build/doom-jpg-shows-ids-art-under-no-recorded-terms.md deleted file mode 100644 index 518f966df5f..00000000000 --- a/issues/build/doom-jpg-shows-ids-art-under-no-recorded-terms.md +++ /dev/null @@ -1,17 +0,0 @@ ---- -status: owner -kind: question -opened: 2026-09-26 ---- - -# `doom.jpg` shows id's art, and nothing records the terms it is under - -`doom.jpg`, the screenshot `README.md` shows, is a picture of this system -running doom, so most of its pixels are `assets/DOOM1.WAD`'s graphics. -Its row in `COMMITTED_FILES` (`src/licence.rs`) declares `NOASSERTION`, -because the tree's `MIT OR Apache-2.0` does not reach id's art, and the -shareware terms in `NOTICE` do not name screenshots. No image -ships the file, so the licence gate does not judge it. - -**Exit**: the owner rules on the terms, and the row says them. Or the -screenshot is replaced by one that shows only this tree's own work. diff --git a/issues/build/the-pcid-negative-control-runs-nowhere.md b/issues/build/the-pcid-negative-control-runs-nowhere.md new file mode 100644 index 00000000000..0cd7cfc69c4 --- /dev/null +++ b/issues/build/the-pcid-negative-control-runs-nowhere.md @@ -0,0 +1,19 @@ +--- +status: open +kind: tooling +opened: 2026-09-29 +--- + +# `toyos-pcid`'s negative control runs nowhere + +`toyos-pcid/Cargo.toml` declares `counting-allocator`, the control that reverts +`PcidPool` to the counter that reissued a live tag. Nothing runs it: +`src/ci.rs`'s `CONTROLS` has no row for it, and `src/build.rs`'s +`declared_model_controls` does not read `toyos-pcid/Cargo.toml`, so +`every_model_control_is_run` cannot notice. The control still has teeth, run by +hand: `cargo test -p toyos-pcid --features counting-allocator` exits 101 with +`tests::two_live_address_spaces_never_share_a_pcid ... FAILED`. + +**Exit:** the control is a `CONTROLS` row demanding that `FAILED` line, and +`declared_model_controls` reads every manifest in the host workspace rather than +a list, so a control declared in a crate the list forgot reds. diff --git a/issues/build/two-builds-of-one-llvm-key-differ-in-their-bytes.md b/issues/build/two-builds-of-one-llvm-key-differ-in-their-bytes.md new file mode 100644 index 00000000000..76a2d235b9c --- /dev/null +++ b/issues/build/two-builds-of-one-llvm-key-differ-in-their-bytes.md @@ -0,0 +1,41 @@ +--- +status: open +kind: defect +opened: 2026-09-30 +--- + +# Two builds of one LLVM key differ in their bytes + +`src/llvm.rs` holds that an LLVM is a function of its key. Two builds of one +key by `build_in_fork`, in two fork checkouts, differed in three things the key +does not name: + +- **The LLVM checkout's `origin`.** LLVM's CMake writes it into + `VCSRevision.h` as `LLVM_REPOSITORY` (`get_source_info` in + `llvm/cmake/modules/VersionFromVCS.cmake`), and clang and LLD put it in their + version (`clang/lib/Basic/Version.cpp`, `lld/Common/Version.cpp`), so in the + guest's bytes: LLD writes it into the `.comment` of the `libstd` a sysroot + links, and clang into that of each C and C++ object of its `c/lib/libc++.a`. + Bootstrap's `update_submodule` sets a checkout's origin from the fork's + `.gitmodules`, `ToyOSOrg`, when it moves the checkout to the gitlink's commit, + and leaves it when the checkout holds that commit already. A worktree's + checkout is cloned from the URL the fork's shared config names, `ToyOSOrg`. +- **The build directory.** `lld`'s `LC_RPATH` and `llvm-config`'s object and + source roots name the `build/toyos-llvm` of the checkout that built it. +- **Archive dates.** The members of every static archive carry the time they + were built (`ar tv`). + +This is not `issues/build/two-checkouts-of-one-tree-build-different-guest-bytes.md`, +which holds the LLVM fixed and varies rustc's inputs: every checkout on a host +links the one LLVM its key names, so that gate sees no origin, and a comparison +of two LLVM installs sees no panic path rustc writes. + +**Exit**: `config_text` sets `LLVM_FORCE_VC_REPOSITORY` and +`LLVM_FORCE_VC_REVISION`, since the repository alone drops the revision, and a +host test that runs `llvm/cmake/modules/GenerateVersionFromVCS.cmake` with +`config_text`'s defines on two checkouts of one commit whose `origin`s differ +writes one header from both, where today each names its own origin. What only +a configure or a build writes, the build directory and the archive dates, a +nightly check sees: it builds one key in two fork checkouts at different paths +whose origins differ, checks after each build that they still do, and finds +every file of the two installs byte-identical. diff --git a/issues/build/two-checkouts-of-one-tree-build-different-guest-bytes.md b/issues/build/two-checkouts-of-one-tree-build-different-guest-bytes.md index 8939f054607..31275116251 100644 --- a/issues/build/two-checkouts-of-one-tree-build-different-guest-bytes.md +++ b/issues/build/two-checkouts-of-one-tree-build-different-guest-bytes.md @@ -10,7 +10,9 @@ Self-hosting is ToyOS rebuilding itself and reproducing the host's bytes, and the host does not yet reproduce its own across checkout paths. The same sources copied to two paths and built with the same sysroot, profile and linker gave three different kernels, bootloaders and `snake`s; the same path -built twice into two target directories gave identical ones. +built twice into two target directories gave identical ones. The LLVM is held +fixed here: two builds of one LLVM key differing is +`issues/build/two-builds-of-one-llvm-key-differ-in-their-bytes.md`. - The kernel and the bootloader carry each path dependency's absolute source path in their panic locations (55 strings in the kernel, 13 in the diff --git a/issues/build/userland-programs-are-never-linted.md b/issues/build/userland-programs-are-never-linted.md new file mode 100644 index 00000000000..848e71af2b4 --- /dev/null +++ b/issues/build/userland-programs-are-never-linted.md @@ -0,0 +1,20 @@ +--- +status: open +kind: tooling +opened: 2026-09-29 +--- + +# No userland program is linted, though the gated ones build for the host + +`src/clippy.rs` lints no userland crate because the `toyos` toolchain ships no +clippy. The crates `src/userlandhost.rs` gates already build for the host with +the stable toolchain, and stable's clippy lints them there. They are not clean: +`cargo clippy --manifest-path userland//Cargo.toml --target +aarch64-apple-darwin --all-targets -- -D warnings` exits 101 with 8 findings +in soundd and 10 in netd. The compositor stops on 1 finding in its dependency +`userland/toyos-window`. + +**Exit:** a shape in `src/clippy.rs`, which `--clippy` and `--ci host` both +run, lints every crate the survey gates, on the host target, with warnings +denied, and those findings are fixed. A clippy finding planted in a gated +userland crate reds `--clippy`. diff --git a/issues/build/whether-doomgeneric-becomes-a-submodule-is-the-owners.md b/issues/build/whether-doomgeneric-becomes-a-submodule-is-the-owners.md deleted file mode 100644 index 82fa7f73378..00000000000 --- a/issues/build/whether-doomgeneric-becomes-a-submodule-is-the-owners.md +++ /dev/null @@ -1,20 +0,0 @@ ---- -status: owner -kind: question -opened: 2026-09-29 ---- - -# Whether doomgeneric becomes a submodule is the owner's - -`userland/doom/build.rs` fetches doomgeneric's commit `fc601639` as a GitHub -archive into the gitignored `userland/doom/doomgeneric/` and compiles it. The -fetch costs five build-dependencies in `userland/doom/Cargo.toml` — `ureq`, -`rustls-rustcrypto`, `webpki-roots`, `flate2`, `tar` — and leaves a ToyOS -change to the C nowhere to live but that untracked directory, which the build -replaces whenever its stamp disagrees with the pin. A fork repository of -doomgeneric as a submodule at that path deletes the fetch and the five. Neither -`ToyOSOrg/doomgeneric` nor `Japabu/doomgeneric` exists (`gh repo view`), and -creating one is the owner's. - -**Exit**: the owner rules — the fetch stays, or a submodule on a repository he -creates replaces it. diff --git a/issues/design-debt/a-boot-start-refused-a-present-device-runs-without-it.md b/issues/design-debt/a-boot-start-refused-a-present-device-runs-without-it.md index 14ee6cc58f9..0702b2c38d7 100644 --- a/issues/design-debt/a-boot-start-refused-a-present-device-runs-without-it.md +++ b/issues/design-debt/a-boot-start-refused-a-present-device-runs-without-it.md @@ -1,26 +1,21 @@ --- -status: owner -kind: question +status: open +kind: defect opened: 2026-09-24 --- # A boot start refused a device the machine has runs without it -`/system/bin/init`'s `start` mints every device a `[programs]` row names. At a -swap or the restart that rolls one back, a device the process being replaced -held and the new one is refused fails the start (`Served::Restart`'s `owed`). -At boot there is no process before it, and a refusal is said in the kernel's -own word (`init: netd: pci:8086:10c9 is on this machine and could not be -handed over`) and the program is started without that device. +A boot start's device refusal is fatal only for a device its `[programs]` row +marks as required; for any other, init logs the refusal loudly and starts the +program without that device (owner, 2026-09-30). No row can mark a device +required, so every boot start refused a device the machine has — +`NotSupported`, `AlreadyExists`, `ResourceExhausted` — runs without it: +`/system/bin/init`'s `start` says the refusal in the kernel's own word +(`init: netd: pci:8086:10c9 is on this machine and could not be handed over`) +and starts the program. -That is not the swap's defect — init claims nothing about a boot start beyond -`started`, and a row names every device its program can drive, so a machine -with a subset is a configuration, not a fault — but a refusal other than -`NotFound` is a fault, and the program then runs on less than the machine -has. `tests/common/faults.rs`'s `refused_claim` and `pci_function_is_exclusive` -assert exactly this behaviour today. - -The question for the owner: should a boot start refused a device the machine -has — `NotSupported`, `AlreadyExists`, `ResourceExhausted` — be fatal? A failed -boot start is an init panic, so the machine would not boot over one device; -and a row cannot yet say which of its devices its program needs. +**Exit**: a `[programs]` row can mark a device required, a host test holds the +mark and init's choice between a fatal refusal and a logged one, and a metal +row or, where none can, a guest test shows a boot start refused a required +device is fatal. diff --git a/issues/design-debt/os-toyos-io-traits-keep-a-posix-name.md b/issues/design-debt/os-toyos-io-traits-keep-a-posix-name.md index d76b3aeaa2e..df9885c522c 100644 --- a/issues/design-debt/os-toyos-io-traits-keep-a-posix-name.md +++ b/issues/design-debt/os-toyos-io-traits-keep-a-posix-name.md @@ -1,37 +1,18 @@ --- -status: owner -kind: question +status: open +kind: defect opened: 2026-08-24 --- -# Does `std::os::toyos::io` get to keep `AsRawFd`/`FromRawFd`? +# `std::os::toyos::io` names its raw-handle traits `AsRawFd` and `FromRawFd` -The owner ruled on 2026-08-19 that **"fds belong only in libc jargon"**: the -kernel has no file descriptors, a process holds typed handles -(`toyos-abi/src/handle.rs`), and `userland/libc` is the one layer whose -interface the word describes. The sweep that followed took the kernel, the ABI, -the SDK, every forced call site and both fork pins. One name it did not settle -is left, and it is the owner's because it is a fork-repo name (`Japabu/rust`) -and not this tree's to rename alone. +`os::toyos::*` is ToyOS's own extension API and speaks ToyOS, and "fds belong +only in libc jargon" (owner, 2026-08-19). The `rust/` fork's +`library/std/src/os/toyos/io.rs` re-exports `std::os::fd`, so +`std::os::toyos::io::{AsRawFd, FromRawFd}` still speak POSIX. +`tests/toyos-rust-tests/src/bin/std_fs.rs` is their one caller in this +repository. -`std::os::toyos::io::{AsRawFd, FromRawFd}`. The naming table ruled that -`os::toyos::*` is ToyOS's own extension API and must speak ToyOS, and it named -every `process` row — but not this one. std's *POSIX* surface keeps the word by -charter; whether an `os::toyos` trait carrying a POSIX name is that surface or -the extension API is what nobody has decided. - -The two answers, both defensible: - -- **It stays**, as a deliberate mirror of std's own - `std::os::unix::io::{AsRawFd, FromRawFd}` convention — the trait names an - `os::toyos` caller reaches for by muscle memory, and diverging costs every - such caller a lookup. -- **It is renamed the next time the trait is touched in the fork**, to whatever - `os::toyos` decides its own non-POSIX vocabulary for a raw handle is, because - the extension API speaking POSIX is precisely what the ruling forbade - everywhere else. - -`tests/toyos-rust-tests/src/bin/std_fs.rs:4` is the only caller in this -repository (`use std::os::toyos::io::{AsRawFd, FromRawFd};`, used at `:59`). -Nothing is blocked on the answer; the word is legal in the two places it is -still spoken here, so this is a naming decision and not a defect. +**Exit**: the next time the trait is touched in the fork, `AsRawFd` and +`FromRawFd` in `std::os::toyos::io` are renamed to `os::toyos`'s own word for +a raw handle, and `std_fs.rs` uses the new names (owner, 2026-09-30). diff --git a/issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md b/issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md new file mode 100644 index 00000000000..cb5d6473474 --- /dev/null +++ b/issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md @@ -0,0 +1,42 @@ +--- +status: open +kind: defect +opened: 2026-10-01 +--- + +# The ABI still keeps retired syscall numbers + +The ABI is completely unstable and a removed syscall's number is free (owner, +2026-10-01): there are no retired numbers and no compatibility shims. The tree +still keeps them: + +- `kernel/src/syscall/dispatch.rs`'s `retired_syscalls!` names deleted calls so + that "an old binary is told which call it was", logging each call of one; + every other unassigned number answers `InvalidArgument` silently. +- `toyos-abi/src/syscall.rs` and `toyos-abi/src/inbox.rs` carry a "formerly …" + or "retired and unused" entry per deleted syscall, `SYS_DEBUG` action and + inbox op, `SYS_INBOX_SETUP`'s doc states the retirement rule, and + `kernel/src/inbox/mod.rs` names op 2 retired. +- `toyos-abi/src/syscall.rs`'s `device_classes!` keeps device classes 3 (`Nic`) + and 4 (`Audio`) "retired rather than reused", and the decode test in + `toyos-abi/src/inventory.rs` refuses them as retired. +- `tests/toyos-rust-tests/src/bin/panic_halts_first.rs` and `tests/toyos.rs` + take syscall 26 as their logged refusal and read `syscall 26 is retired`. +- Plans still follow the old rule: + `issues/kernel/sys-clock-realtime-is-now-a-format-of-sys-clock-epoch.md`, + `issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md`, + `issues/kernel/the-kernel-still-parses-what-userland-writes.md`, + `issues/kernel/the-capability-end-state-is-twelve-answers.md`, + `issues/kernel/sys-debug-actions-and-two-loader-words-that-nothing-calls.md`, + `issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md`, + `issues/diagnostics/the-kernel-keeps-nothing-it-enumerates.md`, + `issues/diagnostics/no-cyclictest.md`, and + `issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md`, whose + "ABI brief" names a gate `.claude/agents/implementer.md` no longer has. + +**Exit**: `retired_syscalls!` and every retirement entry are gone, a deleted +number answers as an unassigned one does, both test sites take their logged +refusal from something live, and no issue plans by retirement. + +Owner: `toyos-abi` and `kernel/src/syscall/dispatch.rs`, whoever next changes +the ABI. diff --git a/issues/hardware/a-tco-base-word-near-all-ones-panics-the-loader-and-the-kernel.md b/issues/hardware/a-tco-base-word-near-all-ones-panics-the-loader-and-the-kernel.md new file mode 100644 index 00000000000..2fe7519243e --- /dev/null +++ b/issues/hardware/a-tco-base-word-near-all-ones-panics-the-loader-and-the-kernel.md @@ -0,0 +1,27 @@ +--- +status: open +kind: defect +opened: 2026-10-01 +--- + +# A TCO base word near all-ones panics the loader and the kernel + +`toyos_tco::Chipset::port` (`toyos-tco/src/lib.rs:275-289`) refuses an +all-ones base word by name, but under Tiger Lake-LP's row (`8086:a0a3`, +`base_mask: !1`) a word from `0xFFFF_FFEE` to `0xFFFF_FFFE` reaches `:284`, +where `base + base_offset + TCO_TMR + 1` overflows `u32`. The loader and the +kernel each read the word out of the PCH's configuration space and hand it to +`port` (`bootloader/src/watchdog.rs:141`, `:73`; +`kernel/src/arch/x86_64/watchdog.rs:51`, `:58`), and both build with +`overflow-checks`. So a boot that arms the watchdog on a machine whose function +answers one of those words panics where it should refuse. + +A host program calling `port` with every `u32` base word and the row's enable +bit, built with `overflow-checks`, panics with "attempt to add with overflow" +at `toyos-tco/src/lib.rs:284:19` for 17 words under `8086:a0a3`, +`0xFFFF_FFEE` the first and `0xFFFF_FFFE` the last, and for none under q35's +`8086:2918`. + +**Exit:** `port` refuses, by name, every base word whose block does not fit +the I/O space, and a host test in `toyos-tco` asks it `0xFFFF_FFFE` under the +`a0a3` row. diff --git a/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md b/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md index 9965b710d98..dbd4f3f6901 100644 --- a/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md +++ b/issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md @@ -84,7 +84,5 @@ nothing. ## Open with the owner -- Before stage 3: the ask's ABI, which is Q6a of - `issues/kernel/the-child-process-track-waits-on-the-owners-rulings.md`; whether a program - started through `launcher` is asked or only stopped; whether - `SYS_SHUTDOWN`/`SYS_REBOOT` change at all. +- Before stage 3: whether a program started through `launcher` is asked or + only stopped; whether `SYS_SHUTDOWN`/`SYS_REBOOT` change at all. diff --git a/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md b/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md index 8a725701db4..f22a4a92526 100644 --- a/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md +++ b/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md @@ -10,8 +10,7 @@ libc cannot start a child: `fork` and `execvp` answer `ENOSYS`, `waitpid` `ECHILD` and `system` `-1` (`userland/libc/src/misc.rs`, `userland/libc/src/stdio.rs`), and there is no `posix_spawn`. M2 and M4 of `issues/build/toyos-builds-itself.md` and the exit of -`issues/kernel/toyos-runs-on-arm64.md` need it. Stages 2, 3, 5, 6 and 7 wait -on `issues/kernel/the-child-process-track-waits-on-the-owners-rulings.md`. +`issues/kernel/toyos-runs-on-arm64.md` need it. **Constraints.** libc is built on ToyOS and never ToyOS on libc (owner, 2026-09-30: "toyos moderness and anti legacy may never be compromised because @@ -30,17 +29,15 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md — exit, crash or kill — kills its whole subtree at once; quitting gracefully is the parent's job before it exits, never the kernel's. The kernel starts init and init starts everything else, with no rule for services: a server - is init's child and lives as long as init. Each login, desktop or SSH, is a - session process under init that parents everything its user starts; an app - uses a server through handles and is never its child; logging out ends the - session. "Keep running after I close this" asks init, through `launcher`, to - be the parent, and nothing else reparents. A kill is cleanup without - cooperation, run on the causing path in bounded steps. init never dies and - stays tiny. + is init's child and lives as long as init. An app uses a server through + handles and is never its child. "Keep running after I close this" asks init, + through `launcher`, to be the parent, and nothing else reparents. A kill is + cleanup without cooperation, run on the causing path in bounded steps. init + never dies and stays tiny. ## Stages -0. **`SYS_PROCESS_OPEN` goes** (ruled). Deleted: `sys_process_open` and every +0. **`SYS_PROCESS_OPEN` goes**. Deleted: `sys_process_open` and every name only it reaches — `process_open`, `SysCap::open_process`, `MANAGE` in init's `SysCap`, the `reopenable` column, `process::process_object`, `reopen_selftest` and `sched::kthread::open_selftest` with their actuator, @@ -54,7 +51,7 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md whose symbols are test data. Negative control: the stage reverted whole, where `sys_process_open` answers 110. Oracle: rustc's name resolution, which fails the build of any caller left. -1. **An end is an event** (ruled). `read_watch` and `has_data` answer for a +1. **An end is an event**. `read_watch` and `has_data` answer for a `Process`, whose watch becomes an `Arc` as an `Acceptor`'s is, and `close_ends_polls` answers `false` for one; init's waiter threads go. *Exit*: children held on their stdin (`process_lifecycle`'s `held` role), @@ -64,33 +61,36 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md handles to a held child closes, a non-blocking submit finds nothing. Negative control: the stage reverted whole (`NotSupported`); mutation: `close_ends_polls` at `true` reds the last arm. Oracle: pidfd_open(2). -2. **An end says how** (Q2, Q6d). *Exit*: exited 137, killed, and each fault +2. **An end says how**. An end reads as an exit, a kill or a fault kind + alike on every architecture, and a bare code reads the last two as failures. + No end reads as a quit's reason. *Exit*: exited 137, killed, and each fault kind the architecture raises, each read from a child that did it. Negative control: the stage reverted whole, where the first two read alike. Oracle: POSIX ``'s classes. -3. **libc starts and waits for children** (Q3a–Q3c; blocked on +3. **libc starts and waits for children** (blocked on `issues/isolation/a-childs-stdio-handle-is-not-the-one-command-named.md`). - `posix_spawn` and `posix_spawnp` with file actions and attributes, routed by - std's launch-or-spawn rule moved into `toyos`; a descriptor table with - close-on-exec, a child getting descriptors 0–2 and exactly what its file - actions name; `waitpid`, `wait`, `wait4`; `ppoll`, `sigpending` and the + libc imitates `SIGCHLD`. A C child starts with descriptors 0–2 and exactly + what its file actions name, a stated departure from POSIX. `posix_spawn` and + `posix_spawnp` with file actions and attributes, routed by std's + launch-or-spawn rule moved into `toyos`; a descriptor table with + close-on-exec; `waitpid`, `wait`, `wait4`; `ppoll`, `sigpending` and the signal-set calls; `environ`, `setenv`, `unsetenv`; `kill` with `SIGKILL` for a child or a `-pgid` of its children, 0 as a probe, `EINVAL` for other signals until stage 6; `system`, `popen` and `pclose` through `/system/bin/shell -c`. No `select`: a descriptor is a handle and passes - `FD_SETSIZE`. A handler for any signal libc imitates runs inside a blocking - call of a thread that leaves the signal unblocked (`ppoll` and `sigsuspend` - under their mask); after it `read`, `recv`, `accept`, `wait`, `waitpid` and - `wait4` restart if it was installed with `SA_RESTART` and answer `EINTR` if - not, and `poll`, `ppoll`, `pause`, `sigsuspend`, `nanosleep` and `usleep` - answer `EINTR`, `sleep` the time left. While no thread is in one, it runs at - once on libc's own thread; while every thread blocks the signal, it stays - pending, as `sigpending` reports, until one unblocks it. libc's own thread - watches at most 255 unwaited children on its own inbox, the 256th watch - being stage 6's notice, and wakes a thread in a blocking call through one - pipe, so `poll` takes at most 255 descriptors and children never count - against them. *Exit*, C corpus files, the mutation that reds an arm after - it: a piped stdout read to EOF through `poll`, then `waitpid`; + `FD_SETSIZE`. A handler for any signal libc imitates runs at once, beside the + program: inside a blocking call of a thread that leaves the signal unblocked + (`ppoll` and `sigsuspend` under their mask), and, while no thread is in one, + on libc's own thread. After it `read`, `recv`, `accept`, `wait`, `waitpid` + and `wait4` restart if it was installed with `SA_RESTART` and answer `EINTR` + if not, and `poll`, `ppoll`, `pause`, `sigsuspend`, `nanosleep` and `usleep` + answer `EINTR`, `sleep` the time left; while every thread blocks the signal, + it stays pending, as `sigpending` reports, until one unblocks it. libc's own + thread watches at most 255 unwaited children on its own inbox, the 256th + watch being stage 6's notice, and wakes a thread in a blocking call through + one pipe, so `poll` takes at most 255 descriptors and children never count + against them. *Exit*, C corpus files, the mutation that reds an arm after it: + a piped stdout read to EOF through `poll`, then `waitpid`; `waitpid(-1)` answering three children once each in the order they end, then `ECHILD` (waiting on the first alone), and `wait` and `wait4` alike (either answering `ECHILD` as `waitpid` does today); a killed child `WIFSIGNALED` @@ -117,7 +117,7 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md the pipe answering (passed to the kernel, which ends the caller). Negative control: today's libc, against which the corpus does not link. Oracle: POSIX's `posix_spawn`, `waitpid`, `poll` and `SA_RESTART`, and signal(7). -4. **A parent takes its children down** (ruled). A spawn's parent is its +4. **A parent takes its children down**. A spawn's parent is its spawner. A launch names its parent in its request, by one of two words: the caller, whose place is a copy of the handle to itself every process starts holding under the label `self` (`WRITE`, `DUP`, `TRANSFER`), carried as one @@ -168,20 +168,21 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md place spawned under init (it starts); std's `NotDeclared` fallback, or its early return for an endowment or extra slot, kept for init (it starts). Oracle: cgroup v2's `cgroup.kill` and `cgroup.max.depth`. -5. **A login is a session under init** (ruled; Q5a–Q5c). init starts a session - process per login — the desktop's at boot, one per SSH connection at sshd's - request — and the compositor and sshd start a login's programs under it. Its - namespace is `issues/filesystem/a-user-is-a-home-tree-and-a-login-row.md`'s - login row. *Exit*: killing the desktop session ends every program the - compositor started and leaves the compositor running; killing the compositor - ends none of them; a dropped SSH connection ends its session and all under - it. Negative control: the stage reverted whole, where the compositor's - programs die with it. Oracle: systemd-logind, whose session scope ends every - process of a login. -6. **A program is asked to quit** (Q3b, Q6a–Q6f). `SYS_PROCESS_QUIT` (124, - never assigned), `(process, reason)`, needs `MANAGE` as the kill does and - reaches the process's subtree by stage 4's walk; the reason is interrupt, - hang-up or terminate. Each process is started holding its notice, a new +5. **A login is a session under init**. Per login init starts a session + program that logout ends and that only parents what its user starts — the + desktop's at boot, one per SSH connection at sshd's request. As it starts + one, init hands the compositor or sshd a right to start programs in that + session alone. sshd's right also quits and kills it. Its namespace is + `issues/filesystem/a-user-is-a-home-tree-and-a-login-row.md`'s login row. + *Exit*: killing the desktop session ends every program the compositor started + and leaves the compositor running; killing the compositor ends none of them; + a dropped SSH connection ends its session and all under it. Negative control: + the stage reverted whole, where the compositor's programs die with it. + Oracle: systemd-logind, whose session scope ends every process of a login. +6. **A program is asked to quit**. `SYS_PROCESS_QUIT` (124, never + assigned), `(process, reason)`, needs `MANAGE` as the kill does, and the + reason is interrupt, hang-up or terminate. A quit reaches the process's + subtree by stage 4's walk. Each process is started holding its notice, a new object kind under the label `quit`, `READABLE` while a reason is asked and not yet read, whose read takes it. A quit kills a process that has never watched its notice or whose notice still holds a reason, whoever asks, in @@ -192,8 +193,8 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md Unix arm maps; rustc stops skipping its handler on ToyOS (`rust/compiler/rustc_driver_impl/src/lib.rs`, `install_ctrlc_handler`). Stage 5's session, `/system/bin/terminal` and `/system/bin/shell` listen and - relay nothing, the walk having asked what they run: each ends once that has, - `-c` with its status, but an interactive shell does not end on an interrupt. + relay nothing, each ending once what it runs has, `-c` with its status, but + an interactive shell does not end on an interrupt. libc takes the three reasons as `SIGINT`, `SIGHUP` and `SIGTERM`, watching the notice once one is handled, ignored or blocked: an ignored reason is dropped and an unhandled one ends the process with exit code 128 plus the @@ -236,25 +237,24 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md watching only from a handler's install (the ignoring child is killed); a spawn carrying no ignore, no mask or no `SETSIGDEF` (each spawned child in turn). Oracle: POSIX's signal actions (XSH 2.4.3) and `posix_spawn`. -7. **Job control** (Q6a–Q6c, Q7). The shell hands its terminal a `MANAGE`-only - duplicate of each process of the foreground line over its `surface` port. - The terminal turns each Ctrl+C into their interrupt, and before it closes it - asks its shell's tree to hang up and kills what is left at its deadline. - sshd does that to a dropped connection's session, and init at shutdown and - test-runner at a test's deadline ask with terminate. `&` is a line the shell - watches. sshd awaits exits through a ToyOS `tokio/src/process` arm in the - tokio fork, shaped as its Linux pidfd arm (`process/unix/pidfd_reaper.rs`). - Stop and continue need a suspend primitive and are not proposed. *Exit*: +7. **Job control**. The shell hands its terminal a `MANAGE`-only + duplicate of each process of the foreground line over its `surface` port. The + terminal turns each Ctrl+C into their interrupt, and a second kills only a + program that has not yet taken the first. Before it closes, the terminal asks + its shell's tree to hang up and kills what is left at its deadline; sshd does + that to a dropped connection's session, and init at shutdown and test-runner + at a test's deadline ask with terminate. `&` is a line the shell watches. + sshd awaits exits through a ToyOS `tokio/src/process` arm in the tokio fork, + shaped as its Linux pidfd arm (`process/unix/pidfd_reaper.rs`). *Exit*: Ctrl+C ends a pipeline whose stage started a child, none of it left and the prompt back, and ends no shell nested on the foreground line; a child that reads each interrupt survives three Ctrl+C, each pressed once it reports its - watch armed or the one before read, and one that reports its watch armed - and never reads dies on the second; a dropped SSH connection whose session - runs a watching child sees it write its file before it ends; tokio's - `Child::wait` returns with no loop polling `try_wait`. Negative control: - today's terminal and shell, where the pipeline runs on. Mutations: a quit - that kills on a reason already read (the reading child dies); a session - that never listens (the file unwritten); a shell that ends on an interrupt - (the nested shell). Oracle: POSIX's INTR, a `SIGINT` "sent to all processes - in the foreground process group" (XBD 11.1.9), and tokio's - `tests/process_smoke.rs`. + watch armed or the one before read, and one that reports its watch armed and + never reads dies on the second; a dropped SSH connection whose session runs a + watching child sees it write its file before it ends; tokio's `Child::wait` + returns with no loop polling `try_wait`. Negative control: today's terminal + and shell, where the pipeline runs on. Mutations: a quit that kills on a + reason already read (the reading child dies); a session that never listens + (the file unwritten); a shell that ends on an interrupt (the nested shell). + Oracle: POSIX's INTR, a `SIGINT` "sent to all processes in the foreground + process group" (XBD 11.1.9), and tokio's `tests/process_smoke.rs`. diff --git a/issues/kernel/a-panic-is-never-an-accident.md b/issues/kernel/a-panic-is-never-an-accident.md new file mode 100644 index 00000000000..7336e8d2e97 --- /dev/null +++ b/issues/kernel/a-panic-is-never-an-accident.md @@ -0,0 +1,60 @@ +--- +status: open +kind: track +opened: 2026-10-01 +--- + +# A panic is never an accident: the owner's rule, not yet enforced + +| Tier | Rule | +|---|---| +| 1. Input boundaries | No panic at all: the set below is forbidden. Where the build allows, each parser's entry point also carries a link-time no-panic proof. | +| 2. The kernel | No implicit panic: the set is denied. A deliberate stop for a broken internal invariant stays, spelled out at its site as an `#[expect(…, reason = "…")]` whose reason names the invariant. | +| 3. System services | The same rule as the kernel. | +| 4. Apps, ports, test code and build tooling | Normal Rust. | + +Tiers 2 to 4 are programs, not crates: a tier holds every crate of this tree +its programs link, and a crate in two tiers is held to the stricter. The loader +is tier 2. A system service is a program any shipped mode's config marks +`service = true`, as `build::shipped` reads the modes, or whose manifest +declares `exempt.owns`, as `src/userlandhost.rs` reads it. Every other shipped +program is an app, `toybox`, `terminal` and the `exempt.manages` tools included. + +**The set**, all `clippy::`: `indexing_slicing`, `string_slice`, +`arithmetic_side_effects`, `unwrap_used`, `expect_used`, `panic`, +`unreachable`, `todo`, `unimplemented` and `panic_in_result_fn`; the lossy +casts `cast_possible_truncation`, `cast_possible_wrap`, `cast_sign_loss` and +`cast_precision_loss`, which for tiers 1 to 3 overrides their rejection in +`issues/build/clippy-stage-two-is-lints-one-at-a-time.md`; `disallowed_methods` +naming `slice::split_at`, `slice::split_at_mut`, `slice::copy_from_slice` and +`slice::clone_from_slice`; and `disallowed_macros` naming `core::assert`, +`core::assert_eq` and `core::assert_ne`. It does not see a shift by a variable +amount, `pow`, `abs`, a standard function it does not name or a callee's panic. + +**Stages, in order.** +1. Tier 1, one crate at a time; an input boundary inside the kernel or the + loader is moved into a crate of its own first. The first exit is one + declaration of the set that every tier-1 crate's library is linted under in + place of its own copy, its tests left at tier 4. `disallowed_methods` and + `disallowed_macros` read the nearest `clippy.toml` alone, and the root's + reaches every crate beneath it, tier 4 included. An area's exit is its crate + forbidding the set under `cargo run -- --clippy`. +2. The link-time proof on the parser entry points. The exit is a step of + `cargo run -- --ci host` whose host build refuses to link an entry point + with a panicking path, or a `rejected` issue with the measurement. +3. The kernel and the loader, one module at a time. The stage ends when + `kernel/src/main.rs`, `bootloader/src/main.rs` and every crate of this tree + `kernel/Cargo.toml` or `bootloader/Cargo.toml` links, unless stage 1 + already forbids the set in it, deny the set and forbid + `clippy::allow_attributes` and `clippy::allow_attributes_without_reason` + under `--clippy`, and a step of `--ci host` refuses an inner `#![allow]` + and an `#[expect]` of the set over more than one finding. + `arithmetic_side_effects` reports an expression once, however many + operators it nests: `x * x + y * y - 1` is one finding, so one `#[expect]` + covers its four overflows. +4. The system services, once `issues/build/userland-programs-are-never-linted.md` + has put userland in `src/clippy.rs`. The stage ends when every crate of this + tree a service links by a normal edge, whatever its `cfg`, its own and + `toyos` included, unless stage 1 already forbids the set in it, is linted + by `--clippy` under stage 3's attributes and passes stage 3's step, and the + step finds those crates itself. diff --git a/issues/kernel/the-child-process-track-waits-on-the-owners-rulings.md b/issues/kernel/the-child-process-track-waits-on-the-owners-rulings.md deleted file mode 100644 index a6d2f66ca93..00000000000 --- a/issues/kernel/the-child-process-track-waits-on-the-owners-rulings.md +++ /dev/null @@ -1,283 +0,0 @@ ---- -status: owner -kind: question -opened: 2026-09-30 ---- - -# The child-process track waits on the owner's rulings - -Stages 2, 3, 5, 6 and 7 of -`issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md` -wait on these, and are written there as recommended here. Each question is one -decision. Its ruling goes into the track as one line, and its entry here is -deleted. - -## Q2. Can a parent tell a failure from a kill or a crash? (stage 2) - -When a child ends, its parent reads one number. Today a child that exits with -137 and one that is killed both read 137, and a crash reads 139 or -1, numbers -a program can exit with too. - -- **Yes.** The number says whether the program ended itself or the kernel - ended it, and if the kernel did, why: killed, or which kind of crash. A crash - reads the same on x86-64 and on ARM64. A program that reads the number the - old way still sees a kill or a crash as a failure, never as a success. -- **No.** A tool that retries a crashed job cannot tell it from one that - failed, and a program that exits with 137 on purpose reads as killed. - -*Recommended: yes.* - -## Q3a. How does a C program learn that its child ended? (stage 3) - -C programs written for Unix learn it from a Unix signal, `SIGCHLD`: Ninja -waits for one (v1.13.1 `src/subprocess-posix.cc`), and so does libuv where it -has no kqueue (v1.53.0 `src/unix/process.c`). ToyOS has no signals, and its -kernel never will. - -- **libc imitates the signal.** libc, the C library, learns of the end from - ToyOS's own event and runs the program's handler for it. Ninja runs - unchanged, and the kernel gains nothing. -- **Each such program is changed for ToyOS.** Ninja, libuv and every other - program that waits for its children carries a ToyOS change of its own, as a - fork, for what one part of libc would serve. - -*Recommended: libc imitates it.* - -## Q3b. When does a C program's signal handler run? (stages 3 and 6) - -On Unix a signal interrupts the program wherever it is, and the handler runs -in its place. ToyOS never does that; the owner ruled it out as legacy. libc can -run a handler only on a thread it controls. - -- **At once, beside the program.** While one of the program's threads is - waiting inside libc — for input, for a child, for time to pass or for a - signal — the handler runs on that thread and cuts the wait short, as on - Unix: a C prompt waiting for a line comes back on Ctrl+C. As on Unix, a - program can ask that a wait for input or for a child carry on after its - handler instead, and a wait for time or for a signal is always cut short. - While no thread is waiting, libc runs the handler on a thread of its own, and - the program's own code goes on running at the same time. Ctrl+C stops clang - at once, and clang deletes the file it was writing. The cost: the handler and - the program's own code can then work on the same thing at once and collide. - GNU make's Ctrl+C handler collects the compiles that have ended, which make's - own code also does; on Windows, the one system where make's handler already - runs beside its code, make stops its own code first for exactly that reason - (make 4.4.1 `src/commands.c:508-510`). On ToyOS it collides only when Ctrl+C - comes while make is working rather than waiting. A handler that jumps back - into the program's main loop, as some C programs' do, lands on the wrong - thread and breaks the program; and clang can delete that file while its main - code is still writing into it. POSIX lets a handler run on any of a program's - threads that does not block the signal; this one is a thread the program - never made. -- **Only when the program next waits inside libc.** Nothing runs beside the - program. A compiler that is computing does not wait, so Ctrl+C does not stop - clang until it is killed, and its half-written file stays behind. - -*Recommended: at once, beside the program.* - -## Q3c. What does a C child start with? (stage 3) - -On Unix a child starts holding every file its parent has open, unless the -parent marked the file not to be passed on. - -- **Only what the parent names**, a stated departure from POSIX. A C child - starts with its input, output and error, and exactly what its parent names - when it starts it. A program started from C then holds what it holds when - started from Rust: what it is declared to hold. The cost: a C program that - hands its child a file by leaving it open, without naming it, loses it. GNU - make does that to share its limit on how many jobs run at once, with a pipe - it leaves open to a make it starts; LLVM's tools, told to share the limit, - take the same pipe (`llvm/lib/Support/Unix/Jobserver.inc`). Without it, a - make that make starts warns and runs one job at a time (make 4.4.1 - `src/main.c:1849`), and an LLVM tool warns and runs as many as it would - alone. make shares the limit by a named pipe instead where the system has - one, and ToyOS has none. -- **Unix's rule.** A C parent with any other file open passes it on, and every - program it starts runs with what that parent holds, where the same program - started from Rust runs with what it is declared to hold. - -*Recommended: only what the parent names.* - -## Q5a. What program is a login? (stage 5) - -Each login, on the desktop or over SSH, has one program that is the parent of -everything its user starts, so that logging out ends all of it. - -- **A program of its own that does nothing else.** init starts a small session - program for each login, and everything the user starts runs under it. It - listens: when the login is asked to end, its programs are asked too, and it - ends once they have. A crash of any program the user started leaves the - login and the rest of its programs running. -- **The login's first program.** An SSH login's session is the shell sshd - starts for it, and the desktop's is the first program the desktop starts for - the user. When that program ends or crashes, everything the user started - ends with it, as on Unix when the login shell exits. One program fewer runs - per login. - -*Recommended: a program of its own.* - -## Q5b. Who lets the compositor and sshd start programs in a login? (stage 5) - -A login's programs are started under its session, never under the compositor -or sshd, so each of them needs a right to start programs there. - -- **init hands it over when it starts the session.** At boot init starts the - desktop's session and gives the compositor the right to start programs in - it; for SSH, sshd asks init for a session per connection and gets that right - back. Each right reaches one session. -- **The session program hands it over itself**, over a connection init sets up - between them. The session program then has a second job besides being a - parent, and each login costs one connection more. - -*Recommended: init hands it over.* - -## Q5c. Can sshd end a login it asked for? (stage 5) - -When an SSH connection drops, its login ends: its programs are asked to hang -up, which reaches them because the session program listens (Q5a), and what is -left is killed. - -- **Yes, only its own.** The right sshd gets for a connection's session also - lets it ask that session to quit and kill it, so a dropped connection ends - its login with nobody else involved. sshd can end no other login. -- **No, it asks init.** Only init can end a login: sshd tells init that the - connection dropped, and init ends it. init does one more job, and stays the - only program that can end a login. - -*Recommended: yes, only its own.* - -## Q6a. Can one program ask another to quit? (stage 6) - -Today the only way one program can end another is a kill, which leaves it no -chance to save anything. The same answer decides how init asks each service to -finish when the machine shuts down -(`issues/isolation/the-supervisor-is-host-tested-and-owns-the-stop.md`). - -- **Yes, with a reason.** A program allowed to end another — its parent, or - the terminal it runs in — can ask it to quit and say why: interrupt - (Ctrl+C), hang-up (a window closed, a connection dropped) or terminate - (shutdown, a deadline). The program hears it among its other events and - decides what to do: an editor saves, a compiler deletes its half-written - output, Ctrl+C in a shell or an editor stops what it is doing and ends - nothing, and a server may take a hang-up as the signal to reload its - settings. The shell, the terminal and a login's session program listen, so - the programs under them are asked too. Whoever asked still kills if the end - it wants does not come. The kernel gains one call and one kind of object. -- **Yes, without a reason.** A server that reloads on a hang-up and exits on - terminate would exit when asked to reload, and no program could tell Ctrl+C - from shutdown. -- **No: a kill only, as today.** Nothing saves its work when its window closes - or the machine shuts down. - -*Recommended: yes, with a reason.* Rejected: a quit channel std and libc would -make at every start, which a program started any other way would not have; and -a handler that interrupts the program's own code wherever it is, the legacy -the owner ruled out. - -## Q6b. What does a quit do to a program that never listens for one? (stage 6) - -Most programs never listen: `ls`, `cat`, a Rust program that does not ask. - -- **It is killed at once, with everything it started**, as a Unix program - that handles nothing dies of Ctrl+C. Ctrl+C stops such a program on the - first press. The cost: a program that listens still dies at once, without - cleaning up, when the program that started it does not listen, because it - dies with that program; on Unix only the one that does not listen would die. - ToyOS's own shell and terminal and each login's session program listen, so - a compiler run from a login cleans up. -- **Nothing, until whoever asked gives up and kills it.** Ctrl+C does nothing - to `cat` until a second press, and shutdown waits out its whole deadline for - every program that never listens. - -*Recommended: killed at once.* - -## Q6c. How far does a quit reach? (stages 6 and 7) - -- **The program and everything it started**, as a kill does. Ctrl+C reaches - `make` and every compile it is running at once, as on Unix, where Ctrl+C - reaches every program of the job in front. A closing terminal asks its shell - and everything the shell started to hang up. The cost: a program cannot keep - a helper it started out of a Ctrl+C meant for itself, as a Unix program can - by putting the helper in a job of its own. Ninja does that with every - compile and passes the Ctrl+C on to them itself (v1.13.1 - `src/subprocess-posix.cc:102`, `:451-457`), so here each compile is asked - twice, and one that has not yet taken the first when Ninja's comes is - killed (Q7), leaving its half-written file. The shell and the session - program pass nothing on, since the quit already reached what they run. -- **The program alone.** Ctrl+C reaches `make` and not its compiles, and make, - which expects its compiles to have been interrupted with it, waits for every - running one to finish before it stops (GNU make's `fatal_error_signal`). The - shell and the session program must then pass each quit on to what they run, - or a closing terminal reaches its shell and not what the shell started. - -*Recommended: the program and everything it started.* - -## Q6d. Does a program that Ctrl+C ended read as interrupted? (stages 2 and 6) - -Under the answers recommended here, no child that Ctrl+C ended ever reads to -its parent as interrupted. One that never listened reads as killed; a C -program that lets the interrupt end it once it has cleaned up, as clang does, -reads as having exited with 130, which is a failure. - -- **No.** The number keeps meaning only whether a program ended itself or the - kernel ended it. A build tool that tells an interrupt from a failure, as - Ninja does, may report a compile the user interrupted as failed; Ninja still - stops the build, because the Ctrl+C reached it too. -- **Yes.** Q2's reasons gain three — interrupted, hung up, terminated — so a - program a quit ended reads as ended for that reason, and a C program that - started it sees what it would on Unix. For clang to read that way, a program - must be able to end itself for one of those reasons, as a Unix program can - by sending the signal to itself, so the number no longer says only what the - kernel did. - -*Recommended: no.* - -## Q6e. Does an unchanged Rust program's Ctrl+C handler run? (stage 6) - -Rust programs that handle Ctrl+C mostly do it through one widely used library, -`ctrlc`; rustc is one of them. - -- **Yes.** Rust's standard library hands a program its quit notice, and the - `ctrlc` fork ToyOS already carries listens to it, so a Rust program's handler - runs on a quit as it does on Unix, with no change to the program. rustc - stops leaving its handler out on ToyOS. -- **No.** A Rust program hears a quit only through a call written for ToyOS; - rustc and every other `ctrlc` user is treated as a program that never - listens (Q6b). - -*Recommended: yes.* - -## Q6f. Does a C program hear a quit as the Unix signal for it? (stage 6) - -- **Yes.** libc turns interrupt into `SIGINT`, hang-up into `SIGHUP` and - terminate into `SIGTERM`, so a C program written for Unix acts as it does - there: clang deletes the file it was writing and stops, a server reloads on - a hang-up, and a program that says only to ignore Ctrl+C runs on through - it. A C program's `kill` with one of those signals asks for a quit. -- **No.** A C program never hears a quit, and is treated as a program that - never listens (Q6b): clang leaves its half-written file behind, and a server - asked to reload ends. - -*Recommended: yes.* - -## Q7. What does a second Ctrl+C do? (stage 7) - -A program that listens may stop what it is doing on Ctrl+C and carry on: a -Python prompt goes back to its prompt, an editor cancels a command. - -- **It kills only a program that has not yet taken the first.** A program - that takes each Ctrl+C and carries on survives any number of them; one stuck - so that it never takes the first dies on the second. The same holds whoever - asks: a program asked to quit again before it took the last ask is killed. - Two quick presses: a C program, or a Rust program using `ctrlc`, takes each - on a thread of its own as it comes, so it survives them unless the machine - is too busy to run that thread between the two; a program that takes its - Ctrl+C in its own loop and is busy across both presses dies, where on Unix - its handler would have run twice. A program that takes a Ctrl+C and then - hangs is not ended by Ctrl+C: closing its window or a kill ends it, as on - Unix. -- **It always kills.** A hung program always dies on the second press, but a - Python prompt or an editor dies on the second Ctrl+C of its life, and loses - what was not saved. - -*Recommended: only a program that has not yet taken the first.* diff --git a/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md b/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md index 582c77b7baf..065269837c1 100644 --- a/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md +++ b/issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md @@ -144,15 +144,14 @@ times: written once as straight-line code. **Exit**: no interrupts-off window longer than a register access, and keyboard input keeps flowing while a stick misbehaves. -6. **The scheduler knows nothing about devices.** A handler posts its +6. A handler posts its device's `Watch` and ends its interrupt; the thread waiting on that watch does the work, and no step creates a kernel thread. `irq_ring`, the driver list in `drain_irqs` and the idle loop's device checks are gone by step 5. Each step measures the kernel's lines, and from step 2 the longest interrupts-off and preemption-off windows, against stage 6's first commit. Steps 3 and 4 do not land alone: they land with #592's i8042 stage and - with usbd. Stage 6 stays open past step 5 until the owner rules on the - panel the dump paints. + with usbd. 1. **Interrupts post.** A post is legal in a handler: the watches a handler posts, and the completions of a ring they complete into, sit behind interrupts-off locks nothing allocates or frees under, and every other @@ -182,9 +181,9 @@ times: them. **Exit**: step 10's. 5. **The pass is the scheduler's.** `drain_irqs` goes: the blocked-task dump and the heartbeat become `pass`'s own, and the TCO feed stays, - since what it proves is that passes run. The dump still paints its - report on the panel and holds it there, a device the pass reaches; - whether that stays is the owner's ruling. **Exit**: `drain_irqs` and the + since what it proves is that passes run. The dump keeps painting its + report on the panel and holding it there, a device the pass reaches + (owner, 2026-09-30). **Exit**: `drain_irqs` and the idle loop's device checks are gone, both windows are measured against stage 6's start, and the exits of `issues/kernel/an-irq-watchs-freeing-cancel-compiles-in-a-handler.md` diff --git a/kernel/Cargo.toml b/kernel/Cargo.toml index 12a6c3c57d1..38bf47fef07 100644 --- a/kernel/Cargo.toml +++ b/kernel/Cargo.toml @@ -340,10 +340,8 @@ heap-lockspin = ["pass-spin"] # The one profile every guest binary is built with. # Optimised, because an unoptimised guest mismeasures everything under TCG; -# debug-assertions and overflow-checks on, because fail-fast beats speed here -# and both crafted-ELF kernel panics in `issues/` were *found* by -# an overflow check. `--release` would turn the last two off silently, so this -# build system no longer has the flag. +# debug-assertions and overflow-checks on, because fail-fast beats speed here. +# `--release` would turn the last two off silently. [profile.toyos] inherits = "dev" opt-level = 2 diff --git a/src/build.rs b/src/build.rs index f37b0342292..c5ddf3a228a 100644 --- a/src/build.rs +++ b/src/build.rs @@ -368,10 +368,7 @@ fn config_crates(root: &Path, config: &SystemConfig) -> Vec { /// it in. /// /// One name, passed to every `cargo build` here and declared by every crate -/// root the image is made of. `--release` used to be a flag on `cargo run`, and -/// it silently turned `debug-assertions` and `overflow-checks` off — the two -/// knobs `issues/`'s crafted-ELF panics were *found* by. There is -/// no longer a second profile to pick, which is why there is no longer a flag. +/// root the image is made of. pub const PROFILE: &str = "toyos"; /// What every guest `cargo` and `rustc` here runs with: the toolchain directory @@ -598,10 +595,7 @@ fn contains_subslice(haystack: &[u8], needle: &[u8]) -> bool { /// /// [`PROFILE`] states them and `--release` is gone from this build system, so /// the way they can still be lost is somebody editing `[profile.toyos]`. This -/// asks the artifact rather than the manifest, which is the only question worth -/// asking: `issues/`'s two crafted-ELF kernel panics were both -/// *found* by an overflow check, and one of them had no configuration in which -/// it was an error return. +/// asks the artifact rather than the manifest. fn assert_overflow_checked(what: &str, image: &[u8]) { let found = contains_subslice(image, OVERFLOW_CHECK_MARKER); assert!( diff --git a/src/licence.rs b/src/licence.rs index 2c402c12d46..b19a2d51b52 100644 --- a/src/licence.rs +++ b/src/licence.rs @@ -436,8 +436,8 @@ pub const COMMITTED_FILES: &[(&str, &str, &str, Terms)] = &[ ( "doom.jpg", "ae22f71dc732580bd4f789937c9fe564969029413fc2092f27bdae8d1ceaf8e3", - "a screenshot of this system running doom, in README.md; \ - issues/build/doom-jpg-shows-ids-art-under-no-recorded-terms.md", + "the owner's own screenshot of this system running doom, in README.md; \ + no licence question applies", Terms::Spdx("NOASSERTION"), ), ( diff --git a/userland/Cargo.toml b/userland/Cargo.toml index 55b8b482f91..c2aa2e8f806 100644 --- a/userland/Cargo.toml +++ b/userland/Cargo.toml @@ -65,10 +65,8 @@ winit = { git = "https://github.com/ToyOSOrg/winit", branch = "toyos-0.30.13" } # The one profile every guest binary is built with. # Optimised, because an unoptimised guest mismeasures everything under TCG; -# debug-assertions and overflow-checks on, because fail-fast beats speed here -# and both crafted-ELF kernel panics in `issues/` were *found* by -# an overflow check. `--release` would turn the last two off silently, so this -# build system no longer has the flag. +# debug-assertions and overflow-checks on, because fail-fast beats speed here. +# `--release` would turn the last two off silently. # # What it replaces here: dependencies at 2 through `package."*"`, doom and the # compositor named individually at 2, and every other program at the dev