Skip to content

The ABI is free to change, and the kernel takes on only what userland cannot - #666

Merged
2 commits merged into
mainfrom
wt/toyos-abirule
Oct 1, 2026
Merged

2 commits merged into
mainfrom
wt/toyos-abirule

Conversation

@Japabu

@Japabu Japabu commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Two owner rulings of 2026-10-01, written into the prompts that govern agents. Root CLAUDE.md goes from 16077 to 16036 bytes.

The ABI is completely unstable

"We can change system calls. We can remove them, add them, change them. I want to have the cleanest, most sustainable ABI."

  • Root CLAUDE.md, "Syscall ABI": "completely unstable, read the code. Never add or change a syscall without discussion; a deleted syscall's number is retired, never reused." (138 characters) becomes "completely unstable. The cleanest, most sustainable ABI beats convenience; a removed number is free." (100). "read the code" goes because the Architecture section's header already says it. Workflow's "An ABI change lands with the work that needs it" stays.
  • implementer.md: the "unless the brief is an ABI brief" gate on toyos-abi/src, toyos/src and userland/libc/src is deleted. The brief's fence already bounds what an implementer touches.
  • reviewer.md, "What no gate reads": the BLOCKER for reusing a retired syscall, SYS_DEBUG action or inbox op number, or for declaring a retired ABI name (SharedToken, services::connect), is deleted. Root CLAUDE.md's Capabilities paragraph ("No registry, no connect-by-name, no pid-as-authority") still bans those designs under any name.
  • The code still keeps retired numbers: issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md.

The kernel takes on only what userland cannot

This comes from the symlink incident. libc's symlink needed to refuse an existing name. The implementer returned ENOSYS to avoid a kernel change, and the reviewer asked the kernel to refuse the name. Neither asked whether the kernel should own symlinks at all.

  • Root CLAUDE.md, Architecture "Kernel": "minimal; new additions are discussed and justified. Resource management, scheduling, process lifecycle, filesystem, device arbitration." (135 characters) becomes "takes on only what userland cannot. Resource management, scheduling, process lifecycle, device arbitration; files: see Capabilities." (132). "Discussed" no longer stands between a clean design and a new syscall. The filesystem leaves the job list because file systems are userland servers by design. The kernel still holds the machine-wide tree (kernel/src/vfs.rs, tmpfs.rs, SYS_SYMLINK), so the line points at the Capabilities paragraph, whose "Not yet true of files" sentence already records that. It adds no second record.
  • implementer.md, a paragraph under "The brief is a fence": before adding kernel behaviour, ask whether userland can own it. When the clean design changes the ABI, change the ABI. A clean design that reaches past the fence blocks the implementer, so this paragraph does not widen the fence. The rule itself is stated once, in root CLAUDE.md. There is no list of kernel jobs here.
  • reviewer.md, "Fit": one BLOCKER for a kernel addition that userland could own. Another for a design made worse to spare the ABI. Both halves of the incident are covered: a kernel refusal is a kernel addition, and the ENOSYS dodge is a worse design. A branch that fixes a defect in kernel code that will leave the kernel later is not blocked for keeping it.
  • "Userland" rather than "a userland server": the owner's loader move puts relocation in a Ring 3 loader inside the target's own address space. That is userland, but it is not a server.

Gates

  • cargo run -- --ci host on ec166f1b4: EXIT=0.

Unsure

  • "takes on only what userland cannot" leaves the verb to the reader ("cannot take on"). "Cannot own" did not fit the 135-character budget once the job list and the pointer were kept.

🤖 Generated with Claude Code

https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L

…el can do

Two owner rulings of 2026-10-01, written into the prompts that govern agents.

The ABI. "We can change system calls. We can remove them, add them, change
them. I want to have the cleanest, most sustainable ABI." Three lines said the
opposite and steered agents away from the ABI:

- Root CLAUDE.md, "Syscall ABI": "Never add or change a syscall without
  discussion; a deleted syscall's number is retired, never reused." It now says
  the cleanest, most sustainable ABI beats convenience and a removed number is
  free. "read the code" goes with it: the Architecture section's header already
  says it. The line is shorter than the one it replaces. Workflow's "An ABI
  change lands with the work that needs it" stays.
- implementer.md: "Never touch toyos-abi/src, toyos/src or
  userland/libc/src unless the brief is an ABI brief." Deleted. The brief's
  fence already bounds what an implementer touches.
- reviewer.md, "What no gate reads": a BLOCKER for reusing a retired syscall,
  SYS_DEBUG action or inbox op number, or declaring a retired ABI name.
  Deleted. Connect-by-name and pid-as-authority stay banned by root
  CLAUDE.md's Capabilities paragraph, which bans the design whatever it is
  called.

Kernel or server. In the symlink incident, libc's symlink needed "refuse an
existing name". The implementer avoided the kernel change and returned ENOSYS,
and the reviewer asked for the kernel to refuse the name. Neither asked whether
the kernel should own symlinks at all; fsd already resolves them. The owner:
"one less thing for the kernel to do… the reviewer and implementer should know
we try to use userland programs instead of the kernel."

- implementer.md: before adding or keeping kernel behaviour, ask whether a
  userland server can own it. When the clean design changes the ABI, change
  the ABI.
- reviewer.md, "Fit": a BLOCKER each for a kernel addition or a kept kernel
  path a userland server could own, and for a design made worse to spare the
  ABI.

The code still keeps retired numbers: retired_syscalls!, the "formerly …"
entries in toyos-abi, four test sites that use syscall 26 as their logged
refusal, and seven issue files that plan by retirement. That is outside this
branch's fence, so it is filed as
issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu
Japabu marked this pull request as ready for review October 1, 2026 07:44
@Japabu

Japabu commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Review of 066e1a028 against .claude/agents/reviewer.md. CI at that head: host completed, conclusion success (run 36832092333). Net: 4 files, +49 −10, all prose, no production or test line. Root CLAUDE.md goes from 16077 to 16071 bytes; the ABI span from 138 to 132 characters. Searched CLAUDE.md, */CLAUDE.md and .claude/agents/*.md: only CLAUDE.md:41 still steers a syscall toward discussion or a job toward the kernel.

BLOCKER

  • CLAUDE.md:41 — the Kernel line is left as it was, and it now contradicts both rulings. First, "new additions are discussed and justified" still puts a new syscall behind the discussion ruling 1 struck from line 47; the PR body's own Unsure says an agent can read it that way. Second, "Resource management, scheduling, process lifecycle, filesystem, device arbitration" makes the filesystem a kernel job, and the filesystem is the symlink incident's own domain. That disagrees with .claude/agents/implementer.md:17-19, while implementer.md:7 makes root CLAUDE.md the law, and the orchestrator who draws the fence reads only the root. This line must change in this PR: the "only what only a kernel can" rule replaces that 135-character span, written as a rule on what the kernel takes on and not as a snapshot, and no longer than the span. The present state needs no new record, because issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md and issues/kernel/every-driver-is-still-in-the-kernel.md already own storage and the drivers leaving the kernel.
  • .claude/agents/implementer.md:17-19 — "(address spaces, scheduling, handles and capabilities, interrupts, device arbitration)" is a closed list the owner never gave, and a second declaration of the kernel's scope beside CLAUDE.md:41. It already contradicts what the owner ruled the kernel keeps: ROOT's in-memory read path and exec (issues/kernel/the-kernel-is-small-interrupts-post-and-threads-wait.md:73), the panic console and its hotkey (same file, :76), and the record ring, console and panel (CLAUDE.md:45). Delete the parenthesis. The rule stays "only what only a kernel can", stated once.
  • .claude/agents/reviewer.md:45-46, .claude/agents/implementer.md:17 — "or a kernel path the branch keeps" and "or keeping" say more than ruling 2. Every branch that fixes a defect in the kernel VFS, block layer, NVMe, xHCI, HDA or virtio-gpu keeps a path those two tracks give to userland, so each becomes a BLOCKER outside its fence. That contradicts reviewer.md:48 ("Nothing outside the brief's fence") and the off-path filing rule at implementer.md:13-14. Both halves of the incident are caught without the clause: a kernel refusal is "a kernel addition", and the ENOSYS dodge is "a design made worse to spare the ABI". Delete "or a kernel path the branch keeps," and "or keeping".
  • .claude/agents/reviewer.md:46, .claude/agents/implementer.md:17 — "a userland server" is narrower than ruling 2's "userland programs". A kernel addition that code in the caller could own passes: the owner's own loader move puts relocation in "a Ring 3 loader inside the target's own address space" (issues/kernel/the-kernel-still-parses-what-userland-writes.md:19-20), and the storage ruling makes the VFS a client library. Replace "a userland server" with "userland" in both files.

NOTE

  • .claude/agents/implementer.md:19 — with the ABI-brief gate deleted, "change the ABI" says nothing about a clean design that lies outside the brief's fence, which is the incident's exact shape (libc inside the fence, fsd outside it). The fence wins and the implementer stops and reports. Say that here, so the paragraph is not read as widening the fence.
  • issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md — it names no owner. reviewer.md:82-83 requires "an owner, evidence and an exit condition", and its neighbours name one, for example issues/design-debt/toyos-abi-decodes-only-the-roster-header-not-its-entries.md:32.
  • issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md:15-19 — it misses the device-class retirement in the file it cites: toyos-abi/src/syscall.rs:1245 ("3 was Nic … Retired rather than reused for PciFunction") and its test toyos-abi/src/inventory.rs:590. Without it, the exit's "every retirement entry" cannot be checked against the list.

REMOVE

  • CLAUDE.md:47 — "add, change or remove syscalls;" — "completely unstable" already says it.
  • PR body — "Root CLAUDE.md is untouched apart from the ABI line. Its Architecture section already says the kernel is minimal." It is false once the first BLOCKER is fixed.
  • PR body — the "Filed, not fixed" bullet list and "All of this is outside this branch's fence." The issue file is the record; naming it is enough.
  • PR body — the PR The owner's rulings of 2026-09-30 written where they govern #645 bullet. It is about another branch and does not belong in main's record.
  • PR body — "(log at the job scratchpad, host2.log)". Nobody can open that path from main.

SEND BACK

Review of 066e1a0 on #666 sent the branch back.

- CLAUDE.md's Kernel line said new additions are "discussed" and named the
  filesystem a kernel job, against both rulings. It now says the kernel
  takes on only what userland cannot, keeps the job list without the
  filesystem, and points at the Capabilities paragraph, whose "Not yet true
  of files" sentence already records the machine-wide tree the kernel still
  holds. 132 characters replace 135.
- implementer.md's closed list of kernel jobs was a second declaration of
  the kernel's scope that the owner never gave; deleted, along with the
  restated rule, which root CLAUDE.md now carries.
- "or keeping" (implementer.md) and "or a kernel path the branch keeps"
  (reviewer.md) made every kernel-side defect fix a BLOCKER outside its
  fence; deleted. A kernel refusal is still "a kernel addition", and an
  ENOSYS dodge is still "a design made worse to spare the ABI".
- "a userland server" became "userland" in both files: the owner's loader
  move puts relocation in a Ring 3 loader in the target's own address space,
  which is userland but no server.
- implementer.md says a clean design that reaches past the fence blocks the
  implementer, so the ABI sentence does not widen the fence.
- The Syscall ABI line drops "add, change or remove syscalls", which
  "completely unstable" already says.
- issues/design-debt/the-abi-still-keeps-retired-syscall-numbers.md names an
  owner and lists the retired device classes 3 and 4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016t9wjdQkB8SH7bmfUoiy6L
@Japabu Japabu changed the title The ABI is free to change, and the kernel keeps only what only a kernel can do The ABI is free to change, and the kernel takes on only what userland cannot Oct 1, 2026
@github-merge-queue github-merge-queue Bot closed this pull request by merging all changes into main in 63cb34a Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant