From 0c9ff75e779c84a86ca3f737777130a4328809e9 Mon Sep 17 00:00:00 2001 From: japabu Date: Thu, 8 Oct 2026 09:57:08 +0200 Subject: [PATCH 1/3] Record four kernel trust-boundary defects an outside audit traced Re-traced each against the current tree and recorded it as an issues/ file with an exit a test can fail; no fix (each fix is its own later change). - SYS-01: sys_thread_spawn does not bound the thread entry, so a noncanonical entry faults the Ring 0 iretq and halt_all_cpus takes the machine. - SYS-02: SYS_INBOX_SETUP and SYS_NAMESPACE_OPEN install a handle, drop the process-data lock, and on a failed fallible step roll back with ops::close(...).expect(...); a sibling close makes that expect panic. - CON-01: sys_chdir holds the match-scrutinee VFS guard into with_process_data (VFS then process-data), the inverse of sys_open (process-data then VFS): a two-thread deadlock. - MEM-01: sys_shm_map clones the Arc and maps with no retirement check while on_zero_handles tears down exactly once, publishing a mapping after the only teardown. A red test cannot be committed: a clean-tree red breaks --ci host and the merge queue. The red tests are patches in the pull request; the three races need a new kernel actuator to stage, a separate change past this branch's fence. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A --- ...y-halts-the-kernel-at-its-ring-0-return.md | 33 +++++++++++++++ ...cs-when-a-sibling-closed-its-new-handle.md | 35 ++++++++++++++++ ...lose-publishes-a-mapping-after-teardown.md | 37 ++++++++++++++++ ...cess-data-lock-open-takes-the-other-way.md | 42 +++++++++++++++++++ 4 files changed, 147 insertions(+) create mode 100644 issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md create mode 100644 issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md create mode 100644 issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md create mode 100644 issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md diff --git a/issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md b/issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md new file mode 100644 index 00000000000..a0f67c63a78 --- /dev/null +++ b/issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md @@ -0,0 +1,33 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# A noncanonical thread entry halts the kernel at its Ring 0 return + +`sys_thread_spawn` (`kernel/src/syscall/proc.rs:132`) validates only +`stack_base <= stack_ptr`. The entry argument is never checked for the user +half or for canonical form. It is carried unchanged through +`process::spawn_thread` and `alloc_kernel_stack` into `r12` of `thread_start` +(`kernel/src/arch/x86_64/entry.rs`), which pushes it as the return `RIP` and +executes `iretq`. + +A noncanonical entry (noncanonical under 48-bit and 57-bit paging alike) makes +the `iretq` fault *before* the privilege transition completes, so the `#GP` +frame names the kernel code segment. `fatal_exception` +(`kernel/src/arch/x86_64/idt/exceptions.rs`, the `ring().is_user()` branch near +the end) ends only Ring 3 faults as the current process and calls +`halt_all_cpus()` for a Ring 0 fault. An unmapped but canonical entry faults in +Ring 3 and is survivable; the noncanonical case is not. One ordinary process +can therefore halt the machine — untrusted input reaching the kernel as a +crash, which the kernel must refuse. + +A kernel-half canonical entry (e.g. the direct-map region) is the same class: +the entry is used as a user `RIP` with no bound. + +**Exit**: a process that calls `SYS_THREAD_SPAWN` with a noncanonical entry, and +one with a canonical kernel-half entry, over a valid mapped stack, is refused at +the syscall or ends alone; a separate witness process keeps running and the +kernel neither halts nor resets. A guest test that spawns such a thread reds +today (the machine halts) and passes once the entry is bounds-checked. diff --git a/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md new file mode 100644 index 00000000000..ce940043661 --- /dev/null +++ b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md @@ -0,0 +1,35 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# An IPC setup rollback panics when a sibling closed its new handle first + +`sys_inbox_setup` (`kernel/src/syscall/ipc.rs:444`) installs the inbox handle +under the process-data lock, releases that lock, and then runs `ctx.copy_out` +to a user address. When the copy fails the rollback closes the handle with +`ops::close(...).expect("the inbox this call installed a moment ago")` +(`:466`). `sys_namespace_open` (`:219`) has the same shape: it installs the +client handle, releases the lock, pushes onto the port's pending queue, and on +`QueueFull`/`Closed` rolls back with +`ops::close(...).expect("the connection this call installed a moment ago")` +(`:265`). + +The install is visible to sibling threads the instant the process-data lock is +dropped, and the fallible step (`copy_out`, or the queue push) runs with no +reservation held. A sibling thread of the same process, on another CPU, can +`SYS_CLOSE` that handle — or move it with `SYS_HANDLE_SEND` — between install +and rollback. The rollback's `HandleTable::remove` +(`kernel/src/object/handle.rs`) then returns `Stale`/an error, and the `expect` +turns a correctly detected concurrent close into a kernel panic. The comments' +premise — "the handle this call installed a moment ago" is still held — is not +an invariant across the dropped lock. Untrusted ordering panics the kernel +where it must refuse. + +**Exit**: with a sibling closing (or sending) the just-installed handle while +`SYS_INBOX_SETUP` is in its failing-`copy_out` rollback, and while +`SYS_NAMESPACE_OPEN` is in its queue-full rollback, both syscalls return their +original refusal and the kernel does not panic. (Staging the window +deterministically needs a kernel actuator that pauses the installing thread +between install and the fallible step; the test is in the pull request body.) diff --git a/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md new file mode 100644 index 00000000000..5fff23b4534 --- /dev/null +++ b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md @@ -0,0 +1,37 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# A shared-memory map racing the last close publishes a mapping after teardown + +`sys_shm_map` (`kernel/src/syscall/ipc.rs:361`) resolves the handle and clones +the object's `Arc` under the process-data lock, releases the lock, and then +calls `SharedMemObject::map_into` (`kernel/src/object/shm.rs:132`). `map_into` +never checks whether the object has retired; it allocates a writable mapping and +appends it to `mapped_in`. + +A cloned `Arc` held by a syscall is not a handle and does not postpone +zero-handle retirement. If a sibling closes the last handle after the clone and +before `map_into`, the object retires and `on_zero_handles` +(`kernel/src/object/shm.rs:172`) runs once: it takes `mapped_in`, finds it empty +and returns, and its queued `Arc` drops. When `map_into` then appends the new +mapping, nothing will ever tear it down — `on_zero_handles` has already run and +will not run again. + +With debug assertions on (the guest/test profile sets `debug-assertions = true` +in `kernel/Cargo.toml`), the final `Arc` drop trips the `debug_assert!` in +`SharedMemObject::drop` (`:186`) and panics the kernel. Without them, the +region's owned pages return to the PMM while the page-table entries stay +present, so the surviving process can read or write pages after another process +or the kernel receives them. This is distinct from the release-latency recorded +in `issues/deferred-release-outlives-its-syscall.md`: there memory is reclaimed +safely but late; here a mapping is created *after* the one teardown, so the +pages are freed while still mapped. + +**Exit**: a `SYS_SHM_MAP` that races the last close of its handle either maps +nothing or is torn down with the object; no shm object's pages are present in +any page table after its zero-handle teardown has run, and the kernel does not +panic. (Deterministic staging needs a kernel actuator that pauses the mapping +thread after it clones the `Arc`; the test is in the pull request body.) diff --git a/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md new file mode 100644 index 00000000000..bfb68fdb606 --- /dev/null +++ b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md @@ -0,0 +1,42 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# A successful chdir reverses the open path's VFS/process-data lock order + +`sys_chdir` (`kernel/src/syscall/fs.rs:84`) is written as + +```rust +match vfs::lock().cd(&cwd, path) { + Ok(new_cwd) => { process::with_process_data(|d| d.cwd = new_cwd); 0 } + Err(e) => e.to_u64(), +} +``` + +The `vfs::lock()` guard is a temporary in the match scrutinee. Under the kernel +crate's edition (2021, `kernel/Cargo.toml`), a scrutinee temporary lives to the +end of the whole `match`, so the VFS lock is still held through the `Ok` arm +while `with_process_data` takes the process-data lock +(`kernel/src/process.rs:853`). Order: **VFS then process-data.** + +`sys_open` (`kernel/src/syscall/fs.rs:35`) takes process-data via +`with_process_data` and calls `ops::open`, which takes `crate::vfs::lock()` +inside that closure (`kernel/src/object/ops.rs`). Order: **process-data then +VFS.** + +Two threads of one process (process-data is a per-process `Arc` shared by its +threads; VFS is global) running `SYS_CHDIR` and `SYS_OPEN` on two CPUs can each +hold the lock the other waits for. The ticket spinlock does not break the cycle; +its spin ceiling panics the kernel and unrelated VFS work stalls behind it. This +is a distinct inversion from the one recorded in +`issues/a-user-copy-demand-pages-under-whatever-its-caller-holds.md`. + +**Exit**: `SYS_CHDIR` and `SYS_OPEN` can run concurrently in one process with no +lock-order inversion — `chdir` does not hold the VFS guard while it takes +process-data (the guard is dropped before the `Ok` arm takes the second lock). +A test that stages chdir past a successful lookup with its guard still alive, +then drives open to VFS acquisition under process-data, completes both without +deadlock. (Deterministic staging needs a kernel actuator; the test is in the +pull request body.) From f34839395a8f07369ca302ebfcac6e4b57f5e461 Mon Sep 17 00:00:00 2001 From: japabu Date: Thu, 8 Oct 2026 10:02:47 +0200 Subject: [PATCH 2/3] The records say what ran: a noncanonical thread entry ends its process under TCG, and three races are traced only The one sequential finding was run as a guest test under QEMU TCG. The spawn is accepted and the thread's fault arrives with a Ring 3 frame (cs=0x23), so the kernel ends that process alone and the machine lives: the halt the audit claims did not happen on this instrument. Intel's SDM puts the fault in IRET itself, in Ring 0, and only the T14 can be asked; the issue is renamed to what was measured, becomes a finding, and its exit is that metal row. The three interleavings say "traced, not executed" and name the actuator each needs. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A --- ...y-halts-the-kernel-at-its-ring-0-return.md | 33 ------------- ...anonical-one-is-measured-only-under-tcg.md | 46 +++++++++++++++++++ ...cs-when-a-sibling-closed-its-new-handle.md | 7 +-- ...lose-publishes-a-mapping-after-teardown.md | 5 +- ...cess-data-lock-open-takes-the-other-way.md | 5 +- 5 files changed, 56 insertions(+), 40 deletions(-) delete mode 100644 issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md create mode 100644 issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md diff --git a/issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md b/issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md deleted file mode 100644 index a0f67c63a78..00000000000 --- a/issues/a-noncanonical-thread-entry-halts-the-kernel-at-its-ring-0-return.md +++ /dev/null @@ -1,33 +0,0 @@ ---- -status: open -kind: defect -opened: 2026-10-08 ---- - -# A noncanonical thread entry halts the kernel at its Ring 0 return - -`sys_thread_spawn` (`kernel/src/syscall/proc.rs:132`) validates only -`stack_base <= stack_ptr`. The entry argument is never checked for the user -half or for canonical form. It is carried unchanged through -`process::spawn_thread` and `alloc_kernel_stack` into `r12` of `thread_start` -(`kernel/src/arch/x86_64/entry.rs`), which pushes it as the return `RIP` and -executes `iretq`. - -A noncanonical entry (noncanonical under 48-bit and 57-bit paging alike) makes -the `iretq` fault *before* the privilege transition completes, so the `#GP` -frame names the kernel code segment. `fatal_exception` -(`kernel/src/arch/x86_64/idt/exceptions.rs`, the `ring().is_user()` branch near -the end) ends only Ring 3 faults as the current process and calls -`halt_all_cpus()` for a Ring 0 fault. An unmapped but canonical entry faults in -Ring 3 and is survivable; the noncanonical case is not. One ordinary process -can therefore halt the machine — untrusted input reaching the kernel as a -crash, which the kernel must refuse. - -A kernel-half canonical entry (e.g. the direct-map region) is the same class: -the entry is used as a user `RIP` with no bound. - -**Exit**: a process that calls `SYS_THREAD_SPAWN` with a noncanonical entry, and -one with a canonical kernel-half entry, over a valid mapped stack, is refused at -the syscall or ends alone; a separate witness process keeps running and the -kernel neither halts nor resets. A guest test that spawns such a thread reds -today (the machine halts) and passes once the entry is bounds-checked. diff --git a/issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md b/issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md new file mode 100644 index 00000000000..4a33605ac3f --- /dev/null +++ b/issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md @@ -0,0 +1,46 @@ +--- +status: open +kind: finding +opened: 2026-10-08 +--- + +# A thread's entry is never checked, and what a noncanonical one does is measured only under TCG + +`sys_thread_spawn` (`kernel/src/syscall/proc.rs`) checks `stack_base <= +stack_ptr` and nothing else. The entry reaches `thread_start` +(`kernel/src/arch/x86_64/entry.rs`) unchanged and is the `RIP` its `iretq` +returns to, so what an entry that is no canonical address does is decided by +the CPU and not by the kernel. + +**Measured under QEMU TCG** (an x86-64 guest on an AArch64 host, `tests/testcases`, +two CPUs, at `b432ed21c`): a guest binary that spawns a thread at such an entry +over a mapped stack. The spawn was accepted, the fault was taken with a Ring 3 +frame, and the kernel ended that process alone: + +``` +FAULT rip=0x0100000000000000 cr2=0x0000000000000000 err=0x0000000000000000 ... tid=1 +SIGBUS tid=1: general protection fault (error_code=0x0) + cs=0x0023 ss=0x001b rflags=0x0000000000010202 +exit: test_rs_spawn_noncanonical_ pid=10 code=-1 cpu=9ms +===TEST_END test_rs_spawn_noncanonical_entry exit=-1=== +``` + +The machine went on and the harness finished its run. So the emulator completes +the return and faults at the fetch. + +**Not measured on hardware.** Intel's SDM (Volume 2, `IRET`) raises `#GP(0)` for +a return `RIP` that is not canonical, from the instruction itself — a Ring 0 +frame, which `fatal_exception` +(`kernel/src/arch/x86_64/idt/exceptions.rs`) answers with `halt_all_cpus()`, +since only a Ring 3 fault ends a process. If the T14, an Intel machine, does +what that document says, one process halts it. Nothing has executed that; TCG +cannot, and the T14 is the orchestrator's. + +The binary and its registration are a patch on the pull request that filed +this. + +**Exit**: a metal row runs that binary on the T14. If the machine halts this is +a `defect` whose exit is that `SYS_THREAD_SPAWN` refuses an entry outside the +canonical user half, by name, on every machine; if the process ends alone +there too, the one line this leaves is `sys_thread_spawn`'s doc comment saying +the entry is the CPU's to refuse. diff --git a/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md index ce940043661..8db971aa6d7 100644 --- a/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md +++ b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md @@ -30,6 +30,7 @@ where it must refuse. **Exit**: with a sibling closing (or sending) the just-installed handle while `SYS_INBOX_SETUP` is in its failing-`copy_out` rollback, and while `SYS_NAMESPACE_OPEN` is in its queue-full rollback, both syscalls return their -original refusal and the kernel does not panic. (Staging the window -deterministically needs a kernel actuator that pauses the installing thread -between install and the fallible step; the test is in the pull request body.) +original refusal and the kernel does not panic. **Traced, not executed.** No model compiles the handle table, and staging the +window needs an actuator the kernel does not have: one that holds the +installing thread between its install and its fallible step until a sibling +releases it. diff --git a/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md index 5fff23b4534..abdc2fe5570 100644 --- a/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md +++ b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md @@ -33,5 +33,6 @@ pages are freed while still mapped. **Exit**: a `SYS_SHM_MAP` that races the last close of its handle either maps nothing or is torn down with the object; no shm object's pages are present in any page table after its zero-handle teardown has run, and the kernel does not -panic. (Deterministic staging needs a kernel actuator that pauses the mapping -thread after it clones the `Arc`; the test is in the pull request body.) +panic. **Traced, not executed.** No model compiles the object layer, and staging the +window needs an actuator the kernel does not have: one that holds `sys_shm_map` +between its clone and `map_into` until a sibling releases it. diff --git a/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md index bfb68fdb606..2dfed975732 100644 --- a/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md +++ b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md @@ -38,5 +38,6 @@ lock-order inversion — `chdir` does not hold the VFS guard while it takes process-data (the guard is dropped before the `Ok` arm takes the second lock). A test that stages chdir past a successful lookup with its guard still alive, then drives open to VFS acquisition under process-data, completes both without -deadlock. (Deterministic staging needs a kernel actuator; the test is in the -pull request body.) +deadlock. **Traced, not executed.** No model compiles these two syscalls, and staging the +window needs an actuator the kernel does not have: one that holds `sys_chdir` +in its `Ok` arm until a sibling releases it. From e34a45e99baa5895a68fa1c72f32030d497c5f63 Mon Sep 17 00:00:00 2001 From: japabu Date: Thu, 8 Oct 2026 10:26:49 +0200 Subject: [PATCH 3/3] Answer review round 1: three race records with owners, and the thread-entry record leaves for the branch that fixes it - The thread-entry record is deleted here. `wt/toyos-threadentry` carries it in its first commit and deletes it with the refusal it asks for, so this branch records the three races and nothing else. - Each of the three names an owner. - The IPC record gives the install, the push and the rollback to `connect_through`, which `sys_namespace_open` calls. - "Traced, not executed." is a paragraph of its own before each exit. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A --- ...anonical-one-is-measured-only-under-tcg.md | 46 ------------------- ...cs-when-a-sibling-closed-its-new-handle.md | 18 +++++--- ...lose-publishes-a-mapping-after-teardown.md | 10 ++-- ...cess-data-lock-open-takes-the-other-way.md | 10 ++-- 4 files changed, 25 insertions(+), 59 deletions(-) delete mode 100644 issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md diff --git a/issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md b/issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md deleted file mode 100644 index 4a33605ac3f..00000000000 --- a/issues/a-thread-entry-is-unchecked-and-a-noncanonical-one-is-measured-only-under-tcg.md +++ /dev/null @@ -1,46 +0,0 @@ ---- -status: open -kind: finding -opened: 2026-10-08 ---- - -# A thread's entry is never checked, and what a noncanonical one does is measured only under TCG - -`sys_thread_spawn` (`kernel/src/syscall/proc.rs`) checks `stack_base <= -stack_ptr` and nothing else. The entry reaches `thread_start` -(`kernel/src/arch/x86_64/entry.rs`) unchanged and is the `RIP` its `iretq` -returns to, so what an entry that is no canonical address does is decided by -the CPU and not by the kernel. - -**Measured under QEMU TCG** (an x86-64 guest on an AArch64 host, `tests/testcases`, -two CPUs, at `b432ed21c`): a guest binary that spawns a thread at such an entry -over a mapped stack. The spawn was accepted, the fault was taken with a Ring 3 -frame, and the kernel ended that process alone: - -``` -FAULT rip=0x0100000000000000 cr2=0x0000000000000000 err=0x0000000000000000 ... tid=1 -SIGBUS tid=1: general protection fault (error_code=0x0) - cs=0x0023 ss=0x001b rflags=0x0000000000010202 -exit: test_rs_spawn_noncanonical_ pid=10 code=-1 cpu=9ms -===TEST_END test_rs_spawn_noncanonical_entry exit=-1=== -``` - -The machine went on and the harness finished its run. So the emulator completes -the return and faults at the fetch. - -**Not measured on hardware.** Intel's SDM (Volume 2, `IRET`) raises `#GP(0)` for -a return `RIP` that is not canonical, from the instruction itself — a Ring 0 -frame, which `fatal_exception` -(`kernel/src/arch/x86_64/idt/exceptions.rs`) answers with `halt_all_cpus()`, -since only a Ring 3 fault ends a process. If the T14, an Intel machine, does -what that document says, one process halts it. Nothing has executed that; TCG -cannot, and the T14 is the orchestrator's. - -The binary and its registration are a patch on the pull request that filed -this. - -**Exit**: a metal row runs that binary on the T14. If the machine halts this is -a `defect` whose exit is that `SYS_THREAD_SPAWN` refuses an entry outside the -canonical user half, by name, on every machine; if the process ends alone -there too, the one line this leaves is `sys_thread_spawn`'s doc comment saying -the entry is the CPU's to refuse. diff --git a/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md index 8db971aa6d7..252b15f9e63 100644 --- a/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md +++ b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md @@ -10,9 +10,9 @@ opened: 2026-10-08 under the process-data lock, releases that lock, and then runs `ctx.copy_out` to a user address. When the copy fails the rollback closes the handle with `ops::close(...).expect("the inbox this call installed a moment ago")` -(`:466`). `sys_namespace_open` (`:219`) has the same shape: it installs the -client handle, releases the lock, pushes onto the port's pending queue, and on -`QueueFull`/`Closed` rolls back with +(`:466`). `connect_through` (`:234`), which `sys_namespace_open` (`:219`) calls, +has the same shape: it installs the client handle, releases the lock, pushes +onto the port's pending queue, and on `QueueFull`/`Closed` rolls back with `ops::close(...).expect("the connection this call installed a moment ago")` (`:265`). @@ -27,10 +27,14 @@ premise — "the handle this call installed a moment ago" is still held — is n an invariant across the dropped lock. Untrusted ordering panics the kernel where it must refuse. -**Exit**: with a sibling closing (or sending) the just-installed handle while -`SYS_INBOX_SETUP` is in its failing-`copy_out` rollback, and while -`SYS_NAMESPACE_OPEN` is in its queue-full rollback, both syscalls return their -original refusal and the kernel does not panic. **Traced, not executed.** No model compiles the handle table, and staging the +**Traced, not executed.** No model compiles the handle table, and staging the window needs an actuator the kernel does not have: one that holds the installing thread between its install and its fallible step until a sibling releases it. + +**Exit**: with a sibling closing (or sending) the just-installed handle while +`SYS_INBOX_SETUP` is in its failing-`copy_out` rollback, and while +`SYS_NAMESPACE_OPEN` is in its queue-full rollback, both syscalls return their +original refusal and the kernel does not panic. + +**Owner**: whoever holds `issues/a-panic-is-never-an-accident.md`. diff --git a/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md index abdc2fe5570..017c71a92f8 100644 --- a/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md +++ b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md @@ -30,9 +30,13 @@ in `issues/deferred-release-outlives-its-syscall.md`: there memory is reclaimed safely but late; here a mapping is created *after* the one teardown, so the pages are freed while still mapped. +**Traced, not executed.** No model compiles the object layer, and staging the +window needs an actuator the kernel does not have: one that holds `sys_shm_map` +between its clone and `map_into` until a sibling releases it. + **Exit**: a `SYS_SHM_MAP` that races the last close of its handle either maps nothing or is torn down with the object; no shm object's pages are present in any page table after its zero-handle teardown has run, and the kernel does not -panic. **Traced, not executed.** No model compiles the object layer, and staging the -window needs an actuator the kernel does not have: one that holds `sys_shm_map` -between its clone and `map_into` until a sibling releases it. +panic. + +**Owner**: the shared-memory object, `kernel/src/object/shm.rs`. diff --git a/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md index 2dfed975732..dfe35565140 100644 --- a/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md +++ b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md @@ -33,11 +33,15 @@ its spin ceiling panics the kernel and unrelated VFS work stalls behind it. This is a distinct inversion from the one recorded in `issues/a-user-copy-demand-pages-under-whatever-its-caller-holds.md`. +**Traced, not executed.** No model compiles these two syscalls, and staging the +window needs an actuator the kernel does not have: one that holds `sys_chdir` +in its `Ok` arm until a sibling releases it. + **Exit**: `SYS_CHDIR` and `SYS_OPEN` can run concurrently in one process with no lock-order inversion — `chdir` does not hold the VFS guard while it takes process-data (the guard is dropped before the `Ok` arm takes the second lock). A test that stages chdir past a successful lookup with its guard still alive, then drives open to VFS acquisition under process-data, completes both without -deadlock. **Traced, not executed.** No model compiles these two syscalls, and staging the -window needs an actuator the kernel does not have: one that holds `sys_chdir` -in its `Ok` arm until a sibling releases it. +deadlock. + +**Owner**: `sys_chdir`, `kernel/src/syscall/fs.rs`.