diff --git a/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md new file mode 100644 index 00000000000..252b15f9e63 --- /dev/null +++ b/issues/an-ipc-setup-rollback-panics-when-a-sibling-closed-its-new-handle.md @@ -0,0 +1,40 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# An IPC setup rollback panics when a sibling closed its new handle first + +`sys_inbox_setup` (`kernel/src/syscall/ipc.rs:444`) installs the inbox handle +under the process-data lock, releases that lock, and then runs `ctx.copy_out` +to a user address. When the copy fails the rollback closes the handle with +`ops::close(...).expect("the inbox this call installed a moment ago")` +(`:466`). `connect_through` (`:234`), which `sys_namespace_open` (`:219`) calls, +has the same shape: it installs the client handle, releases the lock, pushes +onto the port's pending queue, and on `QueueFull`/`Closed` rolls back with +`ops::close(...).expect("the connection this call installed a moment ago")` +(`:265`). + +The install is visible to sibling threads the instant the process-data lock is +dropped, and the fallible step (`copy_out`, or the queue push) runs with no +reservation held. A sibling thread of the same process, on another CPU, can +`SYS_CLOSE` that handle — or move it with `SYS_HANDLE_SEND` — between install +and rollback. The rollback's `HandleTable::remove` +(`kernel/src/object/handle.rs`) then returns `Stale`/an error, and the `expect` +turns a correctly detected concurrent close into a kernel panic. The comments' +premise — "the handle this call installed a moment ago" is still held — is not +an invariant across the dropped lock. Untrusted ordering panics the kernel +where it must refuse. + +**Traced, not executed.** No model compiles the handle table, and staging the +window needs an actuator the kernel does not have: one that holds the +installing thread between its install and its fallible step until a sibling +releases it. + +**Exit**: with a sibling closing (or sending) the just-installed handle while +`SYS_INBOX_SETUP` is in its failing-`copy_out` rollback, and while +`SYS_NAMESPACE_OPEN` is in its queue-full rollback, both syscalls return their +original refusal and the kernel does not panic. + +**Owner**: whoever holds `issues/a-panic-is-never-an-accident.md`. diff --git a/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md new file mode 100644 index 00000000000..017c71a92f8 --- /dev/null +++ b/issues/an-shm-map-racing-the-last-close-publishes-a-mapping-after-teardown.md @@ -0,0 +1,42 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# A shared-memory map racing the last close publishes a mapping after teardown + +`sys_shm_map` (`kernel/src/syscall/ipc.rs:361`) resolves the handle and clones +the object's `Arc` under the process-data lock, releases the lock, and then +calls `SharedMemObject::map_into` (`kernel/src/object/shm.rs:132`). `map_into` +never checks whether the object has retired; it allocates a writable mapping and +appends it to `mapped_in`. + +A cloned `Arc` held by a syscall is not a handle and does not postpone +zero-handle retirement. If a sibling closes the last handle after the clone and +before `map_into`, the object retires and `on_zero_handles` +(`kernel/src/object/shm.rs:172`) runs once: it takes `mapped_in`, finds it empty +and returns, and its queued `Arc` drops. When `map_into` then appends the new +mapping, nothing will ever tear it down — `on_zero_handles` has already run and +will not run again. + +With debug assertions on (the guest/test profile sets `debug-assertions = true` +in `kernel/Cargo.toml`), the final `Arc` drop trips the `debug_assert!` in +`SharedMemObject::drop` (`:186`) and panics the kernel. Without them, the +region's owned pages return to the PMM while the page-table entries stay +present, so the surviving process can read or write pages after another process +or the kernel receives them. This is distinct from the release-latency recorded +in `issues/deferred-release-outlives-its-syscall.md`: there memory is reclaimed +safely but late; here a mapping is created *after* the one teardown, so the +pages are freed while still mapped. + +**Traced, not executed.** No model compiles the object layer, and staging the +window needs an actuator the kernel does not have: one that holds `sys_shm_map` +between its clone and `map_into` until a sibling releases it. + +**Exit**: a `SYS_SHM_MAP` that races the last close of its handle either maps +nothing or is torn down with the object; no shm object's pages are present in +any page table after its zero-handle teardown has run, and the kernel does not +panic. + +**Owner**: the shared-memory object, `kernel/src/object/shm.rs`. diff --git a/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md new file mode 100644 index 00000000000..dfe35565140 --- /dev/null +++ b/issues/chdir-holds-the-vfs-lock-into-the-process-data-lock-open-takes-the-other-way.md @@ -0,0 +1,47 @@ +--- +status: open +kind: defect +opened: 2026-10-08 +--- + +# A successful chdir reverses the open path's VFS/process-data lock order + +`sys_chdir` (`kernel/src/syscall/fs.rs:84`) is written as + +```rust +match vfs::lock().cd(&cwd, path) { + Ok(new_cwd) => { process::with_process_data(|d| d.cwd = new_cwd); 0 } + Err(e) => e.to_u64(), +} +``` + +The `vfs::lock()` guard is a temporary in the match scrutinee. Under the kernel +crate's edition (2021, `kernel/Cargo.toml`), a scrutinee temporary lives to the +end of the whole `match`, so the VFS lock is still held through the `Ok` arm +while `with_process_data` takes the process-data lock +(`kernel/src/process.rs:853`). Order: **VFS then process-data.** + +`sys_open` (`kernel/src/syscall/fs.rs:35`) takes process-data via +`with_process_data` and calls `ops::open`, which takes `crate::vfs::lock()` +inside that closure (`kernel/src/object/ops.rs`). Order: **process-data then +VFS.** + +Two threads of one process (process-data is a per-process `Arc` shared by its +threads; VFS is global) running `SYS_CHDIR` and `SYS_OPEN` on two CPUs can each +hold the lock the other waits for. The ticket spinlock does not break the cycle; +its spin ceiling panics the kernel and unrelated VFS work stalls behind it. This +is a distinct inversion from the one recorded in +`issues/a-user-copy-demand-pages-under-whatever-its-caller-holds.md`. + +**Traced, not executed.** No model compiles these two syscalls, and staging the +window needs an actuator the kernel does not have: one that holds `sys_chdir` +in its `Ok` arm until a sibling releases it. + +**Exit**: `SYS_CHDIR` and `SYS_OPEN` can run concurrently in one process with no +lock-order inversion — `chdir` does not hold the VFS guard while it takes +process-data (the guard is dropped before the `Ok` arm takes the second lock). +A test that stages chdir past a successful lookup with its guard still alive, +then drives open to VFS acquisition under process-data, completes both without +deadlock. + +**Owner**: `sys_chdir`, `kernel/src/syscall/fs.rs`.