Skip to content

A claim keeps where its BARs are and no object over them: a BAR asked for again after its handle closed no longer panics the kernel, and two answers held at once map apart - #761

Merged
Japabu merged 7 commits into
mainfrom
wt/toyos-barremap
Oct 8, 2026

Conversation

@Japabu

@Japabu Japabu commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

The defect

A process holding a live PCI claim panicked the kernel with three ordinary calls: SYS_DEVICE_BAR_MAP, close the handle, SYS_DEVICE_BAR_MAP again.

PANIC: panicked at src/object/handle.rs:108:9:
a handle to a retired SharedMem (koid 197)

pcidev::bar_object made one SharedMemObject per memory BAR on the first ask and kept it in the claim's Bound (bars) to answer every later ask with the same object. A kernel object's life is its handles': the last one to close retires it (HandleEntry's drop), and HandleEntry::new asserts that nobody installs a retired object again. The claim outlived the object it cached, so the second ask handed the retired object to ops::install. An ask refused for a full handle table retired the cached object the same way with no close at all: the entry built for it is dropped.

The fix, and what the second ask answers

The object's lifetime belongs to each ask, not to the claim. The claim keeps what it already kept, bar_at and bar_bytes, and bar_object makes a fresh SharedMemObject over that window every time; Bound::bars is gone. The panic site is untouched: its assertion was right.

The second ask answers a fresh mapping, not a refusal.

  • The tree already states this rule for the other long-lived owner of device memory: device::Screen "stores Regions rather than SharedMemObjects because each claim mints its own object, and a SharedMemObject is retired once its handle count reaches zero". The BAR cache was the one site that broke it.
  • The claim handle is the authority over the window (pci_slot: "the handle is the authority and the slot is what it names"). A SharedMem handle derived from it is closed without spending that authority, as a SYS_DEVICE_DMA_ALLOC after a closed grant handle is still answered.
  • A refusal needs a per-BAR "was asked" bit kept only to say no, and leaves a driver that unmapped its registers no way back but dropping the claim, which resets the function.
  • Keeping one object and reviving it was not an option either: asking whether the cached object is retired and then installing it is a check-then-act race, since a SharedMem handle carries TRANSFER and its last close can run under another process's lock.

What changes for a caller: two asks while the first handle is still held are now two objects and two mappings of the same uncacheable window, where they were two handles to one object. The holder already drives those registers, and each object costs it a handle slot. No caller in the tree asks twice (netstack's two drivers and diskserver's NVMe each call map_bar once). The ABI's doc comment said "Idempotent per BAR: a second call answers the same object" and now says what is true; a comment under toyos-abi/src changes no identity.

Production: kernel/ and toyos-abi/ are 11 insertions, 13 deletions.

The same shape, checked

Every site that turns an object into a handle (HandleEntry::new, ops::install, install_all), and every kernel-side holder of an Arc to a handle-counted object:

site what it installs verdict
syscall/device.rs sys_device_bar_map the claim's cached BAR object the defect, fixed
syscall/device.rs sys_device_dma_alloc dma_alloc's fresh region; the claim's Grant::memory keeps the Arc for the pages and never hands it out again not the shape
pcidev::dma_map installs nothing; holds a region resolved from a live handle not the shape
object/device.rs DeviceInfo::mint (framebuffer, HDA, virtio-sound buffers) objects held in the claim's description, minted once per claim behind described.bytes; install_all reserves room before any HandleEntry::new, so a refusal retires nothing; remint is given fresh objects from device::framebuffer_info not the shape
device::Screen, drivers::hda::info, drivers::virtio_sound::info Regions, a fresh object per claim not the shape
object/namespace.rs holds Arc<Connector>s and never installs one: sys_namespace_open makes a new Connection not the shape
process.rs ProcessEntry::object installed only by loader/start.rs, as the object's first handle and a second made while the first is held not the shape
inbox/mod.rs Completions::shm, file_backing.rs SharedImage::object never installed not the shape
syscall/ipc.rs (pipe, join, port create, port mint, namespace build, connect, accept, shm create, inbox setup), inbox/mod.rs accept, loader/mod.rs supervisor's console and SysCap, loader/start.rs minted consoles fresh objects not the shape
sys_handle_recv, HandleTable::transfer/duplicate live entries moved or duplicated under the table's lock not the shape
interrupt objects none exist: a claim's interrupt is a static Interrupt and IrqWatch per slot, read through the claim handle not the shape

Found on the way and filed, not fixed: issues/a-bars-mapping-outlives-the-claim-it-was-asked-on.md (read from the code, not run). It was true before this change and is not made worse by it.

The test

bar_map_again is restored from ef41f0ad1 and strengthened to the decision: it no longer accepts a refusal. The job claims QEMU's virtio NIC on tests/testcases and reads the function's device_feature dword through each answer's mapping:

  1. ask twice and hold both answers: their addresses differ, both read the same non-degenerate dword; drop the earlier, and the later still reads it. This is the one kernel state the fix newly makes reachable — two live objects over one window in one address space, and one's zero-handle hook running while the other is mapped.
  2. ask, map, read, close; ask again, map, read: the same dword.
  3. fill the handle table, ask: ResourceExhausted; empty it, ask again, map, read: the same dword. This is the arm the issue had read and not run.

Why a guest test: the behaviour is a syscall sequence against a live claim on a real PCI function: a handle table, an object's retirement through the deferred queue, a BAR window mapped into a process. No type reaches it without changing the object model (HandleEntry::new taking a not-yet-retired witness is a larger design than this defect), the kernel library's host tests have no handle table over a claimed function, and the T14's one claimable function is its I219, where a red is the bench's machine in a panic.

Checks (high-risk: devices, a capability boundary)

The head is 00b7a0c74: 8a6dc0be3 with origin/main at fb00abfa5 merged in (#749, #757, #759, #762; a clean merge, #762's hunks in kernel/src/object/mod.rs and toyos-abi/src/syscall.rs disjoint from this branch's). Every row with an exit is measured at that head, by one script the orchestrator ran alone and serially from a clean tree; exit codes are the commands' own. Logs under the orchestrator's scratchpad, orch/barremap-r2/, their deciding lines in the round 2 comment below; the three control patches are byte for byte the ones in the round 1 comment, which apply to the merged head unchanged. Each control therefore reverts this branch's whole kernel change onto the base the green arm ran on, #762 included.

check command exit
the test, green cargo test --test toyos-build -- bar_map_again (green.log) 0, all three arms' lines read
negative control, two answers held: the whole kernel change reverted (control-kernel.patch), the job as it stands same command (control-held.log) 1, the job's assert_ne! at bar_map_again.rs:63: two answers held at once are one mapping, both 0xfffec00000; no kernel panic
negative control 1: the same reverted kernel, and the held-together arm removed (control-heldarm.patch) so the ask after a close is reached same command (control-1.log) 1, PANIC: panicked at src/object/handle.rs:108:9: a handle to a retired SharedMem (koid 197), through ops::install from sys_device_bar_map
negative control 2: on top of control 1, the close-then-ask arm removed (control-arm2.patch) so the first ask is the one a full table refuses same command (control-2.log) 1, the same panic, same backtrace
the other guest test the change reaches (netstack maps the virtio NIC's BAR) cargo test --test toyos-build -- iommu_virtio_platform (guest-iommu.log) 0
host suites cargo run -- --ci host CI run 37762030186 at 00b7a0c74: host success (job 113260522116), [ci] Host: 78 step(s), all green
the whole guest suite cargo run -- --ci guest the same run: guest / suite success (job 113261645563), test result: ok. 32 passed, 32 total, PASS bar_map_again, [ci] Guest: 5 step(s), all green; toolchain / build success (113260522448)

The script applied each control with git apply --check, stacked, ran the test, and reversed each with git apply -R (each exit 0; git status --porcelain --ignore-submodules=none empty after). The first comment's patches were against 72a18fa81 and are superseded.

Where this meets #762

#762 made a region's list of mappings a Held its zero-handle hook takes for good, so a retired region can never be mapped again and must never be handed out twice. pcidev::bar_object makes an object per ask and Bound keeps none; sys_device_bar_map is its only caller. The two fixes do not replace each other: the controls above run #762's kernel with this branch's change reverted, and still panic.

What a green bar_map_again shows of the two together: two objects over one frame in one address space map apart and both read the device; #762's hook (take().expect("the zero-handle hook runs once")) runs on the earlier object while the later is mapped, and on an object that was never mapped, the one the full table refused; and no object reaches SharedMemObject's drop with a mapping left (debug_assert!, on in the kernel's one profile). Any of those failing is a kernel panic, which the harness reads as red.

Independent oracle: the recorded real failure. The audit measured this panic at 06bbc236d with a test it wrote before any fix existed, and control 1 reproduces it message for message, koid included. The value read through each mapping is the device's own register, and the kernel's placement probe reads the same dword (toyos_pci::probe::reference) to prove the BAR decodes where it was put.

The issue

issues/a-claims-bar-asked-for-again-after-its-handle-closed-panics-the-kernel.md is deleted: its exit is met on both arms it names. Nothing else in the tree cited its slug (git grep at the head). Its durable rule is one sentence in kernel/src/object/mod.rs's header: an object ends with its last handle and is never handed out again, so a holder that answers more than once keeps what the object is made over. (DeviceInfo keeps objects and is sound because it answers once per claim.)

issues/a-bars-mapping-outlives-the-claim-it-was-asked-on.md says what this fix changes for its exit: the claim holds no reference to an object it minted, so taking a window back needs the claim to learn what it handed out.

Not run, and unsure of

  • No metal row. The T14 rows that reach bar_object's first ask are the ones that boot netstack on the I219 (userland/netstack/src/i219.rs): lan_lease_report, lan_dhcp_lease and lan_message_delivery. That path's behaviour is unchanged by reading (one object over the same region either way); it is not measured on the T14 here.
  • cargo run -- --ci host and the whole guest suite were not run locally; both are read from CI run 37762030186 at this head (the rows above). That run merged this head into main as it was before One workspace, one lock: the root, the kernel, the loader, userland and the SDK resolve together #746 landed; the merge queue runs both again on the tree main becomes.
  • Control 2's job output does not survive the panic, so that the refused ask is the one that retired the object is what its patch leaves possible, not a line in the log.
  • That the later mapping survives the earlier's unmap is not observed on every interleaving. The read through the later answer follows the close of the earlier one, whose hook runs at that syscall's exit (syscall/dispatch.rs) unless another CPU's drain took the batch first (drain_zero_handles), in which case the read can precede the unmap; the job does not wait on that, and has no event to wait on. What it rests on by reading: the list the hook takes is its own object's field (self.mapped_in.take()), free_and_unmap goes by virtual address, and a BAR's region has pages: None, so no frame is freed under the other window.
  • A map of a BAR object through an Arc that outlived its last handle answering Gone is A shared region's last handle takes its list of mappings, and a map that finds none is refused #762's, held by its type; this job does not reach it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A

Japabu and others added 3 commits October 8, 2026 10:18
… red on a kernel panic

A guest job claims QEMU's virtio NIC on tests/testcases, asks for one of its
memory BARs, closes the handle it is given and asks for the same BAR again on
the same live claim. The kernel owes it a handle or a refusal. At b432ed2 it
panics instead, `a handle to a retired SharedMem`, at
kernel/src/object/handle.rs's constructor: the claim's cached BAR object was
retired when its only handle closed, and the second request installs it again.

`cargo test --test toyos-build -- bar_map_again` exits 1 on that panic. The
test is red by design and is deleted by a later commit of this branch, with
its issue naming the revert that restores it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
… asked for as often as its holder likes

`pcidev::bar_object` made one `SharedMemObject` per memory BAR on the first
`SYS_DEVICE_BAR_MAP` and kept it in the claim's `Bound`, to answer every later
ask with "the same object". An object's life is its handles': the last one to
close retires it, and `HandleEntry::new` asserts nobody installs it again. The
claim outlived that, so ask, close, ask handed the retired object to
`ops::install` and panicked the kernel from an unprivileged holder of a live
claim. An ask a full handle table refused retired the object the same way
without a close.

The claim now keeps only what it already kept, `bar_at` and `bar_bytes`, and
every ask makes an object of its own over that window, the rule
`device::Screen` already states for the framebuffer's buffers. The second ask
answers a fresh mapping and not a refusal: the claim handle is the authority
over the window, a `SharedMem` handle derived from it is closed without
spending that authority, and a refusal would be a per-BAR "was asked" bit kept
only to say no. Two live objects over one window are two uncacheable mappings
of registers the holder already drives.

`bar_map_again` comes back and now reads the function's `device_feature`
through each answer's mapping: after a close, and after an ask refused at a
full table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
Found reading `pcidev::tear_down` for the BAR object's owner: read, not run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Negative-control patches, applied onto 72a18fa81 with git apply --check, run with cargo test --test toyos-build -- bar_map_again, and restored with git checkout -- kernel toyos-abi tests (git status --porcelain empty after).

Control 1 — the whole kernel change reverted (control-kernel.patch). Exit 1, PANIC: panicked at src/object/handle.rs:108:9: a handle to a retired SharedMem (koid 197).

diff --git a/kernel/src/object/mod.rs b/kernel/src/object/mod.rs
index 50f57f508..205029806 100644
--- a/kernel/src/object/mod.rs
+++ b/kernel/src/object/mod.rs
@@ -2,8 +2,6 @@
 //!
 //! Objects are plain `Arc<T>`: no custom refcounting, no `Weak`, no `dyn` hierarchy.
 //!
-//! An object ends with its last handle and is never handed out again: a holder that outlives the handles it answers with keeps what an object is made over (`device::Screen`, a claim's BAR windows) and makes a fresh object for each.
-//!
 //! `handle_count`, not the Arc strong count, is what userland-visible lifecycle rides: a syscall's `Arc` can be stranded on a killed thread's kernel stack, so release is deferred through [`ZERO_QUEUE`] — see `issues/deferred-release-outlives-its-syscall.md`.
 
 // `warn` here gates via CI's `-D warnings`; the rest of the kernel is not yet swept.
diff --git a/kernel/src/pcidev/mod.rs b/kernel/src/pcidev/mod.rs
index f6de66efc..1590f7a85 100644
--- a/kernel/src/pcidev/mod.rs
+++ b/kernel/src/pcidev/mod.rs
@@ -238,6 +238,9 @@ struct Bound {
     /// advertises; 0 bytes is a slot with no BAR this claim may map.
     bar_at: [u64; BARS],
     bar_bytes: [u64; BARS],
+    /// Minted on the first map of that BAR, so a second answers the same
+    /// object rather than a second handle to one window.
+    bars: [Option<Arc<SharedMemObject>>; BARS],
     grants: Vec<Grant>,
     /// The [`RESIDUE`] this claim took over and has placed no grant at: mapped
     /// to nothing and counted against nothing, and where a grant of a range's
@@ -804,6 +807,7 @@ fn bring_up(pci: PciDevice, id: PciId, slot: usize) -> Result<Bound, Refusal> {
                 id,
                 bar_at,
                 bar_bytes,
+                bars: [const { None }; BARS],
                 grants: Vec::new(),
                 residue: Vec::new(),
                 mastering: false,
@@ -1493,25 +1497,26 @@ fn with_bound<T>(
     f(bound)
 }
 
-/// One memory BAR as an object to map, a fresh one for every ask.
-///
-/// **The claim keeps where the window is and never an object over it**: an
-/// object ends with its last handle, which its holder closes whenever it
-/// likes, and the claim outlives that.
+/// One memory BAR as an object to map; the same object every time.
 pub fn bar_object(slot: usize, index: u64) -> Result<Arc<SharedMemObject>, SyscallError> {
     with_bound(slot, |bound| {
         let index = usize::try_from(index).map_err(|_| SyscallError::InvalidArgument)?;
         if index >= BARS || bound.bar_bytes[index] == 0 {
             return Err(SyscallError::InvalidArgument);
         }
-        Ok(SharedMemObject::over(Region {
+        if let Some(object) = bound.bars[index].as_ref() {
+            return Ok(Arc::clone(object));
+        }
+        let object = SharedMemObject::over(Region {
             phys: DirectMap::from_phys(bound.bar_at[index]),
             size: align_2m(bound.bar_bytes[index] as usize) as u64,
             // Registers, and the memory type is not firmware's to decide.
             cache: CachePolicy::Uncacheable,
             // The kernel owns no pages here: this is a device's aperture.
             pages: None,
-        }))
+        });
+        bound.bars[index] = Some(Arc::clone(&object));
+        Ok(object)
     })
 }
 
diff --git a/toyos-abi/src/syscall.rs b/toyos-abi/src/syscall.rs
index 53ca1e35f..3e3260d78 100644
--- a/toyos-abi/src/syscall.rs
+++ b/toyos-abi/src/syscall.rs
@@ -2312,8 +2312,7 @@ pub fn pipe_map(handle: RawHandle) -> Result<*mut u8, SyscallError> {
 /// kernel's, and a process that could write it could aim the device's interrupt
 /// at any address the LAPIC decodes.
 ///
-/// Every call answers an object of its own over the same window, so a BAR
-/// whose handle was closed can be asked for again.
+/// Idempotent per BAR: a second call answers the same object.
 ///
 /// A window alone drives nothing: the function masters the bus from its first
 /// [`device_dma_alloc`] and not before.

Control 2 — applied on top of control 1: the job's close-then-ask arm removed, so the first ask is the one a full handle table refuses (control-arm2.patch). Exit 1, the same panic.

diff --git a/tests/toyos-rust-tests/src/bin/bar_map_again.rs b/tests/toyos-rust-tests/src/bin/bar_map_again.rs
index 5823bed50..a92c76f64 100644
--- a/tests/toyos-rust-tests/src/bin/bar_map_again.rs
+++ b/tests/toyos-rust-tests/src/bin/bar_map_again.rs
@@ -50,10 +50,7 @@ fn main() {
     let bytes = info.bar_bytes[bar];
     let bar = bar as u32;
 
-    let first = ask(&nic, bar, bytes);
     println!("bar_map_again: BAR {bar} asked for and its handle closed; asking again");
-    let again = ask(&nic, bar, bytes);
-    assert_eq!(again, first, "bar_map_again: the second ask's mapping is another window");
     println!("bar_map_again: answered with a handle that maps");
 
     // No room for the handle: the ask is refused, and the object made for it
@@ -78,7 +75,6 @@ fn main() {
         "bar_map_again: an ask from a full table was not refused for room"
     );
     println!("bar_map_again: an ask from a full handle table refused; asking again");
-    let after = ask(&nic, bar, bytes);
-    assert_eq!(after, first, "bar_map_again: the ask after the refusal maps another window");
+    let _ = ask(&nic, bar, bytes);
     println!("bar_map_again: answered after the refusal with a handle that maps");
 }

Japabu and others added 2 commits October 8, 2026 10:54
…he kernel

Its exit is met. A BAR asked for again on a live claim, after every handle to
an earlier answer has gone, is answered with a handle that maps, and
`bar_map_again` is green and reads both arms: the close, and the install a
full handle table refused, which the issue had read and not run. With the
kernel change reverted the test is red on the recorded panic, `a handle to a
retired SharedMem (koid 197)` at `src/object/handle.rs:108:9`, on either arm.

The rule it carried is in `kernel/src/object/mod.rs`'s header: an object ends
with its last handle and is never handed out again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 1 of wt/toyos-barremap at 6f05a4941 against origin/main 2e781a49d. Read only: no test or build was run by this review.

Net lines (git diff --shortstat origin/main...6f05a4941: 7 files, +169 -57). Production (kernel/, toyos-abi/): +11 -13. Tests (tests/toyos.rs, the job): +122 -0. issues/: +36 -44.

BLOCKER

  • tests/toyos-rust-tests/src/bin/bar_map_again.rs:53 — no arm ever holds two answers at once, and that is the one kernel state this change newly makes reachable — the body says so itself ("two asks while the first handle is still held are now two objects and two mappings of the same uncacheable window") and toyos-abi/src/syscall.rs:2315 now promises it ("Every call answers an object of its own"). Every ask drops its SharedMemory before the next begins, so the job is green on any kernel that answers sequential asks, whatever two live windows over one device frame do in one address space and whatever one's zero-handle hook does to the other. High-risk code, a claim the change makes, no test that can fail on it. Owed, in the same job and boot: two map_bar answers held together, their as_ptr()s unequal, the register read through both, the first dropped, the register read through the second again. Its negative control is control-kernel.patch with this arm first: one cached object and map_into idempotent per process answer one address, so it must red on the pointer assertion, not on the panic. Post that log.
  • PR body, "Checks" — iommu_virtio_platform is measured at 72a18fa81, not at the head. I confirmed git diff --name-only 72a18fa81 6f05a4941 is issues/ alone, so the kernel it booted is this one; the rule is still the head, and the arm above moves the head anyway. At the next head: bar_map_again, both existing controls, the new control, iommu_virtio_platform and cargo run -- --ci host, each with command, exit code and log at that head.

NOTE

  • PR body, "Not run" — pci_reclaim and the two claim_* rows do not reach bar_object: tests/toyos-rust-tests/src/bin/pci_reclaim.rs:16-25 claims the I219 and drops it, maps no BAR and starts no netstack. The T14 rows that reach the first ask are the ones that boot netstack on the I219 (lan_lease_report on tests/lanleasecase, lan_dhcp_lease and lan_message_delivery on tests/lantalkcase), through userland/netstack/src/i219.rs:170.
  • PR body, "Checks" — git diff --name-only 72a18fa81 6f05a4941 is four files under issues/, not five: the closed issue arrived by the merge and left by the next commit, so it is in neither tree.

What the brief asked to be held hardest

  1. Several live objects over one window. Frames: Region::pages is None for a BAR, SharedMemObject::over allocates nothing, and free_and_unmap (kernel/src/arch/x86_64/paging.rs:631) removes the address region and clears the PDEs; no path frees or scrubs the frames. Hooks: on_zero_handles takes its own object's mapped_in and unmaps those addresses only. That is by reading, and it is what the first BLOCKER asks to be run. Bound: an object lives exactly as long as a handle to it, so live objects are bounded by handle slots (table slots and in-flight transfers), as SYS_SHM_CREATE's are while costing no pages; a refused install drops the object in the same call. Revocation: issues/a-bars-mapping-outlives-the-claim-it-was-asked-on.md is true of main as well — the claim's cached Arc was dropped by tear_down with nothing unmapped, and that one object's handle already carried DUP and TRANSFER, so the set of processes that can keep a window past the claim is the same before and after. Not made worse in reach, so filing it is within "stay on the task"; its owner exists, its exit is one a guest test can read, and its citation of issues/a-moved-handle-is-always-re-movable.md says what that file says. What does change is the fix: the claim now holds no reference to anything it minted, so that issue's exit needs the claim to learn what it handed out. Worth one sentence in the issue; not a finding.
  2. The same-shape table. Checked by my own search of every HandleEntry::new, ops::install, install_all, install_at and every KObjectRef::*( construction in kernel/src. It stands. The only kernel-side holders of a handle-counted Arc that is ever installed are DeviceInfo's buffers (minted once behind described.bytes, install_all checking room before any HandleEntry::new, remint given fresh objects), Grant::memory (installed once, by the call that made it) and ProcessEntry::object (own_handle, one caller). Completions::shm, SharedImage::object and the namespace's connectors are never installed. No second site.
  3. PR A shared region's last handle takes its list of mappings, and a map that finds none is refused #762. The two compose: git merge-tree --write-tree 6f05a4941 <wt/toyos-shmretire 83408305e> exits 0, the kernel/src/object/mod.rs and toyos-abi/src/syscall.rs hunks are disjoint, and neither fixes the other's defect — A shared region's last handle takes its list of mappings, and a map that finds none is refused #762 alone still installs the cached retired object and panics; this alone still leaves a map through a stranded Arc after the last close, which on a BAR object is a register window with no handle. Together a BAR window lives exactly as long as a handle. Land this first: it is the deterministic one. A shared region's last handle takes its list of mappings, and a map that finds none is refused #762 then merges main and owes bar_map_again at its merged head, since it rewrites the map_into and zero-handle path every ask in that job runs. For A shared region's last handle takes its list of mappings, and a map that finds none is refused #762's reviewer: its new issues/a-device-call-acts-on-whoever-holds-its-claims-slot-by-then.md is a sibling of issues/a-pci-call-in-flight-names-its-function-by-a-slot-the-next-claim-reuses.md, which main has had since Three audit claims traced: a BAR asked for twice panics the kernel, netstack keeps a dead client's socket, and a PCI call's slot is not its claim #755.
  4. A T14 row. Not owed for landing. The change targets the object model, not the device: the first ask builds the same Region in the same order and no caller asks twice, and the removed Arc changed no teardown (the handle count, not the Arc, runs the unmap on main too). All three callers reachable under QEMU have a green first ask in the body's logs: diskserver's NVMe in the bar_map_again boot, the job's own, netstack's virtio NIC in iommu_virtio_platform. If a reading is wanted on the next bench pass, lan_lease_report is the row, not pci_reclaim (see NOTE). The reason given for a guest test over a metal row is accepted: the red is a kernel panic, and on the shared tests/testcases metal boot that voids every row unioned into the boot.
  5. The header sentence in kernel/src/object/mod.rs. It is a contract, cites no issue, and carries no story. Its second clause is stated wider than the tree: DeviceInfo (kernel/src/object/device.rs:27-42) is a holder that outlives the handles it answered with and keeps objects, and is sound because it answers once. A source comment's wording is not a finding under this prompt.

Whole guest suite: not run at this head, as the body says. I do not hold that as its own BLOCKER: the function's only behavioural change is the second ask, and each QEMU-reachable caller's first ask is measured.

SEND BACK

…r outlives the earlier

Review round 1 of #761. The fix makes one kernel state newly reachable — two
live objects over one BAR window in one address space — and no arm of the job
held two answers together: every ask dropped its mapping before the next. The
job now opens with that arm: two `map_bar` answers held, their addresses
unequal, the register read through both, the earlier dropped, the register
read through the later again.

The object header's sentence named every holder that outlives its handles;
`DeviceInfo` is one that keeps objects and is sound because it answers once.
The sentence now names the holder that answers more than once.

The issue on a BAR mapping outliving its claim says what the fix changes for
its exit: the claim holds no reference to what it minted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Round 1 controls and logs, all at 8a6dc0be3. One script ran every gate serially from a clean tree: each control patch git apply --checked, applied, cargo test --test toyos-build -- bar_map_again run, and the tree restored with git checkout -- kernel toyos-abi tests (git status --porcelain --ignore-submodules=none empty after). These patches replace the earlier comment's, which were against 72a18fa81. Each log below runs from the harness's first line to its test result line, with the two host: lines and cargo's lines around them, which carry local paths, left out; EXIT= is the command's own exit code.

The new control: the whole kernel change reverted, the job as it stands (control-kernel.patch) — exit 1, on the pointer assertion

The held-together arm runs first, so one cached object and map_into idempotent per process answer one address, and the job ends before any ask that would panic the kernel.

09:09:56 running 1 tests, 12 wide

09:09:56   RUN   bar_map_again
09:09:56   BUILD x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again
09:09:57   BUILT x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again  (2s)
09:09:57   BUILD x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again
09:10:05   BUILT x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again  (8s)
09:10:08 FAIL bar_map_again: the job ended Some(101):

thread 'main' (1) panicked at src/bin/bar_map_again.rs:63:5:
assertion `left != right` failed: bar_map_again: two answers held at once are one mapping
  left: 0xfffec00000
 right: 0xfffec00000
stack backtrace:
   0:      0x1000005b3b0 - _Unwind_Backtrace
   1:      0x10000046173 - <<std[98a7212de726a26d]::sys::backtrace::BacktraceLock>::print::DisplayBacktrace as core[73b405bef30eeff1]::fmt::Display>::fmt
   2:      0x1000005f1f7 - core[73b405bef30eeff1]::fmt::write
   3:      0x1000004996b - <std[98a7212de726a26d]::sys::stdio::toyos::Stderr as core[73b405bef30eeff1]::io::write::Write>::write_fmt
   4:      0x1000002cf53 - std[98a7212de726a26d]::panicking::default_hook::{closure#0}
   5:      0x1000004119e - std[98a7212de726a26d]::panicking::default_hook
   6:      0x10000041359 - std[98a7212de726a26d]::panicking::panic_with_hook
   7:      0x1000002d00d - std[98a7212de726a26d]::panicking::panic_handler::{closure#0}
   8:      0x10000025709 - std[98a7212de726a26d]::sys::backtrace::__rust_end_short_backtrace::<std[98a7212de726a26d]::panicking::panic_handler::{closure#0}, !>
   9:      0x1000002d778 - __rustc[d0c72ce299f394e5]::rust_begin_unwind
  10:      0x1000005f91b - core[73b405bef30eeff1]::panicking::panic_fmt
  11:      0x1000005f88f - core[73b405bef30eeff1]::panicking::assert_failed_inner
  12:      0x10000019833 - core[73b405bef30eeff1]::panicking::assert_failed::<*mut u8, *mut u8>
  13:      0x10000019575 - bar_map_again[ed935c42b0c7202e]::main
  14:      0x10000019ca6 - std[98a7212de726a26d]::sys::backtrace::__rust_begin_short_backtrace::<fn(), ()>
  15:      0x100000198c0 - std[98a7212de726a26d]::rt::lang_start::<()>::{closure#0}
  16:      0x10000040be6 - std[98a7212de726a26d]::rt::lang_start_internal
  17:      0x10000019791 - main
  18:      0x10000044cd3 - std[98a7212de726a26d]::sys::pal::toyos::start_rust
  19:      0x1000001a6de - _start

09:10:08   FAIL  bar_map_again  (3s)
09:10:08   --- 1 guests, 0 of them not the shipping kernel, 1 kernel build(s): [""]
09:10:08   --- irq census: 1 guest(s) reported, 444 interrupt(s), 357 of them on cpu0 (80.4%); per guest cpu0's share is median 80.4% p90 80.4% max 80.4%
09:10:08       timer            18 (4.1% of all), 55.6% of them on cpu0
09:10:08       kick             77 (17.3% of all), 39.0% of them on cpu0
09:10:08       xhci            200 (45.0% of all), 100.0% of them on cpu0
09:10:08       userdev         113 (25.5% of all), 100.0% of them on cpu0
09:10:08       tlb              36 (8.1% of all), 11.1% of them on cpu0
09:10:08       widest guest reported 2 cpu(s)

09:10:08 failures:
09:10:08     bar_map_again: the job ended Some(101):

09:10:08 test result: FAILED. 0 passed, 1 failed, 0 invalidated, 1 total (12.7s; workers: 10s building, 3s testing)
EXIT=1
diff --git a/kernel/src/object/mod.rs b/kernel/src/object/mod.rs
index d2441c247..205029806 100644
--- a/kernel/src/object/mod.rs
+++ b/kernel/src/object/mod.rs
@@ -2,8 +2,6 @@
 //!
 //! Objects are plain `Arc<T>`: no custom refcounting, no `Weak`, no `dyn` hierarchy.
 //!
-//! An object ends with its last handle and is never handed out again: a holder that answers more than once keeps what an object is made over (`device::Screen`, a claim's BAR windows) and makes a fresh object for each answer.
-//!
 //! `handle_count`, not the Arc strong count, is what userland-visible lifecycle rides: a syscall's `Arc` can be stranded on a killed thread's kernel stack, so release is deferred through [`ZERO_QUEUE`] — see `issues/deferred-release-outlives-its-syscall.md`.
 
 // `warn` here gates via CI's `-D warnings`; the rest of the kernel is not yet swept.
diff --git a/kernel/src/pcidev/mod.rs b/kernel/src/pcidev/mod.rs
index f6de66efc..1590f7a85 100644
--- a/kernel/src/pcidev/mod.rs
+++ b/kernel/src/pcidev/mod.rs
@@ -238,6 +238,9 @@ struct Bound {
     /// advertises; 0 bytes is a slot with no BAR this claim may map.
     bar_at: [u64; BARS],
     bar_bytes: [u64; BARS],
+    /// Minted on the first map of that BAR, so a second answers the same
+    /// object rather than a second handle to one window.
+    bars: [Option<Arc<SharedMemObject>>; BARS],
     grants: Vec<Grant>,
     /// The [`RESIDUE`] this claim took over and has placed no grant at: mapped
     /// to nothing and counted against nothing, and where a grant of a range's
@@ -804,6 +807,7 @@ fn bring_up(pci: PciDevice, id: PciId, slot: usize) -> Result<Bound, Refusal> {
                 id,
                 bar_at,
                 bar_bytes,
+                bars: [const { None }; BARS],
                 grants: Vec::new(),
                 residue: Vec::new(),
                 mastering: false,
@@ -1493,25 +1497,26 @@ fn with_bound<T>(
     f(bound)
 }
 
-/// One memory BAR as an object to map, a fresh one for every ask.
-///
-/// **The claim keeps where the window is and never an object over it**: an
-/// object ends with its last handle, which its holder closes whenever it
-/// likes, and the claim outlives that.
+/// One memory BAR as an object to map; the same object every time.
 pub fn bar_object(slot: usize, index: u64) -> Result<Arc<SharedMemObject>, SyscallError> {
     with_bound(slot, |bound| {
         let index = usize::try_from(index).map_err(|_| SyscallError::InvalidArgument)?;
         if index >= BARS || bound.bar_bytes[index] == 0 {
             return Err(SyscallError::InvalidArgument);
         }
-        Ok(SharedMemObject::over(Region {
+        if let Some(object) = bound.bars[index].as_ref() {
+            return Ok(Arc::clone(object));
+        }
+        let object = SharedMemObject::over(Region {
             phys: DirectMap::from_phys(bound.bar_at[index]),
             size: align_2m(bound.bar_bytes[index] as usize) as u64,
             // Registers, and the memory type is not firmware's to decide.
             cache: CachePolicy::Uncacheable,
             // The kernel owns no pages here: this is a device's aperture.
             pages: None,
-        }))
+        });
+        bound.bars[index] = Some(Arc::clone(&object));
+        Ok(object)
     })
 }
 
diff --git a/toyos-abi/src/syscall.rs b/toyos-abi/src/syscall.rs
index 53ca1e35f..3e3260d78 100644
--- a/toyos-abi/src/syscall.rs
+++ b/toyos-abi/src/syscall.rs
@@ -2312,8 +2312,7 @@ pub fn pipe_map(handle: RawHandle) -> Result<*mut u8, SyscallError> {
 /// kernel's, and a process that could write it could aim the device's interrupt
 /// at any address the LAPIC decodes.
 ///
-/// Every call answers an object of its own over the same window, so a BAR
-/// whose handle was closed can be asked for again.
+/// Idempotent per BAR: a second call answers the same object.
 ///
 /// A window alone drives nothing: the function masters the bus from its first
 /// [`device_dma_alloc`] and not before.

Control 1: the same reverted kernel, and the held-together arm removed (control-heldarm.patch on top) — exit 1, the kernel panic on the ask after a close

diff --git a/tests/toyos-rust-tests/src/bin/bar_map_again.rs b/tests/toyos-rust-tests/src/bin/bar_map_again.rs
index 5d835f00d..38097c3d7 100644
--- a/tests/toyos-rust-tests/src/bin/bar_map_again.rs
+++ b/tests/toyos-rust-tests/src/bin/bar_map_again.rs
@@ -56,22 +56,7 @@ fn main() {
     let bytes = info.bar_bytes[bar];
     let bar = bar as u32;
 
-    // Two answers held at once are two mappings, and letting the earlier one
-    // go takes nothing from the later.
-    let earlier = nic.map_bar(bar, bytes).expect("bar_map_again: the BAR, asked for and mapped");
-    let later = nic.map_bar(bar, bytes).expect("bar_map_again: the BAR, asked for while held");
-    assert_ne!(
-        earlier.as_ptr(),
-        later.as_ptr(),
-        "bar_map_again: two answers held at once are one mapping"
-    );
-    let first = read(&earlier);
-    assert_eq!(read(&later), first, "bar_map_again: the two mappings are two windows");
-    drop(earlier);
-    assert_eq!(read(&later), first, "bar_map_again: the later mapping changed with the earlier's end");
-    drop(later);
-    println!("bar_map_again: two answers held at once map apart, and the later outlives the earlier");
-
+    let first = ask(&nic, bar, bytes);
     println!("bar_map_again: BAR {bar} asked for and its handle closed; asking again");
     let again = ask(&nic, bar, bytes);
     assert_eq!(again, first, "bar_map_again: the second ask's mapping is another window");
09:10:09 running 1 tests, 12 wide

09:10:09   RUN   bar_map_again
09:10:09   BUILD x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again
09:10:09   BUILT x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again  (879ms)
09:10:09   BUILD x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again
09:10:11   BUILT x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again  (1s)
09:10:29 FAIL bar_map_again: kernel panic: PANIC: panicked at src/object/handle.rs:108:9: — the guest went quiet because every CPU is halted, not because it was still working. The panic is the finding and the guard never got to be one.
--- what the kernel said as it died ---
[ 3.056 cpu0 kernel alert] PANIC: panicked at src/object/handle.rs:108:9:
a handle to a retired SharedMem (koid 197)
[ 3.056 cpu0 kernel]   Backtrace:
[ 3.058 cpu0 kernel]     0xffff80007af8ea3c  core::panicking::panic_fmt+0x2c
[ 3.058 cpu0 kernel]     0xffff80007ae9a80f  kernel::object::ops::install+0xff
[ 3.059 cpu0 kernel]     0xffff80007af06313  kernel::process::with_process_data::<u64, kernel::syscall::device::sys_device_bar_map::{closure#0}>+0x43
[ 3.060 cpu0 kernel]     0xffff80007af68635  kernel::syscall::dispatch::syscall_dispatch+0x95
[ 3.061 cpu0 kernel]     0xffff80007ae67ca4  kernel::arch::x86_64::syscall::syscall_handler+0x34
[ 3.061 cpu0 kernel]     0xffff80007ae67900  kernel::arch::x86_64::syscall::syscall_entry+0x70
[ 3.062 cpu0 kernel]   Contexts: cpu0 crashed at sp=0xffff800013620af0, asking about ctx 0xffff80000026a060
[ 3.062 cpu0 kernel]   cpu0 is on ctx 0xffff80000026a060 pid=10 tid=0 stack_top=0xffff800013621000 saved_sp=0xffff800013620fc0
[ 3.062 cpu0 kernel]   cpu1 is on ctx 0xffff800000200df0 (its idle context) stack_top=0x0000000000000000 saved_sp=0xffff800000641ef0
[ 3.062 cpu0 kernel]   Running: pid=10 tid=Some(Tid(0))
[ 3.062 cpu0 kernel]   Process: test_rs_bar_map_again pid=10 state=Live
[ 3.062 cpu0 kernel]   Syscall: num=117 user_rip=0x10000019f8c user_rsp=0xfffffffd00
[ 3.062 cpu0 kernel]   User backtrace:
[ 3.063 cpu0 kernel]     0x10000019f8c  /system/bin/test_rs_bar_map_again+0x19f8c id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000019bee  /system/bin/test_rs_bar_map_again+0x19bed id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000018b5a  /system/bin/test_rs_bar_map_again+0x18b59 id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000018dfe  /system/bin/test_rs_bar_map_again+0x18dfd id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000019716  /system/bin/test_rs_bar_map_again+0x19715 id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000019330  /system/bin/test_rs_bar_map_again+0x1932f id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000040646  /system/bin/test_rs_bar_map_again+0x40645 id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000019241  /system/bin/test_rs_bar_map_again+0x19240 id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x10000044733  /system/bin/test_rs_bar_map_again+0x44732 id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.063 cpu0 kernel]     0x1000001a13e  /system/bin/test_rs_bar_map_again+0x1a13d id=ecf5de2553f31383bfd12e2979d4846f2dd8da1f
[ 3.064 cpu0 kernel alert] panic: rebooting in 60 s, timed by the calibrated clock
the job said:

09:10:29   FAIL  bar_map_again  (18s)
09:10:29   --- 1 guests, 0 of them not the shipping kernel, 1 kernel build(s): [""]

09:10:29 failures:
09:10:29     bar_map_again: kernel panic: PANIC: panicked at src/object/handle.rs:108:9: — the guest went quiet because every CPU is halted, not because it was still working. The panic is the finding and the guard never got to be one.

09:10:29 test result: FAILED. 0 passed, 1 failed, 0 invalidated, 1 total (20.4s; workers: 2s building, 18s testing)
EXIT=1

Control 2: on top of control 1, the close-then-ask arm removed (control-arm2.patch), so the first ask is the one a full table refuses — exit 1, the same panic

--- a/tests/toyos-rust-tests/src/bin/bar_map_again.rs
+++ b/tests/toyos-rust-tests/src/bin/bar_map_again.rs
@@ -56,10 +56,7 @@
     let bytes = info.bar_bytes[bar];
     let bar = bar as u32;
 
-    let first = ask(&nic, bar, bytes);
     println!("bar_map_again: BAR {bar} asked for and its handle closed; asking again");
-    let again = ask(&nic, bar, bytes);
-    assert_eq!(again, first, "bar_map_again: the second ask's mapping is another window");
     println!("bar_map_again: answered with a handle that maps");
 
     // No room for the handle: the ask is refused, and the object made for it
@@ -84,7 +81,6 @@
         "bar_map_again: an ask from a full table was not refused for room"
     );
     println!("bar_map_again: an ask from a full handle table refused; asking again");
-    let after = ask(&nic, bar, bytes);
-    assert_eq!(after, first, "bar_map_again: the ask after the refusal maps another window");
+    let _ = ask(&nic, bar, bytes);
     println!("bar_map_again: answered after the refusal with a handle that maps");
 }
09:10:29 running 1 tests, 12 wide

09:10:29   RUN   bar_map_again
09:10:29   BUILD x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again
09:10:30   BUILT x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again  (1s)
09:10:30   BUILD x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again
09:10:31   BUILT x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again  (971ms)
09:10:49 FAIL bar_map_again: kernel panic: PANIC: panicked at src/object/handle.rs:108:9: — the guest went quiet because every CPU is halted, not because it was still working. The panic is the finding and the guard never got to be one.
--- what the kernel said as it died ---
[ 2.928 cpu0 kernel alert] PANIC: panicked at src/object/handle.rs:108:9:
a handle to a retired SharedMem (koid 197)
[ 2.928 cpu0 kernel]   Backtrace:
[ 2.930 cpu0 kernel]     0xffff80007af8ea3c  core::panicking::panic_fmt+0x2c
[ 2.930 cpu0 kernel]     0xffff80007ae9a80f  kernel::object::ops::install+0xff
[ 2.930 cpu0 kernel]     0xffff80007af06313  kernel::process::with_process_data::<u64, kernel::syscall::device::sys_device_bar_map::{closure#0}>+0x43
[ 2.931 cpu0 kernel]     0xffff80007af68635  kernel::syscall::dispatch::syscall_dispatch+0x95
[ 2.931 cpu0 kernel]     0xffff80007ae67ca4  kernel::arch::x86_64::syscall::syscall_handler+0x34
[ 2.932 cpu0 kernel]     0xffff80007ae67900  kernel::arch::x86_64::syscall::syscall_entry+0x70
[ 2.932 cpu0 kernel]   Contexts: cpu0 crashed at sp=0xffff8000071f2af0, asking about ctx 0xffff8000071c6c10
[ 2.932 cpu0 kernel]   cpu0 is on ctx 0xffff8000071c6c10 pid=10 tid=0 stack_top=0xffff8000071f3000 saved_sp=0xffff8000071f2c60
[ 2.932 cpu0 kernel]   cpu1 is on ctx 0xffff800000200df0 (its idle context) stack_top=0x0000000000000000 saved_sp=0xffff800000641ef0
[ 2.932 cpu0 kernel]   Running: pid=10 tid=Some(Tid(0))
[ 2.932 cpu0 kernel]   Process: test_rs_bar_map_again pid=10 state=Live
[ 2.932 cpu0 kernel]   Syscall: num=117 user_rip=0x10000019d4c user_rsp=0xfffffffd60
[ 2.932 cpu0 kernel]   User backtrace:
[ 2.932 cpu0 kernel]     0x10000019d4c  /system/bin/test_rs_bar_map_again+0x19d4c id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x100000199ae  /system/bin/test_rs_bar_map_again+0x199ad id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x10000018ce9  /system/bin/test_rs_bar_map_again+0x18ce8 id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x100000194d6  /system/bin/test_rs_bar_map_again+0x194d5 id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x100000190f0  /system/bin/test_rs_bar_map_again+0x190ef id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x10000040406  /system/bin/test_rs_bar_map_again+0x40405 id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x10000019001  /system/bin/test_rs_bar_map_again+0x19000 id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x100000444f3  /system/bin/test_rs_bar_map_again+0x444f2 id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.932 cpu0 kernel]     0x10000019efe  /system/bin/test_rs_bar_map_again+0x19efd id=461f9ca37eb154429e69aa76f37f23259fdadfe6
[ 2.934 cpu0 kernel alert] panic: rebooting in 60 s, timed by the calibrated clock
the job said:

09:10:49   FAIL  bar_map_again  (18s)
09:10:49   --- 1 guests, 0 of them not the shipping kernel, 1 kernel build(s): [""]

09:10:49 failures:
09:10:49     bar_map_again: kernel panic: PANIC: panicked at src/object/handle.rs:108:9: — the guest went quiet because every CPU is halted, not because it was still working. The panic is the finding and the guard never got to be one.

09:10:49 test result: FAILED. 0 passed, 1 failed, 0 invalidated, 1 total (20.2s; workers: 2s building, 18s testing)
EXIT=1

Green at the same head

cargo test --test toyos-build -- bar_map_again:

09:09:34 running 1 tests, 12 wide

09:09:34   RUN   bar_map_again
09:09:34   BUILD x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again
09:09:35   BUILT x86_64 bar_map_again of tests/toyos-rust-tests, for bar_map_again  (2s)
09:09:35   BUILD x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again
09:09:38   BUILT x86_64 kernel, loader, ROOT of tests/testcases, for bar_map_again  (3s)
09:09:41   [claims] bar_map_again: two answers held at once map apart, and the later outlives the earlier
09:09:41   [claims] bar_map_again: answered with a handle that maps
09:09:41   [claims] bar_map_again: answered after the refusal with a handle that maps
09:09:41   PASS  bar_map_again  (3s)
09:09:41   --- 1 guests, 0 of them not the shipping kernel, 1 kernel build(s): [""]
09:09:41   --- irq census: 1 guest(s) reported, 507 interrupt(s), 419 of them on cpu0 (82.6%); per guest cpu0's share is median 82.6% p90 82.6% max 82.6%
09:09:41       timer            18 (3.6% of all), 77.8% of them on cpu0
09:09:41       kick             87 (17.2% of all), 42.5% of them on cpu0
09:09:41       xhci            254 (50.1% of all), 100.0% of them on cpu0
09:09:41       userdev         113 (22.3% of all), 100.0% of them on cpu0
09:09:41       tlb              35 (6.9% of all), 2.9% of them on cpu0
09:09:41       widest guest reported 2 cpu(s)

09:09:41 test result: ok. 1 passed, 1 total (7.6s; workers: 5s building, 3s testing)
EXIT=0

cargo test --test toyos-build -- iommu_virtio_platform:

09:09:41 running 1 tests, 12 wide

09:09:41   RUN   iommu_virtio_platform
09:09:41   BUILD x86_64 kernel, loader, ROOT of tests/netcase, for iommu_virtio_platform
09:09:42   BUILT x86_64 kernel, loader, ROOT of tests/netcase, for iommu_virtio_platform  (953ms)
09:09:45   [iommu] headless: 3 virtio function(s) behind a unit = true, the audio function 00:04.0 among them
09:09:48   [iommu] headless-no-iommu: 2 virtio function(s) behind a unit = false, the audio function 00:04.0 among them; the NIC's claim refused for want of a unit
09:09:48   BUILD x86_64 kernel boot-actuators,test-actuators, ROOT of tests/testcases, for iommu_virtio_platform
09:09:50   BUILT x86_64 kernel boot-actuators,test-actuators, ROOT of tests/testcases, for iommu_virtio_platform  (2s)
09:09:53   [iommu] declined: [ 2.486 cpu0 kernel] virtio-sound: NOT INITIALISED — PCI 00:04.0 refused the feature set 0x100000000 the driver accepted, leaving DEVICE_STATUS=0x3 without FEATURES_OK
09:09:53   PASS  iommu_virtio_platform  (9s)
09:09:53   --- 3 guests, 1 of them not the shipping kernel, 2 kernel build(s): ["", "boot-actuators,test-actuators"]
09:09:53   --- irq census: 1 guest(s) reported, 215 interrupt(s), 193 of them on cpu0 (89.8%); per guest cpu0's share is median 89.8% p90 89.8% max 89.8%
09:09:53       timer            10 (4.7% of all), 80.0% of them on cpu0
09:09:53       kick             17 (7.9% of all), 41.2% of them on cpu0
09:09:53       xhci            177 (82.3% of all), 100.0% of them on cpu0
09:09:53       tlb              11 (5.1% of all), 9.1% of them on cpu0
09:09:53       widest guest reported 2 cpu(s)

09:09:53 test result: ok. 1 passed, 1 total (11.7s; workers: 3s building, 9s testing)
EXIT=0

cargo run -- --ci host: EXIT=0; its last line is [ci] Host: 77 step(s), all green.

@Japabu Japabu changed the title A claim keeps where its BARs are and no object over them: a BAR asked for again after its handle closed no longer panics the kernel A claim keeps where its BARs are and no object over them: a BAR asked for again after its handle closed no longer panics the kernel, and two answers held at once map apart Oct 8, 2026
Brings in #749, #757, #759 and #762. #762 changes kernel/src/object/mod.rs
and toyos-abi/src/syscall.rs in hunks disjoint from this branch's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Round 2 gates at the merged head 00b7a0c74 (8a6dc0be3 + origin/main fb00abfa5). One script, run alone and serially from a clean tree by the orchestrator; each EXIT= is the command's own. The three control patches are the ones in the round 1 comment, unchanged: each git apply --checked, stacked, run, and reversed with git apply -R (three reversals exit 0, git status --porcelain --ignore-submodules=none empty after). The deciding lines of each log follow; cargo's lines and the host: lines, which carry local paths, are left out.

cargo test --test toyos-build -- bar_map_again — EXIT=0

10:06:38   [claims] bar_map_again: two answers held at once map apart, and the later outlives the earlier
10:06:38   [claims] bar_map_again: answered with a handle that maps
10:06:38   [claims] bar_map_again: answered after the refusal with a handle that maps
10:06:38   PASS  bar_map_again  (3s)
10:06:38 test result: ok. 1 passed, 1 total (34.5s; workers: 31s building, 3s testing)

cargo test --test toyos-build -- iommu_virtio_platform — EXIT=0

10:06:47   [iommu] headless: 3 virtio function(s) behind a unit = true, the audio function 00:04.0 among them
10:06:50   [iommu] headless-no-iommu: 2 virtio function(s) behind a unit = false, the audio function 00:04.0 among them; the NIC's claim refused for want of a unit
10:07:07   [iommu] declined: [ 2.542 cpu0 kernel] virtio-sound: NOT INITIALISED — PCI 00:04.0 refused the feature set 0x100000000 the driver accepted, leaving DEVICE_STATUS=0x3 without FEATURES_OK
10:07:07   PASS  iommu_virtio_platform  (9s)
10:07:07 test result: ok. 1 passed, 1 total (28.7s; workers: 19s building, 9s testing)

Control, two answers held (control-kernel.patch: this branch's whole kernel change reverted onto the merged base, #762 in it; the job as it stands) — EXIT=1, on the job's pointer assertion, no kernel panic

10:07:29 FAIL bar_map_again: the job ended Some(101):

thread 'main' (1) panicked at src/bin/bar_map_again.rs:63:5:
assertion `left != right` failed: bar_map_again: two answers held at once are one mapping
  left: 0xfffec00000
 right: 0xfffec00000
10:07:29 test result: FAILED. 0 passed, 1 failed, 0 invalidated, 1 total (16.1s; workers: 13s building, 3s testing)

Control 1, close then ask (control-heldarm.patch on top) — EXIT=1, the recorded panic

10:07:51 FAIL bar_map_again: kernel panic: PANIC: panicked at src/object/handle.rs:108:9: — the guest went quiet because every CPU is halted, not because it was still working. The panic is the finding and the guard never got to be one.
[ 2.810 cpu0 kernel alert] PANIC: panicked at src/object/handle.rs:108:9:
a handle to a retired SharedMem (koid 197)
[ 2.812 cpu0 kernel]     0xffff80007aec4d5f  kernel::object::ops::install+0xff
[ 2.812 cpu0 kernel]     0xffff80007aedef93  kernel::process::with_process_data::<u64, kernel::syscall::device::sys_device_bar_map::{closure#0}>+0x43
10:07:51 test result: FAILED. 0 passed, 1 failed, 0 invalidated, 1 total (21.6s; workers: 3s building, 18s testing)

Control 2, refused then ask (control-arm2.patch on top) — EXIT=1, the same panic

10:08:12 FAIL bar_map_again: kernel panic: PANIC: panicked at src/object/handle.rs:108:9: — the guest went quiet because every CPU is halted, not because it was still working. The panic is the finding and the guard never got to be one.
[ 2.658 cpu0 kernel alert] PANIC: panicked at src/object/handle.rs:108:9:
a handle to a retired SharedMem (koid 197)
[ 2.659 cpu0 kernel]     0xffff80007aec4d5f  kernel::object::ops::install+0xff
[ 2.660 cpu0 kernel]     0xffff80007aedef93  kernel::process::with_process_data::<u64, kernel::syscall::device::sys_device_bar_map::{closure#0}>+0x43
10:08:12 test result: FAILED. 0 passed, 1 failed, 0 invalidated, 1 total (21.0s; workers: 3s building, 18s testing)

Not run at this head: cargo run -- --ci host and the whole guest suite. Both are CI's, on the ready pull request.

@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review round 2 of wt/toyos-barremap at 00b7a0c74, against origin/main fb00abfa5 (fetched now; it is also the merge base). Read only: the diff, the changed files whole at the head, the merge commit, the round 2 logs on disk (orch/barremap-r2/), #762's body and reviews. No test, build or cargo command was run by this review.

Net lines (git diff --shortstat origin/main...00b7a0c74: 7 files, +194 -57). Production (kernel/, toyos-abi/): +11 -13. Tests (tests/toyos.rs, the job): +145 -0. issues/: +38 -44.

Round 1's BLOCKERs

  1. No arm held two answers at once: CLOSED. tests/toyos-rust-tests/src/bin/bar_map_again.rs:59-72 holds two map_bar answers together, asserts their addresses differ, reads the register through both, drops the earlier and reads through the later again. Measured at 00b7a0c74 (gates-head.txt): green.log exit 0 with all three arms' lines; control-held.log, the whole kernel change reverted onto the merged base and the job as it stands, exit 1 on src/bin/bar_map_again.rs:63:5, left: 0xfffec00000 right: 0xfffec00000, the job ended Some(101) and no kernel panic. That is the red asked for, on the pointer assertion and not on the panic. control-1.log and control-2.log exit 1 on src/object/handle.rs:108:9: a handle to a retired SharedMem (koid 197) through ops::install from sys_device_bar_map, on a kernel that carries A shared region's last handle takes its list of mappings, and a map that finds none is refused #762: A shared region's last handle takes its list of mappings, and a map that finds none is refused #762 alone does not fix this. The three patches are byte for byte round 1's (cmp), git apply --check of control-kernel.patch still passes at the head (offsets only), the three reversals exit 0 and control-status.txt is empty.
  2. Evidence not at the head: OPEN, in one row. At 00b7a0c74: bar_map_again 0, iommu_virtio_platform 0, the three controls 1, each with command, exit and log. cargo run -- --ci host has no result at 00b7a0c74; the body says so itself. Its one reading (orch/barremap-r1/ci-host.log, [ci] Host: 77 step(s), all green, EXIT=0) is at 8a6dc0be3, and the merge since then moved 67 files under it.

BLOCKER

  • PR body, "Checks" — cargo run -- --ci host is not measured at 00b7a0c74 — the rule is the head, round 1 asked for it at the next head by name, and neither parent's green is the merged tree's: tests/toyos.rs is edited by both sides (RUST_SKIP, MACHINE_TESTS, run_machine_test), and the host suite's tests/checks.rs reads exactly those lists against the discovered binaries. Owed: one run at 00b7a0c74, command, exit code and log in the body. The orchestrator's own run closes it, and so does the host check concluding success on the ready pull request at this head, which runs that command; the three checks the rollup shows at 00b7a0c74 are SKIPPED for the draft and are no result. Nothing in the code is asked for, and the next round is that one line.

NOTE

None.

What the brief asked to be ruled on

  • The merge. Clean and nothing of A shared region's last handle takes its list of mappings, and a map that finds none is refused #762 undone: git merge-tree --write-tree 8a6dc0be3 fb00abfa5 and 00b7a0c74^{tree} are the same tree (d9aeba0f8), so the merge commit is the automatic merge with no hand edit; git diff fb00abfa5 00b7a0c74 is this branch's seven files and nothing else; kernel/src/object/shm.rs at the head is A shared region's last handle takes its list of mappings, and a map that finds none is refused #762's.
  • "Where this meets A shared region's last handle takes its list of mappings, and a map that finds none is refused #762", by the code. True. pcidev::bar_object (kernel/src/pcidev/mod.rs:1513) builds SharedMemObject::over on every ask and Bound has no field that could keep one; sys_device_bar_map is its one caller. Each object has its own Held list (shm.rs:60), on_zero_handles takes self.mapped_in and unmaps those addresses, and free_and_unmap (paging.rs:631) removes an address region and its entries; a BAR's Region::pages is None, so nothing is freed. A refused install drops the entry it built, which queues the never-mapped object, whose hook takes an empty list and returns; control 2 is the measurement that a refused ask retires the object. The kernel's profile has debug-assertions = true (kernel/Cargo.toml:485), so the Drop assertion is live in the kernel the green ran.
  • The implementer's doubt: accepted, nothing more is owed. SharedMemory's drop is a close and no unmap (toyos/src/shm.rs), so the earlier mapping goes by the deferred hook: at that close's own exit (dispatch.rs:684) unless the other CPU's pass takes the batch between the enqueue and that drain, and only then can the read through the later answer precede the unmap. The job has no event to wait on and a wait without one is a flat wait. The only cheap strengthening, holding later to the job's end, puts a live handle under the close-then-ask arm and so removes that arm's red on the cached kernel; it is refused. What the unobserved interleaving rests on is a field of self and an unmap by address, which a reader checks, in code this branch does not change. No mutation is named: one that makes one object's hook reach another's addresses needs the list shared between objects, which a reader of the diff catches.
  • No T14 row: the ruling holds at the merged head. A shared region's last handle takes its list of mappings, and a map that finds none is refused #762's reviewer asked whether a member of the shared, shared-2 and shared-debug boots maps a BAR. None does. The three callers of map_bar in the tree are userland/diskserver/src/nvme.rs:230, userland/netstack/src/i219.rs:170 and userland/netstack/src/virtio_net.rs:396. Those boots are tests/testcases (shared_metal, tests/toyos.rs): it starts no netstack, and its diskserver row names pci:1b36:0010, QEMU's controller, which the T14 does not have, so the supervisor endows no claim and diskserver takes its Drive::Absent arm before Controller::open. Of the jobs, only pci_reclaim claims a PCI function, and it maps nothing; bar_map_again is on RUST_SKIP and rides no shared boot. So the combined kernel differs from the one booted only in bar_object, which those boots never enter, and in Bound losing a field that was None for the whole of every claim they made. The four region rows' reading at 58b768a03 stands for the combined kernel on their path. The rows that do reach the first ask on the T14 are the ones the body names (lan_lease_report, lan_dhcp_lease, lan_message_delivery); the first ask builds the same Region either way and they are not owed for landing.
  • Round 1's NOTEs. Both closed: the body names the lan_* rows and no longer names pci_reclaim or the claim_* rows, and the count of files under issues/ is gone with the sentence that carried it.
  • The issues. The closed issue's exit is met on both arms it names, by green.log and controls 1 and 2; git grep at the head finds its slug nowhere. issues/a-bars-mapping-outlives-the-claim-it-was-asked-on.md has an owner that exists, an exit a guest test reads, and its two citations resolve at the head.

What the landing rests on

host and guest / suite (with toolchain) concluding success on the ready pull request at the head that lands, and again in the merge queue. The whole guest suite has no reading at any head of this branch and is CI's alone; a SKIPPED conclusion is not one. No T14 reading. #763 (wt/toyos-pcislot) conflicts with this head in kernel/src/pcidev/mod.rs alone (git merge-tree), lands after, and owes bar_map_again whole at its merged head.

SEND BACK

@Japabu
Japabu marked this pull request as ready for review October 8, 2026 10:13
@Japabu

Japabu commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 3 — head 00b7a0c74 (unchanged since round 2), origin/main at 1084ddc9a

Read only: git, gh. No cargo, rustc, QEMU or test command run. The change was not re-read; round 2 read it at this head.

Round 2's BLOCKER

CLOSED — cargo run -- --ci host unmeasured at 00b7a0c74. gh run view 37762030186: event pull_request, headSha 00b7a0c74e14d8bd06fa042b39ec2460bcbecfe6, conclusion success.

  • host, job 113260522116, success; step Run cargo run -- --ci host success, 10:14:06 to 10:26:46. Its log, fetched this round: HEAD is now at f77c6ad Merge 00b7a0c74e14d8bd06fa042b39ec2460bcbecfe6 into fb00abfa5a362e0c49925d4891e2fd191a71a81d, and 10:26:46 [ci] Host: 78 step(s), all green.
  • toolchain / build, job 113260522448, success.
  • guest / suite, job 113261645563, success; step Run cargo run -- --ci guest success. The saved log (911 lines) checks out the same f77c6ad, and reads RUN bar_map_again, the three [claims] lines the harness arm demands, PASS bar_map_again (3s), PASS iommu_virtio_platform (8s), test result: ok. 32 passed, 32 total, [ci] Guest: 5 step(s), all green. No PANIC, no FAIL.

The tree CI measured is the head's own: f77c6ad merges the head into fb00abfa5, which is the head's merge base with main and already an ancestor of it, so the merge adds nothing to the head's tree.

The body against the head and the record

True of both. Its gate table's two CI rows name the run, the jobs and the lines read above; its "not run locally" bullet says where the run's base stood. gh pr view: head 00b7a0c74, not a draft, MERGEABLE, CLEAN, auto-merge not armed. The three SKIPPED checks in the rollup are run 37761513021, the draft's; the required names each carry a SUCCESS from 37762030186.

The ruling: main has moved by #758 and #746

git merge-tree --write-tree origin/main 00b7a0c74 exits 0, tree 795a1a5a9: clean.

tests/toyos.rs is the one path both sides changed since fb00abfa5, and the two do not share a hand in it. main's whole change to that file is one line in run_debug_mode (kernel/target/… to target/…, at 4249). The branch's three hunks are at 172 (RUST_SKIP), 275 (MACHINE_TESTS) and 2796 (run_machine_test's arm and bar_map_again). #746 touched none of RUST_SKIP, MACHINE_TESTS or the machine-test arm; the merged file is the head's file plus that one line, and nothing else.

What the branch's new code reaches of #746's, by reading the merged tree:

It may be armed as it stands. The evidence rule asks for the gates green at the head, and they are. The branch touches no manifest, lock, .cargo/, src/ or .github/, so it cannot move what #746 folded; what #746 could move under this branch is one caller of a builder the suite already calls three times. Whether that fourth call builds and passes on the tree main becomes is exactly what the queue's host and guest / suite read, on that tree, before main moves. A merge of main into the branch and one more local reading would measure the same tree the queue measures and buy nothing but a compile. No merge and no further reading are owed.

BLOCKER

None.

NOTE

None.

LAND

@Japabu
Japabu added this pull request to the merge queue Oct 8, 2026
Japabu added a commit that referenced this pull request Oct 8, 2026
…ted per ask

One conflict, pcidev::bar_object: #761's body and doc, an object per ask and
no cache, under this branch's signature, `binding: &Binding` and
`with_bound(binding, ..)`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
Merged via the queue into main with commit d78350c Oct 8, 2026
6 checks passed
@Japabu
Japabu deleted the wt/toyos-barremap branch October 8, 2026 10:58
Japabu added a commit that referenced this pull request Oct 8, 2026
Changes no byte: the branch already carried #761's content, and
`git diff 5d14696 HEAD` is empty.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
Japabu added a commit that referenced this pull request Oct 8, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant