diff --git a/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md b/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md index cae76ce5baa..9d51c4fa901 100644 --- a/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md +++ b/issues/kernel/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md @@ -37,16 +37,15 @@ directory is `issues/isolation/every-program-sees-only-the-files-it-was-given.md ## Stages -1. **An end is an event**. `read_watch` and `has_data` answer for a - `Process`, whose watch becomes an `Arc` as an `Acceptor`'s is, and - `close_ends_polls` answers `false` for one; init's waiter threads go. - *Exit*: children held on their stdin (`process_lifecycle`'s `held` role), - watched in one poller and released one at a time, each completion naming the - child just released, whose code `try_wait` then reads; a watch on a child - already gone completes at once; a kill completes one; after one of two - handles to a held child closes, a non-blocking submit finds nothing. - Negative control: the stage reverted whole (`NotSupported`); mutation: - `close_ends_polls` at `true` reds the last arm. Oracle: pidfd_open(2). +1. **init's waiter threads go**. A child's end is readiness on its handle, + and init still parks one thread per service on `SYS_PROCESS_WAIT` + (`close_when_it_ends`, `userland/init/src/main.rs`). *Exit*: init starts no + thread to wait for a service's end, and `close_when_it_ends` is gone with + all it does kept: a `restart` row that ends is started again on the ports + it kept; a row without one has its acceptors closed, so a client's next + connect is `Gone`; a swap's expected end leaves them open for the binary + after it. `src/metalswap.rs`'s `judge` reads the third, on the T14 through + `toyos-metal --swap`; no test reads the first two. 2. **An end says how**. An end reads as an exit, a kill or a fault kind alike on every architecture, and a bare code reads the last two as failures. No end reads as a quit's reason. *Exit*: exited 137, killed, and each fault diff --git a/issues/kernel/a-watch-outlives-the-close-of-the-process-handle-it-was-made-through.md b/issues/kernel/a-watch-outlives-the-close-of-the-process-handle-it-was-made-through.md new file mode 100644 index 00000000000..9b8af636d7b --- /dev/null +++ b/issues/kernel/a-watch-outlives-the-close-of-the-process-handle-it-was-made-through.md @@ -0,0 +1,29 @@ +--- +status: open +kind: defect +opened: 2026-10-02 +--- + +# A watch outlives the close of the Process handle it was made through + +`ops::close` ends no poll for a `Process` (`close_ends_polls`): closing one +handle may end no other handle's watch, and a close cannot tell the polls made +through its own handle from another's. So a watch made through a handle that +then closes stays armed on the process's watch and kept by its ring, where it +counts against `MAX_PENDING_WATCHES`, until the process ends. The look then +finds the handle stale and answers `-NotFound` under the watch's token. + +Measured at 81e68535c in one QEMU guest, with an arm that is in no commit: a +ring holding one watch on a duplicate of a held child's handle answers nothing +at the submit after the duplicate's close, and answers token 7 with result -1 +once the child ends. + +By `close_ends_polls`, a console's, the log's and a keyboard claim's polls are +kept the same way. + +**Exit**: a handle's close answers the polls its own process made through it, +and no other handle's; a test watches a held child through a duplicate, closes +the duplicate, and reads `-NotFound` before the child ends. That change +answers `issues/kernel/a-close-of-one-handle-ends-every-rings-poll-on-its-object.md` +too, whose first exit, a poll ending only when the last handle to its source +closes, would keep the very poll this exit ends at its own handle's close. diff --git a/kernel/src/object/ops.rs b/kernel/src/object/ops.rs index 8eb9d79f709..d734e59602f 100644 --- a/kernel/src/object/ops.rs +++ b/kernel/src/object/ops.rs @@ -272,6 +272,7 @@ pub fn read_watch(object: &KObjectRef) -> Option { KObjectRef::PipeRead(r) => pipe::read_watch(r.id()).map(WatchRef::Shared), KObjectRef::Connection(c) => pipe::read_watch(c.rx()).map(WatchRef::Shared), KObjectRef::Acceptor(a) => Some(WatchRef::Shared(a.watch().clone())), + KObjectRef::Process(p) => Some(WatchRef::Shared(p.watch().clone())), KObjectRef::Console(_) => Some(WatchRef::Static(&keyboard::WATCH)), KObjectRef::Device(d) => match d.class() { device_registry::DeviceType::Keyboard => Some(WatchRef::Static(&keyboard::WATCH)), @@ -290,7 +291,7 @@ pub fn read_watch(object: &KObjectRef) -> Option { KObjectRef::SysCap(_) => Some(WatchRef::Static(&crate::log::user::WATCH)), KObjectRef::PipeWrite(_) | KObjectRef::File(_) | KObjectRef::Inbox(_) | KObjectRef::Connector(_) | KObjectRef::Namespace(_) - | KObjectRef::SharedMem(_) | KObjectRef::Process(_) => None, + | KObjectRef::SharedMem(_) => None, } } @@ -311,11 +312,14 @@ pub fn write_watch(object: &KObjectRef) -> Option { /// Whether closing one handle to this object ends what its watches watch, so /// every poll on them — in any ring — is answered as gone. `false` for the log /// and the keyboard, which the machine ends on its own and which other handles -/// share: a console closing is not every console's keyboard going away. +/// share: a console closing is not every console's keyboard going away. `false` +/// for a process, which only its own end ends: closing one handle to it ends no +/// other's watch. fn close_ends_polls(object: &KObjectRef) -> bool { match object { KObjectRef::SysCap(_) => false, KObjectRef::Console(_) => false, + KObjectRef::Process(_) => false, KObjectRef::Device(d) => match d.class() { device_registry::DeviceType::Keyboard => false, device_registry::DeviceType::Mouse @@ -328,7 +332,7 @@ fn close_ends_polls(object: &KObjectRef) -> bool { KObjectRef::PipeRead(_) | KObjectRef::PipeWrite(_) | KObjectRef::Connection(_) | KObjectRef::Acceptor(_) | KObjectRef::File(_) | KObjectRef::Inbox(_) | KObjectRef::Connector(_) | KObjectRef::Namespace(_) - | KObjectRef::SharedMem(_) | KObjectRef::Process(_) => true, + | KObjectRef::SharedMem(_) => true, } } @@ -755,6 +759,7 @@ pub fn has_data(object: &KObjectRef) -> bool { KObjectRef::Connection(c) => pipe::has_data(c.rx()), KObjectRef::Console(_) => serial::has_data(), KObjectRef::Acceptor(a) => a.has_pending(), + KObjectRef::Process(p) => p.finished(), KObjectRef::File(_) => true, KObjectRef::Device(d) => match d.class() { device_registry::DeviceType::Keyboard => keyboard::has_data(), @@ -773,7 +778,7 @@ pub fn has_data(object: &KObjectRef) -> bool { }, KObjectRef::PipeWrite(_) | KObjectRef::Inbox(_) | KObjectRef::SysCap(_) | KObjectRef::Connector(_) | KObjectRef::Namespace(_) - | KObjectRef::SharedMem(_) | KObjectRef::Process(_) => false, + | KObjectRef::SharedMem(_) => false, } } diff --git a/kernel/src/object/process.rs b/kernel/src/object/process.rs index 3b651631146..45ec53263db 100644 --- a/kernel/src/object/process.rs +++ b/kernel/src/object/process.rs @@ -2,8 +2,9 @@ //! //! The exit code lives on the object, not a table entry: no zombie, no reap, //! no orphan adoption. A wait after the fact reads a value; a wait before it -//! parks and is woken by the publish. A process nobody holds a handle to -//! disappears. +//! parks and is woken by the publish. An `OP_WATCH` reads the same fact at its +//! look, so a handle is readable from the publish on. A process nobody holds a +//! handle to disappears. use alloc::sync::Arc; use core::sync::atomic::{AtomicBool, Ordering}; @@ -30,7 +31,8 @@ pub struct ProcessObject { /// The same fact, without the lock, for a waiter's per-wake predicate. finished: AtomicBool, /// What `SYS_PROCESS_WAIT` arms on; holding the `Arc` across the park keeps the watch from outliving its subject. - watch: Watch, + /// An `Arc` of its own, as a port's is, for the share `ops::read_watch` answers. + watch: Arc, } impl ProcessObject { @@ -40,7 +42,7 @@ impl ProcessObject { pid, exit: Lock::new(None), finished: AtomicBool::new(false), - watch: Watch::new(), + watch: Arc::new(Watch::new()), }) } @@ -61,7 +63,7 @@ impl ProcessObject { self.exit.lock().as_ref().map(|e| e.stats) } - pub fn watch(&self) -> &Watch { + pub fn watch(&self) -> &Arc { &self.watch } diff --git a/tests/toyos-rust-tests/src/bin/process_lifecycle.rs b/tests/toyos-rust-tests/src/bin/process_lifecycle.rs index 5a63f02dd6f..454b969689d 100644 --- a/tests/toyos-rust-tests/src/bin/process_lifecycle.rs +++ b/tests/toyos-rust-tests/src/bin/process_lifecycle.rs @@ -16,6 +16,10 @@ //! child's exit while the wait is parked on it, which used to return the wait //! and panic the kernel on the exit code that was not there yet. //! +//! **An end is also an event**: an `OP_WATCH` on the handle completes once the +//! exit is published, so one poller waits for any number of children beside +//! whatever else it watches. +//! //! Two roles besides the test. `held` exits with a code of the parent's //! choosing, but not until its stdin closes — which is what lets every arm here //! order the exit against the wait without a clock deciding anything. `waiter` @@ -24,12 +28,17 @@ use std::io::Read; use std::os::toyos::process::{ChildExt, CommandExt}; use std::process::{Child, ChildStdin, Command, Stdio}; -use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::atomic::{AtomicBool, AtomicU32, Ordering}; use std::sync::OnceLock; use toyos::endow::{Endowments, SYSCAP_LABEL}; +use toyos::poller::{Poller, READABLE}; use toyos::process::Process; use toyos::syscap::SysCap; +use toyos_abi::inbox::{ + Completion, RingHeader, Submission, COMPLETION_RING_OFF, OP_WATCH, RING_TAIL_OFF, SUBMISSIONS_OFF, + SUBMISSION_RING_OFF, +}; use toyos_abi::syscall::{self, SyscallError}; use toyos_abi::RawHandle; @@ -60,6 +69,10 @@ fn test() { an_unrelated_wake_does_not_end_the_wait(); two_handles_answer_the_same(); a_kill_publishes_like_an_exit(); + each_end_completes_its_own_watch(); + a_watch_on_an_ended_child_completes_at_once(); + a_kill_completes_a_watch(); + closing_one_handle_ends_no_other_watch(); a_handle_is_the_whole_of_the_right(); an_undefined_wait_flag_bit_is_refused(); println!("a process is a handle: the code is read, not claimed"); @@ -241,6 +254,116 @@ fn a_kill_publishes_like_an_exit() { println!(" a kill publishes {KILLED} once, and a second kill changes nothing"); } +/// Three held children in one poller, each watched once under its index, let +/// go one at a time and out of order. Before each release nothing is ready; +/// after it the one completion names that child, and its code is there without +/// a wait. +fn each_end_completes_its_own_watch() { + const CODES: [i32; 3] = [21, 22, 23]; + let mut held: Vec<(Child, Option)> = + CODES.iter().map(|&code| start(code)).map(|(child, stdin)| (child, Some(stdin))).collect(); + let poller = Poller::new(CODES.len() as u32); + for (i, (child, _)) in held.iter().enumerate() { + poller.watch_raw(RawHandle(child.as_raw_handle()), READABLE, i as u64); + } + for released in [1, 2, 0] { + poller.wait(0, 0, |token| panic!("child {token} is held, and its watch completed")); + drop(held[released].1.take()); + let mut ended = Vec::new(); + poller.wait(1, u64::MAX, |token| ended.push(token)); + assert_eq!(ended, [released as u64], "child {released} was let go"); + let status = held[released].0.try_wait().expect("try_wait"); + assert_eq!(status.and_then(|s| s.code()), Some(CODES[released]), "child {released}'s code"); + } + println!(" one poller: each end completes the watch of the child that ended, and only it"); +} + +/// The registration finds the end already published and completes in the +/// submit that makes it; a held child watched beside it does not. +fn a_watch_on_an_ended_child_completes_at_once() { + const GONE: u64 = 0; + const HELD: u64 = 1; + let (mut gone, release) = start(13); + drop(release); + assert_eq!(gone.wait().expect("wait").code(), Some(13), "the child that ends first"); + let (mut held, keep) = start(0); + let poller = Poller::new(2); + poller.watch_raw(RawHandle(gone.as_raw_handle()), READABLE, GONE); + poller.watch_raw(RawHandle(held.as_raw_handle()), READABLE, HELD); + let mut ready = Vec::new(); + poller.wait(0, 0, |token| ready.push(token)); + assert_eq!(ready, [GONE], "a non-blocking submit of both watches"); + drop(keep); + assert_eq!(held.wait().expect("wait").code(), Some(0), "the held child"); + println!(" a watch on a child already gone completes at once"); +} + +/// A kill ends the child without its cooperation, and its end answers a watch +/// as an exit's does: as readable, and not as a watch whose source is gone. +fn a_kill_completes_a_watch() { + let (mut child, _release) = start(0); + let handle = RawHandle(child.as_raw_handle()); + let result = watch_result_across(handle, || child.kill().expect("kill")); + assert_eq!(result, READABLE as i32, "the killed child's watch"); + let status = child.try_wait().expect("try_wait"); + assert_eq!(status.and_then(|s| s.code()), Some(KILLED), "the killed child's code"); + println!(" a kill completes a watch, as readable"); +} + +/// One `OP_WATCH` `READABLE` on `handle` in a ring of its own: nothing +/// completes before `end` runs, and the answer is the result word of the one +/// completion after it, which `Poller` does not hand out. +fn watch_result_across(handle: RawHandle, end: impl FnOnce()) -> i32 { + const TOKEN: u64 = 0x5EED; + // SAFETY: the page is this function's own ring, mapped until the close below. + let (inbox, base) = unsafe { syscall::inbox_setup(1) }.expect("inbox_setup"); + // SAFETY: slot 0 of a fresh ring's submission array and its tail word, both + // inside the page and aligned; the kernel reads the slot only once the tail + // publishes it, and the tail is reached as the atomic it is. + unsafe { + (base.add(SUBMISSIONS_OFF as usize) as *mut Submission).write(Submission { + op: OP_WATCH, + handle, + op_flags: READABLE, + token: TOKEN, + ..Submission::default() + }); + AtomicU32::from_ptr(base.add(SUBMISSION_RING_OFF as usize + RING_TAIL_OFF) as *mut u32) + .store(1, Ordering::Release); + } + assert_eq!(syscall::inbox_submit(inbox, 1, 0, 0), Ok(0), "the held child's watch completed"); + end(); + assert_eq!(syscall::inbox_submit(inbox, 0, 1, u64::MAX), Ok(1), "the watch's one completion"); + // SAFETY: entry 0 of the completion ring, which the kernel wrote before it + // answered the submit above; copied out whole. + let completion = unsafe { + (base.add(COMPLETION_RING_OFF as usize + core::mem::size_of::()) as *const Completion) + .read_volatile() + }; + syscall::close(inbox); + assert_eq!(completion.token, TOKEN, "the completion's token"); + completion.result +} + +/// A second handle closed while the first is watched: the child did not end, +/// so nothing completes, and the watch still answers the end when it comes. +fn closing_one_handle_ends_no_other_watch() { + let (mut child, release) = start(17); + let second = syscall::dup(RawHandle(child.as_raw_handle())).expect("a Process handle duplicates"); + let poller = Poller::new(1); + poller.watch_raw(RawHandle(child.as_raw_handle()), READABLE, 0); + poller.wait(0, 0, |token| panic!("the held child's watch completed ({token})")); + syscall::close(second); + poller.wait(0, 0, |token| panic!("closing a second handle answered the first's watch ({token})")); + drop(release); + let mut ended = Vec::new(); + poller.wait(1, u64::MAX, |token| ended.push(token)); + assert_eq!(ended, [0], "the watch that outlived the close"); + let status = child.try_wait().expect("try_wait"); + assert_eq!(status.and_then(|s| s.code()), Some(17), "the child's code"); + println!(" closing one handle to a child ends no other handle's watch"); +} + /// **The arm the pid-keyed shape could not have.** The waiter did not spawn the /// subject, is not its parent by any spelling, and holds nothing but a handle /// somebody moved into its table — and that is enough. diff --git a/userland/init/src/main.rs b/userland/init/src/main.rs index 3e55817b735..1901f662fde 100644 --- a/userland/init/src/main.rs +++ b/userland/init/src/main.rs @@ -692,10 +692,6 @@ impl std::fmt::Display for StartError { /// Wait for one start of a service to end, and close its ports if nothing was /// expecting it to, or wake the loop to start it again if its row says so. -/// -/// **A thread, because the kernel answers a process's end to a wait and to -/// nothing a poll can watch.** It parks in the kernel for the process's life -/// and costs nothing until then. fn close_when_it_ends(kept: &Mutex, generation: u64, process: toyos::RawHandle) { let _ = toyos_abi::syscall::process_wait(process); toyos_abi::syscall::close(process);