Repository navigation
The node keeps the pipe order std reads a stream's end by: a FIN after a shutdown reads as a FIN, a failure as a reset - #803
Conversation
… reads as a FIN A stream reaches its client as two pipes, and the receive pipe's end says only that it ended. std (`ended`) reads the kind off the send pipe: one with no reader behind a receive pipe at its end is a failure. The node broke that order for an orderly end. It let the client's send pipe go as soon as the client's FIN was queued, and again once the connection reached CLOSED or TIME-WAIT, so a client that shut its sending half down and read its peer to the end read `ConnectionReset` where the peer had sent a FIN. A failure dropped the receive pipe first and the send pipe later in the same pass, and a stream dropped whole dropped them in declaration order, receive first: a client that read the end between the two read a reset as a FIN. The node now keeps a from-client end whose last byte and FIN are queued, reads it no more, watches its writer and lets it go when the writer leaves, or when the client closes the stream or its reader is gone (nobody then reads the end). A failure lets the send pipe go before the receive pipe, in the pass and in a whole drop, `from_client` being declared first. The connection's reaching CLOSED or TIME-WAIT in order no longer ends the send pipe; only a failure does. A kept end is not on the cut's clock: it holds no byte to send. libc read every end as 0 and every refused write as EIO. It now asks the send pipe what std asks at a read of 0 (`streamend.rs`, host-tested by `toyos-libc-copies`), answers ECONNRESET for a failure and for a write the send pipe refuses for its reader's leaving, and keeps `shutdown`'s halves: a send after SHUT_WR is EPIPE without reaching a pipe the node reads no more, and a recv after SHUT_RD is 0, as std's are. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
… end `a-netstack-client-cannot-tell-a-reset-from-the-peers-fin` described netd's bridge of the smoltcp era. It now says what holds: the node keeps the order std and libc read, netstack on smoltcp closes the two pipes in one pass for every end, and a guest on main read a reset, a reset after a half-close and both FIN orders as a host's TCP does, in two runs. Its exit is the own stack's guest test, which only a netstack on the node can run. `a-shutdown-of-the-sending-half-drops-what-the-send-pipe-still-holds` was read from the code; it now carries the count measured on main: 65,536 of 1,048,576 bytes back, then `ConnectionReset`. The order carries one bit, and std and libc answer every failure ECONNRESET where Linux answers a timeout ETIMEDOUT and an ICMP error its own word: `a-streams-failure-reaches-its-client-as-a-reset-whatever-it-was`, whose clean carrier is a word on a pipe's end, the kernel's to add. The track carries it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Mutation patches and the negative control, each against m1-failure-drops-to-client-first --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -256,8 +256,8 @@
// The connection's failure: the client's send pipe goes first, so that the end it
// reads finds no reader behind it.
(Err(Error::Failed(_)), _) => {
- self.from_client = None;
self.to_client = None;
+ self.from_client = None;
}
(Err(Error::WouldBlock), _) => break,
(Err(refusal), _) => unreachable!("[tcp] refused the holder of a connection a read: {refusal:?}"),m2-fields-to-client-first --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -157,8 +157,8 @@
remote: Ipv4Addr,
/// Before `to_client`, so a stream dropped whole lets its send pipe go first: the order a
/// failure owes its client.
- from_client: Option<Box<dyn FromClient>>,
to_client: Option<Box<dyn ToClient>>,
+ from_client: Option<Box<dyn FromClient>>,
/// The connect is not answered yet, and no byte moves.
connecting: bool,
/// The connect's, while connecting; then the cut's, once the client can see the stream nom3-over-drops-send-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -301,7 +301,7 @@
// A connection that failed takes no more of the client's bytes: its writes fail rather
// than fill a pipe nobody drains, and its end reads as a failure. One that ended in order
// ended after the client's FIN, so its writing was over already.
- if status.failure.is_some() {
+ if matches!(status.state, State::Closed | State::TimeWait) {
self.from_client = None;
}
// A finished send pipe is kept only for a client that can still read the end and lookm4-fin-drops-send-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -287,9 +287,7 @@
if read == 0 {
stack.tcp_shutdown_write(now, conn);
self.fin_queued = true;
- if writer_gone {
- self.from_client = None;
- }
+ self.from_client = None;
break;
}
let Some(bytes) = space.get(..read) else { unreachable!("a pipe read {read} bytes into room for {room}") };m5-kept-pipe-unwatched --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -470,7 +470,7 @@
let (from, to) = (stream.from_client.is_some(), stream.to_client.is_some());
(*id, Watch {
readable: from && stream.room && !stream.fin_queued,
- writer: from && (!stream.done_writing || stream.fin_queued),
+ writer: from && !stream.done_writing,
writable: to && stream.held,
reader: to,
})m6-close-keeps-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -408,7 +408,6 @@
}
stream.to_client = None;
stream.done_writing = true;
- stream.left = true;
self.bridge(now);
}
m7-kept-pipe-timed --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -316,7 +316,7 @@
}
// Nothing of the stream reaches its client again, alive or not, and its pipe still holds
// what may be bytes to send.
- if self.to_client.is_none() && self.done_writing && !self.fin_queued {
+ if self.to_client.is_none() && self.done_writing {
if !self.extended {
let kept = extended.entry(self.remote).or_insert(0);
if *kept < OWNERLESS_PER_PEER {m8-reader-gone-keeps-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -438,10 +438,7 @@
pub fn pipe_gone(&mut self, now: Instant, id: StreamId, end: PipeEnd) {
let Some(stream) = self.streams.live.get_mut(&id).filter(|stream| !stream.connecting) else { return };
match end {
- PipeEnd::ToClient => {
- stream.to_client = None;
- stream.left = true;
- }
+ PipeEnd::ToClient => stream.to_client = None,
// A finished send pipe was kept for its writer alone.
PipeEnd::FromClient if stream.fin_queued => stream.from_client = None,
PipeEnd::FromClient => stream.done_writing = true,control: |
|
The guest case for the move ( The guest: This host's TCP (macOS): diff --git a/tests/toyos-rust-tests/src/bin/stream_ends.rs b/tests/toyos-rust-tests/src/bin/stream_ends.rs
new file mode 100644
index 000000000..293c7fccd
--- /dev/null
+++ b/tests/toyos-rust-tests/src/bin/stream_ends.rs
@@ -0,0 +1,22 @@
+//! `stream_ends.rs`'s report, in a guest: one line each, then `done`, and an
+//! exit of 1 where it is not the one a host's TCP gives.
+//!
+//! argv: the address and port of the harness's peer.
+
+#[path = "../stream_ends.rs"]
+mod stream_ends;
+
+fn main() {
+ let args: Vec<String> = std::env::args().collect();
+ let [_, host, port] = &args[..] else { panic!("usage: stream_ends <host> <peer port>") };
+ let port = port.parse().unwrap_or_else(|e| panic!("the port {port}: {e}"));
+ let mut report = Vec::new();
+ stream_ends::run(host, port, |line| {
+ println!("stream_ends: {line}");
+ report.push(line);
+ });
+ println!("stream_ends: done");
+ if report != stream_ends::EXPECTED {
+ std::process::exit(1);
+ }
+}
diff --git a/tests/toyos-rust-tests/src/stream_ends.rs b/tests/toyos-rust-tests/src/stream_ends.rs
new file mode 100644
index 000000000..893d83fce
--- /dev/null
+++ b/tests/toyos-rust-tests/src/stream_ends.rs
@@ -0,0 +1,135 @@
+//! Each way a stream can end, as std reports it to a client: the same code
+//! runs in a guest against netstack and on the harness's host against the host
+//! kernel's TCP, and the two reports are compared line for line.
+//!
+//! The peer is the harness's: a connection's first byte names what it does.
+//! Every end it makes waits on something the client did, so no line depends
+//! on how long anything took.
+//!
+//! Std alone: the harness compiles this file too.
+
+use std::io::{ErrorKind, Read, Write};
+use std::net::{Shutdown, TcpStream};
+
+/// What a client half-closes after, and reads back.
+pub const BULK: usize = 1 << 20;
+/// What the peer sends before the reset of `reset_mid_stream`.
+pub const AHEAD: usize = 1 << 16;
+
+/// The peer reads to the client's end, sends every byte back, then its FIN.
+pub const HALF_CLOSE: u8 = b'H';
+/// The peer sends its FIN before the client writes.
+pub const PEER_FIN_FIRST: u8 = b'F';
+/// The peer sends [`AHEAD`] bytes, reads the client's one, then resets.
+pub const RESET_MID_STREAM: u8 = b'R';
+/// The peer reads to the client's end, then resets.
+pub const RESET_AFTER_HALF_CLOSE: u8 = b'D';
+/// The peer sends three bytes and its FIN; the client then sends its own.
+pub const BOTH_CLOSE: u8 = b'B';
+
+/// The report a host's TCP gives, and so the one netstack owes.
+pub const EXPECTED: [&str; 7] = [
+ "half_close: wrote 1048576 bytes, then shut the sending half: Ok",
+ "half_close: 1048576 bytes came back as sent, then Ok(0)",
+ "peer_fin_first: the read Ok(0), then a write Ok(5)",
+ "reset_mid_stream: 65536 bytes came as sent, the answer Ok(1), then Err(ConnectionReset)",
+ "reset_after_half_close: shut the sending half: Ok, then Err(ConnectionReset)",
+ "both_close: 3 bytes came as sent, then Ok(0)",
+ "both_close: shut the sending half: Ok, then Ok(0)",
+];
+
+/// Byte `i` of anything sent here: a pattern no shift or cut reproduces.
+pub fn pattern(i: usize) -> u8 {
+ (i % 251) as u8
+}
+
+fn kind<T: std::fmt::Debug>(result: std::io::Result<T>) -> String {
+ match result {
+ Ok(_) => "Ok".to_string(),
+ Err(e) => format!("Err({:?})", e.kind()),
+ }
+}
+
+/// What one more read answers: the end, or a byte that is not one.
+fn the_end(stream: &mut TcpStream) -> String {
+ let mut byte = [0u8; 1];
+ match stream.read(&mut byte) {
+ Ok(n) => format!("Ok({n})"),
+ Err(e) => format!("Err({:?})", e.kind()),
+ }
+}
+
+/// Reads until `want` bytes, the end or a refusal: how many matched
+/// `expect`, and what stopped it, if anything did before `want`.
+fn take(stream: &mut TcpStream, want: usize, expect: impl Fn(usize) -> u8) -> (String, Option<String>) {
+ let mut buf = vec![0u8; 65536];
+ let mut got = 0;
+ while got < want {
+ let room = buf.len().min(want - got);
+ match stream.read(&mut buf[..room]) {
+ Ok(0) => return (format!("{got} bytes came"), Some("Ok(0)".to_string())),
+ Ok(n) => {
+ if let Some(at) = (0..n).find(|&i| buf[i] != expect(got + i)) {
+ return (format!("byte {} differs after {got} bytes", got + at), None);
+ }
+ got += n;
+ }
+ Err(e) if e.kind() == ErrorKind::Interrupted => {}
+ Err(e) => return (format!("{got} bytes came"), Some(format!("Err({:?})", e.kind()))),
+ }
+ }
+ (format!("{got} bytes came"), None)
+}
+
+fn dial(host: &str, port: u16, what: u8) -> TcpStream {
+ let mut stream = TcpStream::connect((host, port)).unwrap_or_else(|e| panic!("the peer at {host}:{port}: {e}"));
+ stream.write_all(&[what]).unwrap_or_else(|e| panic!("naming the end to the peer: {e}"));
+ stream
+}
+
+/// Every end, in order: one line each as [`EXPECTED`] spells them.
+pub fn run(host: &str, port: u16, mut say: impl FnMut(String)) {
+ let mut stream = dial(host, port, HALF_CLOSE);
+ let bulk: Vec<u8> = (0..BULK).map(pattern).collect();
+ let wrote = kind(stream.write_all(&bulk));
+ say(format!("half_close: wrote {BULK} bytes, then shut the sending half: {}", match wrote.as_str() {
+ "Ok" => kind(stream.shutdown(Shutdown::Write)),
+ _ => format!("the write {wrote}"),
+ }));
+ let (came, end) = take(&mut stream, BULK, pattern);
+ let end = end.unwrap_or_else(|| the_end(&mut stream));
+ say(format!("half_close: {} as sent, then {end}", came.replace("came", "came back")));
+
+ let mut stream = dial(host, port, PEER_FIN_FIRST);
+ let read = the_end(&mut stream);
+ let write = match stream.write(b"after") {
+ Ok(n) => format!("Ok({n})"),
+ Err(e) => format!("Err({:?})", e.kind()),
+ };
+ say(format!("peer_fin_first: the read {read}, then a write {write}"));
+
+ let mut stream = dial(host, port, RESET_MID_STREAM);
+ let (came, end) = take(&mut stream, AHEAD, pattern);
+ let line = match end {
+ Some(end) => format!("{came}, then {end}"),
+ None => {
+ let answer = match stream.write(b"k") {
+ Ok(n) => format!("Ok({n})"),
+ Err(e) => format!("Err({:?})", e.kind()),
+ };
+ format!("{came} as sent, the answer {answer}, then {}", the_end(&mut stream))
+ }
+ };
+ say(format!("reset_mid_stream: {line}"));
+
+ let mut stream = dial(host, port, RESET_AFTER_HALF_CLOSE);
+ let shut = kind(stream.shutdown(Shutdown::Write));
+ say(format!("reset_after_half_close: shut the sending half: {shut}, then {}", the_end(&mut stream)));
+
+ let mut stream = dial(host, port, BOTH_CLOSE);
+ let (came, end) = take(&mut stream, 3, |i| b"bye"[i]);
+ let end = end.unwrap_or_else(|| the_end(&mut stream));
+ say(format!("both_close: {came} as sent, then {end}"));
+ let shut = kind(stream.shutdown(Shutdown::Write));
+ say(format!("both_close: shut the sending half: {shut}, then {}", the_end(&mut stream)));
+}
diff --git a/tests/toyos.rs b/tests/toyos.rs
index ed832b7f6..ae7e1f79c 100644
--- a/tests/toyos.rs
+++ b/tests/toyos.rs
@@ -2,6 +2,8 @@
extern crate toyos_build;
mod common;
+#[path = "toyos-rust-tests/src/stream_ends.rs"]
+mod stream_ends;
use std::collections::{BTreeMap, BTreeSet};
use std::fs;
@@ -137,6 +139,9 @@ const RUST_SKIP: &[&str] = &[
// It needs a host that dials its listeners when it says they wait:
// `libc_sockets` runs it on `tests/netcase`.
"nodelay_accepted",
+ // It needs a host peer that ends each stream as the stream asks:
+ // `netstack_stream_ends` runs it on `tests/netcase`.
+ "stream_ends",
// It asserts nothing at all: it holds `dump_nmi_probe`'s boot open for
// twenty seconds. On a shared boot it would be twenty seconds of nothing.
"lan_hold",
@@ -270,6 +275,11 @@ const MACHINE_TESTS: &[&str] = &[
// peer that answers: the calls are requests on netstack's port, netstack
// has no host build, and the T14's peer is the bench's network.
"libc_sockets",
+ // How std reads each way a stream ends, against the host kernel's TCP
+ // running the same code: the end is what netstack, std and the pipes
+ // between them say together, netstack has no host build, and the T14's
+ // bench has no peer that resets on cue.
+ "netstack_stream_ends",
// The nested-NMI report is a raw write to the 16550, which the T14 does not
// have.
"nested_nmi_is_loud",
@@ -3079,6 +3089,97 @@ fn libc_sockets() -> Result<(), String> {
Ok(())
}
+/// Every way `stream_ends.rs` ends a stream, read by std in a guest through
+/// netstack and by the same code on this host through its kernel's TCP, both
+/// against one peer here: the guest's report is the host's, line for line, and
+/// the host's is the one the job spells.
+fn netstack_stream_ends() -> Result<(), String> {
+ const JOB: &str = "stream_ends";
+ let peer = stream_ends_peer()?;
+ let mut host = Vec::new();
+ stream_ends::run("127.0.0.1", peer, |line| host.push(line));
+ if host != stream_ends::EXPECTED {
+ return Err(format!("the oracle: this host's TCP reported\n{}\nand not\n{}", host.join("\n"), stream_ends::EXPECTED.join("\n")));
+ }
+
+ let bin = qemu::build_toyos_bin(qemu::SUITE_ARCH, &compile::repo_root().join("tests/toyos-rust-tests"), JOB);
+ let mut qemu = boot_netcase(&[], &[(JOB.to_string(), bin)], BootOptions::default())?;
+ let result = qemu.run_test(&format!("test_rs_{JOB} 10.0.2.2 {peer}"), Duration::from_secs(120));
+ if let Some(why) = &result.error {
+ return Err(format!("{why}\nthe job said:\n{}", result.stdout));
+ }
+ let guest: Vec<&str> = result.stdout.lines().filter_map(|l| l.trim_end().strip_prefix("stream_ends: ")).collect();
+ if guest.last() != Some(&"done") || guest[..guest.len() - 1] != host {
+ return Err(format!("the guest reported\n{}\nwhere this host's TCP reported\n{}", guest.join("\n"), host.join("\n")));
+ }
+ if result.exit_code != Some(0) {
+ return Err(format!("the job ended {:?}:\n{}", result.exit_code, result.stdout));
+ }
+ Ok(())
+}
+
+/// The peer `stream_ends.rs` dials, for as long as the process lives: its
+/// port. Each connection's first byte names how it ends, and each end waits
+/// on what the client did before it.
+fn stream_ends_peer() -> Result<u16, String> {
+ use std::io::{Read, Write};
+ use std::os::fd::AsRawFd;
+ /// A close that sends a reset (RFC 9293 §3.10.4's ABORT): a zero linger.
+ fn reset(stream: std::net::TcpStream) {
+ let linger = libc::linger { l_onoff: 1, l_linger: 0 };
+ // SAFETY: a live socket, and an option of the size its type says.
+ let set = unsafe {
+ libc::setsockopt(
+ stream.as_raw_fd(),
+ libc::SOL_SOCKET,
+ libc::SO_LINGER,
+ (&raw const linger).cast(),
+ size_of::<libc::linger>() as libc::socklen_t,
+ )
+ };
+ assert_eq!(set, 0, "a zero linger: {}", std::io::Error::last_os_error());
+ }
+ fn serve(mut stream: std::net::TcpStream) {
+ let mut what = [0u8; 1];
+ if stream.read_exact(&mut what).is_err() {
+ return;
+ }
+ match what[0] {
+ stream_ends::HALF_CLOSE => {
+ let mut all = Vec::new();
+ if stream.read_to_end(&mut all).is_ok() {
+ let _ = stream.write_all(&all);
+ }
+ }
+ stream_ends::PEER_FIN_FIRST => {}
+ stream_ends::RESET_MID_STREAM => {
+ let ahead: Vec<u8> = (0..stream_ends::AHEAD).map(stream_ends::pattern).collect();
+ let mut answer = [0u8; 1];
+ if stream.write_all(&ahead).is_ok() && stream.read_exact(&mut answer).is_ok() {
+ reset(stream);
+ }
+ }
+ stream_ends::RESET_AFTER_HALF_CLOSE => {
+ if stream.read_to_end(&mut Vec::new()).is_ok() {
+ reset(stream);
+ }
+ }
+ stream_ends::BOTH_CLOSE => {
+ let _ = stream.write_all(b"bye");
+ }
+ _ => {}
+ }
+ }
+ let peer = std::net::TcpListener::bind(("127.0.0.1", 0)).map_err(|e| format!("the peer: {e}"))?;
+ let port = peer.local_addr().map_err(|e| format!("the peer's port: {e}"))?.port();
+ thread::spawn(move || {
+ for stream in peer.incoming().flatten() {
+ thread::spawn(move || serve(stream));
+ }
+ });
+ Ok(port)
+}
+
/// Run the machine-shape test, which owns its QEMU: the machine shape *is* the
/// test.
fn run_machine_test(name: &str, test_config: &Path) -> Result<(), String> {
@@ -3086,6 +3187,7 @@ fn run_machine_test(name: &str, test_config: &Path) -> Result<(), String> {
"iommu_virtio_platform" => common::iommu::iommu_virtio_platform(test_config),
"netstack_socket_churn" => netstack_socket_churn(),
"libc_sockets" => libc_sockets(),
+ "netstack_stream_ends" => netstack_stream_ends(),
"nested_nmi_is_loud" => faults::nested_nmi_is_loud(test_config),
"machine_shutdown" => power::machine_shutdown(test_config),
"acpi_power_button" => power::acpi_power_button(test_config), |
|
Review, round 1, head Net: Evidence read at the cited heads, each with its log and exit: The questions in the brief:
BLOCKER
NOTE
What
|
|
Round 2, step 1: the row the review asks for, measured at Log The measurement patch (applies to `96f48560f`; not landed)diff --git a/tests/toyos.rs b/tests/toyos.rs
index ed832b7f6..90eb1472d 100644
--- a/tests/toyos.rs
+++ b/tests/toyos.rs
@@ -2,6 +2,8 @@
extern crate toyos_build;
mod common;
+#[path = "toyos-rust-tests/src/stream_ends.rs"]
+mod stream_ends;
use std::collections::{BTreeMap, BTreeSet};
use std::fs;
@@ -137,6 +139,9 @@ const RUST_SKIP: &[&str] = &[
// It needs a host that dials its listeners when it says they wait:
// `libc_sockets` runs it on `tests/netcase`.
"nodelay_accepted",
+ // It needs a host peer that ends each stream as the stream asks:
+ // `netstack_stream_ends` runs it on `tests/netcase`.
+ "stream_ends",
// It asserts nothing at all: it holds `dump_nmi_probe`'s boot open for
// twenty seconds. On a shared boot it would be twenty seconds of nothing.
"lan_hold",
@@ -270,6 +275,11 @@ const MACHINE_TESTS: &[&str] = &[
// peer that answers: the calls are requests on netstack's port, netstack
// has no host build, and the T14's peer is the bench's network.
"libc_sockets",
+ // How std reads each way a stream ends, against the host kernel's TCP
+ // running the same code: the end is what netstack, std and the pipes
+ // between them say together, netstack has no host build, and the T14's
+ // bench has no peer that resets on cue.
+ "netstack_stream_ends",
// The nested-NMI report is a raw write to the 16550, which the T14 does not
// have.
"nested_nmi_is_loud",
@@ -3079,6 +3089,108 @@ fn libc_sockets() -> Result<(), String> {
Ok(())
}
+/// Every way `stream_ends.rs` ends a stream, read by std in a guest through
+/// netstack and by the same code on this host through its kernel's TCP, both
+/// against one peer here: the guest's report is the host's, line for line, and
+/// the host's is the one the job spells.
+fn netstack_stream_ends() -> Result<(), String> {
+ const JOB: &str = "stream_ends";
+ let peer = stream_ends_peer()?;
+ let mut host = Vec::new();
+ stream_ends::run("127.0.0.1", peer, |line| host.push(line));
+ if host != stream_ends::EXPECTED {
+ return Err(format!("the oracle: this host's TCP reported\n{}\nand not\n{}", host.join("\n"), stream_ends::EXPECTED.join("\n")));
+ }
+
+ let bin = qemu::build_toyos_bin(qemu::SUITE_ARCH, &compile::repo_root().join("tests/toyos-rust-tests"), JOB);
+ let case = compile::repo_root().join("tests/netcase");
+ let c_bins = vec![("shut_first".to_string(), compile::link_toyos(&compile::compile_own_c(&case, "shut_first"), "shut_first"))];
+ let mut qemu = boot_netcase(&c_bins, &[(JOB.to_string(), bin)], BootOptions::default())?;
+ let c = qemu.run_test(&format!("test_c_shut_first 10.0.2.2 {peer}"), Duration::from_secs(120));
+ println!("the C job, exit {:?}, error {:?}:\n{}", c.exit_code, c.error, c.stdout);
+ let result = qemu.run_test(&format!("test_rs_{JOB} 10.0.2.2 {peer}"), Duration::from_secs(120));
+ println!("the Rust job, exit {:?}:\n{}", result.exit_code, result.stdout);
+ println!("this host's TCP:\n{}", host.join("\n"));
+ if let Some(why) = &result.error {
+ return Err(format!("{why}\nthe job said:\n{}", result.stdout));
+ }
+ let guest: Vec<&str> = result.stdout.lines().filter_map(|l| l.trim_end().strip_prefix("stream_ends: ")).collect();
+ if guest.last() != Some(&"done") || guest[..guest.len() - 1] != host {
+ return Err(format!("the guest reported\n{}\nwhere this host's TCP reported\n{}", guest.join("\n"), host.join("\n")));
+ }
+ if result.exit_code != Some(0) {
+ return Err(format!("the job ended {:?}:\n{}", result.exit_code, result.stdout));
+ }
+ Ok(())
+}
+
+/// The peer `stream_ends.rs` dials, for as long as the process lives: its
+/// port. Each connection's first byte names how it ends, and each end waits
+/// on what the client did before it.
+fn stream_ends_peer() -> Result<u16, String> {
+ use std::io::{Read, Write};
+ use std::os::fd::AsRawFd;
+ /// A close that sends a reset (RFC 9293 §3.10.4's ABORT): a zero linger.
+ fn reset(stream: std::net::TcpStream) {
+ let linger = libc::linger { l_onoff: 1, l_linger: 0 };
+ // SAFETY: a live socket, and an option of the size its type says.
+ let set = unsafe {
+ libc::setsockopt(
+ stream.as_raw_fd(),
+ libc::SOL_SOCKET,
+ libc::SO_LINGER,
+ (&raw const linger).cast(),
+ size_of::<libc::linger>() as libc::socklen_t,
+ )
+ };
+ assert_eq!(set, 0, "a zero linger: {}", std::io::Error::last_os_error());
+ }
+ fn serve(mut stream: std::net::TcpStream) {
+ let mut what = [0u8; 1];
+ if stream.read_exact(&mut what).is_err() {
+ return;
+ }
+ match what[0] {
+ stream_ends::HALF_CLOSE => {
+ let mut all = Vec::new();
+ if stream.read_to_end(&mut all).is_ok() {
+ let _ = stream.write_all(&all);
+ }
+ }
+ stream_ends::PEER_FIN_FIRST => {}
+ stream_ends::RESET_MID_STREAM => {
+ let ahead: Vec<u8> = (0..stream_ends::AHEAD).map(stream_ends::pattern).collect();
+ let mut answer = [0u8; 1];
+ if stream.write_all(&ahead).is_ok() && stream.read_exact(&mut answer).is_ok() {
+ reset(stream);
+ }
+ }
+ stream_ends::RESET_AFTER_HALF_CLOSE => {
+ if stream.read_to_end(&mut Vec::new()).is_ok() {
+ reset(stream);
+ }
+ }
+ stream_ends::BOTH_CLOSE => {
+ let _ = stream.write_all(b"bye");
+ }
+ stream_ends::SHUT_FIRST => {
+ if stream.write_all(b".").is_ok() && stream.read_to_end(&mut Vec::new()).is_ok() {
+ let _ = stream.write_all(b"late");
+ }
+ }
+ _ => {}
+ }
+ }
+ let peer = std::net::TcpListener::bind(("127.0.0.1", 0)).map_err(|e| format!("the peer: {e}"))?;
+ let port = peer.local_addr().map_err(|e| format!("the peer's port: {e}"))?.port();
+ thread::spawn(move || {
+ for stream in peer.incoming().flatten() {
+ thread::spawn(move || serve(stream));
+ }
+ });
+ Ok(port)
+}
+
/// Run the machine-shape test, which owns its QEMU: the machine shape *is* the
/// test.
fn run_machine_test(name: &str, test_config: &Path) -> Result<(), String> {
@@ -3086,6 +3198,7 @@ fn run_machine_test(name: &str, test_config: &Path) -> Result<(), String> {
"iommu_virtio_platform" => common::iommu::iommu_virtio_platform(test_config),
"netstack_socket_churn" => netstack_socket_churn(),
"libc_sockets" => libc_sockets(),
+ "netstack_stream_ends" => netstack_stream_ends(),
"nested_nmi_is_loud" => faults::nested_nmi_is_loud(test_config),
"machine_shutdown" => power::machine_shutdown(test_config),
"acpi_power_button" => power::acpi_power_button(test_config),
diff --git a/tests/netcase/shut_first.c b/tests/netcase/shut_first.c
new file mode 100644
index 000000000..8a9f568f0
--- /dev/null
+++ b/tests/netcase/shut_first.c
@@ -0,0 +1,32 @@
+/* A client that shuts its sending half down with nothing pending, then reads
+ the peer's four bytes and its FIN. argv: the address and the port of the
+ stream_ends peer. */
+#include <arpa/inet.h>
+#include <errno.h>
+#include <netinet/in.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/socket.h>
+
+int main(int argc, char **argv) {
+ if (argc != 3) return 2;
+ struct sockaddr_in to;
+ memset(&to, 0, sizeof to);
+ to.sin_family = AF_INET;
+ to.sin_port = htons((uint16_t)atoi(argv[2]));
+ inet_pton(AF_INET, argv[1], &to.sin_addr);
+ int fd = socket(AF_INET, SOCK_STREAM, 0);
+ if (connect(fd, (struct sockaddr *)&to, sizeof to) != 0) { printf("shut_first_c: connect failed\n"); return 1; }
+ char c = 'S';
+ printf("shut_first_c: send %ld\n", (long)send(fd, &c, 1, 0));
+ printf("shut_first_c: the answer %ld\n", (long)recv(fd, &c, 1, 0));
+ printf("shut_first_c: shutdown(SHUT_WR) %d\n", shutdown(fd, SHUT_WR));
+ char buf[64];
+ long got = 0, r;
+ while ((r = recv(fd, buf, sizeof buf, 0)) > 0) got += r;
+ int e = errno;
+ printf("shut_first_c: %ld bytes came, then recv %ld%s%s\n", got, r, r < 0 ? ", errno " : "",
+ r < 0 ? (e == ECONNRESET ? "ECONNRESET" : strerror(e)) : "");
+ return r == 0 && got == 4 ? 0 : 1;
+}
diff --git a/tests/toyos-rust-tests/src/stream_ends.rs b/tests/toyos-rust-tests/src/stream_ends.rs
new file mode 100644
index 000000000..e03cc6a21
--- /dev/null
+++ b/tests/toyos-rust-tests/src/stream_ends.rs
@@ -0,0 +1,146 @@
+//! Each way a stream can end, as std reports it to a client: the same code
+//! runs in a guest against netstack and on the harness's host against the host
+//! kernel's TCP, and the two reports are compared line for line.
+//!
+//! The peer is the harness's: a connection's first byte names what it does.
+//! Every end it makes waits on something the client did, so no line depends
+//! on how long anything took.
+//!
+//! Std alone: the harness compiles this file too.
+
+use std::io::{ErrorKind, Read, Write};
+use std::net::{Shutdown, TcpStream};
+
+/// What a client half-closes after, and reads back.
+pub const BULK: usize = 1 << 20;
+/// What the peer sends before the reset of `reset_mid_stream`.
+pub const AHEAD: usize = 1 << 16;
+
+/// The peer reads to the client's end, sends every byte back, then its FIN.
+pub const HALF_CLOSE: u8 = b'H';
+/// The peer sends its FIN before the client writes.
+pub const PEER_FIN_FIRST: u8 = b'F';
+/// The peer sends [`AHEAD`] bytes, reads the client's one, then resets.
+pub const RESET_MID_STREAM: u8 = b'R';
+/// The peer reads to the client's end, then resets.
+pub const RESET_AFTER_HALF_CLOSE: u8 = b'D';
+/// The peer sends three bytes and its FIN; the client then sends its own.
+pub const BOTH_CLOSE: u8 = b'B';
+/// The peer answers one byte; the client then shuts its sending half with
+/// nothing pending, and the peer, at that FIN, sends four bytes and its own.
+pub const SHUT_FIRST: u8 = b'S';
+
+/// The report a host's TCP gives, and so the one netstack owes.
+pub const EXPECTED: [&str; 8] = [
+ "half_close: wrote 1048576 bytes, then shut the sending half: Ok",
+ "half_close: 1048576 bytes came back as sent, then Ok(0)",
+ "peer_fin_first: the read Ok(0), then a write Ok(5)",
+ "reset_mid_stream: 65536 bytes came as sent, the answer Ok(1), then Err(ConnectionReset)",
+ "reset_after_half_close: shut the sending half: Ok, then Err(ConnectionReset)",
+ "both_close: 3 bytes came as sent, then Ok(0)",
+ "both_close: shut the sending half: Ok, then Ok(0)",
+ "shut_first: the answer Ok(1), shut the sending half: Ok, 4 bytes came as sent, then Ok(0)",
+];
+
+/// Byte `i` of anything sent here: a pattern no shift or cut reproduces.
+pub fn pattern(i: usize) -> u8 {
+ (i % 251) as u8
+}
+
+fn kind<T: std::fmt::Debug>(result: std::io::Result<T>) -> String {
+ match result {
+ Ok(_) => "Ok".to_string(),
+ Err(e) => format!("Err({:?})", e.kind()),
+ }
+}
+
+/// What one more read answers: the end, or a byte that is not one.
+fn the_end(stream: &mut TcpStream) -> String {
+ let mut byte = [0u8; 1];
+ match stream.read(&mut byte) {
+ Ok(n) => format!("Ok({n})"),
+ Err(e) => format!("Err({:?})", e.kind()),
+ }
+}
+
+/// Reads until `want` bytes, the end or a refusal: how many matched
+/// `expect`, and what stopped it, if anything did before `want`.
+fn take(stream: &mut TcpStream, want: usize, expect: impl Fn(usize) -> u8) -> (String, Option<String>) {
+ let mut buf = vec![0u8; 65536];
+ let mut got = 0;
+ while got < want {
+ let room = buf.len().min(want - got);
+ match stream.read(&mut buf[..room]) {
+ Ok(0) => return (format!("{got} bytes came"), Some("Ok(0)".to_string())),
+ Ok(n) => {
+ if let Some(at) = (0..n).find(|&i| buf[i] != expect(got + i)) {
+ return (format!("byte {} differs after {got} bytes", got + at), None);
+ }
+ got += n;
+ }
+ Err(e) if e.kind() == ErrorKind::Interrupted => {}
+ Err(e) => return (format!("{got} bytes came"), Some(format!("Err({:?})", e.kind()))),
+ }
+ }
+ (format!("{got} bytes came"), None)
+}
+
+fn dial(host: &str, port: u16, what: u8) -> TcpStream {
+ let mut stream = TcpStream::connect((host, port)).unwrap_or_else(|e| panic!("the peer at {host}:{port}: {e}"));
+ stream.write_all(&[what]).unwrap_or_else(|e| panic!("naming the end to the peer: {e}"));
+ stream
+}
+
+/// Every end, in order: one line each as [`EXPECTED`] spells them.
+pub fn run(host: &str, port: u16, mut say: impl FnMut(String)) {
+ let mut stream = dial(host, port, HALF_CLOSE);
+ let bulk: Vec<u8> = (0..BULK).map(pattern).collect();
+ let wrote = kind(stream.write_all(&bulk));
+ say(format!("half_close: wrote {BULK} bytes, then shut the sending half: {}", match wrote.as_str() {
+ "Ok" => kind(stream.shutdown(Shutdown::Write)),
+ _ => format!("the write {wrote}"),
+ }));
+ let (came, end) = take(&mut stream, BULK, pattern);
+ let end = end.unwrap_or_else(|| the_end(&mut stream));
+ say(format!("half_close: {} as sent, then {end}", came.replace("came", "came back")));
+
+ let mut stream = dial(host, port, PEER_FIN_FIRST);
+ let read = the_end(&mut stream);
+ let write = match stream.write(b"after") {
+ Ok(n) => format!("Ok({n})"),
+ Err(e) => format!("Err({:?})", e.kind()),
+ };
+ say(format!("peer_fin_first: the read {read}, then a write {write}"));
+
+ let mut stream = dial(host, port, RESET_MID_STREAM);
+ let (came, end) = take(&mut stream, AHEAD, pattern);
+ let line = match end {
+ Some(end) => format!("{came}, then {end}"),
+ None => {
+ let answer = match stream.write(b"k") {
+ Ok(n) => format!("Ok({n})"),
+ Err(e) => format!("Err({:?})", e.kind()),
+ };
+ format!("{came} as sent, the answer {answer}, then {}", the_end(&mut stream))
+ }
+ };
+ say(format!("reset_mid_stream: {line}"));
+
+ let mut stream = dial(host, port, RESET_AFTER_HALF_CLOSE);
+ let shut = kind(stream.shutdown(Shutdown::Write));
+ say(format!("reset_after_half_close: shut the sending half: {shut}, then {}", the_end(&mut stream)));
+
+ let mut stream = dial(host, port, BOTH_CLOSE);
+ let (came, end) = take(&mut stream, 3, |i| b"bye"[i]);
+ let end = end.unwrap_or_else(|| the_end(&mut stream));
+ say(format!("both_close: {came} as sent, then {end}"));
+ let shut = kind(stream.shutdown(Shutdown::Write));
+ say(format!("both_close: shut the sending half: {shut}, then {}", the_end(&mut stream)));
+
+ let mut stream = dial(host, port, SHUT_FIRST);
+ let answer = the_end(&mut stream);
+ let shut = kind(stream.shutdown(Shutdown::Write));
+ let (came, end) = take(&mut stream, 4, |i| b"late"[i]);
+ let end = end.unwrap_or_else(|| the_end(&mut stream));
+ say(format!("shut_first: the answer {answer}, shut the sending half: {shut}, {came} as sent, then {end}"));
+}
diff --git a/tests/toyos-rust-tests/src/bin/stream_ends.rs b/tests/toyos-rust-tests/src/bin/stream_ends.rs
new file mode 100644
index 000000000..293c7fccd
--- /dev/null
+++ b/tests/toyos-rust-tests/src/bin/stream_ends.rs
@@ -0,0 +1,22 @@
+//! `stream_ends.rs`'s report, in a guest: one line each, then `done`, and an
+//! exit of 1 where it is not the one a host's TCP gives.
+//!
+//! argv: the address and port of the harness's peer.
+
+#[path = "../stream_ends.rs"]
+mod stream_ends;
+
+fn main() {
+ let args: Vec<String> = std::env::args().collect();
+ let [_, host, port] = &args[..] else { panic!("usage: stream_ends <host> <peer port>") };
+ let port = port.parse().unwrap_or_else(|e| panic!("the port {port}: {e}"));
+ let mut report = Vec::new();
+ stream_ends::run(host, port, |line| {
+ println!("stream_ends: {line}");
+ report.push(line);
+ });
+ println!("stream_ends: done");
+ if report != stream_ends::EXPECTED {
+ std::process::exit(1);
+ }
+} |
|
Round 2, step 1, second arm: the same measurement with Log |
|
Round 2, step 2: libc's half of round 1, removed from this branch, for the move to carry. It reads right only on the node: on diff --git a/tests/libc-arch/src/lib.rs b/tests/libc-arch/src/lib.rs
index a6fa47e5d..fa2c1def0 100644
--- a/tests/libc-arch/src/lib.rs
+++ b/tests/libc-arch/src/lib.rs
@@ -9,7 +9,7 @@
//! `long double` widening against compiler-builtins', the errno codes against
//! `include/errno.h`, IPv4 addresses and their texts against the host C
//! library's, and the socket option rule against the host's `setsockopt` and
-//! `getsockopt`.
+//! `getsockopt`, and the rule that reads a stream's end.
#[cfg(test)]
extern crate alloc;
@@ -48,6 +48,9 @@ mod sigmask;
#[path = "../../../userland/libc/src/sockopt.rs"]
mod sockopt;
#[cfg(test)]
+#[path = "../../../userland/libc/src/streamend.rs"]
+mod streamend;
+#[cfg(test)]
#[path = "../../../userland/libc/src/strtonum.rs"]
mod strtonum;
#[cfg(test)]
@@ -88,6 +91,8 @@ mod signal_masks;
#[cfg(test)]
mod socket_options;
#[cfg(test)]
+mod stream_ends;
+#[cfg(test)]
mod strtonum_differential;
#[cfg(test)]
mod text_differential;
diff --git a/tests/libc-arch/src/stream_ends.rs b/tests/libc-arch/src/stream_ends.rs
new file mode 100644
index 000000000..d25480b32
--- /dev/null
+++ b/tests/libc-arch/src/stream_ends.rs
@@ -0,0 +1,30 @@
+//! libc's rule for a stream's end (`streamend.rs`): a read of 0 is the peer's FIN only while the
+//! send pipe still has its reader, a write the send pipe refuses for its reader's leaving is a
+//! reset, and nothing else either pipe says is read as one.
+
+use toyos_abi::syscall::SyscallError;
+
+use crate::streamend::{read_end, write_refused, Refusal};
+
+/// Every refusal but the two a pipe answers on purpose.
+const OTHERS: [SyscallError; 4] =
+ [SyscallError::InvalidArgument, SyscallError::PermissionDenied, SyscallError::BadAddress, SyscallError::Io];
+
+#[test]
+fn a_read_of_zero_is_the_peers_fin_only_while_the_send_pipe_is_read() {
+ assert_eq!(read_end(Ok(0)), Ok(()));
+ assert_eq!(read_end(Err(SyscallError::WouldBlock)), Ok(()), "a full send pipe still has its reader");
+ assert_eq!(read_end(Err(SyscallError::Gone)), Err(Refusal::Reset));
+ for other in OTHERS {
+ assert_eq!(read_end(Err(other)), Err(Refusal::Other), "{other:?}");
+ }
+}
+
+#[test]
+fn a_write_the_send_pipe_refuses_for_its_reader_is_a_reset() {
+ assert_eq!(write_refused(SyscallError::Gone), Refusal::Reset);
+ assert_eq!(write_refused(SyscallError::WouldBlock), Refusal::Again);
+ for other in OTHERS {
+ assert_eq!(write_refused(other), Refusal::Other, "{other:?}");
+ }
+}
diff --git a/userland/libc/src/lib.rs b/userland/libc/src/lib.rs
index 94e87be16..ef7401847 100644
--- a/userland/libc/src/lib.rs
+++ b/userland/libc/src/lib.rs
@@ -28,6 +28,7 @@ mod sigmask;
mod socket;
mod sockopt;
mod stdio;
+mod streamend;
mod string;
mod strtonum;
mod text;
diff --git a/userland/libc/src/socket.rs b/userland/libc/src/socket.rs
index 1a5986a94..0992bf218 100644
--- a/userland/libc/src/socket.rs
+++ b/userland/libc/src/socket.rs
@@ -8,9 +8,10 @@ use toyos_abi::syscall;
use toyos::net::{NetError, TcpOptions, TcpSocketId, UdpSocketId, OPT_BROADCAST, OPT_NODELAY};
use crate::errno::{
- EACCES, EADDRINUSE, EAFNOSUPPORT, EBADF, ECONNREFUSED, ECONNRESET, EFAULT, EINVAL, EIO, ENOMEM, ENOPROTOOPT, ENOSPC,
- ENOTCONN, EOPNOTSUPP, ETIMEDOUT,
+ EACCES, EADDRINUSE, EAFNOSUPPORT, EAGAIN, EBADF, ECONNREFUSED, ECONNRESET, EFAULT, EINVAL, EIO, ENOMEM, ENOPROTOOPT,
+ ENOSPC, ENOTCONN, EOPNOTSUPP, EPIPE, ETIMEDOUT,
};
+use crate::streamend::{self, Refusal};
use crate::inaddr::{self, SockaddrIn, AF_INET};
use crate::sockopt::{self, Kept};
@@ -69,6 +70,9 @@ struct SocketEntry {
// after its `connect`.
nodelay: bool,
broadcast: bool,
+ /// `shutdown` was asked for this half of a stream.
+ read_shut: bool,
+ write_shut: bool,
}
const MAX_SOCKETS: usize = 128;
@@ -178,6 +182,8 @@ pub unsafe extern "C" fn socket(domain: i32, sock_type: i32, _protocol: i32) ->
notify_fd: 0,
nodelay: false,
broadcast: false,
+ read_shut: false,
+ write_shut: false,
};
let fd = alloc_socket(entry);
if fd < 0 {
@@ -319,6 +325,8 @@ pub unsafe extern "C" fn accept(
notify_fd: 0,
nodelay: accepted.options.nodelay(),
broadcast: false,
+ read_shut: false,
+ write_shut: false,
};
let new_fd = alloc_socket(new_entry);
if new_fd < 0 {
@@ -346,10 +354,14 @@ pub unsafe extern "C" fn send(fd: i32, buf: *const u8, len: usize, _flags: i32)
match entry.kind {
SocketKind::Tcp => {
+ if entry.write_shut {
+ set_errno(EPIPE);
+ return -1;
+ }
let data = core::slice::from_raw_parts(buf, len);
match syscall::write(RawHandle(entry.tx_fd as u32), data) {
Ok(n) => n as isize,
- Err(_) => { set_errno(EIO); -1 }
+ Err(e) => { set_errno(stream_errno(streamend::write_refused(e))); -1 }
}
}
SocketKind::Udp => {
@@ -380,10 +392,17 @@ pub unsafe extern "C" fn recv(fd: i32, buf: *mut u8, len: usize, _flags: i32) ->
match entry.kind {
SocketKind::Tcp => {
+ if entry.read_shut {
+ return 0;
+ }
let data = core::slice::from_raw_parts_mut(buf, len);
- match syscall::read(RawHandle(entry.rx_fd as u32), data) {
+ let read = syscall::read(RawHandle(entry.rx_fd as u32), data).map_err(|_| Refusal::Other).and_then(|n| match n {
+ 0 => streamend::read_end(syscall::write_nonblock(RawHandle(entry.tx_fd as u32), &[])).map(|()| 0),
+ n => Ok(n),
+ });
+ match read {
Ok(n) => n as isize,
- Err(_) => { set_errno(EIO); -1 }
+ Err(refusal) => { set_errno(stream_errno(refusal)); -1 }
}
}
SocketKind::Udp => {
@@ -504,7 +523,7 @@ pub unsafe extern "C" fn shutdown(fd: i32, how: i32) -> i32 {
Some(s) => s,
None => { set_errno(EBADF); return -1; }
};
- let entry = match slot.as_ref() {
+ let entry = match slot.as_mut() {
Some(e) => e,
None => { set_errno(EBADF); return -1; }
};
@@ -514,10 +533,24 @@ pub unsafe extern "C" fn shutdown(fd: i32, how: i32) -> i32 {
set_errno(net_err_to_errno(e));
return -1;
}
+ entry.read_shut |= matches!(how, SHUT_RD | SHUT_RDWR);
+ entry.write_shut |= matches!(how, SHUT_WR | SHUT_RDWR);
}
0
}
+const SHUT_RD: i32 = 0;
+const SHUT_WR: i32 = 1;
+const SHUT_RDWR: i32 = 2;
+
+fn stream_errno(refusal: Refusal) -> i32 {
+ match refusal {
+ Refusal::Reset => ECONNRESET,
+ Refusal::Again => EAGAIN,
+ Refusal::Other => EIO,
+ }
+}
+
// close (for socket fds)
diff --git a/userland/libc/src/streamend.rs b/userland/libc/src/streamend.rs
new file mode 100644
index 000000000..0831218f8
--- /dev/null
+++ b/userland/libc/src/streamend.rs
@@ -0,0 +1,38 @@
+//! How a stream socket's end reaches its caller. A stream is two of netstack's pipes, and the
+//! receive pipe says only that it ended: netstack lets the send pipe go first when the connection
+//! failed, and keeps it after an orderly end for as long as its writer holds it, so what a
+//! zero-byte write into it answers after the end is the end's kind. It reads nothing but what it
+//! is handed, so the host tests it (`toyos-libc-copies`).
+
+use toyos_abi::syscall::SyscallError;
+
+/// What a call on a stream answers instead of bytes.
+#[derive(Clone, Copy, Debug, PartialEq, Eq)]
+pub(crate) enum Refusal {
+ /// `ECONNRESET`: the connection failed.
+ Reset,
+ /// `EAGAIN`.
+ Again,
+ /// `EIO`: the handle is not the pipe end the socket was made with.
+ Other,
+}
+
+/// A read of 0 from the receive pipe, given what a zero-byte write into the send pipe then
+/// answered: `Ok` is the peer's FIN, after its last byte.
+pub(crate) fn read_end(probe: Result<usize, SyscallError>) -> Result<(), Refusal> {
+ match probe {
+ Ok(_) | Err(SyscallError::WouldBlock) => Ok(()),
+ Err(SyscallError::Gone) => Err(Refusal::Reset),
+ Err(_) => Err(Refusal::Other),
+ }
+}
+
+/// A write the send pipe refused. One after a shutdown of the sending half is refused before it
+/// reaches the pipe, which netstack reads no more but keeps.
+pub(crate) fn write_refused(refused: SyscallError) -> Refusal {
+ match refused {
+ SyscallError::Gone => Refusal::Reset,
+ SyscallError::WouldBlock => Refusal::Again,
+ _ => Refusal::Other,
+ }
+} |
… it reads a half-close's FIN as a reset libc's recv probed the send pipe at a read of 0, as std does, and answered ECONNRESET where it found no reader. That reads right only on the node. On main, netstack on smoltcp closes both pipes in one pass once a connection is done in both directions, so a client that shut its sending half down reads its peer's FIN with the send pipe already gone. Measured in a guest at 96f4856 on tests/netcase: a C client that shuts its sending half down with nothing pending and reads the peer's four bytes and FIN read recv -1 ECONNRESET with this libc and recv 0 with main's; std read ConnectionReset in both, where the host's TCP read Ok(0). So userland/libc goes back to main's, with streamend.rs and its host test in toyos-libc-copies. The removed diff is posted on pull request #803 for the move to carry with the guest C case that the issue's exit now names. The issues say what main does by that measurement: the reset after a half-close is bridge_piped's, not the dropped bytes'; the race between the two closes was met (two runs of four); libc reads no end's kind; and libc's EPIPE without SIGPIPE is recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
…d a writer that leaves after the shutdown is tested [tcp]'s status answers writable 0 once the FIN is queued, so the pass's read loop already breaks on room == 0 and the watch's readable is already false: the two fin_queued conjuncts held nothing, and go. a_writer_that_leaves_after_the_shutdown_leaves_the_stream_to_its_reader plays the order no test played: the client shuts its sending half down, its writer leaves before the peer's FIN, the send pipe goes first and the stream lives on its to-client pipe until the peer's text and FIN. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Round 2: the mutations and the negative control at the final head Log m1-failure-drops-to-client-first --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -256,8 +256,8 @@
// The connection's failure: the client's send pipe goes first, so that the end it
// reads finds no reader behind it.
(Err(Error::Failed(_)), _) => {
- self.from_client = None;
self.to_client = None;
+ self.from_client = None;
}
(Err(Error::WouldBlock), _) => break,
(Err(refusal), _) => unreachable!("[tcp] refused the holder of a connection a read: {refusal:?}"),m10-writer-leaving-unheard diff --git a/userland/netstack/node/src/streams.rs b/userland/netstack/node/src/streams.rs
index 6e5bc08f4..bb9a6ea2c 100644
--- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -443,7 +443,6 @@ impl Node {
stream.left = true;
}
// A finished send pipe was kept for its writer alone.
- PipeEnd::FromClient if stream.fin_queued => stream.from_client = None,
PipeEnd::FromClient => stream.done_writing = true,
}
self.bridge(now);m2-fields-to-client-first --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -157,8 +157,8 @@
remote: Ipv4Addr,
/// Before `to_client`, so a stream dropped whole lets its send pipe go first: the order a
/// failure owes its client.
- from_client: Option<Box<dyn FromClient>>,
to_client: Option<Box<dyn ToClient>>,
+ from_client: Option<Box<dyn FromClient>>,
/// The connect is not answered yet, and no byte moves.
connecting: bool,
/// The connect's, while connecting; then the cut's, once the client can see the stream nom3-over-drops-send-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -301,7 +301,7 @@
// A connection that failed takes no more of the client's bytes: its writes fail rather
// than fill a pipe nobody drains, and its end reads as a failure. One that ended in order
// ended after the client's FIN, so its writing was over already.
- if status.failure.is_some() {
+ if matches!(status.state, State::Closed | State::TimeWait) {
self.from_client = None;
}
// A finished send pipe is kept only for a client that can still read the end and lookm4-fin-drops-send-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -287,9 +287,7 @@
if read == 0 {
stack.tcp_shutdown_write(now, conn);
self.fin_queued = true;
- if writer_gone {
- self.from_client = None;
- }
+ self.from_client = None;
break;
}
let Some(bytes) = space.get(..read) else { unreachable!("a pipe read {read} bytes into room for {room}") };m5-kept-pipe-unwatched --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -470,7 +470,7 @@
let (from, to) = (stream.from_client.is_some(), stream.to_client.is_some());
(*id, Watch {
readable: from && stream.room,
- writer: from && (!stream.done_writing || stream.fin_queued),
+ writer: from && !stream.done_writing,
writable: to && stream.held,
reader: to,
})m6-close-keeps-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -408,7 +408,6 @@
}
stream.to_client = None;
stream.done_writing = true;
- stream.left = true;
self.bridge(now);
}
m7-kept-pipe-timed --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -316,7 +316,7 @@
}
// Nothing of the stream reaches its client again, alive or not, and its pipe still holds
// what may be bytes to send.
- if self.to_client.is_none() && self.done_writing && !self.fin_queued {
+ if self.to_client.is_none() && self.done_writing {
if !self.extended {
let kept = extended.entry(self.remote).or_insert(0);
if *kept < OWNERLESS_PER_PEER {m8-reader-gone-keeps-pipe --- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -438,10 +438,7 @@
pub fn pipe_gone(&mut self, now: Instant, id: StreamId, end: PipeEnd) {
let Some(stream) = self.streams.live.get_mut(&id).filter(|stream| !stream.connecting) else { return };
match end {
- PipeEnd::ToClient => {
- stream.to_client = None;
- stream.left = true;
- }
+ PipeEnd::ToClient => stream.to_client = None,
// A finished send pipe was kept for its writer alone.
PipeEnd::FromClient if stream.fin_queued => stream.from_client = None,
PipeEnd::FromClient => stream.done_writing = true,m9-writer-leaving-takes-the-stream diff --git a/userland/netstack/node/src/streams.rs b/userland/netstack/node/src/streams.rs
index 6e5bc08f4..1e3f090b8 100644
--- a/userland/netstack/node/src/streams.rs
+++ b/userland/netstack/node/src/streams.rs
@@ -443,7 +443,10 @@ impl Node {
stream.left = true;
}
// A finished send pipe was kept for its writer alone.
- PipeEnd::FromClient if stream.fin_queued => stream.from_client = None,
+ PipeEnd::FromClient if stream.fin_queued => {
+ stream.from_client = None;
+ stream.to_client = None;
+ }
PipeEnd::FromClient => stream.done_writing = true,
}
self.bridge(now);Runner#!/bin/zsh
# Each patch: checked apply, node tests, reverse; the tree is checked clean after each.
cd <home path> || exit 9
M=<home path>
echo "head $(git rev-parse HEAD)"
for p in $M/m1-*.patch $M/m2-*.patch $M/m3-*.patch $M/m4-*.patch $M/m5-*.patch $M/m6-*.patch $M/m7-*.patch $M/m8-*.patch $M/m9-*.patch $M/m10-*.patch control; do
if [[ $p == control ]]; then name=control; git apply -R --check $M/control-node-change.patch || { echo "$name: does not apply"; exit 8; }; git apply -R $M/control-node-change.patch
else name=$(basename $p .patch); git apply --check $p || { echo "$name: does not apply"; exit 8; }; git apply $p; fi
log=$M/$name.log
cargo test -p toyos-net-node --no-fail-fast > $log 2>&1; ex=$?
if [[ $p == control ]]; then git apply $M/control-node-change.patch; else git apply -R $p; fi
reds=$(grep -E '^test .* FAILED$' $log | sed -e 's/^test //' -e 's/ \.\.\. FAILED$//' | tr '\n' ' ')
echo "$name EXIT=$ex red: $reds"
[[ -z $(git status --porcelain) ]] || { echo "tree not clean after $name"; git status --porcelain; exit 7; }
done
echo DONE |
|
Round 2 gates at r2-node.logr2-tcp.logr2-libc-copies.logr2-clippy.log |
|
Round 2 gates at |
|
Round 2 gates at r2-build.logr2-netguest.log |
|
Round 2 gates at |
|
Answer to the round-1 review (comment 6082008828), head BLOCKER 1 — fixed by removal, after measuringThe row you asked for was measured at
Your reading of What changed:
BLOCKER 2 — moved with the code it testedThe libc wiring is no longer on this branch, so nothing here is left unreached. The exit of
It must be red under each of your three mutations. The track's line says the same. NOTEs
Gates at
|
| gate | exit | result | comment |
|---|---|---|---|
| node tests | 0 | 157 tests | 6082190310 |
| [tcp] | 0 | 385 tests | 6082190310 |
toyos-libc-copies |
0 | 46 tests | 6082190310 |
| clippy | 0 | 24 invocations clean | 6082190310 |
--ci host |
0 | 78 steps | 6082273747 |
--build-only |
0 | 6082291835 | |
netstack_socket_churn and libc_sockets |
0 | 2 of 2 | 6082291835 |
| whole guest suite | 0 | 38 of 38, load 27.64 at start | 6082313797 |
The control has exit 101 with the six reds of round 1. m1 to m10 each have exit 101 (comment 6082159813).
|
Two secret gists linked above held the round's long host and suite logs. Both are deleted: their build lines carried this machine's per-user temporary-directory path, which identifies the machine and may not be posted. The pull request rests on the exits and lines quoted in its body and comments; the full logs stay on the development machine. |
|
Review, round 2, head Net:
Round 1's BLOCKERs
Round 1's NOTEs:
Evidence at
|
| check | exit | result |
|---|---|---|
| node | 0 | 157 tests (1+9+1+20+31+19+33+3+40) |
| [tcp] | 0 | 385 tests |
toyos-libc-copies |
0 | 46 tests |
| clippy | 0 | 24 invocations clean |
--ci host |
0 | Host: 78 step(s), all green; it runs userland/netstack/node's tests, the new one among them (r2-host.log:7767) |
--build-only |
0 | |
netstack_socket_churn and libc_sockets |
0 | 2 passed, 2 total |
| the whole guest suite | 0 | 38 passed, 38 total; no FAIL line |
The control exits 101 with the six named reds. m1 to m10 each exit 101 at the tests the body names, with the tree checked clean after each.
BLOCKER
None.
NOTE
- The run counts in two issues rest on a run whose log is gone. In
issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md: "Ok(0)in two runs", "in each of four runs". Inissues/a-shutdown-of-the-sending-half-drops-what-the-send-pipe-still-holds.md: "65,536 bytes ... in two runs". One of each count is round 1's run at the base, whose log was lost, as the body says; three runs are logged. Either state three, or re-run and post the run. Every conclusion stands on the three logged runs. - In the body, "each log posted on this pull request as it was made" is no longer true of
--ci hostand the guest suite. Comments 6082273747 and 6082313797 link gists that were deleted. Their exits and their lines stand quoted.
What host and guest / suite must show at the head that lands
main has moved to a944746d5, past this head's merge of 558283168; #799 changed the suite's boots. The head that lands is that merge, the merge queue's group commit.
hostmust be green at that exact head. It must runuserland/netstack/node's tests (157 here, 40 of them instreams.rs) witha_writer_that_leaves_after_the_shutdown_leaves_the_stream_to_its_readeramong them. This check is the one gate that carries the change's claims.guest / suitemust be green at that exact head. No guest boots the node until the move, so the suite reaches nothing of this change. Its green says only that the merged tree regresses nothing.- The pull request is a draft, so ci.yml's
hosthas not run on it. It must be marked ready first.
The orchestrator may land on reading those two checks green at that head. No BLOCKER is open, and both NOTEs concern records he corrects himself.
LAND
|
CI at |
…m end's order (#803) and the .local name's re-probe (#800), into the app grants Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
…nd's order (#803) and the .local name's probing (#800), into the HTTPS client ring_kat's line in DRIVEN_AND_SHARED met random_draws' there; the merge keeps main's, and the next commit decides ring_kat's place. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
… the stream end's order (#803), into the stop's hold of the console wire Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
The node keeps the pipe order std reads a stream's end by: a client that shut its sending half down reads its peer's FIN as an end, and a failure still reads as a reset.
The defect
A stream reaches its client as two pipes; the receive pipe's end says only that it ended. std's
ended(sdk/std/sys/net/connection.rs) reads the kind off the send pipe: no reader behind a receive pipe at its end is a failure. The node broke that order on an orderly end: it let the client's send pipe go as soon as the client's FIN was queued (streams.rs, theread == 0arm) and again when the connection reached CLOSED or TIME-WAIT. Soshutdown(Write)then read to end answeredConnectionReset(measured on the move, #801, both cards, 4,194,304 bytes). It also broke it for failures in a narrower way: a failure dropped the receive pipe first and the send pipe later in the same pass, and a stream dropped whole droppedto_clientfirst by declaration order, so a client reading the end between the two could read a reset as a FIN.Decision: the node keeps the order; std and the ABI are unchanged
Weighed, from the code:
kernel/src/pipe.rs,PIPE_SIZE; a connection is two pipes,kernel/src/object/service.rs). A reason the node keeps for the client to ask about after both pipes are gone is kept for ever for a client that died: with both ends let go nothing tells the node it left. The clean carrier is a word the writer sets on its pipe end as it lets go, which is a kernel change outside this fence. Recorded:issues/a-streams-failure-reaches-its-client-as-a-reset-whatever-it-was.md, on the track; connect's ICMP refusal (ERR_OTHERon the move) belongs to the same word set and is named there, not done here.What the node does now (
userland/netstack/node/src/streams.rs, its header states the contract):Watch::writer), and let go when its writer leaves, or when the client closes the stream or its reader is gone (nobody then reads the end:Stream::left).from_clientgo beforeto_client: in the pass's failure arm, and in any whole drop (pipe_broken, a broken pipe's abort),from_clientbeing declared first.status.failure). An orderly end of the connection comes after the client's FIN, so its writing was over already.OWNERLESS_LIFE): it holds no byte to send. It holds its place for as long as its client holds the stream, as any stream does.libc is not changed here: its reading of an end goes with the move
Round 1 also taught libc to read an end's kind as std does (a zero-byte write into the send pipe at a read of 0;
ECONNRESETwhere it finds no reader;EPIPEafterSHUT_WR, 0 afterSHUT_RD). That reads right only on the node. Onmainthe shipped netstack is the smoltcp shell, whosebridge_piped(userland/netstack/src/main.rs) closes both pipes in one pass once a connection is done in both directions, so a client that shut its sending half down finds the send pipe gone at the peer's FIN. Measured (below): with that libc, a C client'sshutdown(SHUT_WR)thenrecvto the end readsECONNRESETonmain; withmain's libc it reads 0. So the libc change is removed from this branch,userland/libcandtests/libc-archaremain's, and the removed diff with its host test is posted on this pull request (comment 6082081208) for the move to carry with the guest C caseissues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md's exit now names.What
maindoes at a stream's end, measuredThe instrument: one std program (
stream_ends) and one C program (shut_first.c) run in a guest ontests/netcaseagainst a peer on this host, and the same std source compiled into the harness and run against the same peer on this Mac's TCP, the two reports compared line for line. The peer is driven by the client's own steps, so no line depends on timing. The case is not landed (below); its patch is in comments 6081887880 (round 1) and 6082065591 (this round'sshut_firstrow and C job). Runs: round 1's two (one at the base, whose log was lost with a wiped scratch directory, one at1e303178b, comment 6081887880) and this round's two at96f48560f, one with round 1's libc (comment 6082065591, exit 1) and one withmain's libc (comment 6082078719, exit 1).main's netstack on smoltcp, stdshutdown(Write), read to endOk(0)Err(ConnectionReset)shut_first: shut down with nothing pending, peer then sends 4 bytes and FINOk(0)Err(ConnectionReset), both runsOk(0), then a writeOk(5)Err(ConnectionReset)Err(ConnectionReset)Ok(0), thenOk(0)Ok(0), thenOk(0)in round 1's two runs andErr(ConnectionReset)in this round's two: the race betweenbridge_piped's two closesC,
shut_first:recv -1, errno ECONNRESETwith round 1's libc;recv 0withmain's.So on
main, std reads the end of a half-closed stream as a reset whether or not a byte was cut: that isbridge_piped's, recorded inissues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md, and the cut count isissues/a-shutdown-of-the-sending-half-drops-what-the-send-pipe-still-holds.md's. Both close at the move, whose netstack runs the node. The smoltcp shell is not touched here: the owner's ruling is that nothing new is built on it.Why no guest test lands here
tests/netcaseboots netstack on smoltcp onmain, where the case is red for the two defects above, which only the move closes, and the owner's ruling is that no new test is built on smoltcp. The case is posted for the move instead; on the move it is also this change's guest negative control (this change reverted, the half-close row readsConnectionReset, as #801 measured). The node's host tests are the tier that reaches the rule itself: the pipe order is the node's, the node has a host build with faked pipe ends, and every rule here is red there under mutation.What the move's shell must do with this: nothing new in the ABI. It already lets a pipe go when the node drops it and reports
Node::pipe_gone(id, PipeEnd::FromClient)when the kernel says the send pipe's writer left; that watch must be asked wheneverWatch::writersays so, which is now also true of a kept end after the FIN (writer: from && (!done_writing || fin_queued)). Itsnetstack_streamscan read to the stream's end again.Recorded here
issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md: whatmaindoes, by the measurement above; the race was met; libc reads no end's kind; exit, on the move, the std case and the guest C case with the libc change, red under each of round 1's review's three libc mutations.issues/a-shutdown-of-the-sending-half-drops-what-the-send-pipe-still-holds.md: the count, and that the reset after it is the other issue's.issues/a-streams-failure-reaches-its-client-as-a-reset-whatever-it-was.md: one bit carries no failure's kind.issues/libc-answers-epipe-and-raises-no-sigpipe.md: libc'swriteanswersEPIPEand raises noSIGPIPE(libc has no signals), as ifMSG_NOSIGNALwere always given; whether it should is a libc design decision, its exit that decision and a guest C case on both paths.issues/toyos-has-its-own-network-stack.md.Checks
High-risk (the pipe order between netstack and every program, and std's reading of it). Every row at
052679491(this round's head,mainat558283168merged in), each log posted on this pull request as it was made except--ci hostand the guest suite, whose logs were too long for a comment and stay on the development machine (their exits and lines are quoted here).cargo test -p toyos-net-node --no-fail-fastcargo test -p toyos-net-tcp --no-fail-fast([tcp], unchanged)cargo test -p toyos-libc-copies --no-fail-fastmain's count (comment 6082190310)cargo run -- --clippyclippy: 24 invocations clean(comment 6082190310)cargo run -- --ci hostHost: 78 step(s), all green(comment 6082273747)cargo run -- --build-onlycargo test --test toyos-build -- netstack_socket_churn libc_sockets2 passed, 2 total(comment 6082291835)cargo test, the whole guest suite38 passed, 38 total; load average 27.64 at its start (comment 6082313797)Negative control: the node's whole change reverted onto the base (
git diff origin/main HEAD -- userland/netstack/node/src/streams.rs, applied reversed as a checked patch, restored after) under this branch's tests: exit 101, redthe_peers_fin_after_the_clients_shutdown_leaves_the_send_pipe_to_its_writer,a_shutdown_after_the_peers_fin_keeps_the_send_pipe_past_the_connections_end,a_shutdown_sends_what_the_pipe_held_and_then_the_fin,a_finished_send_pipe_goes_with_its_reader,a_reset_ends_both_pipes_and_the_stream,a_pipe_that_is_no_pipe_resets_its_connection. Green on the base too:a_reset_after_the_clients_shutdown_lets_the_send_pipe_go_first,a_reset_after_the_peers_fin_ends_the_send_pipeand this round'sa_writer_that_leaves_after_the_shutdown_leaves_the_stream_to_its_reader: those orders were right before, and the tests hold that they still are.Mutations at
052679491, each a checked patch, built, run with--no-fail-fast, reversed, the tree checked clean after; patches, runner and log in comment 6082159813.to_clientbeforefrom_clienta_reset_ends_both_pipes_and_the_stream,a_reset_after_the_clients_shutdown_lets_the_send_pipe_go_firstto_clientdeclared first againa_pipe_that_is_no_pipe_resets_its_connectiona_shutdown_after_the_peers_fin_keeps_the_send_pipe_past_the_connections_end,the_peers_fin_after_the_clients_shutdown_leaves_the_send_pipe_to_its_writera_finished_send_pipe_goes_with_its_reader,a_shutdown_after_the_peers_fin_keeps_the_send_pipe_past_the_connections_end,a_shutdown_sends_what_the_pipe_held_and_then_the_fin,the_peers_fin_after_the_clients_shutdown_leaves_the_send_pipe_to_its_writera_shutdown_sends_what_the_pipe_held_and_then_the_fin,the_peers_fin_after_the_clients_shutdown_leaves_the_send_pipe_to_its_writera_close_sends_what_the_pipe_held_and_then_the_fin,a_close_with_text_unread_resets,a_wake_is_owed_only_for_a_connection_there_is_a_place_forthe_peers_fin_after_the_clients_shutdown_leaves_the_send_pipe_to_its_writer(cut at 100 s past TIME-WAIT)a_finished_send_pipe_goes_with_its_readera_writer_that_leaves_after_the_shutdown_leaves_the_stream_to_its_readera_writer_that_leaves_after_the_shutdown_leaves_the_stream_to_its_reader,a_shutdown_after_the_peers_fin_keeps_the_send_pipe_past_the_connections_end,the_peers_fin_after_the_clients_shutdown_leaves_the_send_pipe_to_its_writerIndependent oracles: this host's TCP running the same program against the same peer (the table above, macOS); RFC 9293 §3.6 and §3.10.7.4 for the orders the node tests script. Not yet: Linux, and the guest case on the node, which is the move's.
Unsure
Net
git diff --shortstat origin/main...HEAD: 7 files, +336 / -49. Production: the node +67 / -21; tests +134 / -5;issues/the rest.🤖 Generated with Claude Code
https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
Known imprecision, corrected in the move (#801), which edits both files:
issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.mdandissues/a-shutdown-of-the-sending-half-drops-what-the-send-pipe-still-holds.mdcount one run whose log was lost with the orchestrator's scratch directory; three runs are logged, and every conclusion stands on those three.