Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
74 changes: 52 additions & 22 deletions issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,25 +4,55 @@ kind: defect
opened: 2026-09-26
---

# A netd client cannot tell a reset connection from the peer's FIN

A piped TCP connection reaches its client as two pipes, and the only thing
the receive pipe can say about the end of the stream is EOF: its write end
closed. `bridge_piped` (`userland/netd/src/main.rs`) closes that end for
every ending alike — the peer's FIN, the peer's RST, and netd resetting the
connection itself through `PipedConnection::refuse` (a pipe whose ring page
could not be allocated, or a handle netd cannot write or read). The client
reads the same zero-byte `read` in all of them, so a stream cut short by a
reset is indistinguishable from one the peer finished, and a protocol with no
length framing of its own takes a truncated stream as a whole one.

The `ResourceExhausted` case is reachable without a hostile client: the ring
page is allocated on a pipe's first write, and a peer that speaks first (an
ssh banner) makes netd's first write the allocation. The only record is
netd's own log line.

Read from the code, not measured.

Exit condition: the pipe protocol carries a reset as something other than
EOF — for instance an error the client's next `read` returns — and a test
whose connection netd resets sees that error rather than EOF.
# A netstack client cannot tell a reset connection from the peer's FIN

A stream reaches its client as two pipes, and the receive pipe's end says only
that it ended. std (`ended`, `sdk/std/sys/net/connection.rs`) reads the kind
off the send pipe: one with no reader behind a receive pipe at its end is a
failure. libc reads no kind: its `recv` answers every end 0 and every refused
read `EIO` (`userland/libc/src/socket.rs`), so a C client reads a reset as the
peer's FIN.

On the node that order holds: a failure lets the send pipe go before the
receive pipe, and an orderly end keeps a send pipe whose last byte and FIN are
queued until its writer leaves (`userland/netstack/node/src/streams.rs`'s
header; its host tests in `userland/netstack/node/tests/streams.rs`, each red
with the order broken).

What `main` ships is netstack on smoltcp, whose `bridge_piped`
(`userland/netstack/src/main.rs`) closes the receive pipe and then, in the same
pass, the send pipe of a connection that is no longer open, whether it failed
or ended in order. A client that reads the end between the two reads a
failure as a FIN, and one that reads it after both reads an orderly end as a
failure. A client that shut its sending half down before the peer's FIN is
the second: its connection is done in both directions at that FIN. Measured
in a guest on `main`, std, against the host kernel's TCP behind QEMU's user
network, the same program on the harness host's TCP (macOS) as the oracle:

- a client that shut its sending half down with nothing pending, then read the
peer's four bytes and its FIN, read the four and then `ConnectionReset`, in
each of two runs, where the host read `Ok(0)`; a C client doing the same
read `recv` 0 with `main`'s libc, which reads every end so, and
`ECONNRESET` with a libc that reads the kind as std does;
- a client that read the peer's FIN and then shut its sending half down read
its next read as `Ok(0)` in two runs and as `ConnectionReset` in two others,
where the host read `Ok(0)`: the race between the two closes;
- a reset mid-stream, a reset after the client's half-close and a FIN before
the client's write read as the host read them, in each of four runs.

**Exit condition**: netstack runs on the node, and on `tests/netcase`:

- a guest std test whose client shuts its sending half down reads its peer's
bytes and then the end as an end, one whose peer resets a stream reads the
reset as a reset, and each line it prints is the host kernel's for the same
program against the same peer;
- libc reads an end's kind as std does, and a guest C case reads `recv` 0 at
the peer's FIN after `shutdown(SHUT_WR)`, `ECONNRESET` on a reset mid-stream
and on one after `SHUT_WR`, `send` `EPIPE` after `SHUT_WR` and `recv` 0
after `SHUT_RD`; and it is red with `recv`'s probe of the send pipe replaced
by a plain 0, with `shutdown` not marking the sending half shut, and with
`recv` not answering 0 after `SHUT_RD`. The libc change and its host test
were written for the node and are posted on pull request #803 for the move
to carry with that case.

**Owner**: whoever holds `issues/toyos-has-its-own-network-stack.md`.
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,15 @@ stream that ends short with a clean FIN.
Dropping the write end instead is the path that works: the bridge reads the
pipe to its end and closes the socket after the last byte.

Read from the code, not measured. `netstack_socket_churn`
Measured in a guest on `main`, against the host kernel's TCP behind QEMU's
user network, with a peer that echoes what it read once the client's FIN
arrives: a client that wrote 1,048,576 bytes and shut its sending half down
at once read 65,536 bytes back as sent in two runs and 65,535 in two others,
where the same program on a host's TCP reads all 1,048,576 and then the end.
What it read after them, `ConnectionReset`, is not this defect but
`issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md`'s: a
client that shut its sending half down with nothing pending reads it too.
`netstack_socket_churn`
(`tests/toyos-rust-tests/src/bin/netstack_socket_churn.rs`) writes into the
pipe after the shutdown, which is the client's own error and not this.

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
---
status: open
kind: defect
opened: 2026-10-09
---

# A stream's failure reaches its client as a reset, whatever it was

The pipe ABI tells a stream's client how the stream ended by the order its two
pipes end in (`userland/netstack/node/src/streams.rs`'s header), which carries
one bit: the peer's FIN, or a failure. The node knows which failure
(`toyos_net_tcp::Failure`: a reset, R2's timeout, an ICMP error), and std
answers every one `ConnectionReset`, where Linux answers a connection [tcp]
gave up on `ETIMEDOUT` and one an ICMP error ended `EHOSTUNREACH` or
`ENETUNREACH`. libc reads no failure yet
(`issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md`), and
the change the move carries for it answers every one `ECONNRESET`.

A pipe's end carries no word, and every other carrier costs what this one does
not: a third pipe or a kept request connection is a 2 MiB page each
(`kernel/src/pipe.rs`, `PIPE_SIZE`), and a reason the node keeps for a client to
ask about after both pipes are gone is kept for ever for a client that died,
since nothing then tells the node it left. The clean carrier is a word the
writer of a pipe sets as it lets its end go and the reader reads at the end,
which is the kernel's to add.

The same word set would answer a connect that an ICMP error ended, which the
move's netstack answers `ERR_OTHER` for want of a word.

Read from the code, not measured.

**Exit condition**: a pipe's end carries its reason, the node writes
`Failure`'s through it, and std and libc answer each as Linux does; a node
host test reads each failure's word, and a guest test whose peer is
unreachable reads it as unreachable.

**Owner**: whoever holds `issues/toyos-has-its-own-network-stack.md`.
36 changes: 36 additions & 0 deletions issues/libc-answers-epipe-and-raises-no-sigpipe.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
---
status: open
kind: defect
opened: 2026-10-09
---

# libc answers EPIPE and raises no SIGPIPE

POSIX's `write` to a pipe or socket whose reader is gone, and its `send` on a
stream that can send no more, fail `EPIPE` and also send the calling thread
`SIGPIPE`, whose default action ends the process; `send` raises none only with
`MSG_NOSIGNAL`. ToyOS's libc fails `EPIPE` and raises nothing
(`userland/libc/src/posix_io.rs`, `SyscallError::Gone`): it has no signals
(`signal`, `sigaction` and `raise` in `userland/libc/src/misc.rs` do nothing),
and none of those
`issues/a-childs-end-is-an-event-and-a-parent-takes-its-children-down.md`
has it imitate is `SIGPIPE`. So every write behaves as if `MSG_NOSIGNAL` were
given, and a C program that relies on the default action to stop writing into a
closed pipe runs on, and stops only if it reads the error. `send` on a stream answers `EIO` for every refusal today; the libc
change the move carries for
`issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md` answers
`EPIPE` after `shutdown(SHUT_WR)`, and raises nothing either.

Whether libc raises `SIGPIPE` on these, or states the departure as it states
others, is a decision of libc's design that nobody has made.

Read from the code, not measured.

**Exit condition**: the decision is made and recorded at `posix_io.rs`, and a
guest C case asserts it on both paths, a `write` into a pipe whose reader is
gone and a `send` after `shutdown(SHUT_WR)`: with `SIGPIPE` raised, the default
action ends the process, an ignored one leaves `EPIPE` and `MSG_NOSIGNAL`
raises none; with the departure kept, each answers `EPIPE` and the process
lives on.

**Owner**: `userland/libc`.
1 change: 1 addition & 0 deletions issues/toyos-has-its-own-network-stack.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ What the node does not yet meet:
- The node counts a wake unspent until an accept arrives, and cannot tell an accept on its way from one that will never be sent: an owner that reads a wake and sends no accept is owed one wake fewer from then on, and the last connection waiting is not announced. That is the node's half of `issues/an-accept-that-never-reaches-netstack-strands-its-listener.md`. Exit: that issue's, by a client that cannot spend a wake without its accept arriving, or by an accept that waits in netd and needs no wake.
- A stream its client can see no more, its reader gone or the peer's FIN read through and its client writing no more, is reset once its pipe has given up no byte for 100 s, and a peer address restarts that clock for at most 16 such streams at once (`OWNERLESS_PER_PEER`, `userland/netstack/node/src/streams.rs`); one past them has 100 s from the pass that found it so, until one of the 16 is done and it takes its place. What holds in this stage, from the code: per address, 16, an estimate no measurement set. There is no floor on the rate: each of the 16 is kept for as long as its pipe gives up one byte per 100 s, which for a full 2 MiB pipe is 2,097,152 x 100 s, over six years, and then for its tail in [tcp] (the line below on a client gone with nothing left in its pipe). The 16 is no bound across addresses: addresses are not counted, and where a connect's address is its client's choice, an accepted stream's is the peer's, which on the link answers for as many addresses as it likes. The bound across them is the node's places (`userland/netstack/node/src/places.rs`), of which every stream holds one until the node lets it go, and there is no second one: in `peers_at_more_addresses_than_there_are_places_hold_no_stream_past_them` (`userland/netstack/node/tests/listeners.rs`) peers at four addresses bring sixteen connections each to a node with three places, which holds two such streams at a time, each kept alive past its 100 s, and refuses every accept and connect past them. `issues/netstack-cuts-a-departed-clients-unsent-tail-at-the-ceiling.md` stays open until netd runs on the node: on a host its exit's second clause is met (`a_departed_clients_tail_arrives_whole_while_its_peer_takes_it`) and its first for 16 streams an address. Exit: the T14 or a guest shows a peer holding one at a byte per 100 s, and the cut takes a floor on the rate, or the owner rules it needs none; and the move closes that issue with its bench row.
- A stream its peer reset is gone from the node at once, so for a client that still holds it `Node::set_nodelay` answers `false` and `Node::nodelay` `None`: `issues/a-stream-its-peer-reset-refuses-the-option-requests-a-host-answers.md`, carried by the node as built. std's `nodelay()` asks netd nothing: it answers from the value its `TcpStream` holds (`sdk/std/sys/net/connection.rs`), and libc's `getsockopt` of `TCP_NODELAY` from its socket's entry (`userland/libc/src/socket.rs`). What the node carries is `set_nodelay` on a reset stream, refused where macOS refuses it under another kind and Linux answers. The node cannot hold the stream for the client: its two pipe ends are all it has of the client, letting both go is how the client learns of the reset, and with both gone nothing tells it when the client leaves, so a kept stream would be kept for ever. Exit, outside the node: that issue's guest test is green, and a `getsockopt` of `TCP_NODELAY` on a stream its peer reset is answered, by libc from a value its socket holds as std's does or by a pipe ABI that gives netd a handle whose other end's leaving the kernel reports.
- How a stream ended reaches its client as the order its two pipes end in, which is one bit: the peer's FIN, or a failure. The node keeps a send pipe whose FIN is queued until its writer leaves and lets a failed stream's send pipe go first (`userland/netstack/node/src/streams.rs`), so std reads a FIN as a FIN and a failure as `ConnectionReset`, whichever failure it was: `issues/a-streams-failure-reaches-its-client-as-a-reset-whatever-it-was.md`. libc reads no end's kind until the move carries its change for `issues/a-netstack-client-cannot-tell-a-reset-from-the-peers-fin.md` (posted on pull request #803) with that issue's guest C case on `tests/netcase`. Exit: those two issues'.
- Every frame and every deadline ends in a pass over every stream, every transmit opportunity walks every stream twice (once for the count each peer address keeps alive, once to pass those whose connect is not answered), and `Node::next_deadline` reads every stream: linear in streams per frame, as netd's `bridge_piped` is today. For an idle stream a pass is one `recv_with`, two `status` and one read of its client's pipe, which in netd is a system call: 1,000 system calls a frame with 1,000 idle streams. A pass's `recv_with` also re-files its connection's deadline and offers it to the transmit round, so one frame makes every idle stream eligible and the next opportunity serves each to learn it has nothing. [tcp] cannot say which connections a segment, a timer, an ICMP error or a failed next hop moved: it reports `drain_eligible` and `drain_gone`, which are the round's. The change: [tcp] marks a held connection in `for_conn`, `tick`, `icmp` and `end` and hands the marks out once (`drain_moved`); the node maps a connection to its stream, keeps its deadlines ordered, and passes only the streams [tcp] moved, the one a pipe's wake names, and those due. Exit: at the move, netd's CPU time per frame on the T14 with 1, 100 and 1,000 idle streams beside one bulk stream is measured, and the change lands if 1,000 idle streams cost more than 12.3 µs a frame over what one costs. The 12.3 µs is the time a full frame takes on a wire of 1 Gbit/s, 1,538 bytes of 8 ns with its preamble and gap: past it the passes alone take longer than the bulk stream's frames take to arrive. It is a threshold by arithmetic and no measurement; by estimate, not measured, 1,000 system calls exceed it.
- A client that is gone with nothing left in its pipe is finished by [tcp]'s rules for a user who let go: a reset once it has been idle 60 s, where netd on smoltcp reset it 100 s after the client left whatever the peer said. A peer that acknowledges a byte a minute holds such a connection for as long as it has bytes to acknowledge. Exit: the T14 or a guest shows such a peer holding one, and [tcp]'s rule takes a bound from the client's leaving; or the owner rules the idle bound is the one.
- `Tcp::ready`, `Tcp::orphans` and `Tcp::options` are additions to [tcp] the specifications do not hold, made for the node's wakes, its places and the copy a stream keeps of its connection's options; their tests (`toyos-net-shard/tcp/tests/held.rs`) name no scenario id. `Shard::listen` refuses an address [ip] does not hold usable, by the rule [udp]'s bind refuses by (`udp.bind-address-not-local`), and counts nothing for it; the node's test of it (`a_listener_is_at_the_address_it_named_and_only_one_the_machine_holds`) names no scenario either. A listener whose address the machine loses stands as a datagram socket bound to it does, with its port and its place and its owner told nothing, and answers again once the address is back (`a_listener_stands_while_its_address_is_lost_and_answers_when_it_is_back`, `userland/netstack/node/tests/lease.rs`); no host was read for it. Exit: the specifications hold the three calls and the refusal, and the tests name their scenarios.
Expand Down
Loading
Loading