Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 8 additions & 1 deletion issues/toyos-has-its-own-network-stack.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,13 +17,19 @@ netd runs smoltcp. Its replacement is ToyOS's own stack, built clean-room: reade
- Then multi-core, netring (blocked on the owner's ABI ruling), TCP and IP hardening, IPv6, offloads, soak.
- This track owns the T14's outbound rows, the machine reaching its router and the internet on the I219, judged from the stick: they arrive on this stack and not on smoltcp (owner: "no smoltcp."), and until they do no T14 row reads the wired card (`issues/the-host-cannot-reach-the-t14-while-it-runs-toyos.md`).

Stage 5 lands in slices, the orchestrator's cut under the owner's words "i want smoltcp out as fast as possible": netd's decisions go into `toyos-net-node` (`userland/netstack/node`), pure and host-tested and shipped in nothing, and then one change moves netd onto it and deletes smoltcp. In the tree: the node, its DHCP lease, its clients' datagram sockets, its mDNS name and its resolver. Still to build on it: streams, listeners; then the move.
Stage 5 lands in slices, the orchestrator's cut under the owner's words "i want smoltcp out as fast as possible": netd's decisions go into `toyos-net-node` (`userland/netstack/node`), pure and host-tested and shipped in nothing, and then one change moves netd onto it and deletes smoltcp. In the tree: the node, its DHCP lease, its clients' datagram sockets, its mDNS name, its resolver and streams. Still to build on it: listeners; then the move.

What the node does not yet meet:

- Of what a third party wrote, only slirp's recorded OFFER and ACK have reached it (`userland/netstack/node/tests/slirp.rs`); every other frame it has answered is the tests' own, from the RFCs' layouts. Exit: netd runs on it against slirp in a guest and a router on the T14.
- It counts a DHCP message [udp] refused as `node.dhcp-unsent`, whatever the rule; the `dhcp.renew-unroutable` scenario owed above is not written. Exit: that scenario names the counter, or the node counts the renewal apart.
- With the link down the DHCP client keeps its timers: every 4 to 64 s the node has a deadline, builds a DISCOVER that [udp] refuses, and counts it in `dhcp.tx.discover` and `node.dhcp-unsent`, from `Node::new` on. Exit: `toyos-dhcp`'s client is told the link went down and waits for it, and the node's first DISCOVER is the one that leaves.
- No second TCP has been on the far end of a stream: the peer of `userland/netstack/node/tests/streams.rs` is the tests' own script, which acknowledges what arrives in order and loses, reorders and repeats nothing; what is not ours there is `etherparse`, which reads every segment the node emits and builds every one it receives. Exit: netd runs on the node against the host kernel's TCP through slirp in a guest.
- The node has no bound on its streams, where netd refuses a connect past `max_piped_connections`, and a connection whose client let go is [tcp]'s alone and counted by nobody; [tcp] has no table limit of its own, and each connection is two 64 KiB buffers of netd's. Exit, before netd runs on the node: a connect past a bound is refused, and the bound counts connecting streams and the connections [tcp] finishes alone.
- A stream its client can see no more, its reader gone or the peer's FIN read through and its client writing no more, is reset once its pipe has given up no byte for 100 s, and a peer address restarts that clock for at most 16 such streams at once (`OWNERLESS_PER_PEER`, `userland/netstack/node/src/streams.rs`); one past them has 100 s from the pass that found it so, until one of the 16 is done and it takes its place. What holds in this stage, from the code: per address, 16, an estimate no measurement set. There is no floor on the rate: each of the 16 is kept for as long as its pipe gives up one byte per 100 s, which for a full 2 MiB pipe is 2,097,152 x 100 s, over six years, and then for its tail in [tcp] (the line below on a client gone with nothing left in its pipe). There is no bound across addresses: addresses are not counted, so the node holds 16 times as many as there are addresses holding one, and where a connect's address is its client's choice, an accepted stream's is the peer's, which on the link answers for as many addresses as it likes. That bound is the stream bound of the line above, built one stage on as the node's places, of which every stream holds one until the node lets it go; this stage builds no second one. `issues/netstack-cuts-a-departed-clients-unsent-tail-at-the-ceiling.md` stays open until netd runs on the node: on a host its exit's second clause is met (`a_departed_clients_tail_arrives_whole_while_its_peer_takes_it`) and its first for 16 streams an address. Exit: a test of the node's places in which peers at more addresses than the node has places for each hold 16 such streams, and the node holds no stream past its places; the T14 or a guest shows a peer holding one at a byte per 100 s, and the cut takes a floor on the rate, or the owner rules it needs none; and the move closes that issue with its bench row.
- A stream its peer reset is gone from the node at once, so for a client that still holds it `Node::set_nodelay` answers `false` and `Node::nodelay` `None`: `issues/a-stream-its-peer-reset-refuses-the-option-requests-a-host-answers.md`, carried by the node as built. std's `nodelay()` asks netd nothing: it answers from the value its `TcpStream` holds (`sdk/std/sys/net/connection.rs`). What the node carries is `set_nodelay` on a reset stream, refused where macOS refuses it under another kind and Linux answers, and libc's `getsockopt` of `TCP_NODELAY`, which asks through `toyos::net::tcp_get_option` (`userland/libc/src/socket.rs`) and is refused where both hosts answer. The node cannot hold the stream for the client: its two pipe ends are all it has of the client, letting both go is how the client learns of the reset, and with both gone nothing tells it when the client leaves, so a kept stream would be kept for ever. Exit, outside the node: that issue's guest test is green, and a `getsockopt` of `TCP_NODELAY` on a stream its peer reset is answered, by libc from a value its socket holds as std's does or by a pipe ABI that gives netd a handle whose other end's leaving the kernel reports.
- Every frame and every deadline ends in a pass over every stream, every transmit opportunity walks every stream twice (once for the count each peer address keeps alive, once to pass those whose connect is not answered), and `Node::next_deadline` reads every stream: linear in streams per frame, as netd's `bridge_piped` is today. For an idle stream a pass is one `recv_with`, two `status` and one read of its client's pipe, which in netd is a system call: 1,000 system calls a frame with 1,000 idle streams. A pass's `recv_with` also re-files its connection's deadline and offers it to the transmit round, so one frame makes every idle stream eligible and the next opportunity serves each to learn it has nothing. [tcp] cannot say which connections a segment, a timer, an ICMP error or a failed next hop moved: it reports `drain_eligible` and `drain_gone`, which are the round's. The change: [tcp] marks a held connection in `for_conn`, `tick`, `icmp` and `end` and hands the marks out once (`drain_moved`); the node maps a connection to its stream, keeps its deadlines ordered, and passes only the streams [tcp] moved, the one a pipe's wake names, and those due. Exit: at the move, netd's CPU time per frame on the T14 with 1, 100 and 1,000 idle streams beside one bulk stream is measured, and the change lands if 1,000 idle streams cost more than 12.3 µs a frame over what one costs. The 12.3 µs is the time a full frame takes on a wire of 1 Gbit/s, 1,538 bytes of 8 ns with its preamble and gap: past it the passes alone take longer than the bulk stream's frames take to arrive. It is a threshold by arithmetic and no measurement; by estimate, not measured, 1,000 system calls exceed it.
- A client that is gone with nothing left in its pipe is finished by [tcp]'s rules for a user who let go: a reset once it has been idle 60 s, where netd on smoltcp reset it 100 s after the client left whatever the peer said. A peer that acknowledges a byte a minute holds such a connection for as long as it has bytes to acknowledge. Exit: the T14 or a guest shows such a peer holding one, and [tcp]'s rule takes a bound from the client's leaving; or the owner rules the idle bound is the one.
- The word each of [udp]'s refusals is answered in on the pipe ABI is `Refused` in `userland/netstack/node/src/datagram.rs`, tested by `each_refusal_is_answered_in_the_pipes_word_for_it` against no reader's scenario; `udp.no-ephemeral-port` has a word and no test that reaches it. Exit: the scenario owed below exists, and the test names its id.
- A client's datagram to a broadcast address is refused `udp.broadcast-not-permitted` and answered invalid input, where smoltcp sends it: the pipe ABI has no request that permits it, and std's `set_broadcast` (`sdk/std/sys/net/connection.rs`) answers `Ok` to a permission nothing grants. No program in the tree sends to a broadcast address or sets the option; the sources of the third-party C programs an image builds from a recipe have not been searched. The owner's ruling: "Of course carry the permission why would you even ask. The premise is and always has been to do things properly". The move of netd onto the node does not land with this open. Exit, both before the move, as a stage of their own: that search, for `SO_BROADCAST` and a `sendto` to a broadcast address, is posted with its command and output; and the pipe ABI carries the permission from std's and libc's setters to `Udp::set_broadcast`, a send refused for want of it answered in a word that says permission and not input.
- A refusal [udp] logs against a client's call waits in the stack until the node's next `receive`, `fire` or `link`, because only those pass over the stack's reports. Exit, at the move and not after it: the datagram calls end in that pass, or netd drains it once per turn of its loop.
Expand Down Expand Up @@ -59,6 +65,7 @@ What stage 4 departs from its specifications:
- [ip]'s `Event::Cleared` and `Event::Room` are not in the specifications. Exit: they name them.
- The readers' stage 4 design has a flow waiting for its next hop keep its place in the round. [shard]'s round lets it go, deficit and all, and [tcp] offers it again when woken, at the round's tail, so a waiting flow costs the round nothing. Exit: a `toyos-net-shard/tests/drr.rs` test finds a woken flow served in the place it held when it began to wait.
- The readers' stage 4 design has a flow leave the round once it has nothing eligible, its deficit reset. [udp] says so as it hands out a sender's last datagram; [tcp] learns it only when it serves the connection, since only its `next_segment` decides what is due, so a connection whose turn ended on its last segment stays in the round with its deficit until its next turn, and what it queues before then is sent against that deficit, debt included, as RFC 8290 §4.2 keeps an emptied queue until it is next selected. Exit: [tcp] reports at hand-off that nothing is left, and a `toyos-net-shard/tests/drr.rs` test finds such a connection out of the round, its deficit reset.
- `Tcp::recv_with`, the read of a reader that may take less than it is shown, is not in the specifications: the node's bridge to a client's pipe needs it, since a byte `recv` copied out has opened the window and cannot be put back. Its tests (`toyos-net-shard/tcp/tests/take.rs`) name no scenario. Exit: the specification names the call and its scenarios, and the tests take their ids.
- The readers' TCP specification disagrees with itself, and [tcp] follows its user timeout: one replaces R2 and runs only while sent data is unacknowledged or a zero window holds data back, so with one set a connection whose data never left, for want of credit or of a next hop, has no give-up, where its rule that give-up clocks run on wall time has a device that never offers credit end its connections. Exit: RFC 9293 §3.10.8 ends a connection in any state once its user timeout expires: a test that sets one and withholds credit, as `s_pl_009_give_up_without_credit` does without one, finds the connection ended.

Open for the owner. On 2026-10-03 he chose "Fair scheduler now": "Build the byte-fair scheduler as the network spec describes, about 280 lines, and decide later on the laptop whether new connections should jump the queue." It is built as one round, in the order flows became eligible, and nobody jumps it. The track reads a queue that new connections jump as RFC 8290's new-flows list, and offers him this evidence: a sparse flow such as a DNS query waits one turn of every flow ahead of it in the round, up to a quantum of 1,514 bytes each, where [udp]'s datagrams and [tcp]'s segments alternated frame by frame before, so a query waited at most one TCP frame: in DRR-02 a 100-byte segment that becomes eligible leaves behind a full frame of each of the three bulk flows ahead of it (`s_shard_drr_002_a_bulk_flow_starves_no_other_beyond_its_quantum`); a T14 measurement at stage 5 of a sparse flow's wait behind bulk flows would add to it. Exit: the owner rules on whether new connections jump the queue.
Expand Down
13 changes: 13 additions & 0 deletions toyos-net-shard/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -489,6 +489,19 @@ impl Shard {
done
}

/// [`Tcp::recv_with`]: `take` is shown the oldest bytes held and answers how many it took.
pub fn recv_with(&mut self, now: Instant, id: ConnId, take: impl FnOnce(&[u8]) -> usize) -> Result<Received, toyos_net_tcp::Error> {
let done = self.tcp.recv_with(now, id, take);
self.settle(now);
done
}

pub fn set_options(&mut self, now: Instant, id: ConnId, options: toyos_net_tcp::Options) -> Result<(), toyos_net_tcp::Error> {
let done = self.tcp.set_options(now, id, options);
self.settle(now);
done
}

pub fn shutdown_write(&mut self, now: Instant, id: ConnId) -> Result<(), toyos_net_tcp::Error> {
let done = self.tcp.shutdown_write(now, id);
self.settle(now);
Expand Down
12 changes: 12 additions & 0 deletions toyos-net-shard/tcp/src/conn.rs
Original file line number Diff line number Diff line change
Expand Up @@ -865,6 +865,18 @@ impl Sync {
}
}

/// [`Self::recv`] for a reader that may take less than it is shown.
pub fn recv_with(&mut self, take: impl FnOnce(&[u8]) -> usize) -> Result<Received, Error> {
if self.rx.unread() == 0 {
return if self.rx.closed { Ok(Received::End) } else { Err(Error::WouldBlock) };
}
let n = self.rx.read_with(take);
if n > 0 {
self.rx.after_read(self.smss());
}
Ok(Received::Data(n))
}

pub fn shutdown_write(&mut self, now: Instant) {
if self.tx.fin.is_some() {
return;
Expand Down
9 changes: 9 additions & 0 deletions toyos-net-shard/tcp/src/rx.rs
Original file line number Diff line number Diff line change
Expand Up @@ -338,6 +338,15 @@ impl Rx {
self.buf.read(out)
}

/// Shows `take` the oldest bytes held, as far as they lie in one piece, and lets go of as
/// many of them as it answers it took.
pub fn read_with(&mut self, take: impl FnOnce(&[u8]) -> usize) -> usize {
let (held, _) = self.buf.slices(0, self.buf.len());
let n = take(held).min(held.len());
self.buf.consume(n);
n
}

/// `shutdown_read`: what is held is dropped, and later text is dropped as if read.
pub fn stop_reading(&mut self) {
self.discard = true;
Expand Down
19 changes: 19 additions & 0 deletions toyos-net-shard/tcp/src/stack.rs
Original file line number Diff line number Diff line change
Expand Up @@ -998,6 +998,25 @@ impl Tcp {
result
}

/// [`Self::recv`] for a reader that may take less than it is shown, such as a pipe that is
/// nearly full: `take` sees the oldest bytes held, as far as they lie in one piece, and answers
/// how many of them it took; no other byte leaves the buffer. With nothing held it is not
/// called, and the answer is [`Self::recv`]'s; `Data(0)` is a reader that took nothing.
pub fn recv_with(&mut self, now: Instant, id: ConnId, take: impl FnOnce(&[u8]) -> usize) -> Result<Received, Error> {
let conn = self.conn(id)?;
let result = match &mut conn.state {
Tcb::SynSent(_) | Tcb::SynRcvd(_) => Err(Error::WouldBlock),
Tcb::Sync(sync) => sync.recv_with(take),
Tcb::Ended(Ended { failure: Some(failure), .. }) => Err(Error::Failed(*failure)),
Tcb::Ended(Ended { failure: None, rx }) => match rx.as_mut() {
Some(rx) if rx.unread() > 0 => Ok(Received::Data(rx.read_with(take))),
_ => Ok(Received::End),
},
};
self.settle(id.index, now);
result
}

pub fn shutdown_write(&mut self, now: Instant, id: ConnId) -> Result<(), Error> {
let conn = self.conn(id)?;
let result = match &mut conn.state {
Expand Down
Loading
Loading