Repository navigation
Conversation
…so ToyOS's TCP offers a scaled window toyos-net-tcp already negotiated RFC 7323 window scaling and timestamps with PAWS; it took the smallest shift that fits the receive buffer, and netstack's 65,535-byte buffer made that shift 0, so a sender never had more than 64 KiB in flight. The T14's download read 35.1 Mb/s, 65,535 bytes per 14.9 ms. - netstack's TCP_BUFFER is 4 MiB each way: gigabit to a 33 ms round trip, at window scale 7. PLACE_BYTES is now a stream's (two pipes and two 4 MiB buffers, 12 MiB) rather than a listener's. - The receive buffer starts at 65,535 (limits::RECEIVE_BUFFER_INITIAL) and, once a round trip has passed, grows to twice what the user read per round trip, up to the configured buffer. Only reads grow it: a peer sending into a connection nobody reads, a listener's unaccepted children included, finds 65,535 bytes of room and no more. A window field at the negotiated shift bounds it, so an unscaled connection never grows past 65,535. - A ring's storage grows, by doubling, with the furthest offset written, so a 4 MiB send buffer costs what it holds. - After a read, a window update leaves when the window could at least double. netstack, like the test network, transmits a batch's ACKs before its reader drains the batch, so the ACKs offer the window less that batch; without the update the sender learned of the room only with the next data's ACK, one window per two round trips, and the growth rule, reading half a window per round trip, never grew the buffer past about 133 KB. Linux's tcp_cleanup_rbuf sends the same update. - A scaled window that would round to a zero field is offered as one unit where the buffer has room, as Linux's __tcp_select_window does. With a hole open the edge holds, and at shift 7 a window under 128 bytes read as zero: the sender persisted and resent one byte of the hole per backed-off probe. A property run at the new buffer found it. s_net_005 runs at netstack's buffer: at 65,535 a window-limited sender's window reopens each round on one update ACK, which its every-second-ACK loss takes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
|
Mutation patches at m1-no-scale-offereddiff --git a/toyos-net-shard/tcp/src/open.rs b/toyos-net-shard/tcp/src/open.rs
index 7322f8503..af20ede15 100644
--- a/toyos-net-shard/tcp/src/open.rs
+++ b/toyos-net-shard/tcp/src/open.rs
@@ -160,7 +160,7 @@ fn syn_options(local: &Local, peer: Option<&Negotiated>, now: Instant) -> SynOpt
mss: Some(local.mss),
sack_permitted: peer.is_none_or(|n| n.sack),
timestamps,
- window_scale: if peer.is_none_or(|n| n.scaled) { WindowShift::new(local.shift).ok() } else { None },
+ window_scale: None,
}
}
m2-absent-option-ignoreddiff --git a/toyos-net-shard/tcp/src/open.rs b/toyos-net-shard/tcp/src/open.rs
index 7322f8503..278ba060c 100644
--- a/toyos-net-shard/tcp/src/open.rs
+++ b/toyos-net-shard/tcp/src/open.rs
@@ -38,7 +38,7 @@ pub fn negotiate(seg: &In<'_>, local: &Local, first_tsval: u32, ctx: &mut Ctx<'_
}
Some(mss) => mss,
};
- let scaled = options.window_scale().is_some();
+ let scaled = true;
let (snd_shift, rcv_shift) = match options.window_scale() {
Some(scale) => {
if scale.raw() > scale.effective() {
@@ -46,7 +46,7 @@ pub fn negotiate(seg: &In<'_>, local: &Local, first_tsval: u32, ctx: &mut Ctx<'_
}
(scale.effective(), local.shift)
}
- None => (0, 0),
+ None => (0, local.shift),
};
let ts = options.timestamps().map(|t| Ts { recent: t.value, recent_at: ctx.now, offset: local.ts_offset, first: first_tsval });
Negotiated { peer_mss, snd_shift, rcv_shift, scaled, sack: options.sack_permitted(), ts }m3-no-growthdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..a12195bc7 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -392,7 +392,7 @@ impl Rx {
let read = u128::try_from(read).unwrap_or(u128::MAX);
let per_rtt = read.saturating_mul(rtt.as_nanos()).checked_div(now.since(began).as_nanos()).unwrap_or(0);
let want = usize::try_from(per_rtt.saturating_mul(2)).unwrap_or(usize::MAX);
- self.buf.grow(want.min(self.max));
+ let _ = want;
self.round = (now, 0);
}
if let Some(candidate) = self.candidate(mss) {m4-fixed-full-bufferdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..54e3b37b2 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -90,7 +90,7 @@ impl Rx {
next,
edge: next.add(window),
shift,
- buf: Ring::new(max.min(initial)),
+ buf: Ring::new(max),
ranges: Vec::new(),
stamp: 0,
fin: None,m5-no-window-updatediff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..1dd7f64cc 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -396,7 +396,7 @@ impl Rx {
self.round = (now, 0);
}
if let Some(candidate) = self.candidate(mss) {
- if self.last_window < self.threshold(mss) || candidate.since(self.next) >= self.window().saturating_mul(2) {
+ if self.last_window < self.threshold(mss) {
self.ack_now = true;
}
}m6-no-round-updiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..a60e8a25b 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -274,7 +274,7 @@ impl Rx {
// room, as Linux does: else a sender owing the text of a hole waits on a window not shut.
let unit = 1u32.checked_shl(u32::from(self.shift)).unwrap_or(u32::MAX);
let window = edge.since(self.next);
- let edge = if window > 0 && window < unit && unit <= self.free() { self.next.add(unit) } else { edge };
+ let _ = (window, unit);
let field = (edge.since(self.next) >> self.shift).min(u32::from(u16::MAX));
(edge, u16::try_from(field).unwrap_or(u16::MAX))
}m7-no-relayoutdiff --git a/toyos-net-shard/tcp/src/ring.rs b/toyos-net-shard/tcp/src/ring.rs
index aaf45daab..e47345fce 100644
--- a/toyos-net-shard/tcp/src/ring.rs
+++ b/toyos-net-shard/tcp/src/ring.rs
@@ -53,7 +53,7 @@ impl Ring {
/// offset, laid out again from the head.
fn reserve(&mut self, end: usize) {
if end > self.bytes.len() {
- self.bytes.rotate_left(self.head);
+ let _ = 0;
self.head = 0;
let len = end.max(self.bytes.len().saturating_mul(2)).min(self.capacity);
self.bytes.resize(len, 0);m8-no-offerable-capdiff --git a/toyos-net-shard/tcp/src/rx.rs b/toyos-net-shard/tcp/src/rx.rs
index a23e09cc2..65693a591 100644
--- a/toyos-net-shard/tcp/src/rx.rs
+++ b/toyos-net-shard/tcp/src/rx.rs
@@ -85,7 +85,7 @@ impl Rx {
pub fn new(next: Seq, max: usize, shift: u8, window: u32, now: Instant) -> Self {
let initial = usize::try_from(crate::limits::RECEIVE_BUFFER_INITIAL).unwrap_or(usize::MAX);
let offerable = usize::from(u16::MAX).checked_shl(u32::from(shift)).unwrap_or(usize::MAX);
- let max = max.min(offerable);
+ let _ = offerable;
Self {
next,
edge: next.add(window), |
|
T14 at |
|
Review of #820 at Net lines, own change: production +102/−37 (net +65: BLOCKER
NOTE
SEND BACK |
Carries #801 (
wt/toyos-move, merged at409a4fa9b) until it lands; this branch's own change isd286ece79,git diff d286ece79~1 d286ece79.What changed and why
toyos-net-tcpalready negotiated RFC 7323 window scaling (both ways, never unless both SYNs offer it, the shift clamped at 14 and countedWscaleClamped), and timestamps with PAWS. It picked the smallest shift that fits the receive buffer, and netstack'sTCP_BUFFERof 65,535 made that shift 0. So the work is the buffer, and what the buffer's growth needs from the receiver.TCP_BUFFERis 4 MiB each way, window scale 7. At 1 Gb/s that covers a 33.5 ms round trip; Ubuntu on the T14 autotunes to 6 MB (tcp_rmemmax).PLACE_BYTESbecomes a stream's: two 2 MiB pipes and two 4 MiB buffers, 12 MiB, up from the listener's 10 MiB (128 × 65,535 + 2 MiB). That is about 170 places on 16 GiB instead of about 200, under the same one-eighth share.rx's module doc). It starts atlimits::RECEIVE_BUFFER_INITIAL= 65,535, the most a SYN offers. Once a round trip (SRTT) has passed, it grows to twice what the user read per round trip, up to the configured buffer, and it never shrinks (RFC 7323 §2.4). Text nobody reads grows nothing: a listener's 128 unaccepted children, or a stalled reader, hold at most 65,535 each. A window field at the negotiated shift bounds the growth, so an unscaled connection stays at 65,535. The base had a latent defect here: with a buffer over 65,535 and a peer that does not scale, the edge it recorded ran past what the 16-bit field had offered (83,955 measured).tcp_cleanup_rbufrule). netstack transmits a received batch's ACKs beforesockets.bridgedrains the batch into the client's pipe, and the test network does the same. Those ACKs offer the window less the batch, and the old rule sent no update unless the window was under one MSS. Measured in the test network: the sender learned of the room one round trip late, and moved one window per two round trips. The growth rule then read half a window per round trip and settled at about 133 KB. With the update, 32 MiB each way over gigabit at 20 ms takes 721 simulated ms against 5,921, and the window reaches 4 MiB. The 64 KiB reading on the T14 (35.1 Mb/s = 65,535 B / 14.9 ms) may therefore be half a window per 7.5 ms rather than a whole one per 14.9; The T14 downloads over HTTPS through ToyOS's own stack inside its list's bound, and netstack's 64 KiB receive window is filed as the ceiling its rate sits on #818'srtt_mswill say which.__tcp_select_windowdoes). With a hole open the edge holds, and at shift 7 a 93-byte window read as zero. The sender then persisted and resent one byte of the hole per backed-off probe: a property run at the new buffer stalled for 600 s. This is reachable only with a shift above 0, so this change is what makes it reachable.conn.rs). At 1 Gb/s the 2^31 sequence half-space wraps in about 17 s, inside the 60 s TIME-WAIT/MSL, so PAWS is what keeps an old duplicate out at these windows. Nothing to add.Tests and their controls
Host tests,
toyos-net-shard/tcp/tests/net.rs(ours against ours, the crate's consistency control):rfc_7323_2_a_scaled_window_carries_more_than_64_kib_in_flight_each_way: gigabit credit, 20 ms, 32 MiB each way. Asserts shift (7,7) on both ends, flight > 65,535 each way, the window grown to the whole 4 MiB, and every byte arrives. A second arm drops 1 data segment in 3,000, so out-of-order text is held across growth.rfc_7323_2_2_no_scaling_unless_both_syns_offer_it: one SYN stripped of its options. Asserts shift (0,0), flight and window ≤ 65,535 at a 4 MiB buffer, and every byte arrives.a_receive_buffer_nobody_reads_never_grows: 10 s of a peer sending into an unread connection leaves unread and window ≤ 65,535 and the sender's flight ≤ 65,535. Once read, the window grows past 65,535 and every byte arrives.props.rs: odd seeds of every property run (loss, reordering, duplication, an adversary, every invariant checked after each event) use 4 MiB.s_net_005_lost_acksruns at netstack's buffer. At 65,535 with the update rule it reds at both drop parities (and passes at both without the rule): a window-limited sender's window reopens each round on one update ACK, and the scenario's every-second-ACK loss takes it, costing an RTO each time. A tail-loss probe is what would cover it; this is recorded here and in the commit.Already in the tree and unchanged: option parsing and the malformed or out-of-range Window Scale (
toyos-net-wires_topt_007,s_topt_019,s_tser_009), the clamp to 14 counted by name (s_op_010), no scaling when the peer offers none (s_op_011), Window Scale outside a SYN ignored (s_op_012), a SYN-ACK window never scaled (s_op_013), and the shift that fits the buffer (s_op_031).Mutations: each applied as a checked patch,
cargo test --offline --no-fail-fastintoyos-net-shard/tcp, thengit apply -R; every one RESTORED=0 with the tree clean. All exit 101. The patches are in the comment below.rfc_7323_2_…in_flight_each_way,a_receive_buffer_nobody_reads_never_grows,s_op_001,s_op_003,s_op_031, 21 handshake testsrfc_7323_2_2_no_scaling…,s_op_002,s_op_011, 5 simultaneous-open testsrfc_7323_2_…in_flight_each_way,a_receive_buffer_nobody_reads_never_growsa_receive_buffer_nobody_reads_never_growsrfc_7323_2_…in_flight_each_ways_prop_004s_net_*, and the three new testsrfc_7323_2_2_no_scaling…The whole change reverted (
toyos-net-shard/tcp/srcat the basece95ca9a8, the new tests kept): exit 101,a_receive_buffer_nobody_reads_never_growsandrfc_7323_2_2_no_scaling…red.rfc_7323_2_…in_flight_each_waystays green there: the base crate already scales once it is given a 4 MiB buffer. That test pins netstack's size, and m5 is its control for this crate.Independent oracle. The host's TCP cannot be the peer here: the node's host tests ask the host kernel about
TCP_NODELAYon loopback and connect nothing of the node to it, and reaching the host kernel's TCP with our frames needs a TAP device and privileges. QEMU's user network cannot be the peer either: libslirp'stcp_dooptionsparses MSS alone (read in libslirp 4.8.0 and 4.9.1 source; 4.9.5 is the installed one), so a guest's TCP peer never offers Window Scale. The oracle is therefore the T14 against the internet, a CDN's TCP (Ubuntu on the same machine and cable negotiated scaling and read 803.6–827.2 Mb/s), plus RFC 7323's text.No new guest test. A QEMU guest's peer is slirp, which never scales, so a guest can show neither a scaled window nor more than 64 KiB in flight. The no-offer path against a third-party TCP is what every existing guest TCP test now runs: each SYN of ours offers WS=7 and slirp ignores it. The whole guest suite is the check of that.
Gates at
d286ece79cargo test --offlineintoyos-net-shard/tcp: EXIT=0, 388 passed.cargo run -- --ci host: EXIT=0,[ci] Host: 78 step(s), all green.cargo run -- --build-only: EXIT=0.cargo test: EXIT=0, harness43 passed, 43 total,netstack_streams,netstack_streams_e1000e,netstack_socket_churn,libc_socketsamong them; host load averages 16.03 before, 12.53 after.downloadboot is built from this branch merged withorigin/wt/toyos-speedat3b196d075(merge751dfc6f6, pushed aswt/toyos-wscale-metal):cargo test --test toyos-build -- --metal --metal-readback <dir> internet_download, exit 2 (staged; the machine was not touched), image sha256e35e2466…cb658539e. The 64 KiB issue on The T14 downloads over HTTPS through ToyOS's own stack inside its list's bound, and netstack's 64 KiB receive window is filed as the ceiling its rate sits on #818 (netstacks-receive-buffer-caps-every-connection-at-one-64-kib-window-per-round-trip.md) closes on that reading ifmbps> 524.3 / min(rtt_ms); Ubuntu reads 803.6–827.2 Mb/s.What I am unsure of
RX_RINGof 256 × 2 KiB descriptors is about 3.1 ms of gigabit frames beforemissed; netstack takesRX_BUDGET64 frames a pass;TX_RINGis 16, so ACKs go out in passes of 15. Then one copy into the client's 2 MiB pipe and one out. The job reads 64 KiB chunks, so its ~1.9 KB per read call at 35 Mb/s was what was there to read, not a cap.🤖 Generated with Claude Code
https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C