Skip to content

A lost .local name is probed for again a minute after each loss: claimed when no host answers, at most one probing a minute whatever a peer sends, and one log line a loss - #800

Merged
Japabu merged 2 commits into
mainfrom
wt/toyos-mdns-retry
Oct 9, 2026

Conversation

@Japabu

@Japabu Japabu commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

A machine that lost its .local name stayed nameless until its link next returned, even after the other machine had left. It now probes for its own name again a minute after each loss, and takes it back when no host answers.

Head d6e9beaaf8aaad289141236c987c1bf9072ebdbe, on origin/main 1621281ae (#793); merging origin/main at round 2 found it already there (git merge origin/main: already up to date).

The rulings

Asked "When two machines on one network have the same ToyOS host name, what should the second one do?", the owner answered "Hold no name, say so (Recommended)". Asked on 2026-10-09 "A ToyOS machine that loses its .local name to a conflict stays nameless until its network link next returns, even after the other machine has left. Should it try for its name again by itself?", he answered "Retry on a slow interval (Recommended)". The option's description gave once a minute as an example.

What changed, per decision

The rule is in toyos-mdns (userland/netstack/mdns/src/lib.rs). Of the callers, only the shipped shell's loss line changes.

  • A lost name is a probing owed a minute out. Claim::Lost, a state nothing left but a link after none, is deleted. A conflicting response under the probe sets Claim::Owed { from_ms: now + RETRY_MS }, the state a start-up probing begins in. So Responder::owed_at carries the retry with no new arm, and both callers already wake on it: the node through Name::next_deadline (userland/netstack/node/src/name.rs), the shipped shell through Responder::wake_in (userland/netstack/src/mdns.rs), each a one-line map of owed_at. Read, not changed.
  • The interval is RETRY_MS = 60_000, my choice. The owner's example was once a minute. It is twelve times RFC 6762 §8.1's five seconds ("wait five seconds after any failed probe attempt before trying again"), so a host that really holds the name answers one probe a minute per host that wants it. A name whose holder left is back within 61 s of the last loss: the minute, at most 250 ms of delay, 750 ms of probing.
  • A probing no host answers claims and announces as at start-up (Event::Claimed); one the holder answers waits a minute from that answer.
  • The loss line says the retry (userland/netstack/src/mdns.rs): "netstack: mDNS: another host answered for {host}.local; this machine answers to no name and asks for {host}.local again every 60 s", the number RETRY_MS / 1000, so RETRY_MS is pub beside PORT, GROUP and TTL. With Lost said once, it is all the log holds while the name stays lost; the silence after it reads as still asking and the claim line as its end. No other line is added. The move (wt/toyos-move, which deletes mdns.rs and carries a copy of this string in its serve.rs) takes the same text, built from the same constant, when it merges this.
  • Event::Lost is said once until the name is claimed again (lost: bool, the one new field). A retry that meets the holder says nothing, so a retry writes no log line. Both callers log an event as it comes, with no logic of their own.
  • A lost name that loses a tiebreak waits the minute, not §8.2's second. With §8.2's second, forged winning probes would buy one probe in five seconds (§8.1) from a host that holds no name, for as long as they are sent. A name not lost keeps §8.2's second.
  • A link after none, and an address after none, still probe at once. Claim is one value, so the probing replaces the retry owed: two at once is unrepresentable.
  • RFC 6762. §9: the loser "MUST cease using the name". A probe is a question (§8.1) and uses none; nothing is announced or answered between a loss and a claim. §8.1 permits another attempt after five seconds and requires none. §8.2 as above. Read from rfc-editor.org; no other implementation's source was read.

What a peer can buy from a host that holds no name

  • Forged conflicts or answers: between two probings no probe is out, so nothing is read. Under a retry's probe a conflicting response puts the next probing a minute later. At most one probing in the interval, which is at most three probes in any minute, no announcement, no log line.
  • Forged probes that win the tiebreak: the same bound, by the rule above.
  • A flapping link: unchanged from toyos-mdns claims its name before it uses it: three probes, the tie-break, a conflict's outcome, RFC 6762 §8.1's bound on what a peer's messages cost, and the link's return in both callers #793. Each return costs its three probes, and five seconds first within ten of a conflict (§8.1), so one probing in five seconds while a holder answers each. Never a retry beside it.
  • State: fixed fields; a conflict writes the claim, the flag and one time.
  • Log: one line when the name is lost, one when it is claimed. A forger that lets each retry claim the name and then takes it again buys two lines a minute.
  • What it cannot close: two packets still take a held name, and one more a minute keeps it. A host that answers every probe is what a holder of the name is.

Tests

No test is built on smoltcp. The shell's part is the deadline it passes on: held by its existing wake_in_asks_for_what_the_name_owes_and_nothing_else (userland/netstack/src/mdns.rs), since a retry is the same Claim::Probing { sent: 0, at_ms } as a start-up delay; no test below drives the shell. No guest test is added or changed: no test reads the loss line, and the netcase boots wait for the claim line, unchanged, by event as before.

toyos-mdns, six new (src/tests.rs):

  • a_lost_name_is_probed_for_again_a_minute_after_the_loss_and_not_before: a pass every millisecond of the minute sends nothing; the probe leaves at the minute plus the drawn delay, for four draws; a new address under a lost name is the one the retry proposes.
  • a_retry_the_holder_answers_loses_again_and_the_next_is_a_minute_after_it: answered after the first probe and in the last 250 ms; no second Lost.
  • a_retry_no_host_answers_claims_and_announces_the_name: three probes, two announcements; and a name claimed is said lost when lost again.
  • a_storm_of_forged_conflicts_at_a_lost_name_costs_one_probe_a_minute: a response, a winning probe and a query each millisecond for five minutes buy four probes and no event.
  • forged_probes_that_win_every_tiebreak_at_a_lost_name_cost_one_probe_a_minute.
  • a_links_return_under_a_lost_name_probes_at_once_in_place_of_the_retry: one probing, none where the retry was owed; with the holder there, the next is a minute from that loss.

Five existing tests asserted that a lost name stays lost for an hour; they now assert nothing for 59,999 ms and the deadline at the minute.

The node, two new (tests/name.rs), driven by Node::next_deadline alone:

  • a_lost_name_is_probed_for_again_a_minute_after_each_loss_and_claimed_once_no_host_answers
  • a_links_return_under_a_lost_name_is_probed_on_at_once_and_no_retry_follows_it

Two existing node tests fired sixty seconds past a loss and expected silence; they fire 59.

Mutations

Seven, each a checked patch applied, run and restored in one script, tree clean after each. Patches, exits and the tests each turned red are in the first comment on this pull request. All seven exit 101 in toyos-mdns.

mutation toyos-mdns node tests/name.rs
m0 the whole rule reverted (negative control: lib.rs as on origin/main) 101, 10 red 101, 1 red
m1 the interval is five seconds 101, 11 red 101, 4 red
m2 Lost said at every loss 101, 3 red 101, 1 red
m3 a claim does not clear lost 101, 1 red 0
m4 a lost tiebreak waits a second 101, 1 red 0
m5 a link's return under a lost name keeps the retry 101, 4 red 101, 2 red
m6 a retry the holder answers is retried at once 101, 3 red 101, 1 red

m3 and m4 are green in the node's file: both are the responder's decisions, held by its own tests.

Gates, at this head (d6e9beaaf), each by its own exit code

command exit
cargo test --locked -p toyos-mdns 0 36 passed
cargo test --locked -p toyos-net-node 0 153 passed over nine targets; tests/name.rs 21
cargo test --locked -p netstack 0 29 passed
cargo run -- --ci host 0 78 steps, all green; the source gates and clippy's shapes among them
cargo run -- --build-only 0
cargo test --test toyos-build -- netstack_socket_churn 0 1 passed, 1 guest (x86-64); boot_netcase awaits the claim line, so this boot read it

The netcase boot ran on a loaded, shared host: uptime load averages 31.87 37.89 35.61 before and 28.51 36.97 35.31 after, 14 cores. The tree was clean after.

At round 1 (129d45af2), before the loss line changed, cargo run -- --clippy exited 0, and the whole guest suite exited 0, 37 passed of 37 over 39 guests, x86-64 and AArch64; neither was re-run at this head.

The oracle is RFC 6762 §8.1, §8.2 and §9, quoted at each test. The negative control is m0.

Size

6 files, +297 −49. Production: lib.rs +63 −28, of which code is +10 −8 and the rest the module header and doc comments; node/src/name.rs +2 −2, a comment; the shell's mdns.rs +5 −2, the loss line. Tests: +226 −16. The track's name line: +1 −1.

Unsure of

  • The retry's delay is drawn at the loss, a minute before it is used. Claim::Owed draws at the next owed, as it already does inside §8.1's five-second wait. The alternative, a state that waits before it draws, is more code for the same distribution.
  • Lost is no longer said for a loss that follows a link's return under a name already said lost. On main each such loss was a line. The log now says a change of what the machine answers to, and nothing else.
  • A lost name's tiebreak departs from §8.2's "waiting one second". A stale probe heard under a retry costs a minute where §8.2 would cost a second. The retry itself is outside what the RFC asks.
  • The shipped shell's module header is untouched, to keep clear of the branch that deletes that file. It says a name another host answers for is not this machine's, which is still true; it does not mention the retry, which the loss line now says.
  • Not measured on metal or between two guests. The track's name line keeps both exits open.

🤖 Generated with Claude Code

https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C

The owner ruled today on the question the network-stack track held open.
Asked "A ToyOS machine that loses its .local name to a conflict stays
nameless until its network link next returns, even after the other machine
has left. Should it try for its name again by itself?", he answered "Retry
on a slow interval (Recommended)".

toyos-mdns: a conflicting response under the probe no longer parks the
name in a state nothing leaves but a link after none. It owes a probing
RETRY_MS, a minute, after the loss: `Claim::Lost` is gone, and a lost name
is `Claim::Owed` a minute out, so `owed_at` carries the retry to both
callers with no change in either. A probing no host answers claims and
announces as at start-up; one the holder answers waits a minute from that
answer. `Event::Lost` is said once until the name is claimed again, so a
retry costs its owner's log nothing. A lost name that loses a tiebreak
(RFC 6762 §8.2) waits the minute and not §8.2's second, which forged
probes would turn into one probe in five seconds from a host that holds no
name. A link after none still probes at once and replaces the retry owed.

The minute: the owner's example was once a minute; it is twelve times
§8.1's five seconds, so a host that really holds the name answers one
probe a minute per host that wants it.

The five tests that held "a lost name stays lost" now hold the minute; six
new ones in toyos-mdns and two in the node's tests/name.rs hold the rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Mutations at 129d45af26070e2f5eb4bbebf79c75887aa57802, each applied as a checked patch (git apply --check, git apply), run against toyos-mdns's suite (cargo test in userland/netstack/mdns) and the node's cargo test --test name, and restored (git apply -R) in one script; the last column is git status --porcelain --ignore-submodules=none | wc -l after the restore.

m0-the-whole-rule-reverted: toyos-mdns EXIT=101, node tests/name.rs EXIT=101, tree after: 0 changed
m1-interval-five-seconds: toyos-mdns EXIT=101, node tests/name.rs EXIT=101, tree after: 0 changed
m2-lost-said-at-every-loss: toyos-mdns EXIT=101, node tests/name.rs EXIT=101, tree after: 0 changed
m3-claim-does-not-clear-lost: toyos-mdns EXIT=101, node tests/name.rs EXIT=0, tree after: 0 changed
m4-lost-tiebreak-waits-a-second: toyos-mdns EXIT=101, node tests/name.rs EXIT=0, tree after: 0 changed
m5-links-return-under-a-lost-name-keeps-the-retry: toyos-mdns EXIT=101, node tests/name.rs EXIT=101, tree after: 0 changed
m6-a-retry-the-holder-answers-is-retried-at-once: toyos-mdns EXIT=101, node tests/name.rs EXIT=101, tree after: 0 changed

Exit 101 is a test failure in each case, not a build failure; the tests each turned red:

== m0-the-whole-rule-reverted.mdns.log
test tests::a_lost_name_is_probed_for_again_a_minute_after_the_loss_and_not_before ... FAILED
test tests::a_links_return_under_a_lost_name_probes_at_once_in_place_of_the_retry ... FAILED
test tests::a_retry_no_host_answers_claims_and_announces_the_name ... FAILED
test tests::a_retry_the_holder_answers_loses_again_and_the_next_is_a_minute_after_it ... FAILED
test tests::an_address_after_none_is_probed_for_and_a_new_one_under_a_held_name_is_announced ... FAILED
test tests::rfc6762_8_1_a_conflicting_response_under_the_probe_takes_the_name_and_nothing_else_does ... FAILED
test tests::a_storm_of_forged_answers_costs_one_probe_and_the_name ... FAILED
test tests::rfc6762_9_a_held_name_another_host_answers_for_under_the_new_probe_is_lost ... FAILED
test tests::forged_probes_that_win_every_tiebreak_at_a_lost_name_cost_one_probe_a_minute ... FAILED
test tests::a_storm_of_forged_conflicts_at_a_lost_name_costs_one_probe_a_minute ... FAILED
test result: FAILED. 26 passed; 10 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
== m0-the-whole-rule-reverted.node.log
test a_lost_name_is_probed_for_again_a_minute_after_each_loss_and_claimed_once_no_host_answers ... FAILED
test result: FAILED. 20 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
== m1-interval-five-seconds.mdns.log
test tests::a_links_return_under_a_lost_name_probes_at_once_in_place_of_the_retry ... FAILED
test tests::a_lost_name_is_probed_for_again_a_minute_after_the_loss_and_not_before ... FAILED
test tests::a_retry_no_host_answers_claims_and_announces_the_name ... FAILED
test tests::a_retry_the_holder_answers_loses_again_and_the_next_is_a_minute_after_it ... FAILED
test tests::a_storm_of_forged_answers_costs_one_probe_and_the_name ... FAILED
test tests::an_address_after_none_is_probed_for_and_a_new_one_under_a_held_name_is_announced ... FAILED
test tests::rfc6762_8_1_a_conflicting_response_received_as_the_last_250_ms_end_takes_the_name ... FAILED
test tests::rfc6762_8_1_a_conflicting_response_under_the_probe_takes_the_name_and_nothing_else_does ... FAILED
test tests::forged_probes_that_win_every_tiebreak_at_a_lost_name_cost_one_probe_a_minute ... FAILED
test tests::rfc6762_9_a_held_name_another_host_answers_for_under_the_new_probe_is_lost ... FAILED
test tests::a_storm_of_forged_conflicts_at_a_lost_name_costs_one_probe_a_minute ... FAILED
test result: FAILED. 25 passed; 11 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s
== m1-interval-five-seconds.node.log
test a_lost_name_is_probed_for_again_a_minute_after_each_loss_and_claimed_once_no_host_answers ... FAILED
test a_conflicting_response_received_as_the_probing_ends_takes_the_name ... FAILED
test a_links_return_under_a_lost_name_is_probed_on_at_once_and_no_retry_follows_it ... FAILED
test another_hosts_answer_under_the_probe_takes_the_name_and_the_node_says_so ... FAILED
test result: FAILED. 17 passed; 4 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
== m2-lost-said-at-every-loss.mdns.log
test tests::a_links_return_under_a_lost_name_probes_at_once_in_place_of_the_retry ... FAILED
test tests::a_retry_the_holder_answers_loses_again_and_the_next_is_a_minute_after_it ... FAILED
test tests::a_storm_of_forged_conflicts_at_a_lost_name_costs_one_probe_a_minute ... FAILED
test result: FAILED. 33 passed; 3 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.02s
== m2-lost-said-at-every-loss.node.log
test a_lost_name_is_probed_for_again_a_minute_after_each_loss_and_claimed_once_no_host_answers ... FAILED
test result: FAILED. 20 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
== m3-claim-does-not-clear-lost.mdns.log
test tests::a_retry_no_host_answers_claims_and_announces_the_name ... FAILED
test result: FAILED. 35 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
== m3-claim-does-not-clear-lost.node.log
test result: ok. 21 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
== m4-lost-tiebreak-waits-a-second.mdns.log
test tests::forged_probes_that_win_every_tiebreak_at_a_lost_name_cost_one_probe_a_minute ... FAILED
test result: FAILED. 35 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
== m4-lost-tiebreak-waits-a-second.node.log
test result: ok. 21 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
== m5-links-return-under-a-lost-name-keeps-the-retry.mdns.log
test tests::an_address_after_none_is_probed_for_and_a_new_one_under_a_held_name_is_announced ... FAILED
test tests::a_links_return_under_a_lost_name_probes_at_once_in_place_of_the_retry ... FAILED
test tests::rfc6762_8_a_link_that_returns_is_probed_on_before_the_name_is_announced_or_answered_with_again ... FAILED
test tests::rfc6762_8_1_a_probe_attempt_within_ten_seconds_of_a_conflict_waits_five_seconds_first ... FAILED
test result: FAILED. 32 passed; 4 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
== m5-links-return-under-a-lost-name-keeps-the-retry.node.log
test a_links_return_under_a_lost_name_is_probed_on_at_once_and_no_retry_follows_it ... FAILED
test another_hosts_answer_under_the_probe_takes_the_name_and_the_node_says_so ... FAILED
test result: FAILED. 19 passed; 2 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s
== m6-a-retry-the-holder-answers-is-retried-at-once.mdns.log
test tests::a_links_return_under_a_lost_name_probes_at_once_in_place_of_the_retry ... FAILED
test tests::a_retry_the_holder_answers_loses_again_and_the_next_is_a_minute_after_it ... FAILED
test tests::a_storm_of_forged_conflicts_at_a_lost_name_costs_one_probe_a_minute ... FAILED
test result: FAILED. 33 passed; 3 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s
== m6-a-retry-the-holder-answers-is-retried-at-once.node.log
test a_lost_name_is_probed_for_again_a_minute_after_each_loss_and_claimed_once_no_host_answers ... FAILED
test result: FAILED. 20 passed; 1 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s

m0 is the negative control: the whole production change reverted (git diff HEAD origin/main -- userland/netstack/mdns/src/lib.rs) under this branch's tests. m3 and m4 are green in the node's file by design: what a claim does to the next loss's word and what a lost tiebreak waits are the responder's decisions, held by its own tests.

m0-the-whole-rule-reverted.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..08dce4865 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -45,27 +45,19 @@
 //! when the last 250 ms end is heard under the probe and takes the name,
 //! where the other order would claim and announce it first.
 //!
-//! **A lost name is not replaced, and is probed for again.** §9 recommends a
-//! responder change its name and probe again; this one holds none and says so
-//! to its owner once. The name was moved in by the owner, who also asks the
+//! **A lost name is not replaced.** §9 recommends a responder change its name
+//! and probe again; this one holds none, says so to its owner and waits for a
+//! link after none. The name was moved in by the owner, who also asks the
 //! network to record it with the lease: a responder that picked another would
 //! answer to a name nothing else on the machine knows, that no storage keeps
 //! for the next boot (§9's third step), and that any host on the link could
-//! move again by answering for it. It probes for the same name again
-//! [`RETRY_MS`] after each loss, and at once on a link after none: a probing
-//! no host answers claims and announces the name as at start-up, and one the
-//! holder answers waits the interval again and says nothing. §9's loser "MUST
-//! cease using the name", and a probe uses none: it is a question, and
-//! nothing is announced or answered between a loss and a claim. §8.1 lets a
-//! failed probe be tried again five seconds later and asks for no retry at
-//! all.
+//! move again by answering for it.
 //!
 //! **What a peer can make this responder do.** Every byte read is a peer's,
 //! from a source on this link (§11); none is trusted, and none is kept.
 //!
 //! - It keeps nothing a message brought: its state is its fixed fields, of
-//!   which a conflict writes the claim, whether its loss was said and one
-//!   time.
+//!   which a conflict writes the claim and one time.
 //! - One conflict costs at most one probing: three probes and two
 //!   announcements.
 //! - §8.1: "If fifteen conflicts occur within any ten-second period, then the
@@ -78,25 +70,10 @@
 //!   the ten seconds before it is probed on at once, as §9 asks. So a peer's
 //!   messages buy at most one probing in five seconds, for as long as it
 //!   sends them, and the name is held again when it stops.
-//! - A lost name's probing begins at least [`RETRY_MS`] after the one
-//!   before, whatever arrives: between two nothing is read, since no probe is
-//!   out for a message to conflict with or tie; under a probe a conflicting
-//!   response and a probe that wins the tiebreak each put the next probing
-//!   the interval later, where §8.2's second would let forged probes buy one
-//!   in five seconds. So a peer's messages buy at most three probes in the
-//!   interval from a host that holds no name, no announcement and no word to
-//!   its owner.
-//! - A link after none is probed on at once, whoever holds the name, in place
-//!   of the retry that was owed and never beside it. Under a lost name it
-//!   costs what it costs under any, its three probes, and §8.1's five seconds
-//!   first within ten of a conflict: one probing in five seconds while a
-//!   holder answers each.
-//! - Two packets from any host on the link take a held name, a response that
-//!   takes it back to probing and a response under the probe, and one more in
-//!   each interval keeps it: a host that answers every probe is what a holder
-//!   of the name is. When it stops, the next probing claims the name. One
-//!   that lets each retry claim the name and then takes it buys the owner two
-//!   words in the interval, claimed and lost.
+//! - A lost name sends nothing and reads nothing until a link after none: it
+//!   stays lost after the other host has left or a forger has stopped. Two
+//!   packets from any host on the link buy that, a response that takes a held
+//!   name back to probing and a response under the probe.
 //!
 //! Where an answer goes follows the question (§5.4, §6.7):
 //!
@@ -294,9 +271,8 @@ pub enum Event {
     /// announced and answered with from here on.
     Claimed,
     /// §8.1: another host answered for the name under the probe. It is not
-    /// this host's, and nothing is announced or answered until it is
-    /// [`Event::Claimed`]. Said once: a later probing that host answers too
-    /// says nothing.
+    /// this host's, and nothing is announced or answered until a link comes
+    /// after none.
     Lost,
 }
 
@@ -320,13 +296,6 @@ const DEFER_MS: u64 = 1_000;
 const CONFLICT_WINDOW_MS: u64 = 10_000;
 const LIMITED_WAIT_MS: u64 = 5_000;
 
-/// How long after losing the name it is probed for again: a minute. A host
-/// that holds the name answers one probe a minute for each host that wants
-/// it, twelve times fewer than §8.1's five seconds would ask of it, and a
-/// name whose holder has left is back within the minute and the second a
-/// probing takes. §8.1's five seconds after a failed attempt are inside it.
-const RETRY_MS: u64 = 60_000;
-
 /// §6: a record is multicast on an interface at most once a second.
 const GROUP_EVERY_MS: u64 = 1_000;
 
@@ -350,16 +319,16 @@ enum Claim {
     /// the name is held. A message conflicts only once `sent` is not zero.
     Probing { sent: u8, at_ms: u64 },
     Held,
+    Lost,
 }
 
 /// This host's one record on its link, on the caller's monotonic clock in
-/// milliseconds: probed for on every link after none and [`RETRY_MS`] after
-/// each loss, then announced twice a second apart (§8, §8.3), answered to
-/// whoever asks, and multicast at most once a second (§6), so no host can
-/// turn the queries it sends into a multicast to every host on the link at
-/// its own rate. A query §6 holds back is answered when the second ends, not
-/// dropped: the caller wakes at [`Responder::owed_at`], for that and for
-/// every probe, a retry's included.
+/// milliseconds: probed for on every link after none, then announced twice a
+/// second apart (§8, §8.3), answered to whoever asks, and multicast at most
+/// once a second (§6), so no host can turn the queries it sends into a
+/// multicast to every host on the link at its own rate. A query §6 holds back
+/// is answered when the second ends, not dropped: the caller wakes at
+/// [`Responder::owed_at`], for that and for every probe.
 ///
 /// A message counts as sent once [`Responder::owed`] has handed it over.
 #[derive(Debug)]
@@ -368,8 +337,6 @@ pub struct Responder<'a> {
     /// The link the address is held on.
     link: Option<Link>,
     claim: Claim,
-    /// [`Event::Lost`] was said, and [`Event::Claimed`] not since.
-    lost: bool,
     /// When the name last met a conflict.
     conflict_ms: Option<u64>,
     last_group_ms: Option<u64>,
@@ -384,7 +351,7 @@ pub struct Responder<'a> {
 
 impl<'a> Responder<'a> {
     pub const fn new(host: Host<'a>) -> Self {
-        Self { host, link: None, claim: Claim::Owed { from_ms: 0 }, lost: false, conflict_ms: None, last_group_ms: None, owed_ms: None, again: false }
+        Self { host, link: None, claim: Claim::Lost, conflict_ms: None, last_group_ms: None, owed_ms: None, again: false }
     }
 
     /// This host is on `link` at `now_ms`, or on none while it holds no
@@ -428,7 +395,6 @@ impl<'a> Responder<'a> {
                 return (Some(probe(self.host, link.addr)), None);
             }
             self.claim = Claim::Held;
-            self.lost = false;
             event = Some(Event::Claimed);
             self.owed_ms = Some(self.group_free_ms(now_ms, GROUP_EVERY_MS));
             self.again = true;
@@ -480,6 +446,7 @@ impl<'a> Responder<'a> {
             Claim::Owed { from_ms } => Some(from_ms),
             Claim::Probing { at_ms, .. } => Some(at_ms),
             Claim::Held => self.owed_ms,
+            Claim::Lost => None,
         }
     }
 
@@ -506,35 +473,33 @@ impl<'a> Responder<'a> {
                 self.rival(message, link, now_ms);
                 (None, None)
             }
-            Claim::Owed { .. } | Claim::Probing { .. } => (None, None),
+            Claim::Owed { .. } | Claim::Probing { .. } | Claim::Lost => (None, None),
         }
     }
 
     /// A response: §6 has one from a port other than 5353 silently ignored,
-    /// and one that conflicts takes a name under probe (§8.1), which is
-    /// probed for again [`RETRY_MS`] later, and sends a held one back to
-    /// probing (§9).
+    /// and one that conflicts takes a name under probe (§8.1) and sends a
+    /// held one back to probing (§9).
     fn response(&mut self, message: &[u8], from: Source, link: Link, now_ms: u64) -> Option<Event> {
         let probing = match self.claim {
             Claim::Probing { sent, .. } if sent > 0 => true,
             Claim::Held => false,
-            Claim::Owed { .. } | Claim::Probing { .. } => return None,
+            Claim::Owed { .. } | Claim::Probing { .. } | Claim::Lost => return None,
         };
         if from.port != PORT || !conflicts(message, self.host, link.addr).unwrap_or(false) {
             return None;
         }
         let wait = self.conflict(now_ms);
         if probing {
-            self.release(Claim::Owed { from_ms: now_ms.saturating_add(RETRY_MS) });
-            return (!core::mem::replace(&mut self.lost, true)).then_some(Event::Lost);
+            self.release(Claim::Lost);
+            return Some(Event::Lost);
         }
         self.release(Claim::Owed { from_ms: now_ms.saturating_add(wait) });
         None
     }
 
     /// A query heard under this host's probe: §8.2's comparison if it is
-    /// another host's probe for the name. A lost name defers for
-    /// [`RETRY_MS`], as it does to an answer.
+    /// another host's probe for the name.
     fn rival(&mut self, query: &[u8], link: Link, now_ms: u64) {
         let Some((kind, _)) = asked(query, self.host) else { return };
         let Some((theirs, more)) = proposed(query, self.host, kind) else { return };
@@ -547,7 +512,7 @@ impl<'a> Responder<'a> {
             Ordering::Less => false,
         };
         if later {
-            let wait = self.conflict(now_ms).max(if self.lost { RETRY_MS } else { DEFER_MS });
+            let wait = self.conflict(now_ms).max(DEFER_MS);
             self.release(Claim::Probing { sent: 0, at_ms: now_ms.saturating_add(wait) });
         }
     }

m1-interval-five-seconds.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..bfac64c82 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -325,7 +325,7 @@ const LIMITED_WAIT_MS: u64 = 5_000;
 /// it, twelve times fewer than §8.1's five seconds would ask of it, and a
 /// name whose holder has left is back within the minute and the second a
 /// probing takes. §8.1's five seconds after a failed attempt are inside it.
-const RETRY_MS: u64 = 60_000;
+const RETRY_MS: u64 = 5_000;
 
 /// §6: a record is multicast on an interface at most once a second.
 const GROUP_EVERY_MS: u64 = 1_000;

m2-lost-said-at-every-loss.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..14f2016b6 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -526,7 +526,8 @@ impl<'a> Responder<'a> {
         let wait = self.conflict(now_ms);
         if probing {
             self.release(Claim::Owed { from_ms: now_ms.saturating_add(RETRY_MS) });
-            return (!core::mem::replace(&mut self.lost, true)).then_some(Event::Lost);
+            self.lost = true;
+            return Some(Event::Lost);
         }
         self.release(Claim::Owed { from_ms: now_ms.saturating_add(wait) });
         None

m3-claim-does-not-clear-lost.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..11954dfe2 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -428,7 +428,6 @@ impl<'a> Responder<'a> {
                 return (Some(probe(self.host, link.addr)), None);
             }
             self.claim = Claim::Held;
-            self.lost = false;
             event = Some(Event::Claimed);
             self.owed_ms = Some(self.group_free_ms(now_ms, GROUP_EVERY_MS));
             self.again = true;

m4-lost-tiebreak-waits-a-second.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..f2d6fe990 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -547,7 +547,7 @@ impl<'a> Responder<'a> {
             Ordering::Less => false,
         };
         if later {
-            let wait = self.conflict(now_ms).max(if self.lost { RETRY_MS } else { DEFER_MS });
+            let wait = self.conflict(now_ms).max(DEFER_MS);
             self.release(Claim::Probing { sent: 0, at_ms: now_ms.saturating_add(wait) });
         }
     }

m5-links-return-under-a-lost-name-keeps-the-retry.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..1836793cd 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -398,6 +398,7 @@ impl<'a> Responder<'a> {
             return;
         };
         match self.link.replace(link) {
+            None if self.lost => {}
             None => self.release(Claim::Owed { from_ms: now_ms.saturating_add(self.limit_ms(now_ms)) }),
             Some(was) if was.addr != link.addr && matches!(self.claim, Claim::Held) => {
                 self.owed_ms = Some(now_ms);

m6-a-retry-the-holder-answers-is-retried-at-once.patch

diff --git a/userland/netstack/mdns/src/lib.rs b/userland/netstack/mdns/src/lib.rs
index eac4f424d..8894f2459 100644
--- a/userland/netstack/mdns/src/lib.rs
+++ b/userland/netstack/mdns/src/lib.rs
@@ -525,7 +525,7 @@ impl<'a> Responder<'a> {
         }
         let wait = self.conflict(now_ms);
         if probing {
-            self.release(Claim::Owed { from_ms: now_ms.saturating_add(RETRY_MS) });
+            self.release(Claim::Owed { from_ms: now_ms.saturating_add(if self.lost { wait } else { RETRY_MS }) });
             return (!core::mem::replace(&mut self.lost, true)).then_some(Event::Lost);
         }
         self.release(Claim::Owed { from_ms: now_ms.saturating_add(wait) });

@Japabu Japabu changed the title A lost .local name is probed for again a minute after each loss A lost .local name is probed for again a minute after each loss: claimed when no host answers, at most one probing a minute whatever a peer sends, and one log line a loss Oct 9, 2026
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

Review, round 1, of 129d45af26070e2f5eb4bbebf79c75887aa57802 against origin/main 1621281ae. One commit; read: the diff, the five changed files whole, both callers (userland/netstack/node/src/name.rs, userland/netstack/src/mdns.rs, the shell's loop in main.rs), the gate and mutation logs. I ran no test and no build.

Net: 5 files, +292 −47. Production userland/netstack/mdns/src/lib.rs +63 −28 (code +10 −8, one field, one constant, Claim::Lost gone), node/src/name.rs +2 −2 (comment). Tests +226 −16. Track +1 −1. The growth is accepted: the rule costs two more lines of code than the state it deletes.

BLOCKER

None.

NOTE

  • userland/netstack/src/mdns.rs:107 — the one line a loss now writes says "this machine answers to no name" and nothing of the retry — with Event::Lost said once, it is all a reader of the log gets for as long as the name stays lost, and it reads as final; it must say the name is asked for again every minute, so that silence after it reads as "still asking" and the claim line as its end. One string, no test reads it (the netcase boots wait on the claim line only). The move's serve.rs carries a copy of the same string and takes the same text.
  • issues/toyos-has-its-own-network-stack.md:41 — the move's branch rewrites this same line; a trial merge of this head with it conflicts there and nowhere else — whichever lands second resolves it by accounting for both sides' sentences, not by taking one.
  • Pull request body, "Tests" — prose: it says the shell's part is "held by reading and by the tests below"; no test below drives the shell. What holds it is the shell's existing wake_in_asks_for_what_the_name_owes_and_nothing_else and the fact that a retry is the same Claim::Probing { sent: 0, at_ms } a start-up delay is.

The rule against RFC 6762

  • A probe a minute after a loss is permitted: §8.1 bounds a new attempt from below at five seconds and asks for none; a probe is a question and uses no name. Stands.
  • A retry no host answers runs the start-up path unchanged (three probes, Claimed, two announcements, §8.3): the same arm of owed, no second path.
  • §9 "cease using the name": between loss and claim the state is Owed or Probing, and heard answers only under Held. Nothing is announced or answered. Held by a_lost_name_is_probed_for_again_a_minute_after_the_loss_and_not_before and the node's test.

The three decisions

  1. A lost name that loses a tiebreak waits the minute: stands. §8.2's "waiting one second" carries no requirement word, and a host that already deferred under §8.1 and probes again only by its own choice may wait longer. It does not make two honest machines collide: the tiebreak is decided by the proposed address, not by the delay, so of two that retry in the same instant the later address goes on and claims within 750 ms, and the other is answered by it at its next retry. One round, never a repeat, whichever of them carries lost. The cost is real and small: a stale probe (§8.2's reason for the second) under a retry costs a minute to a machine that already has no name. Held by m4.
  2. The delay drawn at the loss: stands. Same distribution, no state added; a draw is spent on a retry a link's return later replaces, which costs nothing. The tests prove no later call draws (undrawn panics).
  3. Lost said once until the next claim: stands, with the NOTE above. The log then carries changes of what the machine answers to; the last name line is always the present state. The line itself is what falls short.

Bounds while nameless

  • Forged responses, queries and winning probes: nothing is read with no probe out; under a retry's probe either puts the next probing a minute on. At most three probes in any minute, no announcement, no log line. Held by the two storm tests.
  • A flapping link replaces the retry (Claim is one value) and costs what toyos-mdns claims its name before it uses it: three probes, the tie-break, a conflict's outcome, RFC 6762 §8.1's bound on what a peer's messages cost, and the link's return in both callers #793 left. Held by m5.
  • No state grows: one bool.
  • A peer keeps the machine nameless forever with one packet a minute: said plainly in the module header and on the track's line.
  • A peer cannot make it claim: a claim needs three probes with no conflicting response read, and a forger can add responses, not remove a holder's. What the retry does add, inherent in the ruling: a holder that misses one 750 ms probing (asleep, off the link) loses the name to the retry, where before it lost it only to a link's return.
  • New and said in the body: a forger that lets each retry claim and then takes the name buys two log lines a minute for as long as it sends. On main a peer alone bought one.

The mechanism

Verified by reading. A loss leaves Owed { now + 60 s }, which the same pass's owed turns into Probing { sent: 0, at_ms }; owed_at returns it. The node: Name::next_deadline maps it and Node::next_deadline takes the minimum; the node's new test advances only by next_deadline, so it fails if the retry is not a deadline. The shell: wake_in maps it and the loop takes the minimum into its wait. A nameless idle machine wakes for the retry with nothing else scheduled.

Tests and mutations

The five rewritten responder tests each assert the deadline at the minute and silence to the millisecond before it; the two node tests moved from 60 s to 59 s of silence and the retry is asserted beside them. None is merely loosened: m1 reds all seven. The seven mutation patches apply to the code as it stands, their logs postdate the commit, each exits 101 on test failures. No test is built on smoltcp. I name no further mutation: a retry is the start-up state, so what breaks one breaks the other's tests.

Evidence at this head

toyos-mdns 36, node 153 over nine targets, netstack 29, clippy 0, cargo run -- --ci host 0 with 78 steps (the three crates' steps among them), --build-only 0, guest suite 0 with 37 of 37 over 39 guests: each log ends in its own exit line, the recorded head is this one and the tree was clean after. No guest test is added or changed, and none reaches the retry; the host tier holds it. Nothing here targets hardware.

Other branches

#796 and #797 touch none of these files. The move collides on the track's line only (NOTE above); userland/netstack/src/mdns.rs is untouched here unless the first NOTE is taken in it, which then meets the move's deletion and is resolved by carrying the string into the move's copy.

Landing

The pull request is a draft, and host, toolchain and guest are skipped on it: a skipped check is not a green one. At the head that lands, after the first NOTE: host concluded success and guest / suite concluded success on that exact commit, read from the check runs' conclusions. The orchestrator may land on reading those two; no further review round is needed for a change confined to that string.

LAND AFTER NAMED CHANGES

…w often

With `Event::Lost` said once until the next claim, that line is all the log
holds for as long as the name stays lost, and "answers to no name" read as
final. It now ends "and asks for <host>.local again every 60 s", so the
silence after it reads as still asking and the claim line as its end.

The number is `toyos_mdns::RETRY_MS / 1000`, which makes the constant `pub`
beside `PORT`, `GROUP` and `TTL`: the interval has one declaration and the
line cannot drift from it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RvnWQFcMuGqTHYhvSnTe8A
@Japabu

Japabu commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator Author

CI at d6e9beaaf, read from each job's log: host 79 steps all green; toolchain restored (llvm a2cc063281d9f279, compiler 6c76c19e5ffcc869, freestanding 0f5a5faab7bbc78c, sysroot 4ff97f02cb7938e0); guest suite 37/37 ok. Merges clean with main 8dbccd3ab. Queued.

@Japabu
Japabu added this pull request to the merge queue Oct 9, 2026
Merged via the queue into main with commit d6298c8 Oct 9, 2026
6 checks passed
@Japabu
Japabu deleted the wt/toyos-mdns-retry branch October 9, 2026 14:56
Japabu added a commit that referenced this pull request Oct 9, 2026
…m end's order (#803) and the .local name's re-probe (#800), into the app grants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
Japabu added a commit that referenced this pull request Oct 9, 2026
…nd's order (#803) and the .local name's probing (#800), into the HTTPS client

ring_kat's line in DRIVEN_AND_SHARED met random_draws' there; the merge keeps
main's, and the next commit decides ring_kat's place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
Japabu added a commit that referenced this pull request Oct 9, 2026
… the stream end's order (#803), into the stop's hold of the console wire

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cSFvbD35xJ2kGANVdm23C
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant