Repository navigation
fix(protocol): a resend is offered to the mesh when no carrier reaches the recipient - #540
Merged
Merged
Conversation
bahdotsh
changed the base branch from
main
to
feat/session-establishment-over-mesh
October 8, 2026 08:39
…s the recipient Only a message's first send was handed to neighbours when no carrier took it (handle_send_failure) or when a carrier took it for a recipient it cannot reach (handle_send_success). The retry queue and both outbox flushes asked neither question. A message whose first attempt went into a direct link that had already died was therefore never carried: by the time its acknowledgement timed out the link was gone, every resend was refused and re-queued, and a neighbour that could reach the recipient was never asked, for the life of the outbox entry. offer_resend_to_mesh applies the first send's rule to every resend: offer when no carrier took the frame, and when one took it for a recipient that can_reach_recipient says it cannot reach. A resend over a link that still reaches the recipient is not offered, so an ordinary retransmission costs one copy. The ACK-timeout path re-queues into the retry queue and the parked-probe sends run through it, so both are covered by the same call. Five simulator cases: refused retry after a dead direct link (delivered and settled), refused flush, swallowed retry, swallowed flush, and a resend over a live link that must not add a mesh copy. Each of the five call sites and the reachability check has a mutation that turns one red.
Several comments and the outbox state machine said Wi-Fi Direct enqueues for any recipient and reports success. It has not since the peer-stream slot replaced it: WiFiDirect::send returns PeerNotReachable for a recipient no stream has proved, as BLE does for one that is not a connected peer. The carriers that accept any recipient are the internet transport and Reticulum. The claims now name those two, in can_reach_without_carrying, can_address_via, handle_send_success, the relay hint pinning rule, route_ack, WiFiDirect::send_to_peer and outbox-and-retries.md. The route_ack gate is described as what it now is: a check before the send, not the only thing standing between an answer and a link that never drains. ADR 0015 is left as the record of its decision.
The five resend cases asserted that the message reached carol. For a message the mesh carries without a pending acknowledgement that proves nothing about the sender: before settle_dm_without_pending_ack dropped the live park counter gate, the refused flush case delivered the message and alice kept it in her outbox, re-offering it until the outbox lifetime failed it seven days later. The test stayed green. The four cases that carry a message across the mesh now also require message_delivered at alice and that the outbox entry is gone: a flush, which re-drives every entry regardless of backoff, sends the id nowhere, and the retry queue is empty. The flush cases step until the acknowledgement arrives rather than a fixed six rounds. Mutations: restoring the park counter gate turns the refused flush case red; keeping the outbox entry on the no-pending-ACK settle turns it red at the settle check; keeping it on the pending-ACK settle turns the other three red. The swallowing-carrier comment no longer claims Wi-Fi Direct accepts any recipient.
offer_resend_to_mesh and send_or_carry_handshake each wrote the same test: offer to the mesh when no carrier took the frame, or when one took it for a recipient can_reach_recipient says no carrier reaches. The two arrived on separate branches. Copies of that rule drifting apart would leave one path swallowing frames the other carries, so both now ask needs_mesh_copy. No behaviour change.
Three places on main still described the gap this branch closes. The local API chapter said a message a direct link took and lost "never crosses the mesh", the offline-first demo listed it as a known gap, and the demo's runner justified its stop-before-disconnect dance with the same claim. That is no longer true, but it is not "fixed" either. While a dead stream still looks open, the carrier accepts the resend and still lists the recipient as reachable, so nothing is offered. Only the first retry after keepalive ends the stream (about 30 s) goes to the neighbours. So the message crosses, late. Say exactly that. The demo keeps stopping C before parting it from A, not because the message is lost any more, but because it would land past the scenario's window for C to receive it. Fixes #541.
bahdotsh
force-pushed
the
fix/retry-offers-to-mesh
branch
from
October 8, 2026 11:42
967083e to
734d990
Compare
bahdotsh
changed the base branch from
feat/session-establishment-over-mesh
to
main
October 8, 2026 11:42
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Only a message's first send is handed to mesh neighbours when no carrier reaches the recipient:
handle_send_failureoffers it when no carrier took the frame, andhandle_send_successoffers it when a carrier took it for a recipient it cannot reach. The resend paths asked neither question:process_retry_queue(protocol/mod.rs), both branches;try_flush_send_via, behindflush_outbox_allandflush_outbox_for_peer_via, both branches.So a message whose first attempt went into a direct link that had already died (a half-open stream accepts frames until its keepalive notices) was never carried. By the time its acknowledgement timed out the link was gone, every resend was refused and went back on the retry queue, and a neighbour that could reach the recipient was never asked, for the whole 7-day outbox lifetime.
Found in the three-host lab run of #539: A sent to C over a direct peer stream, C was then reachable only through B, and minutes later A had 1 pending ACK, 1 queued retry and 0 mesh transmissions.
Change
offer_resend_to_mesh(message, accepted)insend.rsapplies the first send's rule to every resend: offer when no carrier took the frame, and when one took it for a recipientcan_reach_recipientsays it cannot reach. It is called from the four resend branches. A resend over a link that still reaches the recipient is not offered, so an ordinary retransmission costs one copy.The whole class, path by path:
flush_outbox_all,flush_outbox_for_peer_via), refused / swallowedWhy re-offering on every resend is safe (same argument as the park path): a neighbour that took the frame refuses further copies of the id for the
RelaySeenCacheretention window (600 s); the retry backoff bounds how often we ask; each offer spends own-send tokens like any other frame. Excluding neighbours already handed the id would be wrong: a neighbour records an id only when it accepts the frame, and one that refused it (rate limit, full queue) is the one a later offer should reach.Not handled: a resend into a link that is half-open and not yet detected
While the carrier still lists the recipient as connected, the resend is accepted and
can_reach_recipientis true, so nothing is offered. Offering there would not help anyway:offer_to_meshhands a frame for a recipient who is a neighbour straight to that recipient, over the same suspect link. Routing around a link we still believe is up needs a "suspect link" state, which this PR does not add. The window is bounded by the carrier's own dead-link detection (the Python peer stream sets TCP keepalive 15 s + 3 x 5 s andTCP_USER_TIMEOUT30 s on Linux), after which resends are refused and this fix carries them.Tests
Five new cases in
crates/offline-protocol/tests/mesh_forwarding.rs:a_message_lost_on_a_dead_direct_link_is_carried_by_the_mesh_on_retry: the first frame is swallowed by the direct link, the link dies, the retry goes through B, and A getsmessage_delivered. Red on main (nothing reaches C).a_flushed_message_no_carrier_takes_is_handed_to_the_neighborsa_resend_a_carrier_swallows_is_still_handed_to_the_neighborsa_flush_a_carrier_swallows_is_still_handed_to_the_neighborsa_resend_over_a_live_link_is_not_also_handed_to_the_mesh(cost side: exactly one resend copy, none to B)Mutations, each turning a test red: removing any one of the four call sites, never offering an accepted resend, and always offering an accepted resend (dropping the reachability check).
Locally:
cargo test --workspace --lib(engine lib 2015 tests),cargo test -p offline-protocol(all binaries),cargo clippy --workspace -- -D warnings,cargo fmt --all -- --check,RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps, all clean. No UDL change, no binding regeneration. Core and sealed untouched.Related
examples/offline-first/run.pystops C before taking it off the shared network, so A sees the stream close. Once this lands that workaround can go: a half-open link now recovers through B on the next retry after the carrier notices.message_delivered. The first test here does not depend on it: the first attempt was accepted, so an ACK was registered and the mesh-carried answer settles it on main.Docs:
docs/state-machines/outbox-and-retries.md,docs/mesh.md(mixed-neighbourhood section),docs/local-api.md,examples/offline-first/,CHANGELOG.mdunder Unreleased / Fixed.Rebased onto main now that #537 has merged (a resend the mesh carries settles through its
settle_dm_without_pending_ack). The last commit updates what main still said about the gap:docs/local-api.mdand the offline-first demo's README andrun.pynow say a lost message crosses the mesh late (on the first retry after keepalive ends the dead stream), not never; the demo keeps stopping C first so the message lands inside its receive window.Fixes #541.