Skip to content

fix(protocol): a resend is offered to the mesh when no carrier reaches the recipient - #540

Merged
bahdotsh merged 5 commits into
mainfrom
fix/retry-offers-to-mesh
Oct 8, 2026
Merged

bahdotsh merged 5 commits into
mainfrom
fix/retry-offers-to-mesh

Conversation

@bahdotsh

@bahdotsh bahdotsh commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Problem

Only a message's first send is handed to mesh neighbours when no carrier reaches the recipient: handle_send_failure offers it when no carrier took the frame, and handle_send_success offers it when a carrier took it for a recipient it cannot reach. The resend paths asked neither question:

  • the retry loop in process_retry_queue (protocol/mod.rs), both branches;
  • try_flush_send_via, behind flush_outbox_all and flush_outbox_for_peer_via, both branches.

So a message whose first attempt went into a direct link that had already died (a half-open stream accepts frames until its keepalive notices) was never carried. By the time its acknowledgement timed out the link was gone, every resend was refused and went back on the retry queue, and a neighbour that could reach the recipient was never asked, for the whole 7-day outbox lifetime.

Found in the three-host lab run of #539: A sent to C over a direct peer stream, C was then reachable only through B, and minutes later A had 1 pending ACK, 1 queued retry and 0 mesh transmissions.

Change

offer_resend_to_mesh(message, accepted) in send.rs applies the first send's rule to every resend: offer when no carrier took the frame, and when one took it for a recipient can_reach_recipient says it cannot reach. It is called from the four resend branches. A resend over a link that still reaches the recipient is not offered, so an ordinary retransmission costs one copy.

The whole class, path by path:

Path Before After
First send refused / swallowed offered unchanged
Retry queue, refused re-queued only offered
Retry queue, accepted for an unreachable recipient counted as sent offered
Outbox flush (flush_outbox_all, flush_outbox_for_peer_via), refused / swallowed re-queued / counted as sent offered
ACK timeout re-queues into the retry queue covered by the retry loop
Parked-DM probe offered at park; probe sends ran through the retry loop without an offer offered at park and on each probe the relay or gateway does not reach

Why re-offering on every resend is safe (same argument as the park path): a neighbour that took the frame refuses further copies of the id for the RelaySeenCache retention window (600 s); the retry backoff bounds how often we ask; each offer spends own-send tokens like any other frame. Excluding neighbours already handed the id would be wrong: a neighbour records an id only when it accepts the frame, and one that refused it (rate limit, full queue) is the one a later offer should reach.

Not handled: a resend into a link that is half-open and not yet detected

While the carrier still lists the recipient as connected, the resend is accepted and can_reach_recipient is true, so nothing is offered. Offering there would not help anyway: offer_to_mesh hands a frame for a recipient who is a neighbour straight to that recipient, over the same suspect link. Routing around a link we still believe is up needs a "suspect link" state, which this PR does not add. The window is bounded by the carrier's own dead-link detection (the Python peer stream sets TCP keepalive 15 s + 3 x 5 s and TCP_USER_TIMEOUT 30 s on Linux), after which resends are refused and this fix carries them.

Tests

Five new cases in crates/offline-protocol/tests/mesh_forwarding.rs:

  • a_message_lost_on_a_dead_direct_link_is_carried_by_the_mesh_on_retry: the first frame is swallowed by the direct link, the link dies, the retry goes through B, and A gets message_delivered. Red on main (nothing reaches C).
  • a_flushed_message_no_carrier_takes_is_handed_to_the_neighbors
  • a_resend_a_carrier_swallows_is_still_handed_to_the_neighbors
  • a_flush_a_carrier_swallows_is_still_handed_to_the_neighbors
  • a_resend_over_a_live_link_is_not_also_handed_to_the_mesh (cost side: exactly one resend copy, none to B)

Mutations, each turning a test red: removing any one of the four call sites, never offering an accepted resend, and always offering an accepted resend (dropping the reachability check).

Locally: cargo test --workspace --lib (engine lib 2015 tests), cargo test -p offline-protocol (all binaries), cargo clippy --workspace -- -D warnings, cargo fmt --all -- --check, RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps, all clean. No UDL change, no binding regeneration. Core and sealed untouched.

Related

Docs: docs/state-machines/outbox-and-retries.md, docs/mesh.md (mixed-neighbourhood section), docs/local-api.md, examples/offline-first/, CHANGELOG.md under Unreleased / Fixed.

Rebased onto main now that #537 has merged (a resend the mesh carries settles through its settle_dm_without_pending_ack). The last commit updates what main still said about the gap: docs/local-api.md and the offline-first demo's README and run.py now say a lost message crosses the mesh late (on the first retry after keepalive ends the dead stream), not never; the demo keeps stopping C first so the message lands inside its receive window.

Fixes #541.

@bahdotsh
bahdotsh changed the base branch from main to feat/session-establishment-over-mesh October 8, 2026 08:39
…s the recipient

Only a message's first send was handed to neighbours when no carrier took
it (handle_send_failure) or when a carrier took it for a recipient it
cannot reach (handle_send_success). The retry queue and both outbox
flushes asked neither question. A message whose first attempt went into a
direct link that had already died was therefore never carried: by the time
its acknowledgement timed out the link was gone, every resend was refused
and re-queued, and a neighbour that could reach the recipient was never
asked, for the life of the outbox entry.

offer_resend_to_mesh applies the first send's rule to every resend: offer
when no carrier took the frame, and when one took it for a recipient that
can_reach_recipient says it cannot reach. A resend over a link that still
reaches the recipient is not offered, so an ordinary retransmission costs
one copy. The ACK-timeout path re-queues into the retry queue and the
parked-probe sends run through it, so both are covered by the same call.

Five simulator cases: refused retry after a dead direct link (delivered
and settled), refused flush, swallowed retry, swallowed flush, and a
resend over a live link that must not add a mesh copy. Each of the five
call sites and the reachability check has a mutation that turns one red.
Several comments and the outbox state machine said Wi-Fi Direct
enqueues for any recipient and reports success. It has not since the
peer-stream slot replaced it: WiFiDirect::send returns
PeerNotReachable for a recipient no stream has proved, as BLE does
for one that is not a connected peer. The carriers that accept any
recipient are the internet transport and Reticulum.

The claims now name those two, in can_reach_without_carrying,
can_address_via, handle_send_success, the relay hint pinning rule,
route_ack, WiFiDirect::send_to_peer and outbox-and-retries.md. The
route_ack gate is described as what it now is: a check before the
send, not the only thing standing between an answer and a link that
never drains. ADR 0015 is left as the record of its decision.
The five resend cases asserted that the message reached carol. For a
message the mesh carries without a pending acknowledgement that
proves nothing about the sender: before settle_dm_without_pending_ack
dropped the live park counter gate, the refused flush case delivered
the message and alice kept it in her outbox, re-offering it until the
outbox lifetime failed it seven days later. The test stayed green.

The four cases that carry a message across the mesh now also require
message_delivered at alice and that the outbox entry is gone: a flush,
which re-drives every entry regardless of backoff, sends the id
nowhere, and the retry queue is empty. The flush cases step until the
acknowledgement arrives rather than a fixed six rounds.

Mutations: restoring the park counter gate turns the refused flush
case red; keeping the outbox entry on the no-pending-ACK settle turns
it red at the settle check; keeping it on the pending-ACK settle turns
the other three red. The swallowing-carrier comment no longer claims
Wi-Fi Direct accepts any recipient.
offer_resend_to_mesh and send_or_carry_handshake each wrote the same
test: offer to the mesh when no carrier took the frame, or when one
took it for a recipient can_reach_recipient says no carrier reaches.
The two arrived on separate branches. Copies of that rule drifting
apart would leave one path swallowing frames the other carries, so
both now ask needs_mesh_copy. No behaviour change.
Three places on main still described the gap this branch closes. The
local API chapter said a message a direct link took and lost "never
crosses the mesh", the offline-first demo listed it as a known gap,
and the demo's runner justified its stop-before-disconnect dance with
the same claim.

That is no longer true, but it is not "fixed" either. While a dead
stream still looks open, the carrier accepts the resend and still
lists the recipient as reachable, so nothing is offered. Only the
first retry after keepalive ends the stream (about 30 s) goes to the
neighbours. So the message crosses, late.

Say exactly that. The demo keeps stopping C before parting it from A,
not because the message is lost any more, but because it would land
past the scenario's window for C to receive it.

Fixes #541.
@bahdotsh
bahdotsh force-pushed the fix/retry-offers-to-mesh branch from 967083e to 734d990 Compare October 8, 2026 11:42
@bahdotsh
bahdotsh changed the base branch from feat/session-establishment-over-mesh to main October 8, 2026 11:42
@bahdotsh
bahdotsh merged commit 5671619 into main Oct 8, 2026
1 check passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 8, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(protocol): a retry that the direct carriers refuse is never offered to the mesh

1 participant