diff --git a/CHANGELOG.md b/CHANGELOG.md index d50b431f..ae26e6fd 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -192,6 +192,18 @@ archived by series under [docs/changelog/](docs/changelog/); see the peripheral when a central is subscribed. The peripheral alone takes nothing while no central is subscribed, where a notification reached no one. No board run has been done yet. +- **A resend is offered to the mesh when no carrier reaches the recipient.** + Only a message's first send was handed to neighbours when no carrier took + it. A message whose first attempt went into a direct link that had already + died (a half-open stream accepts frames until its keepalive notices) was + therefore never carried: by the time its acknowledgement timed out the link + was gone, every resend was refused and went back on the retry queue, and a + neighbour that could reach the recipient was never asked, for the life of + the outbox entry. The retry queue and both outbox flushes now make the same + offer the first send does, when no carrier takes the frame and when a + carrier takes it for a recipient it cannot reach. A resend over a link that + still reaches the recipient is not offered, so an ordinary retransmission + costs one copy. ## [0.28.0] — 2026-10-06 diff --git a/crates/offline-protocol-transport/src/wifi_direct.rs b/crates/offline-protocol-transport/src/wifi_direct.rs index 42ea9876..dea50571 100644 --- a/crates/offline-protocol-transport/src/wifi_direct.rs +++ b/crates/offline-protocol-transport/src/wifi_direct.rs @@ -371,10 +371,9 @@ impl Transport for WifiDirectTransport { /// Queues `message` for a specific connected peer. /// - /// Unlike [`Transport::send`], which accepts any recipient and lets the - /// platform discover it cannot be reached, this refuses a peer with no - /// live link — a forwarding caller needs the failure synchronously so it - /// can pick another neighbor instead. + /// Refuses a peer with no live link, as [`Transport::send`] does: a + /// forwarding caller needs the failure synchronously so it can pick + /// another neighbor instead. fn send_to_peer(&self, peer_id: &str, message: &Message) -> Result<()> { if self.layer_status() != TransportStatus::Available { return Err(crate::Error::TransportNotAvailable( diff --git a/crates/offline-protocol/src/protocol/mod.rs b/crates/offline-protocol/src/protocol/mod.rs index 07dd6e1d..50ee18c4 100644 --- a/crates/offline-protocol/src/protocol/mod.rs +++ b/crates/offline-protocol/src/protocol/mod.rs @@ -2805,10 +2805,12 @@ impl OfflineProtocol { current_transport, Some(attempt_count.saturating_add(1)), ); + self.offer_resend_to_mesh(&message, true); debug!(message_id = %message.id, "Flush send succeeded"); current_transport } Err(e) => { + self.offer_resend_to_mesh(&message, false); let _ = self.retry_queue.enqueue(message.clone(), attempt_count); debug!(message_id = %message.id, error = %e, "Flush send failed, re-enqueued"); None @@ -3885,6 +3887,7 @@ impl OfflineProtocol { if let Some(transport) = current_transport { self.transport_manager.reset_retry_count(transport); } + self.offer_resend_to_mesh(&entry.message, true); debug!( message_id = %entry.message.id, @@ -3894,6 +3897,7 @@ impl OfflineProtocol { ); } Err(e) => { + self.offer_resend_to_mesh(&entry.message, false); // Re-enqueue with incremented retry count for backoff let next_retry_at = self .retry_queue diff --git a/crates/offline-protocol/src/protocol/send.rs b/crates/offline-protocol/src/protocol/send.rs index f0b50e39..7c34678f 100644 --- a/crates/offline-protocol/src/protocol/send.rs +++ b/crates/offline-protocol/src/protocol/send.rs @@ -588,9 +588,9 @@ impl OfflineProtocol { /// 2. **Pinned to [`TransportType::Internet`].** These frames are /// self-addressed, and DORS demotes Internet below every mesh transport /// (`INTERNET_FALLBACK_DEMOTION`), so ordinary routing hands them to - /// BLE/Wi-Fi Direct first. BLE fails closed (self is never a connected - /// peer), but Wi-Fi Direct and Reticulum enqueue unconditionally and - /// return `Ok` — swallowing the frame while reporting success, which on + /// the mesh carriers first. BLE and Wi-Fi Direct fail closed (self is + /// never a connected peer), but Reticulum enqueues unconditionally and + /// returns `Ok`, swallowing the frame while reporting success, which on /// the broadcast path means the group message is delivered to nobody. /// /// Errors propagate to the caller rather than routing through @@ -3606,11 +3606,11 @@ impl OfflineProtocol { self.ensure_ack_registration(message)?; // A transport reporting success is not always evidence the frame can - // arrive. Wi-Fi Direct and Reticulum accept any recipient and return - // `Ok`, so a send to someone we hold no link to is queued for a link - // that never drains — and because the mesh hand-off used to hang off - // the *failure* path, a device with either carrier up would swallow the - // frame instead of asking its neighbors to carry it. That is precisely + // arrive. The internet transport and Reticulum accept any recipient + // and return `Ok`, so a send to someone neither can reach is reported + // as sent and never arrives, and because the mesh hand-off used to hang + // off the *failure* path, a device with either carrier up would swallow + // the frame instead of asking its neighbors to carry it. That is precisely // the out-of-range case forwarding exists for, so the question is asked // directly rather than inferred from an error that never comes. // @@ -3677,6 +3677,57 @@ impl OfflineProtocol { Ok(next_retry_at) } + /// The mesh offer for a **resend**: the retry queue and both outbox + /// flushes. A resend asks the same question a first send does, so it gets + /// the same answer as [`Self::handle_send_success`] and + /// [`Self::handle_send_failure`]: a send no carrier took is offered to the + /// neighbors, and so is one a carrier took for a recipient it cannot + /// reach. + /// + /// Without this a message whose first attempt a carrier accepted and lost + /// (a stream that was already dead and had not noticed) never reached the + /// mesh at all: by the time its acknowledgement timed out the direct link + /// was gone, every resend was refused, and each refusal only went back on + /// the retry queue while a neighbor that could reach the recipient sat + /// unused for the life of the outbox entry. + /// + /// Re-offering on every resend is cheap for the reasons the park path + /// gives: a neighbor that took the frame refuses another copy of the id + /// for the whole `RelaySeenCache` retention window, the retry backoff + /// bounds how often we ask, and each offer spends the own-send tokens any + /// other frame of ours does. Excluding neighbors we already handed it to + /// would be wrong, because a neighbor records an id only when it *accepts* + /// the frame, and one that refused it (rate limit, full queue) is exactly + /// the one a later offer should reach. + pub(super) fn offer_resend_to_mesh(&mut self, message: &Message, accepted: bool) -> usize { + if !self.needs_mesh_copy(message.recipient.as_str(), accepted) { + return 0; + } + let handed_to_mesh = self.offer_to_mesh(message); + if handed_to_mesh > 0 { + debug!( + message_id = %message.id, + recipient = %message.recipient, + accepted, + handed_to_mesh, + "Resend could not reach the recipient directly; handed it to neighbors" + ); + } + handed_to_mesh + } + + /// Whether a frame for `recipient` that a carrier did (`accepted`) or did + /// not take must also be handed to the neighbours. + /// + /// The one rule for a resend and a handshake frame: yes when no carrier + /// took it, and yes when one took it for a recipient that + /// [`Self::can_reach_recipient`] says no carrier reaches, because a + /// carrier's `Ok` is not delivery. Two copies of it drifting apart would + /// leave one path swallowing frames the other carries. + fn needs_mesh_copy(&self, recipient: &str, accepted: bool) -> bool { + !accepted || !self.can_reach_recipient(recipient) + } + /// Hands a locally-originated frame to nearby devices so it can travel /// toward a recipient we cannot reach ourselves. /// @@ -3825,12 +3876,11 @@ impl OfflineProtocol { /// accepted the frame and no neighbour took it. pub(super) fn send_or_carry_handshake(&mut self, message: &Message) -> Result { let direct = self.transport_manager.send(message); - let handed_to_mesh = - if direct.is_err() || !self.can_reach_recipient(message.recipient.as_str()) { - self.offer_to_mesh(message) - } else { - 0 - }; + let handed_to_mesh = if self.needs_mesh_copy(message.recipient.as_str(), direct.is_ok()) { + self.offer_to_mesh(message) + } else { + 0 + }; if handed_to_mesh > 0 { debug!( recipient = %message.recipient, @@ -4947,14 +4997,14 @@ impl OfflineProtocol { /// ack handling already absorbs. When nothing reaches the sender directly /// the mesh remains the whole answer and step 3 is skipped, as before. /// - /// Step 1 is *gated on addressability* rather than on a send error, and - /// that gate is load-bearing. A transport returning `Ok` is not evidence - /// the frame can arrive: Wi-Fi Direct enqueues for any recipient, and it is - /// the preferred mesh carrier — so on the last hop of a forwarded frame the - /// answer would be queued for a link that never drains, reported as sent, - /// and step 2 would never run. The sender's retransmissions would take the - /// same path every time, ending in a failure report for a delivered - /// message. + /// Step 1 is *gated on addressability* as well as on the send result. A + /// transport returning `Ok` is not evidence the frame can arrive: a mesh + /// carrier that enqueued for any recipient (Wi-Fi Direct did, before it + /// refused a recipient no stream had proved) would, on the last hop of a + /// forwarded frame, report the answer as sent toward a sender it holds no + /// link to, and step 2 would never run. The sender's retransmissions would + /// take the same path every time, ending in a failure report for a + /// delivered message. /// /// An acknowledgement no route takes is held, not dropped /// ([`Self::hold_unrouted_ack`]): it was owed the moment the message was diff --git a/crates/offline-protocol/src/transport_manager.rs b/crates/offline-protocol/src/transport_manager.rs index a64fc6c4..c37d5cc8 100644 --- a/crates/offline-protocol/src/transport_manager.rs +++ b/crates/offline-protocol/src/transport_manager.rs @@ -1143,10 +1143,10 @@ impl TransportManager { /// by other devices. /// /// This exists because **a send returning `Ok` is not evidence of - /// reachability**. Only BLE refuses a recipient it holds no link to; Wi-Fi - /// Direct and Reticulum enqueue for any recipient and report success, so a - /// frame handed to them for someone out of range is queued for a link that - /// will never drain — reported as sent, silently swallowed. Anything that + /// reachability**. BLE and Wi-Fi Direct refuse a recipient they hold no + /// link to, but the internet transport and Reticulum enqueue for any + /// recipient and report success, so a frame handed to them for someone + /// neither can reach is reported as sent and never arrives. Anything that /// decides "the mesh has to carry this" from a send failure therefore never /// fires on a device where one of those carriers is up. Asking this instead /// keeps that decision on facts we can check. @@ -1214,9 +1214,9 @@ impl TransportManager { /// The narrower question [`Self::can_reach_without_carrying`] answers across /// every carrier, asked of one. It exists for the same reason: a transport /// returning `Ok` is not evidence the frame can arrive. A **mesh** carrier - /// can only address a peer it holds a live link to, and Wi-Fi Direct - /// enqueues for any recipient regardless — so a frame handed to it for - /// someone several hops away is queued for a link that never drains. + /// can only address a peer it holds a live link to: BLE and Wi-Fi Direct + /// both refuse a send to anyone else, and asking first lets a caller choose + /// another route before it spends a send, rather than after the refusal. /// **Infrastructure** carriers do their own routing, so for them the answer /// is yes whenever they are available. /// diff --git a/crates/offline-protocol/tests/mesh_forwarding.rs b/crates/offline-protocol/tests/mesh_forwarding.rs index b5ec5bab..d4e2da8c 100644 --- a/crates/offline-protocol/tests/mesh_forwarding.rs +++ b/crates/offline-protocol/tests/mesh_forwarding.rs @@ -1004,3 +1004,336 @@ fn a_device_with_custody_off_still_strips_the_request_and_holds_nothing() { assert_eq!(net.node("bob").custody_stats().held, 0); assert_eq!(net.node("bob").custody_stats().accepted, 0); } + +/// Records every `message_delivered` a device reports. +fn watch_deliveries(net: &mut Neighborhood, id: &str) -> Arc>> { + let delivered = Arc::new(Mutex::new(Vec::::new())); + let seen = Arc::clone(&delivered); + net.node(id).on_event(move |event| { + if let Event::MessageDelivered { message_id, .. } = event { + seen.lock().unwrap().push(message_id); + } + }); + delivered +} + +/// Steps the network until `delivered` holds something, at most `rounds` +/// times, without sleeping. +fn step_until_delivered( + net: &mut Neighborhood, + delivered: &Arc>>, + rounds: usize, +) { + for _ in 0..rounds { + net.step(); + if !delivered.lock().unwrap().is_empty() { + return; + } + } +} + +/// Asserts that `sender` has settled `message_id`: its outbox no longer holds +/// it, so a flush, which re-drives every entry whatever its backoff, puts it +/// on no carrier, and nothing is left on the retry queue. +/// +/// A message the mesh carried without a pending acknowledgement reaches its +/// recipient whether or not the sender settles it, so the recipient's inbox +/// proves nothing about this. Unsettled, it stays in the outbox, re-offered on +/// every retry, until the outbox lifetime fails it seven days later. +fn assert_settled( + net: &mut Neighborhood, + sender: &str, + message_id: &str, + other_carriers: &[&MockTransport], +) { + for carrier in other_carriers { + carrier.clear_sent_messages(); + } + let before = net + .transmissions + .iter() + .filter(|(from, _, id)| from == sender && id == message_id) + .count(); + net.node(sender).flush_outbox_all(); + net.step(); + let after = net + .transmissions + .iter() + .filter(|(from, _, id)| from == sender && id == message_id) + .count(); + assert_eq!( + after, before, + "{sender} still holds {message_id} in its outbox: a flush sent it again" + ); + for carrier in other_carriers { + assert!( + !carrier + .sent_messages() + .iter() + .any(|m| m.id.as_str() == message_id), + "{sender} still holds {message_id} in its outbox: a flush put it on a carrier" + ); + } + assert_eq!( + net.node(sender).retry_queue_size(), + 0, + "{sender} still has {message_id} on its retry queue" + ); +} + +#[test] +fn a_message_lost_on_a_dead_direct_link_is_carried_by_the_mesh_on_retry() { + // The first send took the direct link, which was already dead: the carrier + // accepted the frame and it never arrived (a half-open stream does exactly + // this until its keepalive notices). By the time the acknowledgement times + // out the link is gone, carol is reachable only through bob, and every + // carrier refuses the resend because none holds a link to her. + // + // A refused first send is offered to the neighbors; a refused resend has to + // be too, or the message waits in the outbox for a direct link that may + // never come back while a path through the crowd sits unused. + let mut alice_config = default_config("alice"); + alice_config.reliability.ack.default_timeout_ms = 50; + alice_config.reliability.retry.initial_delay_ms = 10; + + let mut net = Neighborhood::new(&["bob", "carol"]); + net.add_node("alice", alice_config); + net.link("alice", "bob"); + net.link("bob", "carol"); + net.link("alice", "carol"); + + let delivered = Arc::new(Mutex::new(Vec::::new())); + let seen = Arc::clone(&delivered); + net.node("alice").on_event(move |event| { + if let Event::MessageDelivered { message_id, .. } = event { + seen.lock().unwrap().push(message_id); + } + }); + + let msg_id = net.send("alice", "carol", "around the dead link"); + + // The direct link swallows the frame. + assert_eq!( + net.radios["alice"].sent_messages().len(), + 1, + "the first send should have gone straight to carol" + ); + assert!(net.radios["alice"].peer_sends().is_empty()); + net.radios["alice"].clear_sent_messages(); + + // Then the link is gone; bob still hears both of them. + net.radios["alice"].remove_connected_peer("carol"); + net.radios["carol"].remove_connected_peer("alice"); + net.node("alice").on_neighbor_lost("carol"); + net.node("carol").on_neighbor_lost("alice"); + net.links.get_mut("alice").unwrap().retain(|p| p != "carol"); + net.links.get_mut("carol").unwrap().retain(|p| p != "alice"); + + // Let the acknowledgement time out and the retry come due. + for _ in 0..50 { + std::thread::sleep(std::time::Duration::from_millis(20)); + net.step(); + if !delivered.lock().unwrap().is_empty() { + break; + } + } + + assert_eq!( + net.inbox("carol"), + vec!["around the dead link".to_string()], + "the resend must be carried through bob" + ); + assert!( + net.transmissions + .iter() + .any(|(from, to, id)| from == "bob" && to == "carol" && id == &msg_id), + "bob must have carried it" + ); + assert_eq!( + *delivered.lock().unwrap(), + vec![msg_id.clone()], + "and alice must learn it was delivered" + ); + assert_settled(&mut net, "alice", &msg_id, &[]); +} + +#[test] +fn a_flushed_message_no_carrier_takes_is_handed_to_the_neighbors() { + // A carrier coming up re-drives the whole outbox at once, ignoring the + // backoff. Alice was alone when she wrote to carol, so nothing took the + // message and nobody was there to carry it. Bob arrives, and he can reach + // carol; the flush that follows must offer him the message rather than + // putting it straight back on the retry queue to wait out its backoff. + let mut net = Neighborhood::new(&["alice", "bob", "carol"]); + net.link("bob", "carol"); + + let delivered = watch_deliveries(&mut net, "alice"); + let msg_id = net.send("alice", "carol", "carried on the flush"); + assert_eq!(net.transmission_count(&msg_id), 0, "alice was alone"); + + net.link("alice", "bob"); + net.node("alice").flush_outbox_all(); + // Rounds only, no sleeping: the retry queue's own timer is a second away + // and must not be what delivers this. + step_until_delivered(&mut net, &delivered, 12); + + assert_eq!( + net.inbox("carol"), + vec!["carried on the flush".to_string()], + "the flush must hand the message to bob" + ); + // Nothing accepted the flushed frame, so it has no pending acknowledgement + // and carol's answer is the only way alice learns it arrived. + assert_eq!( + *delivered.lock().unwrap(), + vec![msg_id.clone()], + "carol's acknowledgement must settle the message the mesh carried" + ); + assert_settled(&mut net, "alice", &msg_id, &[]); +} + +#[test] +fn a_resend_a_carrier_swallows_is_still_handed_to_the_neighbors() { + // A carrier can take a frame for a recipient it cannot reach and report + // success (the internet transport does, for one the relay has said it + // cannot reach; the mock here stands in for it on a mesh slot, so that no + // relay verdict is needed). Alice's first attempt went that way while she + // was alone; + // bob has arrived since, and he can reach carol. The resend after the + // acknowledgement times out is swallowed the same way, and it has to be + // offered to bob as a first send would be, not counted as sent. + let mut alice_config = default_config("alice"); + alice_config.reliability.ack.default_timeout_ms = 50; + alice_config.reliability.retry.initial_delay_ms = 10; + + let mut net = Neighborhood::new(&["bob", "carol"]); + net.add_node("alice", alice_config); + let swallowing = MockTransport::new(TransportType::WiFiDirect); + swallowing.start().unwrap(); + net.node("alice") + .transport_manager_mut() + .add_transport(TransportType::WiFiDirect, Box::new(swallowing.clone())); + net.link("bob", "carol"); + + let delivered = watch_deliveries(&mut net, "alice"); + let msg_id = net.send("alice", "carol", "past the swallowing carrier"); + assert!( + swallowing + .sent_messages() + .iter() + .any(|m| m.id.as_str() == msg_id), + "the first attempt should have been swallowed" + ); + assert_eq!(net.transmission_count(&msg_id), 0, "alice was alone"); + + net.link("alice", "bob"); + for _ in 0..50 { + std::thread::sleep(std::time::Duration::from_millis(20)); + net.step(); + if !delivered.lock().unwrap().is_empty() { + break; + } + } + + assert_eq!( + net.inbox("carol"), + vec!["past the swallowing carrier".to_string()], + "the swallowed resend must also be handed to bob" + ); + assert_eq!( + *delivered.lock().unwrap(), + vec![msg_id.clone()], + "carol's acknowledgement must settle the message the mesh carried" + ); + assert_settled(&mut net, "alice", &msg_id, &[&swallowing]); +} + +#[test] +fn a_flush_a_carrier_swallows_is_still_handed_to_the_neighbors() { + // The flush counterpart of the case above: the carrier that comes up and + // triggers the flush is one that takes any recipient, so the flushed + // message "succeeds" into it. Bob, who can reach carol, must still be + // offered it. + let mut net = Neighborhood::new(&["alice", "bob", "carol"]); + net.link("bob", "carol"); + + let delivered = watch_deliveries(&mut net, "alice"); + let msg_id = net.send("alice", "carol", "flushed past the swallowing carrier"); + assert_eq!(net.transmission_count(&msg_id), 0, "alice was alone"); + + let swallowing = MockTransport::new(TransportType::WiFiDirect); + swallowing.start().unwrap(); + net.node("alice") + .transport_manager_mut() + .add_transport(TransportType::WiFiDirect, Box::new(swallowing.clone())); + net.link("alice", "bob"); + net.node("alice").flush_outbox_all(); + assert!( + swallowing + .sent_messages() + .iter() + .any(|m| m.id.as_str() == msg_id), + "the flush should have gone to the carrier that accepts anyone" + ); + step_until_delivered(&mut net, &delivered, 12); + + assert_eq!( + net.inbox("carol"), + vec!["flushed past the swallowing carrier".to_string()], + "the swallowed flush must also be handed to bob" + ); + assert_eq!( + *delivered.lock().unwrap(), + vec![msg_id.clone()], + "carol's acknowledgement must settle the message the mesh carried" + ); + assert_settled(&mut net, "alice", &msg_id, &[&swallowing]); +} + +#[test] +fn a_resend_over_a_live_link_is_not_also_handed_to_the_mesh() { + // The cost side of the resend offer. Carol is in range and her first + // acknowledgement is lost, so alice resends over the link she still has. + // That resend reaches carol; offering it to the mesh as well would put a + // second copy on the air for every ordinary retransmission. + let mut alice_config = default_config("alice"); + alice_config.reliability.ack.default_timeout_ms = 50; + alice_config.reliability.retry.initial_delay_ms = 10; + + let mut net = Neighborhood::new(&["bob", "carol"]); + net.add_node("alice", alice_config); + net.link("alice", "bob"); + net.link("alice", "carol"); + + let delivered = Arc::new(Mutex::new(Vec::::new())); + let seen = Arc::clone(&delivered); + net.node("alice").on_event(move |event| { + if let Event::MessageDelivered { message_id, .. } = event { + seen.lock().unwrap().push(message_id); + } + }); + + let msg_id = net.send("alice", "carol", "say again"); + net.step(); + assert_eq!(net.inbox("carol"), vec!["say again".to_string()]); + // Carol's acknowledgement is lost on the way back. + net.node("carol").process().unwrap(); + net.radios["carol"].clear_sent_messages(); + net.radios["carol"].clear_peer_sends(); + + for _ in 0..50 { + std::thread::sleep(std::time::Duration::from_millis(20)); + net.step(); + if !delivered.lock().unwrap().is_empty() { + break; + } + } + + assert_eq!(*delivered.lock().unwrap(), vec![msg_id.clone()]); + assert_eq!( + net.deliveries_to("carol", &msg_id), + 2, + "the first copy and one resend, nothing extra" + ); + assert_eq!(net.deliveries_to("bob", &msg_id), 0); +} diff --git a/docs/local-api.md b/docs/local-api.md index c8ae4f9f..3e1785ff 100644 --- a/docs/local-api.md +++ b/docs/local-api.md @@ -214,12 +214,11 @@ Two devices that reach each other only through a third need not have met: while a message waits for a peer no carrier reaches directly, the key package and the Welcome cross the mesh like every frame after them, and the recipient's acknowledgement comes back the same way, so `message_delivered` -on the sender proves the crossing end to end. One gap remains: only a send -no carrier takes is handed to the mesh. A message a direct link took and -then lost (a stream to a device that went away without closing it, until -keepalive ends it in about 30 seconds) is retried over direct carriers -only, so it never crosses the mesh; it waits for a direct link to the -recipient. +on the sender proves the crossing end to end. A message a direct link took +and then lost (a stream to a device that went away without closing it) +crosses the mesh too, but late: while the stream still looks open its +retries go down it, and only once keepalive ends it (about 30 seconds) is +the next retry handed to the neighbours. ## The policy file diff --git a/docs/mesh.md b/docs/mesh.md index 8c73d316..9e0e38dc 100644 --- a/docs/mesh.md +++ b/docs/mesh.md @@ -570,6 +570,8 @@ The catch is that the question being asked there is mostly about this device's o A verdict is also **remembered**, not just acted on. It is recorded as a fact about that recipient on that carrier, so the *next* send does not have to repeat the failure to learn the same thing: a carrier that has said it cannot reach someone stops counting as a way to reach them until the fact ages out (ten minutes) or a presence answer supersedes it. Facts decay on purpose. A remembered "unreachable" that never expired would keep a path shut long after the recipient came back, and since nothing is ever settled by a claim, the worst a stale one can cost is latency. - **How the message arrived.** An acknowledgement for a message that reached this device across the mesh goes back the way it came, whatever this device's own carriers say. Otherwise an online recipient answers over the relay, where an offline sender cannot see it — and that sender retransmits a message that was delivered and read, eventually reporting it failed. When this device is also online, the answer goes both ways; the duplicate costs one frame. +The question is asked again on every **resend**, not only on the first send: a retry or an outbox flush that no carrier takes, or that a carrier takes for a recipient it cannot reach, is handed to the neighbors too. That is what recovers a message whose first attempt went into a direct link that had already died (a half-open stream reports success until its keepalive notices): by the time the acknowledgement times out the link is gone and every resend is refused, and a neighbor may be the only way left. A resend over a link that still reaches the recipient is not offered, so an ordinary retransmission puts one copy on the air, not two. + Because parking removes the pending acknowledgement, a parked message that is then delivered across the mesh is settled from the acknowledgement alone. So is a message no carrier took directly, which was handed to the neighbors at once and never had a pending acknowledgement. Apps see the ordinary `message_delivered` event either way. What is still not covered: a device whose only infrastructure is **Nostr** never receives an unreachable verdict at all (a broadcast relay reports no per-recipient delivery), so nothing contradicts the initial "reachable" answer and no mesh fallback fires for it. That gap is permanent for Nostr rather than unfinished: there is no verdict to be had. Reticulum is no longer in that position. Its managers speak [the gateway contract](spec/gateway-contract.md), so a gateway's `recipient_unreachable` verdict reaches the same parking machinery the relay's does (it was always keyed to the verdict rather than to the relay), and a device attached to a gateway gets mesh fallback for a recipient that gateway cannot reach. What remains carrier-specific is only that a zone with no gateway has no verdicts to receive, which is the same as having no infrastructure at all. Note also that carrier status is reported by the platform bridge and means "this carrier is up", not "the relay connection is authenticated"; a bridge that reports a connection it never authenticates produces no verdicts either, and its messages settle by acknowledgement timeout as they always did. diff --git a/docs/state-machines/outbox-and-retries.md b/docs/state-machines/outbox-and-retries.md index e18e9c32..22f506f0 100644 --- a/docs/state-machines/outbox-and-retries.md +++ b/docs/state-machines/outbox-and-retries.md @@ -76,7 +76,14 @@ should be: - **No transport is not terminal.** A send with nowhere to go persists to the outbox, offers the frame to the mesh, schedules a retry, and emits a non-terminal deferral event. Terminal failure comes only from retry-budget - exhaustion or expiry. + exhaustion or expiry. **A resend asks the same question.** The retry queue + and both outbox flushes offer the frame to the mesh when no carrier takes + it, and when a carrier takes it for a recipient it cannot reach (the + internet transport and Reticulum accept any recipient; BLE and Wi-Fi Direct + refuse one they hold no link to). Without that, a message whose first + attempt went into a link that was already dead waited out its whole outbox + lifetime for that link to return, while a neighbour that could reach the + recipient was never asked. - **Parking is entered from `Pending`, not from `Queued`.** The unreachable verdict is an asynchronous relay report about a message the transport already accepted, which is why there is a pending acknowledgement for the park to diff --git a/examples/offline-first/README.md b/examples/offline-first/README.md index e6737056..2d9c1e8e 100644 --- a/examples/offline-first/README.md +++ b/examples/offline-first/README.md @@ -113,9 +113,11 @@ each: ## Known gaps -- **A message a direct link took and then lost never crosses the mesh** - (#541): a message sent down a stream to a device that went away without - closing it waits for a direct link to that device. +- **A message a direct link took and then lost crosses the mesh late**: + a message sent down a stream to a device that went away without closing + it is retried down that stream until keepalive ends it (about 30 s), and + only the next retry is handed to the neighbours. That is why scenario 4 + stops C before taking it off the network A shares with it. - **One phone per Linux box over Bluetooth LE.** The box's peripheral learns which phone wrote to it, but it maps a phone to its user id only while that phone is the one central connected, so with two a reply has no route back diff --git a/examples/offline-first/run.py b/examples/offline-first/run.py index 01bbba49..a767593e 100644 --- a/examples/offline-first/run.py +++ b/examples/offline-first/run.py @@ -277,8 +277,8 @@ def part_a_and_c(self) -> None: # Stopped while still on net-ab, so A sees the stream close. Taken off # the network first, C would leave A a stream that looks open until # keepalive ends it (about 30 s), and a message A sends in that window - # goes down it, is retried over direct carriers only, and never - # reaches the mesh (#541). + # goes down it and is retried down it, reaching the mesh only on the + # first retry after keepalive, past HOP_RECEIVED_S. self._docker("stop", f"{PROJECT}-c") self._docker("network", "disconnect", f"{PROJECT}_net-ab", f"{PROJECT}-c") self._docker("start", f"{PROJECT}-c")