Skip to content

feat(bindings): offline-protocol-verify, and scenario tests for the networking properties - #538

Merged
bahdotsh merged 10 commits into
mainfrom
feat/offline-protocol-verify
Oct 8, 2026
Merged

bahdotsh merged 10 commits into
mainfrom
feat/offline-protocol-verify

Conversation

@bahdotsh

@bahdotsh bahdotsh commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

What

offline-protocol-verify, a command for the device a service runs on, and scenario tests that run the networking properties end to end through the local API.

The engine already reports every networking fact as an event: message_deferred while a recipient is away, message_relayed on a device that carries a frame for someone else, message_received with hop_count and transport, and message_delivered, the recipient's own acknowledgement naming the carrier it arrived over. The verifier sends through the service and waits for those events, so a check passes or fails on what the engine says.

Commands: state, send, await ID --until delivered|received, pair, ping, watch. JSON lines on stdout, a summary on stderr, exit 0 pass / 2 timeout / 1 terminal failure. await matches events by the id they carry, never by the server's correlation, so it works on a sender restarted since the send (and does not depend on the message_undeliverable terminal-tag fix).

The verifier has its own small client (verify/client.py) rather than sharing the HTTP front's persistent one: a command is short-lived, and the verifier must not depend on the front's package. Only watch reconnects, around it.

Scenario tests (tests/scenarios/)

Loopback peer streams, file stores with a store key, encryption on and required, as in the container image. A device is switched off (server and engine stop, stores released) and on again over the same stores and port.

  • Store and forward: B off, A sends, message_deferred; B on, B gets the message, A gets message_delivered{transport: wifiDirect}. Also through the commands (await started while B is off).
  • Reboot: B off, A sends, A restarts, B on, the restarted A delivers it (and its router never saw the id). Control: the same restart over an empty protocol-state directory delivers nothing.
  • Through the middle: A and C meet once directly (the key exchange and Welcome never cross a hop), then C hears only B; A sends to C, C receives at hop_count == 1, B emits message_relayed.

Mutation checks, each red, then restored: dropping A's restart (reboot), dropping the pre-pair (C never receives: first contact does not cross a hop), keeping A and C linked (hop_count 0), keeping B on (no message_deferred).

Engine defect found: the sender's receipt never crosses the mesh

In the through-the-middle scenario C's acknowledgement is carried back by B (message_relayed from C to A) and reaches A, which drops it. A message whose only route is the mesh is refused by every carrier of the sender (handle_send_failure), so no pending acknowledgement is registered, and an acknowledgement with no pending record settles only a relay-parked DM (settle_parked_dm_from_ack). A keeps re-offering, C keeps re-acknowledging, and the outbox reports the delivered message failed after 7 days. The Rust simulator's the_answer_finds_its_way_back asserts only that B transmits toward A.

This is the bug #537 fixes. test_the_receipt_comes_back_through_the_middle_device records it as a strict expected failure: whichever of this PR and #537 merges second will see it XPASS and fail, and the marker comes off then. The guide and CHANGELOG say so.

Not covered

  • A receipt the engine emits while no client of the application is connected is not held (only received messages are), so await/watch on the sender must start before the recipient can answer. Documented.
  • Scenario tests skip on Windows (Unix socket); the TCP carrier has its own test.

Verified

  • Python suite: 998 passed + 1 xfailed locally (macOS, 3.14).
  • cargo test -p offline-protocol-uniffi --lib: 205 passed (the guard that walks every .py in the package).

Second commit (53cfad2)

Found by running #539 in containers: state and pair connected under the verifier's application id, so on the recipient they took the message the service was holding for that application and await --until received then timed out. They now connect as <app id>.observer; a test pins it (red with the old id). state no longer reports get_topology, which is empty on a running service. The guide also records a second engine gap seen there: a message that a direct stream took and then lost is retried over direct carriers only and never handed to the mesh.

The documented check, `send` and then `await <id> --until delivered`,
timed out whenever the recipient was reachable. `send` closes its
connection once `send_message` answers, and `await` opens a new one; a
receipt the engine emits in between reaches no client of the
application, and the local API holds only inbound tags. Over loopback
the receipt lands in that gap every time (5 of 5 runs exited 2).

`send --await [--timeout S]` keeps the connection `send_message` was
called on, which `VerifyClient.open` subscribed before the call, and
waits on it exactly as `await --until delivered` does, so a receipt
that beats the call's own answer is queued rather than lost. The guide's
example uses it; a separate `await` stays for a recipient that is away,
started before it returns.
`watch` promises to span a restart of the service, but it retried only
an `OSError`, a closed connection or a refused `hello`. A service that
is stopping or starting can accept the connection and close it before
the WebSocket handshake answers; websockets raises `InvalidMessage`
for that (a `WebSocketException`, not an `OSError`), and the watcher
died with status 1 in the very gap it exists to cover. An opening
handshake that times out raises `asyncio.TimeoutError`, which is not
an `OSError` on Python 3.10. Both are now retried.

A `watch --duration` that never reached the service also exited 0
with an empty log, which an orchestrator cannot tell from a quiet
network. It now says once that it is waiting, and exits 1 if the
duration ends without it ever having connected.
`test_without_the_saved_state_the_message_is_gone` passed when no
receipt for the queued message came within five seconds of the relink.
That also holds if wiping the protocol-state directory broke sending
altogether (a lost session, a refused carrier), so the control could
pass without saying anything about the saved queue.

It now sends a fresh message after the restart and waits for both its
receipt on A and its arrival on B before asserting the queued one
neither settles on A nor arrives on B. Keeping the saved state (the
mutation the control exists for) still turns it red.
`await` and `watch` printed nothing between connecting and the event
they wait for, so a script that starts one in the background and then
triggers the event (brings the recipient back, sends from another
device) could only sleep and hope the subscription was in place. A
receipt or a relay report emitted before it is not held, so a short
sleep loses it and a long one only makes that less likely.

Both now print the line `subscribed`, alone, on standard error once
`subscribe` has answered (`watch` again after each reconnect). The
string is `commands.READY_LINE`, documented in the guide and pinned in
a test, since a script matches it.
asyncio has no Unix socket client on Windows: create_unix_connection
raises NotImplementedError, so the test failed in the Windows wheel
job before the watcher ever got to retry. Every other test that
needs a Unix socket already skips there through the harness, and
the verifier reaches a Windows service over --tcp. This one built
its target by hand and so missed the skip.
#537 is on main now, so the carried receipt settles the sender's
message. The strict expected failure on the receipt test would XPASS
and turn red the moment this branch merged main, which is exactly
what the marker was set up to do. It comes off.

The guide, pair's docstring and the scenario module still said two
devices must have met directly before reaching each other through a
third, because the key package and the Welcome never crossed a hop.
That stopped being true with the same merge: while a message waits
for a peer no carrier reaches, both ride the mesh. Telling operators
to pre-pair devices that don't need it is not great.

Add the scenario that proves it: A and C never link, B sits between
them, A sends to C, C receives at one hop and A gets
message_delivered. The guide now says the receipt proves a crossing
end to end, and keeps the one gap still real: a message a direct
stream took and then lost is retried over direct carriers only.
@bahdotsh
bahdotsh merged commit ace5dd8 into main Oct 8, 2026
24 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 8, 2026
@bahdotsh
bahdotsh deleted the feat/offline-protocol-verify branch October 8, 2026 11:09
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant