Repository navigation
feat(bindings): an offline-first demo on three devices, and OP_GATEWAY in the image - #539
Merged
Merged
Conversation
…etworking properties
This was referenced Oct 8, 2026
Merged
Adding OP_GATEWAY to the comment block that lists the entrypoint's variables took a space from the OP_SOCKET line, so its description no longer lined up with every other entry. Restore it.
Under --ssh, run.py starts `await --until delivered` on A in the background and then prompts the operator to switch B back on. The ssh client inherited the terminal's stdin and forwards whatever it reads to the remote command, so it competed with input() for the operator's Enter: the keypress could be swallowed by the remote side, leaving the prompt waiting while the receipt window ran out. Run every verifier, foreground and background, with stdin from /dev/null. Nothing the verifier does reads it.
The hardware runbook has each host run the service image under host networking, but run.py --ssh ran `offline-protocol-verify --socket /run/offline-protocol/api.sock` on the host itself. The local API socket is inside the container and the host has no verifier, so every command failed and no scenario could start. A host running the service directly fared no better: its default socket is under XDG_RUNTIME_DIR or ~/.offline-protocol, never the image's path. Add --remote-exec, a command the verifier runs through on each host (`docker exec <container>` for the image), and --socket for a service started with another path. The README's hardware section says which to use.
Two ways a scenario that stopped part-way poisoned what ran after it: - Scenario 1 or 2 raising between switching B off and on (a refused send, a socket that never answered) left B stopped. Every later scenario then began with `pair` toward B and failed after 90 s on a device that was simply off, which reads as a networking failure. - Scenario 4 stopping between joining C to net-ab and parting it left C on net-ab, and `compose up` does not take it off. Every later run then failed scenario 4 at once: `docker network connect` refuses a container that is already on the network. After a scenario fails, switch every device back on (Docker: `docker start` all three, a no-op for a running one; ssh: ask the operator) and wait for each socket before the next. Join C to net-ab only when it is not already there. A failure raised by a scenario now carries its name in the table instead of its bare number, and the docstring no longer names a --docker flag the script never had. Checked against the lab: with C left on net-ab, scenario 4 failed before and passes now; with scenario 1 failing while B is off, scenario 2 passes straight after.
The offline-first runbook had the relay leg run scenario 1 with config-relay.json, which keeps wifi_direct_enabled on, on boxes the other legs put on one LAN. The engine tries a live direct mesh link to the recipient before transport selection (send.rs, the direct_mesh branch), so once B is back on the LAN the outbox retry goes down the peer stream, the receipt names wifiDirect, and the leg passes without the relay having carried anything. The gateway leg has the same hole. Say that both legs need no peer stream between the boxes (OP_LAN=0 and no OP_PEERS, or two networks), and that the receipt names internet when the relay carried it.
The documented check, `send` and then `await <id> --until delivered`, timed out whenever the recipient was reachable. `send` closes its connection once `send_message` answers, and `await` opens a new one; a receipt the engine emits in between reaches no client of the application, and the local API holds only inbound tags. Over loopback the receipt lands in that gap every time (5 of 5 runs exited 2). `send --await [--timeout S]` keeps the connection `send_message` was called on, which `VerifyClient.open` subscribed before the call, and waits on it exactly as `await --until delivered` does, so a receipt that beats the call's own answer is queued rather than lost. The guide's example uses it; a separate `await` stays for a recipient that is away, started before it returns.
`watch` promises to span a restart of the service, but it retried only an `OSError`, a closed connection or a refused `hello`. A service that is stopping or starting can accept the connection and close it before the WebSocket handshake answers; websockets raises `InvalidMessage` for that (a `WebSocketException`, not an `OSError`), and the watcher died with status 1 in the very gap it exists to cover. An opening handshake that times out raises `asyncio.TimeoutError`, which is not an `OSError` on Python 3.10. Both are now retried. A `watch --duration` that never reached the service also exited 0 with an empty log, which an orchestrator cannot tell from a quiet network. It now says once that it is waiting, and exits 1 if the duration ends without it ever having connected.
`test_without_the_saved_state_the_message_is_gone` passed when no receipt for the queued message came within five seconds of the relink. That also holds if wiping the protocol-state directory broke sending altogether (a lost session, a refused carrier), so the control could pass without saying anything about the saved queue. It now sends a fresh message after the restart and waits for both its receipt on A and its arrival on B before asserting the queued one neither settles on A nor arrives on B. Keeping the saved state (the mutation the control exists for) still turns it red.
`await` and `watch` printed nothing between connecting and the event they wait for, so a script that starts one in the background and then triggers the event (brings the recipient back, sends from another device) could only sleep and hope the subscription was in place. A receipt or a relay report emitted before it is not held, so a short sleep loses it and a long one only makes that less likely. Both now print the line `subscribed`, alone, on standard error once `subscribe` has answered (`watch` again after each reconnect). The string is `commands.READY_LINE`, documented in the guide and pinned in a test, since a script matches it.
run.py started each background await or watch and then slept two seconds before switching B on or sending. Neither a receipt nor a relay report is held for a client that is not listening yet, so on a slow box (or over ssh) the event could fire before the verifier had subscribed, and the scenario failed 120 s later for no real reason. The verifier now prints a `subscribed` line on stderr once its subscription is confirmed. Wait for that line instead of guessing. Both pipes are drained on threads from the start, so reading stderr early never loses the tail of it and a long watch never blocks on a full stdout pipe. A verifier that exits before subscribing fails the scenario at once with its error, rather than passing as ready.
asyncio has no Unix socket client on Windows: create_unix_connection raises NotImplementedError, so the test failed in the Windows wheel job before the watcher ever got to retry. Every other test that needs a Unix socket already skips there through the harness, and the verifier reaches a Windows service over --tcp. This one built its target by hand and so missed the skip.
…pology # Conflicts: # CHANGELOG.md # bindings/python/offline_protocol_sdk/verify/commands.py # bindings/python/tests/scenarios/test_through_the_middle.py # docs/local-api.md
The demo opens by saying the recipient's acknowledgement comes back as message_delivered, and then scenario 4 passed without it. That was defensible while #537 was open. #537 is merged now, so a missing receipt across the hop is a regression, and the scenario has to fail on it. --require-hop-receipt becomes --allow-missing-hop-receipt, for an image built from a 0.28.0 or earlier wheel whose engine never sends one. The known gaps were stale too. #536 taught the BlueZ peripheral which phone wrote to it. What is left is narrower: a phone is mapped to its user id only while it is the one central connected. The retry-path gap now has an issue (#541), so the README points at it instead of describing the mechanism a third time. The run record says that its lab run used an engine without #537, which is why that run has no receipt. restore() switched devices back on but left C on net-ab after a scenario 4 that failed half-way. Every later scenario then ran with A and C linked, and the lab exists to keep them apart. It now parts them the same way part_a_and_c does, with C stopped first so A sees the stream close. While at it, --ssh without '=' is a usage error instead of a ValueError traceback, and the image README names the verifier check as a second reason the PyPI build fails.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Based on
main(#538 and #537 merged in).What
An offline-first demo on three devices, and the image changes it needs.
examples/offline-first/compose.yml: three services frombindings/python/dockeron two bridge networks (a - net-ab - b - net-bc - c), so b is the only way between a and c. No HTTP front, static peers (the topology is the file's, not multicast DNS's), one identity per volume.examples/offline-first/run.py(standard library only): runs the scenarios withoffline-protocol-verifyinside each device (docker exec, orsshwith the operator switching devices off and on), prints a table, exits non-zero on a failure.SIGKILLed and restarted between the send and B's return;docker network connect, then share no network: C receives athop_count1 within 30 s, B emitsmessage_relayed, and A gets C's receipt back across the hop. A missing receipt fails the scenario;--allow-missing-hop-receiptis for an image built from a 0.28.0 or earlier wheel, whose engine never sends it.examples/offline-first/README.md: the runbook, including the hardware legs (LAN between hosts, the carrier changing under aping, relay, gateway daemon, phone) and a "What has been run" table.OP_GATEWAY->--gateway(+ two entrypoint test cases),config-ble.jsonandconfig-relay.jsonbaked beside the default (one field different each), and the build now also fails when the installed package has nooffline-protocol-verify.What ran
In Docker 28.3 on one arm64 machine,
run.py --freshthenrun.pyagain, on an image whose native library was built frommainafter #537 (overlaid with this branch's Python sources): scenarios 1, 2 and 4 passed both times, scenario 4 with A's receipt required and back at hop 1 (receipts in 33 to 36 ms overwifiDirectin 1 and 2). An earlier pair of runs on the 0.28.0 native library passed 1, 2 and 4 without the receipt, as that engine predicts. The image was never built from the checkout. Nothing has run between separate hosts, over Bluetooth LE, a relay, a gateway daemon, or with a phone; the README says so.scripts/tests/test-docker-entrypoint.shpasses (13 cases) and both scripts are shellcheck-clean; CI runs the test under/bin/shon Ubuntu (dash).Found while running it (engine, not fixed here)
The first lab run failed scenario 4 and the cause is real: a message that a direct stream took and then lost never crosses the mesh. I took C off net-ab before restarting it, so A kept a half-open stream to C for ~30 s; A's send went down it (pending ACK registered), and after keepalive ended the stream every retry was refused by the direct carriers and re-queued (
protocol/mod.rs, the retry loop'sErrarm) withoutoffer_to_mesh, which only the first send's paths call. State minutes later on A:pending_ack_count 1,retry_queue_size 1, meshtransmissions 0.run.pynow stops C while it is still on net-ab so A sees the close (the realistic case is the other one, and the README and the guide list it as a known gap). #537 does not cover it (it never touches the retry loop): filed as #541.Also fixed on #538 while running this:
state/pairconnected under the verifier's application id and so took the message the recipient's service was holding for it; they now connect as<app id>.observer.