Skip to content

feat(bindings): an offline-first demo on three devices, and OP_GATEWAY in the image - #539

Merged
bahdotsh merged 19 commits into
mainfrom
feat/offline-first-topology
Oct 8, 2026
Merged

bahdotsh merged 19 commits into
mainfrom
feat/offline-first-topology

Conversation

@bahdotsh

@bahdotsh bahdotsh commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Based on main (#538 and #537 merged in).

What

An offline-first demo on three devices, and the image changes it needs.

  • examples/offline-first/compose.yml: three services from bindings/python/docker on two bridge networks (a - net-ab - b - net-bc - c), so b is the only way between a and c. No HTTP front, static peers (the topology is the file's, not multicast DNS's), one identity per volume.
  • examples/offline-first/run.py (standard library only): runs the scenarios with offline-protocol-verify inside each device (docker exec, or ssh with the operator switching devices off and on), prints a table, exits non-zero on a failure.
    1. store and forward: B stopped, A sends, B started: B receives, A's receipt within 60 s of B's return;
    2. the same with A SIGKILLed and restarted between the send and B's return;
    3. A and C pair over a temporary docker network connect, then share no network: C receives at hop_count 1 within 30 s, B emits message_relayed, and A gets C's receipt back across the hop. A missing receipt fails the scenario; --allow-missing-hop-receipt is for an image built from a 0.28.0 or earlier wheel, whose engine never sends it.
  • examples/offline-first/README.md: the runbook, including the hardware legs (LAN between hosts, the carrier changing under a ping, relay, gateway daemon, phone) and a "What has been run" table.
  • Image: OP_GATEWAY -> --gateway (+ two entrypoint test cases), config-ble.json and config-relay.json baked beside the default (one field different each), and the build now also fails when the installed package has no offline-protocol-verify.

What ran

In Docker 28.3 on one arm64 machine, run.py --fresh then run.py again, on an image whose native library was built from main after #537 (overlaid with this branch's Python sources): scenarios 1, 2 and 4 passed both times, scenario 4 with A's receipt required and back at hop 1 (receipts in 33 to 36 ms over wifiDirect in 1 and 2). An earlier pair of runs on the 0.28.0 native library passed 1, 2 and 4 without the receipt, as that engine predicts. The image was never built from the checkout. Nothing has run between separate hosts, over Bluetooth LE, a relay, a gateway daemon, or with a phone; the README says so.

scripts/tests/test-docker-entrypoint.sh passes (13 cases) and both scripts are shellcheck-clean; CI runs the test under /bin/sh on Ubuntu (dash).

Found while running it (engine, not fixed here)

The first lab run failed scenario 4 and the cause is real: a message that a direct stream took and then lost never crosses the mesh. I took C off net-ab before restarting it, so A kept a half-open stream to C for ~30 s; A's send went down it (pending ACK registered), and after keepalive ended the stream every retry was refused by the direct carriers and re-queued (protocol/mod.rs, the retry loop's Err arm) without offer_to_mesh, which only the first send's paths call. State minutes later on A: pending_ack_count 1, retry_queue_size 1, mesh transmissions 0. run.py now stops C while it is still on net-ab so A sees the close (the realistic case is the other one, and the README and the guide list it as a known gap). #537 does not cover it (it never touches the retry loop): filed as #541.

Also fixed on #538 while running this: state/pair connected under the verifier's application id and so took the message the recipient's service was holding for it; they now connect as <app id>.observer.

Adding OP_GATEWAY to the comment block that lists the entrypoint's
variables took a space from the OP_SOCKET line, so its description
no longer lined up with every other entry. Restore it.
Under --ssh, run.py starts `await --until delivered` on A in the
background and then prompts the operator to switch B back on. The
ssh client inherited the terminal's stdin and forwards whatever it
reads to the remote command, so it competed with input() for the
operator's Enter: the keypress could be swallowed by the remote
side, leaving the prompt waiting while the receipt window ran out.

Run every verifier, foreground and background, with stdin from
/dev/null. Nothing the verifier does reads it.
The hardware runbook has each host run the service image under host
networking, but run.py --ssh ran `offline-protocol-verify --socket
/run/offline-protocol/api.sock` on the host itself. The local API
socket is inside the container and the host has no verifier, so
every command failed and no scenario could start. A host running
the service directly fared no better: its default socket is under
XDG_RUNTIME_DIR or ~/.offline-protocol, never the image's path.

Add --remote-exec, a command the verifier runs through on each host
(`docker exec <container>` for the image), and --socket for a
service started with another path. The README's hardware section
says which to use.
Two ways a scenario that stopped part-way poisoned what ran after it:

- Scenario 1 or 2 raising between switching B off and on (a refused
  send, a socket that never answered) left B stopped. Every later
  scenario then began with `pair` toward B and failed after 90 s on
  a device that was simply off, which reads as a networking failure.
- Scenario 4 stopping between joining C to net-ab and parting it
  left C on net-ab, and `compose up` does not take it off. Every
  later run then failed scenario 4 at once: `docker network connect`
  refuses a container that is already on the network.

After a scenario fails, switch every device back on (Docker: `docker
start` all three, a no-op for a running one; ssh: ask the operator)
and wait for each socket before the next. Join C to net-ab only when
it is not already there. A failure raised by a scenario now carries
its name in the table instead of its bare number, and the docstring
no longer names a --docker flag the script never had.

Checked against the lab: with C left on net-ab, scenario 4 failed
before and passes now; with scenario 1 failing while B is off,
scenario 2 passes straight after.
The offline-first runbook had the relay leg run scenario 1 with
config-relay.json, which keeps wifi_direct_enabled on, on boxes the
other legs put on one LAN. The engine tries a live direct mesh link
to the recipient before transport selection (send.rs, the
direct_mesh branch), so once B is back on the LAN the outbox retry
goes down the peer stream, the receipt names wifiDirect, and the
leg passes without the relay having carried anything. The gateway
leg has the same hole.

Say that both legs need no peer stream between the boxes (OP_LAN=0
and no OP_PEERS, or two networks), and that the receipt names
internet when the relay carried it.
The documented check, `send` and then `await <id> --until delivered`,
timed out whenever the recipient was reachable. `send` closes its
connection once `send_message` answers, and `await` opens a new one; a
receipt the engine emits in between reaches no client of the
application, and the local API holds only inbound tags. Over loopback
the receipt lands in that gap every time (5 of 5 runs exited 2).

`send --await [--timeout S]` keeps the connection `send_message` was
called on, which `VerifyClient.open` subscribed before the call, and
waits on it exactly as `await --until delivered` does, so a receipt
that beats the call's own answer is queued rather than lost. The guide's
example uses it; a separate `await` stays for a recipient that is away,
started before it returns.
`watch` promises to span a restart of the service, but it retried only
an `OSError`, a closed connection or a refused `hello`. A service that
is stopping or starting can accept the connection and close it before
the WebSocket handshake answers; websockets raises `InvalidMessage`
for that (a `WebSocketException`, not an `OSError`), and the watcher
died with status 1 in the very gap it exists to cover. An opening
handshake that times out raises `asyncio.TimeoutError`, which is not
an `OSError` on Python 3.10. Both are now retried.

A `watch --duration` that never reached the service also exited 0
with an empty log, which an orchestrator cannot tell from a quiet
network. It now says once that it is waiting, and exits 1 if the
duration ends without it ever having connected.
`test_without_the_saved_state_the_message_is_gone` passed when no
receipt for the queued message came within five seconds of the relink.
That also holds if wiping the protocol-state directory broke sending
altogether (a lost session, a refused carrier), so the control could
pass without saying anything about the saved queue.

It now sends a fresh message after the restart and waits for both its
receipt on A and its arrival on B before asserting the queued one
neither settles on A nor arrives on B. Keeping the saved state (the
mutation the control exists for) still turns it red.
`await` and `watch` printed nothing between connecting and the event
they wait for, so a script that starts one in the background and then
triggers the event (brings the recipient back, sends from another
device) could only sleep and hope the subscription was in place. A
receipt or a relay report emitted before it is not held, so a short
sleep loses it and a long one only makes that less likely.

Both now print the line `subscribed`, alone, on standard error once
`subscribe` has answered (`watch` again after each reconnect). The
string is `commands.READY_LINE`, documented in the guide and pinned in
a test, since a script matches it.
run.py started each background await or watch and then slept two
seconds before switching B on or sending. Neither a receipt nor a
relay report is held for a client that is not listening yet, so on a
slow box (or over ssh) the event could fire before the verifier had
subscribed, and the scenario failed 120 s later for no real reason.

The verifier now prints a `subscribed` line on stderr once its
subscription is confirmed. Wait for that line instead of guessing.
Both pipes are drained on threads from the start, so reading stderr
early never loses the tail of it and a long watch never blocks on a
full stdout pipe. A verifier that exits before subscribing fails the
scenario at once with its error, rather than passing as ready.
asyncio has no Unix socket client on Windows: create_unix_connection
raises NotImplementedError, so the test failed in the Windows wheel
job before the watcher ever got to retry. Every other test that
needs a Unix socket already skips there through the harness, and
the verifier reaches a Windows service over --tcp. This one built
its target by hand and so missed the skip.
Base automatically changed from feat/offline-protocol-verify to main October 8, 2026 11:09
…pology

# Conflicts:
#	CHANGELOG.md
#	bindings/python/offline_protocol_sdk/verify/commands.py
#	bindings/python/tests/scenarios/test_through_the_middle.py
#	docs/local-api.md
The demo opens by saying the recipient's acknowledgement comes back
as message_delivered, and then scenario 4 passed without it. That
was defensible while #537 was open. #537 is merged now, so a
missing receipt across the hop is a regression, and the scenario
has to fail on it. --require-hop-receipt becomes
--allow-missing-hop-receipt, for an image built from a 0.28.0 or
earlier wheel whose engine never sends one.

The known gaps were stale too. #536 taught the BlueZ peripheral
which phone wrote to it. What is left is narrower: a phone is
mapped to its user id only while it is the one central connected.
The retry-path gap now has an issue (#541), so the README points
at it instead of describing the mechanism a third time. The run
record says that its lab run used an engine without #537, which
is why that run has no receipt.

restore() switched devices back on but left C on net-ab after a
scenario 4 that failed half-way. Every later scenario then ran
with A and C linked, and the lab exists to keep them apart. It now
parts them the same way part_a_and_c does, with C stopped first so
A sees the stream close.

While at it, --ssh without '=' is a usage error instead of a
ValueError traceback, and the image README names the verifier
check as a second reason the PyPI build fails.
@bahdotsh
bahdotsh merged commit b1e4580 into main Oct 8, 2026
24 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 8, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant