Skip to content

fix(bindings): Android and the Python service start a session over Bluetooth LE - #542

Merged
bahdotsh merged 12 commits into
mainfrom
fix/ble-android-to-python-peripheral
Oct 8, 2026
Merged

bahdotsh merged 12 commits into
mainfrom
fix/ble-android-to-python-peripheral

Conversation

@bahdotsh

@bahdotsh bahdotsh commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Why

An Android phone and a Mac running the Python service found and verified each other over Bluetooth LE and never formed a session, so with encryption required no message ever crossed. Found on a lab of two Macs and a Galaxy M36 (Android 16). The chain, each step in the logs:

  1. macOS rotates its Bluetooth address. The Android central dialled both addresses, both verified as the same peer, and it kept both links.
  2. The Mac's peripheral then saw two centrals. bless drops the CBCentral behind a CoreBluetooth write, so the peripheral attributed every write to the placeholder ble-peer, and it served no Hello characteristic for the phone to announce itself.
  3. The core requires a hop-0 control frame's sender to match the link's identity, so it refused the phone's key package (TRANSPORT_IDENTITY_MISMATCH). No session, no message.
  4. Two links per Mac filled the phone's four connection slots, so the cap also refused the Macs' own centrals, closing the other route.

Two more faults turned up in the same lab:

  • A Mac that may not use Bluetooth wedged the whole service. bless's CoreBluetooth constructor waits, with no deadline, for a power-on that never comes when the process is not authorised (a service started over ssh). It blocked the event loop, so no API socket opened and the 30 s transport start deadline never fired.
  • A dead link stayed "connected". CoreBluetooth kept a link to a Mac whose service had restarted for over five minutes, and every frame written into it vanished.

What

Android (CentralGattClient, MeshConnectionRegistry, BleTransportFacade)

  • A client link that verifies as a peer already held at another address is closed unannounced, and that address is left undialed for a minute. A server link to the same peer is not a duplicate.

Python peripheral (ble_peripheral.py)

  • Builds bless's server with a subclass of its delegate that passes each write's central, offset and value through, and answers with the ATT code the peripheral returns. BlueZ already names the writer through the bus hook.
  • Serves Hello and follows the chapter: refuse an offset write (0x07) or a value outside 96 to 512 bytes (0x0D) before verifying, verify with the core's verifier, bind the central to the derived address, never rebind, announce only when no other link has. That central's fragments enter the core under the proved address.
  • On macOS, notifies a bound peer's central alone, and retries a notification CoreBluetooth refuses because its queue is full. bless ignored that refusal and dropped fragments of every burst.
  • Builds the server off the loop on macOS with a 20 s deadline, and refuses first when macOS reports Bluetooth denied or restricted. BlueZ still builds on the loop thread, where its setup task must be scheduled.

Python central (ble_manager.py)

  • Writes its hello to a peer that serves one.
  • Closes a second client link to a peer it already holds.
  • Reads each link's Device id back every 15 s, and drops a link whose read fails, hangs past 5 s, or names another address.

iOS central (BleManager.swift, new BleHelloPolicy.swift)

  • Writes its hello, as Android has since Android BLE: writes from an unmapped central address wait for a resolution dial (seconds to ~40 s) #512: Hello is discovered on every path, and the assertion it serves is written with response once per connection, after the Message subscription and before the handshake reads, so ahead of any Message write. Without it an iPhone could verify a Python box and never start a session through the link it opened. A refused hello is a diagnostic, never a dropped link. react_native_ios_central_writes_its_hello pins the call site.

Docs: bridge rule P15 states the invariant (a peripheral link carries the address its hello proved, or it cannot start a session). Also the CHANGELOG, and the offline-first demo's known gap.

Verification

Check Result
Python suite 1067 passed, 1 skipped
FFI crate unit tests and source guards 208 passed; the new iOS guard fails under each of 6 mutations
SwiftPM tests 461 passed, including 5 new BleHelloPolicyTests; CI's iOS bridge typecheck passes
New test_ble_hello.py 34 tests; each of 10 mutations of the fixes turns at least one red, a no-op control stays green
Android JVM unit tests new DuplicateClientLinkTest 5/5; the rest pass except PeerStreamFramingTest, which cannot find its vectors file when run from a copy outside the repo (unrelated to this change)
Phone and this Mac, Bluetooth only session forms; Mac to phone delivered over ble with receipt in 0.185 s; phone to Mac delivered (encrypted, hop 0), phone shows the receipt
Phone and the second Mac session in 1 s; delivered over ble in 0.174 s
Mac to Mac, Bluetooth only each central's hello binds on the other's peripheral; delivery after a restart in 0.11 s
Service started over ssh with Bluetooth on socket up in 21 s, peripheral fails with a reason, other transports run (was: never)

Not exercised on hardware: the stale-link liveness drop (macOS reported every disconnect promptly in this round; covered by unit tests), BlueZ (no Linux box in the lab), and any iPhone: the iOS hello write is typechecked, unit-tested and pinned, but no iPhone was in the lab.

Not fixed here: four INVALID_CIPHERTEXT decrypt failures seen once between the two Macs, cause unknown.

A macOS peripheral rotates its Bluetooth address and advertises under the
old and the new one at once. The Android central dialled both, both
verified as the same peer, and both were kept. Two such peers then filled
the facade's four connection slots, and the cap refused every inbound
central, which is how such a peripheral's own central would have learned
who we are.

A client link that verifies as a peer we already hold a client link to at
another address is now closed unannounced, and its address is left
undialed for a minute so the scanner does not redial it on every
advertisement. A server link to the same peer is not a duplicate: a peer
may hold one link per direction.
An Android phone and a Mac running the service found and verified each
other over Bluetooth LE and never formed a session. With two centrals
connected (the phone held a link to each of the Mac's rotating
addresses) the peripheral attributed every write to the placeholder
ble-peer, so the core refused the phone's key package as a transport
identity mismatch: a hop-0 control frame must name the link's identity.

- bless drops the CBCentral behind a CoreBluetooth write. The server is
  built with a subclass of bless's delegate that passes each request's
  central, offset and value through and answers with the ATT code the
  peripheral returns. On BlueZ the bus hook already names the writer.
- The peripheral serves Hello and runs the chapter's order: refuse an
  offset write or a value outside 96..512 bytes before verifying, verify
  with the core's verifier, bind the central to the derived address,
  never rebind, and announce only when no other link has. That central's
  fragments then enter the core under the proved address.
- The central writes its hello to a peer that serves one, and closes a
  second client link to a peer it already holds.
- On macOS a fragment for a bound peer is notified to that central alone,
  and a NO from updateValue (a full transmit queue), which bless ignored,
  waits for the stack's ready signal and is retried.
- The central reads each link's Device id back every 15 s and drops a
  link whose read fails, hangs or names another address. CoreBluetooth
  kept a link to a restarted peer connected for over five minutes while
  every write into it vanished.
- bless's CoreBluetooth constructor waits, with no deadline, for a
  power-on that never comes when the process may not use Bluetooth. On
  the loop thread that wedged the whole service past its start deadline.
  The server is now built on a daemon thread with a 20 s deadline on
  macOS, and a process macOS reports denied or restricted is refused
  first. BlueZ still builds on the loop thread, where its setup task must
  be scheduled.

Run between two Macs and a Galaxy M36 on Android 16: sessions form and
messages are delivered over BLE in both directions.
P15 states the invariant the peripheral fixes depend on: a link carries
the address its hello proved, or it cannot start a session, and names the
failure each piece prevents. The offline-first demo's known gap narrows
to a phone that writes no hello (iOS today).
An iPhone writing into the Python service's peripheral is named only by
its connection id there, and the core refuses its key package, whose
sender must be the link's identity. So an iPhone could find and verify a
Python box and never start a session through the link it opened. Android
has written the Bluetooth LE Hello since #512; iOS never did.

The iOS central now discovers Hello with the handshake characteristics on
every path, and writes the identity assertion it serves, with response,
once per connection: after the Message subscription and before the
handshake reads, so ahead of any Message write in CoreBluetooth's GATT
order. BleHelloPolicy decides it (the peer serves Hello, nothing written
yet this connection, the value within 96..512 bytes and one write), and a
refused hello is a diagnostic, never a dropped link. The per-connection
record is cleared wherever announcedPeripherals is.

react_native_ios_central_writes_its_hello pins the call site, since no
Swift test reaches the CoreBluetooth delegate; six mutations of it each
fail the guard. No iPhone was run.
The iOS central checked that its hello fits one write by asking
CoreBluetooth for maximumWriteValueLength(for: .withResponse).
On iOS that answer is 512 on every link, whatever the MTU.

It turns out the with-response figure is not "what fits in one
packet". It is "what CoreBluetooth is willing to send for you",
and anything longer than one packet goes out as a prepared write.
The chapter forbids a prepared hello, and the Python peripheral
refuses one with an invalid-offset error before verifying it. So
on any link whose MTU is below the assertion length plus three,
the iPhone sends a hello that can never bind, and the link falls
back to exactly the attribution failure the hello exists to fix.

Ask for the without-response figure instead, which is the
payload of one ATT packet, and keep writing with response. That
is the same bound Android already uses from its negotiated MTU.

The source guard now pins the argument, and every discover call
that asks for the handshake characteristics must ask for Hello
too, not only the three-element literal it used to count. A
four-element list with App tag and no Hello sailed straight
through the old check.
The Python peripheral binds a central to the address its hello
proves, and announces that peer to the core. It only ever forgot
a binding when the monitor saw the central leave.

The monitor tracks subscribers. A central that writes a hello and
disconnects without subscribing was never "known", so it never
"left", and its binding sat there for the life of the process,
along with a route in the core to the address it proved. Any
device holding a valid assertion could grow that map by saying
hello and walking away. This is not great.

Give a bound central five seconds to show up as a subscriber.
A hello legitimately lands just before the subscription does,
so it gets the grace; after that the binding goes, and the peer
is reported lost unless another central of ours or our own
client link still reaches it.

While at it, two small things in the same file:

- The per-recipient notify lookup tested membership in the
  subscriber map and then indexed it. CoreBluetooth's queue
  removes an unsubscribing central between the two, and the
  KeyError dropped a fragment. Use .get.

- A bless constructor that never returned holds the delegate
  swap lock forever, so a later start in the same process sat on
  the lock until the 20 s init deadline and then blamed the
  adapter. Wait two seconds for the lock and say what actually
  happened.
When a central refuses a second client link to a peer it already
holds (a macOS peripheral advertising under an old and a new
address), it leaves the refused address alone for a minute so it
is not redialed, verified and refused again on every
advertisement.

Android did that with a bare deadline. If the kept link died ten
seconds into the window, the peer sat undialed at the address it
still advertises for the other fifty. The Python central did not
suppress at all, and redialed the duplicate every cooldown.

Record which peer the refused address duplicated, and treat the
refusal as live only while a client link to that peer is still
held somewhere else. The kept link going away ends it on the
next dial attempt. The Python central now does the same.

While at it, the earlier commit inserted the refusal helper
between rejectPeer and its KDoc, which left that doc attached to
the wrong function. Move it back where it belongs.

The Android call order (refuse on the verified arm before
announcing, check the refusal before the connection cap) is not
reachable from a JVM test, so a source guard pins it.
The chapter says a peripheral refuses a prepared (long) hello and a
hello at a non-zero offset *before* it verifies anything. The macOS
delegate did not quite do that.

CoreBluetooth hands a long write to the delegate as one batch of
requests at rising offsets, and the delegate walked the batch request
by request. The first chunk sits at offset 0 with a perfectly
plausible length, so it went straight into the verifier. Only the
second chunk's offset produced the refusal. The outcome was still a
refusal, but the verifier ran on bytes it should never have seen.

Judge the batch as a whole first: a hello in a batch of more than one
request, or at any offset, is answered with 0x07 and nothing in the
batch reaches the handler. The scan runs on CoreBluetooth's queue, so
a raise inside it falls back to the old path rather than skipping the
response and leaving the central hanging.

While at it, build the delegate class under the swap lock too. An
Objective-C class name registers once per process, and two
constructions racing to build it would both try.
Two things kept the Python central's hello from doing its one job,
which is letting the peer's peripheral name the link before the key
package arrives.

First, on Linux it never wrote one at all. The bound came from the
client's mtu_size, and it turns out bleak's BlueZ client reports 23
there until someone explicitly acquires the MTU, which nothing does.
That is a 20-byte payload against a 96-byte minimum, so every Linux
hello took the "does not fit" exit. Silently. Take the bound from the
Hello characteristic's max_write_without_response_size instead: it
reads BlueZ's negotiated MTU property, and on CoreBluetooth it is the
same one-packet figure the iOS central already uses.

Second, the central announced the peer *before* the hello. The
announce is what makes the core push a key package, and the drain it
schedules runs during the very next await, so the key package was
written without response ahead of the hello. A Python peripheral
then attributes it to a bare central UUID and the core refuses it as
a transport identity mismatch. Session start waits for the re-push.
This is exactly the failure the hello exists to prevent.

Subscribe, say hello, then announce. The core accepts fragments from
a peer it has not been told about yet, so subscribing first loses
nothing. The hello is written with response under a 3 second
deadline, matching Android's hello watchdog, so a peer that never
answers cannot hold the link unannounced.
The peripheral already asked the central before reporting a peer
lost. The central never asked back. So when our client link to a peer
went away (a liveness drop, an address rotation, the peer's cap
evicting us) while that peer's central was still bound on our own
peripheral, the central told the core the peer was gone. The core
dropped its route, even though a notification addressed to that peer
would still have reached it.

The new liveness read makes this fire on its own, which is how it
stopped being theoretical.

Give the peripheral a holds_peer of its own (a *subscribed* central
carries the address: one that said hello and never subscribed is not
a route), and have the central consult it on disconnect and in the
stale sweep. Android's registry has asked both directions all along.

The two probes are weak references. Bound methods in both directions
made the roles a cycle, and a dropped ProtocolManager then kept the
core, and the stores it holds open, alive until a GC pass. The
store-release tests caught that one. A role that is gone reaches
nothing, which is the right answer anyway.

While at it, stop calling a link that was already torn down by a
liveness drop "unverified" when the stack's own disconnect callback
arrives for it afterwards.
The Windows wheel job went red on the test that drops a hello whose
central never subscribes. The code was fine. The test was not.

The monitor measures the grace with time.monotonic(), and the tests
shrank the grace to 10 ms and then slept 2 ms at a time, expecting
the clock to move. On Windows it ticks every 15.6 ms, and asyncio
runs any timer due within one tick at once, so two hundred "sleeps"
finished inside a single tick. The grace never elapsed and the
binding was never dropped. The sibling test passed by luck.

Give the monitor loop a clock attribute, real by default, and have
the tests hand it one that advances a fixed step per reading. No
test depends on wall time anymore, on any platform.
The hello's bounds (96 to 512 bytes) are written down in Python,
Swift and Kotlin, and nothing compared them. A central and a
peripheral that disagree fail silently: the peripheral refuses a
value under its floor before verifying it, and a central with a wider
bound just spends a write that can only be refused.

Pin all of them. The floor is the sealed codec's own minimum
assertion length; the ceiling is the Android server's, which served
Hello first. Moving any one of them now fails here instead of on a
board run.
@bahdotsh
bahdotsh merged commit d7ea462 into main Oct 8, 2026
24 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 8, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant