Repository navigation
fix(bindings): Android and the Python service start a session over Bluetooth LE - #542
Merged
Merged
Conversation
A macOS peripheral rotates its Bluetooth address and advertises under the old and the new one at once. The Android central dialled both, both verified as the same peer, and both were kept. Two such peers then filled the facade's four connection slots, and the cap refused every inbound central, which is how such a peripheral's own central would have learned who we are. A client link that verifies as a peer we already hold a client link to at another address is now closed unannounced, and its address is left undialed for a minute so the scanner does not redial it on every advertisement. A server link to the same peer is not a duplicate: a peer may hold one link per direction.
An Android phone and a Mac running the service found and verified each other over Bluetooth LE and never formed a session. With two centrals connected (the phone held a link to each of the Mac's rotating addresses) the peripheral attributed every write to the placeholder ble-peer, so the core refused the phone's key package as a transport identity mismatch: a hop-0 control frame must name the link's identity. - bless drops the CBCentral behind a CoreBluetooth write. The server is built with a subclass of bless's delegate that passes each request's central, offset and value through and answers with the ATT code the peripheral returns. On BlueZ the bus hook already names the writer. - The peripheral serves Hello and runs the chapter's order: refuse an offset write or a value outside 96..512 bytes before verifying, verify with the core's verifier, bind the central to the derived address, never rebind, and announce only when no other link has. That central's fragments then enter the core under the proved address. - The central writes its hello to a peer that serves one, and closes a second client link to a peer it already holds. - On macOS a fragment for a bound peer is notified to that central alone, and a NO from updateValue (a full transmit queue), which bless ignored, waits for the stack's ready signal and is retried. - The central reads each link's Device id back every 15 s and drops a link whose read fails, hangs or names another address. CoreBluetooth kept a link to a restarted peer connected for over five minutes while every write into it vanished. - bless's CoreBluetooth constructor waits, with no deadline, for a power-on that never comes when the process may not use Bluetooth. On the loop thread that wedged the whole service past its start deadline. The server is now built on a daemon thread with a 20 s deadline on macOS, and a process macOS reports denied or restricted is refused first. BlueZ still builds on the loop thread, where its setup task must be scheduled. Run between two Macs and a Galaxy M36 on Android 16: sessions form and messages are delivered over BLE in both directions.
P15 states the invariant the peripheral fixes depend on: a link carries the address its hello proved, or it cannot start a session, and names the failure each piece prevents. The offline-first demo's known gap narrows to a phone that writes no hello (iOS today).
An iPhone writing into the Python service's peripheral is named only by its connection id there, and the core refuses its key package, whose sender must be the link's identity. So an iPhone could find and verify a Python box and never start a session through the link it opened. Android has written the Bluetooth LE Hello since #512; iOS never did. The iOS central now discovers Hello with the handshake characteristics on every path, and writes the identity assertion it serves, with response, once per connection: after the Message subscription and before the handshake reads, so ahead of any Message write in CoreBluetooth's GATT order. BleHelloPolicy decides it (the peer serves Hello, nothing written yet this connection, the value within 96..512 bytes and one write), and a refused hello is a diagnostic, never a dropped link. The per-connection record is cleared wherever announcedPeripherals is. react_native_ios_central_writes_its_hello pins the call site, since no Swift test reaches the CoreBluetooth delegate; six mutations of it each fail the guard. No iPhone was run.
The iOS central checked that its hello fits one write by asking CoreBluetooth for maximumWriteValueLength(for: .withResponse). On iOS that answer is 512 on every link, whatever the MTU. It turns out the with-response figure is not "what fits in one packet". It is "what CoreBluetooth is willing to send for you", and anything longer than one packet goes out as a prepared write. The chapter forbids a prepared hello, and the Python peripheral refuses one with an invalid-offset error before verifying it. So on any link whose MTU is below the assertion length plus three, the iPhone sends a hello that can never bind, and the link falls back to exactly the attribution failure the hello exists to fix. Ask for the without-response figure instead, which is the payload of one ATT packet, and keep writing with response. That is the same bound Android already uses from its negotiated MTU. The source guard now pins the argument, and every discover call that asks for the handshake characteristics must ask for Hello too, not only the three-element literal it used to count. A four-element list with App tag and no Hello sailed straight through the old check.
The Python peripheral binds a central to the address its hello proves, and announces that peer to the core. It only ever forgot a binding when the monitor saw the central leave. The monitor tracks subscribers. A central that writes a hello and disconnects without subscribing was never "known", so it never "left", and its binding sat there for the life of the process, along with a route in the core to the address it proved. Any device holding a valid assertion could grow that map by saying hello and walking away. This is not great. Give a bound central five seconds to show up as a subscriber. A hello legitimately lands just before the subscription does, so it gets the grace; after that the binding goes, and the peer is reported lost unless another central of ours or our own client link still reaches it. While at it, two small things in the same file: - The per-recipient notify lookup tested membership in the subscriber map and then indexed it. CoreBluetooth's queue removes an unsubscribing central between the two, and the KeyError dropped a fragment. Use .get. - A bless constructor that never returned holds the delegate swap lock forever, so a later start in the same process sat on the lock until the 20 s init deadline and then blamed the adapter. Wait two seconds for the lock and say what actually happened.
When a central refuses a second client link to a peer it already holds (a macOS peripheral advertising under an old and a new address), it leaves the refused address alone for a minute so it is not redialed, verified and refused again on every advertisement. Android did that with a bare deadline. If the kept link died ten seconds into the window, the peer sat undialed at the address it still advertises for the other fifty. The Python central did not suppress at all, and redialed the duplicate every cooldown. Record which peer the refused address duplicated, and treat the refusal as live only while a client link to that peer is still held somewhere else. The kept link going away ends it on the next dial attempt. The Python central now does the same. While at it, the earlier commit inserted the refusal helper between rejectPeer and its KDoc, which left that doc attached to the wrong function. Move it back where it belongs. The Android call order (refuse on the verified arm before announcing, check the refusal before the connection cap) is not reachable from a JVM test, so a source guard pins it.
The chapter says a peripheral refuses a prepared (long) hello and a hello at a non-zero offset *before* it verifies anything. The macOS delegate did not quite do that. CoreBluetooth hands a long write to the delegate as one batch of requests at rising offsets, and the delegate walked the batch request by request. The first chunk sits at offset 0 with a perfectly plausible length, so it went straight into the verifier. Only the second chunk's offset produced the refusal. The outcome was still a refusal, but the verifier ran on bytes it should never have seen. Judge the batch as a whole first: a hello in a batch of more than one request, or at any offset, is answered with 0x07 and nothing in the batch reaches the handler. The scan runs on CoreBluetooth's queue, so a raise inside it falls back to the old path rather than skipping the response and leaving the central hanging. While at it, build the delegate class under the swap lock too. An Objective-C class name registers once per process, and two constructions racing to build it would both try.
Two things kept the Python central's hello from doing its one job, which is letting the peer's peripheral name the link before the key package arrives. First, on Linux it never wrote one at all. The bound came from the client's mtu_size, and it turns out bleak's BlueZ client reports 23 there until someone explicitly acquires the MTU, which nothing does. That is a 20-byte payload against a 96-byte minimum, so every Linux hello took the "does not fit" exit. Silently. Take the bound from the Hello characteristic's max_write_without_response_size instead: it reads BlueZ's negotiated MTU property, and on CoreBluetooth it is the same one-packet figure the iOS central already uses. Second, the central announced the peer *before* the hello. The announce is what makes the core push a key package, and the drain it schedules runs during the very next await, so the key package was written without response ahead of the hello. A Python peripheral then attributes it to a bare central UUID and the core refuses it as a transport identity mismatch. Session start waits for the re-push. This is exactly the failure the hello exists to prevent. Subscribe, say hello, then announce. The core accepts fragments from a peer it has not been told about yet, so subscribing first loses nothing. The hello is written with response under a 3 second deadline, matching Android's hello watchdog, so a peer that never answers cannot hold the link unannounced.
The peripheral already asked the central before reporting a peer lost. The central never asked back. So when our client link to a peer went away (a liveness drop, an address rotation, the peer's cap evicting us) while that peer's central was still bound on our own peripheral, the central told the core the peer was gone. The core dropped its route, even though a notification addressed to that peer would still have reached it. The new liveness read makes this fire on its own, which is how it stopped being theoretical. Give the peripheral a holds_peer of its own (a *subscribed* central carries the address: one that said hello and never subscribed is not a route), and have the central consult it on disconnect and in the stale sweep. Android's registry has asked both directions all along. The two probes are weak references. Bound methods in both directions made the roles a cycle, and a dropped ProtocolManager then kept the core, and the stores it holds open, alive until a GC pass. The store-release tests caught that one. A role that is gone reaches nothing, which is the right answer anyway. While at it, stop calling a link that was already torn down by a liveness drop "unverified" when the stack's own disconnect callback arrives for it afterwards.
The Windows wheel job went red on the test that drops a hello whose central never subscribes. The code was fine. The test was not. The monitor measures the grace with time.monotonic(), and the tests shrank the grace to 10 ms and then slept 2 ms at a time, expecting the clock to move. On Windows it ticks every 15.6 ms, and asyncio runs any timer due within one tick at once, so two hundred "sleeps" finished inside a single tick. The grace never elapsed and the binding was never dropped. The sibling test passed by luck. Give the monitor loop a clock attribute, real by default, and have the tests hand it one that advances a fixed step per reading. No test depends on wall time anymore, on any platform.
The hello's bounds (96 to 512 bytes) are written down in Python, Swift and Kotlin, and nothing compared them. A central and a peripheral that disagree fail silently: the peripheral refuses a value under its floor before verifying it, and a central with a wider bound just spends a write that can only be refused. Pin all of them. The floor is the sealed codec's own minimum assertion length; the ceiling is the Android server's, which served Hello first. Moving any one of them now fails here instead of on a board run.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
An Android phone and a Mac running the Python service found and verified each other over Bluetooth LE and never formed a session, so with encryption required no message ever crossed. Found on a lab of two Macs and a Galaxy M36 (Android 16). The chain, each step in the logs:
CBCentralbehind a CoreBluetooth write, so the peripheral attributed every write to the placeholderble-peer, and it served no Hello characteristic for the phone to announce itself.TRANSPORT_IDENTITY_MISMATCH). No session, no message.Two more faults turned up in the same lab:
What
Android (
CentralGattClient,MeshConnectionRegistry,BleTransportFacade)Python peripheral (
ble_peripheral.py)0x07) or a value outside 96 to 512 bytes (0x0D) before verifying, verify with the core's verifier, bind the central to the derived address, never rebind, announce only when no other link has. That central's fragments enter the core under the proved address.Python central (
ble_manager.py)iOS central (
BleManager.swift, newBleHelloPolicy.swift)react_native_ios_central_writes_its_hellopins the call site.Docs: bridge rule P15 states the invariant (a peripheral link carries the address its hello proved, or it cannot start a session). Also the CHANGELOG, and the offline-first demo's known gap.
Verification
BleHelloPolicyTests; CI's iOS bridge typecheck passestest_ble_hello.pyDuplicateClientLinkTest5/5; the rest pass exceptPeerStreamFramingTest, which cannot find its vectors file when run from a copy outside the repo (unrelated to this change)blewith receipt in 0.185 s; phone to Mac delivered (encrypted, hop 0), phone shows the receiptblein 0.174 sNot exercised on hardware: the stale-link liveness drop (macOS reported every disconnect promptly in this round; covered by unit tests), BlueZ (no Linux box in the lab), and any iPhone: the iOS hello write is typechecked, unit-tested and pinned, but no iPhone was in the lab.
Not fixed here: four
INVALID_CIPHERTEXTdecrypt failures seen once between the two Macs, cause unknown.