Skip to content

core: give every framework component one owner — a shared I/O loop, executor-scoped affinity, and Bridge/RemoteServer on their owners #847

Description

@Yaraslaut

Summary

morph gives models one threading rule: a single owner (a strand), and state touched only there. The framework's own components do not follow it. They expose public APIs callable from any thread, run private threads, and receive replies on threads they do not own. The visible cost is 57 mutex members across 27 headers, 6 private std::thread members plus one per server client, 10 condition variables, about 10 lock acquisitions per local execute before any handler runs, and an 800-line concurrency spec that mostly documents how those pieces avoid deadlocking each other.

This issue applies the model rule to the framework. Every component takes an executor at construction, touches its state only in tasks running on it, hands that executor (never inlineExecutor()) to every completion or bind it awaits so replies arrive as posted tasks, and enumerates its cross-thread surface in its spec. Debug builds assert it; the ThreadSanitizer leg is a backstop. The mutex count is reported on every PR as a side measure. It is not the goal: a lock count is a control that also passes when locks are hidden in wrappers or moved into core-cpp.

Three mechanisms do almost all of the work: one owner per component with owner-delivered replies (~24 locks), one I/O loop (~9), and one-hop remote dispatch on a server strand (~9). A Completion redesign removes the one lock on every call path. Five steps, each landing on its own; steps 1–2 are non-breaking; steps 3–5 amount to a new major version; step 6 is parked.

Verification status

Measured on master @ 9d001052 (macOS arm64, AppleClang 17; the earlier revision figures on Linux, clang 22):

  • The 57 declarations, per file, with the grep below. Also 54 at 2407ffdf, 55 at a020e69c: the count rises with every feature that adds a runtime-settable field.
  • Every declaration printed with its declaration comment (script below), which is what the cause table rests on.
  • The lock sites of every Bridge, LocalBackend and RemoteServer lock; the lock acquisitions on one local execute.
  • Every caller of Bridge off its owning thread in examples/ and tests/.
  • The framework-owned std::thread members and the condition-variable count.
  • The cancellation gap (completion_awaiter.hpp, wire.hpp, both transports), by reading.
  • What core-cpp v0.5.0 ships, read from the pinned checkout .cache/cpm/core-cpp/a7096e8d…, not run.
  • Two -fsyntax-only probes (under Left in place).

Inferred, not measured: each lock's cause, from its declaration comment and lock sites. Treat each cause's count as ±3. The one-hop dispatch ordering argument (step 5) is inferred; tests/test_remote_execute_ordering.cpp is the check.

Not measured at all: any performance effect of any step. Every performance statement here is an expectation, and the #843 baseline exists to turn it into a measurement.

Evidence

Count. A data-member declaration of any standard mutex type:

grep -rcE '^\s*(mutable\s+)?(inline\s+)?(static\s+)?(std::)?(shared_mutex|mutex|recursive_mutex|timed_mutex|shared_timed_mutex)\s+[_A-Za-z][_A-Za-z0-9]*\s*(\{\s*\}\s*)?;' include/ src/
→ 57 in 27 files (0 in src/)

core/bridge.hpp        9  _backendMtx _mtx _attachMtx _sessionMtx _principalMtx _executeDeadlineMtx registrationMtx mtx(:402 handoff) mtx(:613 BridgeLifetime, shared_mutex)
core/remote.hpp        7  _regMtx _logProviderMtx _limitsMtx _healthMtx _drainMtx _pendingMtx _sessionMtx
net/socket_backend.hpp 7  _socketMtx _connectMtx _reconnectMtx _syncMtx _reconnectHandlerMtx _sessionMtx _handlerMtx
core/backend.hpp       5  _pendingMtx(SynchronousBackendAdapter) _mtx _regMtx _pendingMtx _taskRunsMtx (LocalBackend)
net/socket_server.hpp  3  writeMtx _closeMtx _clientsMtx
2 each: journal/journal.hpp, core/observability.hpp, core/executor.hpp, core/detail/execute_order_gate.hpp
1 each: completion.hpp, detail/reply_router.hpp, detail/subscription_registry.hpp, logger.hpp, strand.hpp,
        timeout_scheduler.hpp, forms/flows.hpp, forms/sections.hpp, journal/action_log.hpp,
        journal/file_action_log.hpp, offline/{file_offline_queue, network_monitor, offline_queue,
        reconnect_coordinator, replay_ledger, sqlite_offline_queue, sync_worker}.hpp, qt/qt_websocket_backend.hpp

Per-declaration listing with comments, for the audit:

python3 - <<'EOF'
import re,glob
pat=re.compile(r'^\s*(mutable\s+)?(inline\s+)?(static\s+)?(std::)?(shared_mutex|mutex|recursive_mutex|timed_mutex|shared_timed_mutex)\s+\w+\s*(\{\s*\}\s*)?;')
n=0
for f in sorted(glob.glob('include/**/*.hpp',recursive=True)):
    L=open(f).read().split('\n')
    for i,l in enumerate(L):
        if pat.match(l): n+=1; print(f"{f}:{i+1} {l.strip()}")
print("TOTAL",n)
EOF

All 57 are live. QtWebSocketBackend::_pendingMtx is declared in the header and locked at ten sites in src/qt/qt_websocket_backend.cpp (:197, :256, :265, :386, :416, :431, :517, :556, :583); a search of the header alone misreads it as dead. Whether that backend is single-thread-only, which would make the lock redundant, is a step 3 question.

Locks on one local execute. From the lock sites: executeVia takes _executeDeadlineMtx (bridge.hpp:2147), _sessionMtx (:2320), BridgeLifetime's gate shared (:2379), _backendMtx (:2384); LocalBackend::execute takes _regMtx (backend.hpp:1441) and _pendingMtx twice (:1521, :1550); CompletionState::mtx on every attach and settle of the three chained completions per call (qt_executor.hpp:30–33). About 10 acquisitions before any handler runs.

Private threads, against docs/spec/concurrency_and_lifetimes.md:30 "morph has no ad-hoc threads scattered through the dispatch path":

net/socket_backend.hpp:1011    std::thread _ioThread;
net/socket_backend.hpp:1012    std::thread _handlerThread;
net/socket_server.hpp:314      std::thread clientThread{...}     (one per client connection)
net/socket_server.hpp:532      std::thread _acceptThread;
core/timeout_scheduler.hpp:325 std::thread _thread;
offline/network_monitor.hpp:197 std::thread _thread;

_handlerThread is not an I/O thread. It exists only because the reconnect handler calls sendSync (a blocking register) and must not do so on the I/O thread (socket_backend.hpp:694–717); bindModel is already non-blocking (sendControlAsync). Once the reconnect is posted to the bridge's owner and registration is always asynchronous, sendSync has no caller and the thread goes with it.

The contract says any thread, in so many words. bridge.md:949 "Bridge is fully thread-safe"; concurrency_and_lifetimes.md:206 teardown is "order-independent, on any thread"; TimeoutScheduler's schedule()/cancel() "may be called from any thread"; SocketBackend "may safely be driven from multiple threads concurrently"; HandlerBinding::registrationMtx exists because "a waiter can be queued or resolved from either the registering thread or the backend's reply-delivering thread"; the bridge.hpp:402 handoff exists for "a backend that replies from another thread while its own dispatch call is still on this stack".

The same spec says the opposite. bridge.md:1031: "The intended usage remains single-GUI-thread affinity: a handler and its subscriptions belong to one GUI thread." Both sentences are in one file.

In practice Bridge already has one owner. Every setDefaultSession/setPrincipal/setExecuteDeadline/switchBackend call outside include/ runs on the app's main thread or a test's main thread. The only off-thread callers are the framework's own reconnect handler, which runs on the transport thread and takes _mtx+_attachMtx, and tests written to the free-threaded contract (test_concurrency_invariants.cpp:269, test_bridge_lifetime.cpp:635, test_switch_backend.cpp:1016). LocalBackend is the same: every _regMtx site (backend.hpp:1271, :1291, :1313, :1324, :1337, :1376, :1441) is on the caller's thread, none on a strand.

The free-threaded contract already needed an escape hatch nobody uses. Bridge takes an optional bridgeExec (bridge.hpp:997) whose documented purpose is to close the late-reply check-then-use window. Callers passing it: one test (tests/test_async_registration.cpp:2500), zero examples. The window the spec describes as "open for [a Bridge] that was not [given an executor]" is open in all seven rungs.

The registration machinery exists because a bind reply may settle anywhere. rebindThroughSurface (bridge.hpp:2597–2620) is the template: bindModel(request, inlineExecutor()), then a continuation that either parkIfInFrames or deliverLates to bridgeExec, then awaitHandoff/claimHandoff on a mayBlock policy. BridgeLifetime, the :402 handoff, deliverLate, registrationMtx, BindWait and whenBound's waiter list all serve that one uncertainty. Pass the owner instead of inlineExecutor() and every branch becomes "post the continuation".

The setter locks come from features, not a design. _sessionMtx/_principalMtx arrived with #34 (ad4f7e1c, "add Principal"). Each runtime-settable field adds a lock because the contract lets any thread set it.

One loop per component is the in-tree anti-pattern. TimeoutScheduler already runs a core::net::PlatformLoop, on its own std::thread natively and host-driven on WASM, and needs _mtx to batch requests across to that thread. A loop per component is a thread and a lock per component.

Remote dispatch takes two hops, and the second hop is the only reason the order gate exists. handleImpl already decodes the model id on the transport thread (remote.hpp:370–388) and then posts to _pool, whose task posts to the model's strand. ExecuteOrderGate (two locks, a condition variable, and awaitTurn()'s cv.wait "with no deadline" on a pool thread, remote.hpp:1012) exists to keep per-model FIFO across that pool hop, and its own doc says it "serialises across every model rather than per model" (execute_order_gate.hpp:203). Post the execute straight to the strand and the strand's FIFO is the guarantee.

The 57, grouped by cause

Cause ≈ Locks
1. Public API callable from any thread ~19 Bridge::_sessionMtx, _principalMtx, _executeDeadlineMtx, _backendMtx, _mtx, _attachMtx; RemoteServer::_logProviderMtx, _limitsMtx, _healthMtx, half of _regMtx (health() "from any thread"); SimulatedRemoteBackend::_sessionMtx; SocketBackend::_sessionMtx; SocketServer::_closeMtx (serialises close() against itself); TimeoutScheduler::_mtx; SyncWorker::_runMtx; ReconnectCoordinator::_mtx; LocalBackend::_regMtx; logger mtx; observability metricMtx, traceMtx
2. A reply or callback lands on a thread the receiver does not own ~15 registrationMtx; the :402 handoff; _pendingMtx ×4 (Local, adapter, Simulated, Qt); PendingCallTable::_mtx; SubscriptionRegistry::_mtx; _taskRunsMtx; LocalBackend::_mtx; forms FlowSession/sections _mtx; SocketBackend::_syncMtx, _handlerMtx, _reconnectHandlerMtx (the sync-register-on-the-I/O-thread problem, not the private-thread problem)
3. The component owns a private thread ~6 SocketBackend::_socketMtx, _connectMtx, _reconnectMtx; SocketServer::writeMtx, _clientsMtx; NetworkMonitor::_mtx
4. Destruction on any thread 1 BridgeLifetime's shared_mutex. A handler and its bridge may be destroyed on different threads in either order, so ~BridgeHandler holds the gate shared while it deregisters and ~Bridge takes it exclusively first. A CallbackToken cannot do this job: its liveness check is advisory across threads.
5. Executor and queue internals 6 executor.hpp ×2, execute_order_gate ×2, strand.hpp's _enrolledMtx, RemoteServer::_drainMtx (guards the drain condition variable)
6. Storage with concurrent users ~10 journal ×2, action log, file action log, offline queues ×3, replay ledger, half of RemoteServer::_regMtx; plus CompletionState::mtx, the cross-executor settle

Causes 1–4 (~41) follow from two choices: free-threaded public objects, and private threads. Causes 5–6 (~16) exist for reasons any design must handle; the gate's two and the settle lock go anyway, for reasons given under steps 4 and 5.

A note on the forms locks: sections.hpp:326 and flows.hpp:487 each say the continuation "runs on whatever thread resolves the BridgeHandler completion". The continuation is _handler.execute(...).then(_callbacks, …) (sections.hpp:294–305), delivered on the handler's guiExec. The lock exists because the forms tests use an InlineExec as guiExec (tests/test_sections.cpp, test_flows_apps.cpp, test_forms_rules.cpp), which makes the continuation run on the settling thread. Naming the owner (step 3) resolves it.

What core-cpp v0.5.0 provides, and what morph is missing to use it

Read on the pinned checkout, not run.

Needed for core-cpp has morph has
"Am I on my owner?" (the assert every affine step relies on) ExecutorScope, currentExecutor(), ExecutorScope::anyInForce; Strand::runningHere(); EventLoop::isOnWorkerThread() Only ModelStrands (strand.hpp:412) and the Task-handler driver (coroutine.hpp:54) state a scope. ThreadPoolExecutor, MainThreadExecutor, QtExecutor, InlineExecutor state none, so on the GUI thread currentExecutor() is null. ExecutorScope takes a core::async::IExecutor& (ExecutorContext.hpp:69); morph's executors implement morph::exec::IExecutor and reach core-cpp only through a per-use CoreExecutorOver adapter (strand.hpp:62), so no morph executor has a stable core identity to name.
One reactor that is also an executor EventLoop : IExecutor with post(), submit(), addTimer, delay, waitReadable/Writable. Its doc: "all scheduler state is touched only on the loop thread. The cross-thread surface is post(), submit(), schedule(), requestCancel() and stop()" — the shape every morph component should be able to state. SocketBackend/SocketServer run raw poll() on private threads over morph's own TcpSocket.
The reactor with no thread (WASM) PlatformLoop is host-driven under single-threaded WASM (PlatformLoop.hpp:20–24: it "must be PUMPED and neither run() nor blockOn() may be called on it") TimeoutScheduler already does this (timeout_scheduler.hpp:16–21). morph_net never builds on WASM (MORPH_BUILD_NET is ON only in linux-everything); the WASM rungs use QtWebSocketBackend on Qt's loop. No Qt host scheduler is needed.
A component as an actor Strand (one, over any executor, StrandReclaim::Never), KeyedStrands ModelStrands wraps KeyedStrands, for models only. A KeyedStrands key is retired when idle (Strand.hpp:669), so a component must own a Strand, not a key.
A cancellable deadline DeadlineTimer TimeoutScheduler reimplements it with a thread and a map.
Coroutine-first client API Task<T>, co_await Already there: Completion<T>::operator co_await (completion.hpp:549).
A settle with no mutex completion_awaiter.hpp's Shared: one atomic outcome, no lock CompletionState::mtx mediates attach-on-consumer-thread against settle-on-producer-thread (completion.hpp:49–140).
Stop propagation StopSource/StopToken/StopCallback, OperationCancelled, whenAll/whenAny parent→child stop, strands with seal()/close() One StopSource per run (LocalRun::stopSource, backend.hpp:118), requested by the client deadline, the server executeTimeout and ~LocalBackend. Not reached from a cancelled co_await (completion_awaiter.hpp:106 onStop only ends the awaiter's own scope), not from cancelPending/switchBackend/~Bridge, and not across the wire (wire.hpp has no cancel kind; neither transport reads call.stopSource). That is #846.

Dependencies

Ticket Provides This issue's relation
#843 Benchmark baseline (morph_bench, morph_bench_alloc, plus a SocketBackend loopback round trip, fan-out across models and connections, Completion settle, registration latency), a per-dispatch lock count, JSON archived as workflow artifacts, and a stated noise band from its best-of-N method Blocks step 2. Every step below reports its numbers against this baseline.
#846 One stop per execute, reached from every cancel path and across the wire (cancel {callId}, negotiated via hello); the G0–G4 cancellation policy in concurrency_and_lifetimes.md, held by tests/test_cancellation_policy.cpp Blocks step 3's scope part. Steps 3–5 change G-levels; that suite is what shows by how much.
#844 (PR) BridgeHandler::executeWhenBound Merge first. Step 3 makes it the behaviour of execute and removes the method. One point to settle there the way step 3 settles it: a handler destroyed before its bind must reject the held Completion, not leave it unsettled (an awaiting coroutine leaks otherwise).

Steps

Every PR reports: morph's mutex count before and after (the grep above); any lock moved into core-cpp, listed separately; the #843 benchmark delta (p50/p99, throughput, allocations per call) against its noise band.

Acceptance for every lock removed: a test that calls the verb from a thread that is not the owner and asserts, inside the verb's body, that the call was posted (ExecutorScope::anyInForce(owner) is true there). It fails deterministically if the post is skipped, with or without a sanitizer. A debug-build assert that trips off-owner is paired with that positive test, not a substitute for it. TSAN is a backstop, not the evidence. A removal no test can tell from the original is not verified.

Step 1 — morph's executors state an ExecutorScope

Non-breaking in API; one pinned behaviour change. Locks removed: 0. Depends on: nothing.

  • Every morph::exec::IExecutor owns a stable core::async::IExecutor adapter (the CoreExecutorOver shape, moved out of strand.hpp), reachable as coreExecutor(). ThreadPoolExecutor, MainThreadExecutor and QtExecutor construct core::async::ExecutorScope{coreExecutor()} around each task they run, the way core-cpp's loop, pool and strand do (ExecutorContext.hpp: "once per turn / once per worker thread / once per batch"). InlineExecutor states no scope: it runs on the caller's thread, and whatever scope is in force there stays in force.
  • Adds morph::exec::runningOn(IExecutor const&) over ExecutorScope::anyInForce, so any component can ask "am I on my owner" in one line.
  • ModelStrands uses the base executor's coreExecutor() instead of its own CoreExecutorOver _base.
  • Behaviour change to pin: completion_awaiter.hpp records ResumeTarget::current() and falls back to cbExec when it is empty. A coroutine started through spawn already resumes on its spawn executor, because ExecutorResumer (coroutine.hpp:42) states a scope of its own. A coroutine run inside an executor's task without spawn has no scope today and resumes on cbExec; after this step it resumes on the executor it was running on. Same thread in practice; pin both cases with tests and state the rule in coroutines.md "Where a coroutine resumes".
  • Why first: without it there is nothing for steps 3–5's asserts to read.
  • Acceptance: for each of the three executors, a test posts a task and observes runningOn(exec) true and core::async::currentExecutor() == &exec.coreExecutor() inside it, false and null outside, and true inside an InlineExecutor task run from inside the posted task. Each test is shown to fail with the scope statement removed. The co_await resumption test above.
  • strand.hpp's _enrolledMtx stays. It guards one unordered_map shared by every key (strand.hpp:299). The around-task hook on key B's strand reads it (:393–394) while enroll on key A's strand inserts (:264–266), so a rehash races the read; and withdraw runs off its key's strand (the handler-end callback calls it before runOnStrand, backend.hpp:1738, remote.hpp:1687). It could go only if _enrolled became per-key state, which is a design change for step 3, not a tidy-up.
  • Spec: docs/spec/core/executor.md gains a "Current executor" section; coroutines.md updated.

Step 2 — one I/O loop per process

Non-breaking. Locks expected to go: ~9, plus 5 thread types. Depends on: #843 (before/after measurement), step 1.

  • One core::net::PlatformLoop owns every socket, the accept loop, every client connection, the timers and the connectivity probe. SocketBackend, SocketServer, TimeoutScheduler (over DeadlineTimer) and NetworkMonitor become flows on that loop; none owns a thread.
  • Natively the loop owns its thread, as _ioThread does today. On single-threaded WASM PlatformLoop is already host-driven, and nothing in morph_net builds there. No Qt host scheduler.
  • Pending-call records (PendingCallTable) are owned by the loop: a record goes out with the request and comes back with the reply. No lock.
  • SocketBackend's "may be driven from multiple threads" becomes: the cross-thread surface is execute/registerModel/deregisterModel, each of which posts to the loop. Enumerate it in backend.md.
  • Locks: SocketBackend::_socketMtx, _connectMtx, _reconnectMtx; SocketServer::writeMtx, _clientsMtx, _closeMtx; NetworkMonitor::_mtx; TimeoutScheduler::_mtx; PendingCallTable::_mtx. (_syncMtx, _handlerMtx, _reconnectHandlerMtx and _handlerThread are step 3's.)
  • Acceptance: grep -nE '<thread>|std::thread|jthread|std::async|detach\(' include/morph/net include/morph/core/timeout_scheduler.hpp include/morph/offline/network_monitor.hpp finds nothing. The SocketBackend loopback and fan-out benchmarks from observability: Tracy zones behind MORPH_ENABLE_TRACY (prefixed MORPH_ZONE macros) and a capture job that proves they fire #843 stay inside its stated noise band; if they do not, see What would change this. The per-lock tests above for each of the nine.
  • Spec: concurrency_and_lifetimes.md's thread-roles table gains "the I/O loop" and loses "the probe thread" and "the transport thread".

Step 3 — Bridge, handlers and LocalBackend have one owner; bind replies are delivered on it

Breaking (contract and API). Locks expected to go: ~24, plus _handlerThread. Depends on: step 1, #844 merged, #846.

Decision to record first, in the spec, by the spec owner: Bridge, every BridgeHandler and LocalBackend belong to the executor given at construction (bridgeExec becomes required; BridgeHandler already takes guiExec). They are created, called and destroyed there. An off-owner call is a contract violation, asserted in debug builds. bridge.md:949 ("fully thread-safe") and concurrency_and_lifetimes.md:206 ("on any thread") become false; bridge.md:1031 becomes the contract. The evidence for the decision is in Evidence above: no consumer calls off-thread, and the free-threaded contract already leaks an optional bridgeExec that nobody passes.

Then, as one change, because every piece below exists for the same uncertainty about where a bind reply settles:

  • bindModel/promoteModel receive the owner as their executor, never inlineExecutor(). Every bind continuation is a posted task on the owner.
  • Registration is asynchronous, without a policy. The rule: the bind is a Completion delivered on the owner; a backend that can settles it before returning; execute dispatches immediately when currentId (already atomic) is set and chains on the bind otherwise. A LocalBackend handler's first execute still dispatches at once. A handler constructed inside a running action on a pool thread (concurrency_and_lifetimes.md:418) is off-owner and posts its registration; that case is why registration is asynchronous, not a cost of it. A failed bind rejects with the bind's error; a handler destroyed mid-bind rejects its held completions. The "handler not bound" failure class disappears, and with it the caller-side gating: 31 trackBound() sites across 16 example files today.
  • Deleted outright: BridgeLifetime, AsyncDispatchHandoff (bridge.hpp:402), parkIfInFrame/awaitHandoff/claimHandoff, deliverLate, registrationMtx, BindWait/bindWaitPolicy(), whenBound(), trackBound(), executeWhenBound(), and the blocking IBackend::attachModel/registerModelShared verbs (attachHandler/ensureBound, bridge.hpp:1108, :1396, are the last blocking register verbs Bridge calls; under one owner they are the same posted bind). This reverses the bridge: opt-in way for BridgeHandler::execute to wait for an in-flight async registration instead of failing "handler not bound" #841 triage ("a separately named method rather than a constructor policy") and says so.
  • _sessionMtx, _principalMtx, _executeDeadlineMtx, _backendMtx, _mtx, _attachMtx, SubscriptionRegistry::_mtx go; the fields are plain members touched on the owner. SubscriptionRegistry::publishResult is already called from the .then continuation on guiExec (bridge.hpp:763).
  • The reconnect handler no longer calls into the bridge from the transport thread; the backend posts "reconnected" to the owner. SocketBackend::sendSync then has no caller, and _handlerThread, _syncMtx, _handlerMtx, _reconnectHandlerMtx go.
  • LocalBackend::_regMtx, _mtx, _pendingMtx, _taskRunsMtx and SynchronousBackendAdapter::_pendingMtx go: every site is owner-side (execute writes _taskRuns/_pending; the destructor and cancelPending read them, on the owner). SimulatedRemoteBackend::_sessionMtx, _pendingMtx likewise (remote.hpp:2201–2229). SocketBackend::_sessionMtx likewise.
  • The forms locks (sections.hpp, flows.hpp) go: the continuation runs on the owner. The forms tests stop using an InlineExec as guiExec.
  • Structured scopes. CallbackScope is extended, not replaced: requestStop() also calls request_stop() on the StopSource of every call the scope owns, and every cancel path (cancelPending, switchBackend, ~Bridge) does the same rather than only settling. The wire cancel envelope is core: a cancellation policy for execute — one stop per call, reached from every cancel path and across the wire #846's; this step only routes the scope's stop into it.
  • Tests written to the free-threaded contract (test_concurrency_invariants.cpp:269, test_bridge_lifetime.cpp:635, test_switch_backend.cpp:1016) are rewritten to post to the owner; each becomes a per-lock test of the shape above.
  • Acceptance: the per-lock tests above for each of the ~24. BridgeLifetime, AsyncDispatchHandoff, BindWait, bindWaitPolicy, whenBound, trackBound, executeWhenBound, attachModel, registerModelShared do not exist in include/morph (symbol absence, not a grep for a name that a rename would defeat). grep trackBound examples finds nothing. tests/test_cancellation_policy.cpp shows each cancel verb moving from G1/G2 to the level the spec now states.
  • Spec: bridge.md "Thread safety" and "Registration readiness" rewritten to the owner model; backend.md "Waiting for a bind" removed; concurrency_and_lifetimes.md loses "teardown is order-independent, on any thread" and the four check-then-call dispositions, and its "Cancellation policy" G-levels are updated.

Step 4 — Completion settles by posting

Breaking only for a consumer that attaches off its executor. Locks removed: 1, on every call path. Depends on: step 3.

  • Today CompletionState::mtx mediates attach on the consumer thread against settle on the producer thread: setValue stores under the lock and drains onOk; attachThen checks ready under the lock and fires now or queues (completion.hpp:49–140). Once every consumer has an owner: settle posts to cbExec and applies there; attach runs on cbExec (asserted). All mutation is on one executor and its queue is the happens-before edge, the shape completion_awaiter.hpp's Shared already has. ready becomes an atomic for co_await's fast path. cbExec == nullptr settles inline, since nothing can be delivered there anyway.
  • Cost: one post per settle even with no handler attached (today setValue posts only when a callback exists, completion.hpp:121). Report it against bench.alloc_budget.
  • Acceptance: the existing completion tests; one that attaches from a pool thread and observes the assert; one that settles from a pool thread with no handler attached and observes the value delivered after a later attach on cbExec. CompletionState has no mutex member.
  • Spec: completion.md and concurrency_and_lifetimes.md "Completion / CompletionState thread-safety" rewritten.

Step 5 — RemoteServer on a Strand, one-hop dispatch, ServerConfig; offline orchestration on a strand

Breaking (API). Locks expected to go: ~9. Depends on: step 1.

  • RemoteServer owns a core::async::Strand over its pool. The registry, the connection directory, the limits and the health state are touched only on it.
  • One hop. handleImpl already decodes the model id on the transport thread; an execute is posted straight to _strands->post(mid, …) and decoded, authorised and run there. Every other envelope is posted to the server's strand. handleInline is unaffected (it already rejects execute, remote.hpp:321–335). ExecuteOrderGate, its two locks, its condition variable and the blocking awaitTurn() are deleted: per-model FIFO is the strand's guarantee. The "synchronous executor re-enters takeAndPost on the same thread" case that forced a recursive_mutex (execute_order_gate.hpp:199) disappears with the gate.
  • setLogProvider, setLimitPolicy, setHealthHandler, setSupportedVersionRange become a ServerConfig given at construction. health() and drainedWithin() return Completions answered on the strand; _drainMtx and its condition variable go.
  • SyncWorker::_runMtx and ReconnectCoordinator::_mtx only serialise run()/onOnline()/onOffline() against themselves (reconnect_coordinator.hpp:238, sync_worker.hpp:295). As posted tasks on the offline strand they are serialised by construction.
  • Locks: ExecuteOrderGate::_enqueueMtx, _mtx; RemoteServer::_drainMtx, _regMtx, _limitsMtx, _healthMtx, _logProviderMtx; SyncWorker::_runMtx; ReconnectCoordinator::_mtx.
  • Acceptance: the per-lock tests above; tests/test_remote_execute_ordering.cpp green with the gate gone; a test calling health() from a pool thread observes the posted answer; ServerConfig has no setter. ExecuteOrderGate does not exist in include/morph.
  • Spec: backend.md RemoteServer section: the setters table becomes a ServerConfig table; cross-thread surface enumerated; "Synchronous re-entry" kept; offline.md thread roles updated.

Step 6 — storage (parked)

Locks in scope: ~9 (journal ×2, action log, file action log, offline queues ×3, replay ledger).

Per-model action logs are already single-writer: model.hpp:257 appends from inside execute, on the model's strand, to the log RemoteServer::attachLogIfConfigured (remote.hpp:704–718) gave that holder. What remains locked is the aggregate Journal (cross-model undoLast/checkpoint, journal.hpp:322–334, read from the app thread) and the offline queues. A shared aggregate with N strand writers and an app reader needs a lock or an owner, and an owner turns an uncontended lock into a queue hop per append. Do not implement until #843's benchmarks show contention on one of these locks.

Left in place, and why

Lock Reason
executor.hpp ×2 Queue internals. Go to zero if morph adopts core::async::ThreadPoolExecutor; separate decision.
logger mtx The serialised sink is a documented contract (concurrency_and_lifetimes.md:640–653).
observability metricMtx, traceMtx Not a one-liner on this toolchain: std::atomic<std::shared_ptr<…>> fails to compile on AppleClang 17 libc++ ("requires that 'T' be a trivially copyable type"), and the std::atomic_load(shared_ptr*) free functions that do compile are deprecated in C++20. Uncontended; leave.
strand.hpp _enrolledMtx One map shared by every key, read on one key's strand while another key's strand inserts, and withdraw runs off-strand (see step 1). Goes only with a per-key _enrolled, in step 3.
storage ×9 Step 6, parked.

Expected floor after steps 1–5: 15, of which 9 are step 6's, 2 are the executor queues, and 1 (_enrolledMtx) is step 3's per-key question.

Effect on work in flight

What would change this

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: coreSubsystem: coreenhancementNew feature or requesttriage: rescopeReal problem, wrong framing; rewrite before building

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions