Problem
A client-side request timeout recovers by acting on the whole host, not on the request that timed out. When one daemon serves several sessions, one slow request ends the others.
- The runner sweep runs on every local timeout.
handleRequestTimeout calls cleanupTimedOutIosRunnerBuilds before it checks the command's timeout policy. The sweep runs pkill -f with patterns that match every agent-device runner xcodebuild on the host: other sessions, other daemons, other workspaces. A timed-out snapshot, whose policy is preserve-daemon, still kills every runner.
- Reset-policy commands kill the shared daemon.
open, back, scroll, settings and the other commands on DEFAULT_TIMEOUT_POLICY SIGKILL the daemon (resetDaemonAfterTimeout), which ends every session it owns.
Consumer impact: tester-army/e2e runs one session per worker slot on one local daemon. A cold iOS runner outlasted the 90 s open envelope, and the reset ended the sibling worker's session (tester-army/e2e#627). The engine now starts each runner with prepare ios-runner and warms devices one at a time only to stay clear of this path.
Required behavior (needs design)
- The sweep targets only the runner of the device or session the timed-out request used, never every runner on the host.
- When the daemon has other live sessions, a timed-out request recovers its own session (cancel the request, retire that session's runner) and leaves the daemon and the other sessions running.
- A full daemon reset stays available for a daemon that does not answer at all.
Open question: the client knows only process-name patterns. Scoping the recovery probably means the daemon owns it, through a cancel or retire call, or through its own request deadline. Decide that before implementation.
Done when
- Two sessions on one daemon: a timed-out
open in session A leaves session B able to run snapshot and press.
- A timed-out
snapshot kills no process outside its own session's runner.
Related: #3105 (timeout reset removes metadata without proving ownership), #3116 (daemon retirement ownership). This issue is about the scope of recovery, not how a reset retires the daemon.
From the tester-army/e2e integration review.
Problem
A client-side request timeout recovers by acting on the whole host, not on the request that timed out. When one daemon serves several sessions, one slow request ends the others.
handleRequestTimeoutcallscleanupTimedOutIosRunnerBuildsbefore it checks the command's timeout policy. The sweep runspkill -fwith patterns that match every agent-device runnerxcodebuildon the host: other sessions, other daemons, other workspaces. A timed-outsnapshot, whose policy ispreserve-daemon, still kills every runner.open,back,scroll,settingsand the other commands onDEFAULT_TIMEOUT_POLICYSIGKILL the daemon (resetDaemonAfterTimeout), which ends every session it owns.Consumer impact: tester-army/e2e runs one session per worker slot on one local daemon. A cold iOS runner outlasted the 90 s
openenvelope, and the reset ended the sibling worker's session (tester-army/e2e#627). The engine now starts each runner withprepare ios-runnerand warms devices one at a time only to stay clear of this path.Required behavior (needs design)
Open question: the client knows only process-name patterns. Scoping the recovery probably means the daemon owns it, through a cancel or retire call, or through its own request deadline. Decide that before implementation.
Done when
openin session A leaves session B able to runsnapshotandpress.snapshotkills no process outside its own session's runner.Related: #3105 (timeout reset removes metadata without proving ownership), #3116 (daemon retirement ownership). This issue is about the scope of recovery, not how a reset retires the daemon.
From the tester-army/e2e integration review.