Skip to content

fix(apple-runner): fence prep spawns after teardown or last-waiter cancellation #3220

Description

@thymikee

Purpose

A runner start must stop creating preparation processes after its execution host is torn down or its last interested waiter cancels. Stopping the currently registered children once is insufficient.

PR #3193 measured this during close with a cold build: close killed the first build-for-testing before obtaining the runner session lock, but the detached prewarm health retry spawned a second build while close waited for that lock. Close reached its timeout; the second build remained until daemon stop. Original evidence: #3193 (comment). This observation is at b6cdd3c49, not a fresh reproduction on main.

Required behavior

  • The in-flight runner start owns whether further preparation is admitted. Device teardown closes that admission before stopping current prep children or waiting for the session lock.
  • Canceling the final interested waiter closes admission for that start. Canceling one waiter must preserve work still needed by another.
  • Check admission at every preparation spawn, including retries after a killed build. A check only when registering an already spawned child is too late.
  • A later independent open may create a fresh start; a retired start cannot resume and publish a runner into it.
  • Prefer one start-owned cancellation/admission mechanism. Reassess whether the request-owner prep filter introduced by fix(daemon): scope client timeout recovery to the timed-out request #3193 remains necessary once the reachable callers and waiter semantics are covered.

Done when

  • A control holds the runner session lock during a cold build, begins non-retained close, and attempts a retry after the first child is stopped. No second prep child is spawned; close settles without waiting for a replacement build.
  • Controls cover cancellation before the first prep spawn, last-waiter cancellation, and another still-interested waiter.
  • Repeat the live close-during-build route with evidence of no respawn, completed close, and cleanup. Retain the original evidence attribution.
  • Focused tests and the affected gate pass; process signaling remains scoped to owned children.

Depends on #3193. This is start admission, not another host-wide timeout sweep or a shorter join timeout.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions