Skip to content

fix(daemon): start a lease's TTL when its allocation completes - #3165

Merged
thymikee merged 2 commits into
callstack:mainfrom
LambdaTest:split/lease-ttl
Oct 3, 2026
Merged

thymikee merged 2 commits into
callstack:mainfrom
LambdaTest:split/lease-ttl

Conversation

@amankansal-lt

@amankansal-lt amankansal-lt commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Split out of #3118. This change affects every provider.

leases.allocate used to start the lease TTL when allocation began. A hosted provider can spend 30–80 s creating its session, so a 60 s lease could expire before the caller got it. The TTL now starts when allocation completes. Two cases are unchanged:

  • If the requester cancels, the lease is not renewed.
  • If allocation fails, the lease is released as before.

A lease allocated without a provider keeps the TTL it was created with. It no longer holds a work pass, which used to renew it on release.

ADR 0007 and remote-proxy.md are updated to match.

Validation

  • pnpm check:affected --run passes.
  • New tests cover:
    • an allocation that outlasts the TTL;
    • a lease without a provider keeping its original TTL;
    • artifacts after a slow allocation and after a cancelled one.

🤖 Generated with Claude Code

The registry records a lease, and stamps its expiry, before the lease
lifecycle provider allocates the session behind it. Hosted providers can
spend 30-80 seconds creating that session, so a lease on the default
one-minute inactivity window was expired, or nearly so, by the time its
client received it: the next command failed with "Lease is not active",
release reported nothing to release, and the paid provider session was
left running. The expiry sweep could also reap the lease while the
provider was still allocating.

Hold a lease work pass for the duration of the provider allocation, the
same protection admitted request work gets. The lease cannot expire
underneath the allocation, and ending the pass while the requester is
still waiting renews the lease for its own TTL from that moment. The
response now carries the renewed lease. A requester that hung up still
protects nothing, and its allocation is released as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 5 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread src/daemon/handlers/__tests__/lease.test.ts Outdated
… a provider

The no-provider allocation test ran on a frozen clock, so it could not tell
a preserved TTL from one renewed when allocation completed. On an advancing
clock it showed the handler did renew: the lease work pass was held, and
its release restarted the TTL, even when no provider allocated anything.

Hold the work pass only while a lease lifecycle provider allocates, and
run the no-provider test on a clock that advances on every read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@thymikee

thymikee commented Oct 3, 2026

Copy link
Copy Markdown
Member

This PR is ready. At e8f387d the lease TTL now starts when allocation completes, and the handler tests through handleLeaseCommands cover both the provider and no-provider routes. The one reported check passes, but it is a fork head, so the Integration and Coverage jobs may not have run; those need a maintainer to approve the CI run before merge. I did not run the tests or pnpm check:affected, and I did not run a live cloud allocation. I judged this daemon-only lease bookkeeping, so the handler tests are the evidence, and I read the pre-change route at c237027 to confirm the tests would catch the old behavior.

Not blocking, and you can take or leave these: remote-proxy.md:48 says "Either window starts when allocation completes", but on the plain proxy route the TTL is stamped when the record is created, so "when a hosted provider finishes allocating" would be more exact. The getLease(...) ?? lease fallback at lease.ts:104 only returns the stale record if the lease is released mid-allocation, which main already did. The retain-and-release pattern at lease.ts:80 now repeats the one in runAdmittedLeaseWork, which is fine for two callers. The lease.test.ts:117 comment could say only the invariant and drop the bug history.

The cubic-dev-ai P2 thread on the no-provider test no longer applies, because the test at head holds the pass without a provider; please resolve it: #3165 (comment)

@thymikee thymikee added the ready-for-human Valid work that needs human implementation, judgment, or maintainer merge label Oct 3, 2026
@thymikee
thymikee merged commit 642da8f into callstack:main Oct 3, 2026
15 of 16 checks passed
thymikee added a commit that referenced this pull request Oct 3, 2026
…token-attach-c41aff

* commit 'd396f3b509d9ed7cddaf170351ea6cf34e02ac16':
  fix(provider-webdriver): harden BrowserStack app references and endpoints (#3169)
  fix(android): back off a timed-out snapshot helper session and bound content re-captures (#3160)
  fix: centralize confirmed daemon retirement (#3126)
  fix: bind daemon registration writes to the acquired owner (#3125)
  fix: return confirmed daemon termination outcomes (#3124)
  refactor(move): share daemon registration and shutdown report modules (#3123)
  fix: preserve process lock exclusion across publication and reclaim (#3122)
  fix(cli): refuse a non-URL install-from-source source up front (#3166)
  fix(daemon): start a lease's TTL when its allocation completes (#3165)
  fix(android): fail doctor when adb is the Windows binary on a POSIX host (#3157)
  0.21.20
  0.21.19

# Conflicts:
#	src/daemon/server/daemon-runtime-metadata-ownership.test.ts
amankansal-lt added a commit to LambdaTest/agent-device that referenced this pull request Oct 5, 2026
Brings feat/testmu-provider-plugin (915e572) up to date with main,
which since landed the split-out callstack#3165 (lease TTL), callstack#3166 (install-from-source
URL refusal), callstack#3169 (BrowserStack app references, hub helpers, materialized
upload) and callstack#3173 (Limrun attached instances).

Main's reviewed versions win for the split-out code: appFileUploadForm with
{provider, service}, resolveHubAppReference with parseReference,
parseBrowserStackAppReference, isHttpUrl in install, the lease work pass only
when a provider allocates, and the upload file check in appFileUploadForm
rather than runtime-deployment. Re-applied on top only what the migration
needs: provider profile-field refusal (connect, WebDriver and Limrun
allocation, refused repeat allocation keeps its lease), optional auth on
fetchProviderVerificationJson for TestMu's public catalog, and the
provider-webdriver/plugin barrel without asRecord.

The TestMu plugin now uses asOptionalRecord from @agent-device/kernel/record,
parses lt:// references itself and uploads a public URL before delegating to
resolveHubAppReference, and passes {provider, service} to appFileUploadForm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-for-human Valid work that needs human implementation, judgment, or maintainer merge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants