fix(scheduler): keep idle fill off a run that is on the retry ramp (AGT-4675) - #815
Merged
Merged
Conversation
…GT-4675) The ramp from AGT-4673 backs a failing run off, but idle fill spends one budget per free slot, in priority order, on RETRY_AT rows whose backoff has not elapsed, so the highest-priority failing run was lifted ahead of 36 READY tasks as soon as a slot freed: AX-1844 failed attempt 8, was backed off at 20:33:01 and was READY and claimed by 20:33:19. Leave a RETRY_AT run alone once its attempt number reaches the ramp, and share the threshold with retryAtFor through RETRY_RAMP_FROM_ATTEMPT.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
The retry backoff from AGT-4673 is bypassed by idle fill.
filterAlreadyProcessedgives each heartbeat anidleFillBudgetequal to the free slots and spends it, in priority order, onRETRY_ATrows whose backoff has not elapsed (markReady). A freed slot goes to the highest-priority failing run, ahead of the READY tasks that are waiting.Evidence (ledger events, 2026-10-03)
failed:20:33:01 PUBLISHING > RETRY_AT(backoff applied),20:33:08 missing_worktree_reconciled clear,20:33:19 RETRY_AT > READY,20:33:19 claimed; attempt 9 was running 7 minutes later.RETRY_ATwith 336 minutes left.superseded).Change
durableRunCoordinator.ts:RETRY_RAMP_FROM_ATTEMPT = 4, used byretryAtFor(same delays as before).autonomousRunner.ts:idleLiftablealso requiresdurableRun.attemptNo < RETRY_RAMP_FROM_ATTEMPT, so a run that has reached the ramp is left to itsretry_at. Operator answers and reopens still readmit it through their own paths.Tests
retry_atstaysRETRY_ATwhen a slot is free; a run at attempt 3 is still lifted to READY (control). Removing the guard fails the first.vitestrunner, coordinator and ledger suites: 20 files, 310 tests pass.Risk
When nothing else is runnable, a slot stays empty instead of re-running a run that has failed three times. READY work is not scarce today.
After deploy
No
RETRY_AT > READYfor a run at attempt 4 or more before itsretry_at; distinct issues per hour and the top-10 share move as described in AGT-4673.