fix(scheduler): back off a task that keeps failing instead of retrying it every 30 minutes (AGT-4673) - #813
Merged
Merged
Conversation
…g it every 30 minutes (AGT-4673) retryAtFor returned now + 30 min for a failed or rejected attempt whatever the attempt number (15 min for infra_error); only superseded backed off. A task that fails every time was re-claimed as soon as its attempt ended and held a slot for good: 11 issues took 74% of 78 attempts in four hours while 80 others waited. Keep today's delay for the first three attempts, then double it per attempt (failed/rejected cap 6 h, infra_error cap 2 h), and pass the run's own attempt count from the failure call site.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
After a failed or rejected attempt
retryAtForreturnednow + 30 minwhatever the attempt number (15 min forinfra_error). Onlysupersededbacked off. A task that fails every time was re-claimed as soon as its attempt ended and kept a slot indefinitely, while the rest of the queue waited.Evidence (ledger, 2026-10-03)
plan.mdoverwrite, fixed separately), 5 unfinished drafts, 1 parked.Change
retryAtFor: unchanged for attempts 1 to 3; from attempt 4 the delay doubles per attempt.failed/rejected: 30 min base, cap 6 h (30, 30, 30, 60, 120, 240, 360).infra_error: 15 min base, cap 2 h.rate_limited,deferred,supersededunchanged.claim.attemptNo; without it the ramp never applied.Tests
New file
durableRunCoordinator.retryBackoff.test.ts(the existing test file is at the 1500-line cap):retryAtForfails the ramp tests.vitestcoordinator, ledger, runner and scheduler suites: 15 files, 255 tests pass before the test move; the two coordinator files pass together (65 tests) after it.Risk
A task that converges slowly, one attempt at a time on its preserved branch, now waits longer between attempts. Each unfinished attempt already publishes a draft PR, so the work stays visible.
After deploy
Distinct issues per hour and the share of attempts taken by the top 10 issues should move: today 11 to 12 issues per hour and 74%.