Skip to content

fix(scheduler): back off a task that keeps failing instead of retrying it every 30 minutes (AGT-4673) - #813

Merged
unohee merged 1 commit into
mainfrom
fix/agt-4673-retry-backoff
Oct 3, 2026
Merged

unohee merged 1 commit into
mainfrom
fix/agt-4673-retry-backoff

Conversation

@unohee

@unohee unohee commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

TL;DR

After a failed or rejected attempt retryAtFor returned now + 30 min whatever the attempt number (15 min for infra_error). Only superseded backed off. A task that fails every time was re-claimed as soon as its attempt ended and kept a slot indefinitely, while the rest of the queue waited.

Evidence (ledger, 2026-10-03)

  • Last 4 h: 78 attempts over 23 issues; the top 11 issues took 58 of them (74%), 4 to 8 attempts each.
  • The same eight issues held all eight slots in two consecutive cycles (17:27 and 18:42 starts), at attempt 6 to 8.
  • Last 2 h: 12 distinct issues got a slot while 28 READY and 54 RETRY_AT rows waited.
  • Last cycle of those eight: 1 approved PR (with a tracked plan.md overwrite, fixed separately), 5 unfinished drafts, 1 parked.

Change

  • retryAtFor: unchanged for attempts 1 to 3; from attempt 4 the delay doubles per attempt. failed/rejected: 30 min base, cap 6 h (30, 30, 30, 60, 120, 240, 360). infra_error: 15 min base, cap 2 h. rate_limited, deferred, superseded unchanged.
  • The failure call site now passes claim.attemptNo; without it the ramp never applied.

Tests

New file durableRunCoordinator.retryBackoff.test.ts (the existing test file is at the 1500-line cap):

  • Ramp values for failed, rejected and infra_error, the unchanged first three attempts, and the rate-limit case.
  • A coordinator-level test drives one run through five failed attempts with fake timers and expects waits of 30, 30, 30, 60, 120 minutes. Removing the call-site argument fails it; flattening retryAtFor fails the ramp tests.
  • vitest coordinator, ledger, runner and scheduler suites: 15 files, 255 tests pass before the test move; the two coordinator files pass together (65 tests) after it.

Risk

A task that converges slowly, one attempt at a time on its preserved branch, now waits longer between attempts. Each unfinished attempt already publishes a draft PR, so the work stays visible.

After deploy

Distinct issues per hour and the share of attempts taken by the top 10 issues should move: today 11 to 12 issues per hour and 74%.

…g it every 30 minutes (AGT-4673)

retryAtFor returned now + 30 min for a failed or rejected attempt whatever the
attempt number (15 min for infra_error); only superseded backed off. A task that
fails every time was re-claimed as soon as its attempt ended and held a slot for
good: 11 issues took 74% of 78 attempts in four hours while 80 others waited.
Keep today's delay for the first three attempts, then double it per attempt
(failed/rejected cap 6 h, infra_error cap 2 h), and pass the run's own attempt
count from the failure call site.
@unohee
unohee merged commit 36569dd into main Oct 3, 2026
7 checks passed
@unohee
unohee deleted the fix/agt-4673-retry-backoff branch October 3, 2026 10:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant