fix(verify): bound concurrent deterministic verifications to two (AGT-4676) - #816
Merged
Merged
Conversation
…-4676) Every attempt that reaches verification runs the repository's full suite in its own sandbox and nothing limited how many did so at once. Eight cgf-portal suites on a machine other sessions also load reached 3% in the 300 s pytest cap (the suite takes 272 s when the machine is calm), so 7 of 7 verdicts were lost and the attempts fell back to the LLM tester, about 16 minutes each. Add a process-wide FIFO semaphore (default 2, OPENSWARM_VERIFY_MAX_CONCURRENT, sized once at startup) around runVerify, so both of its callers are bounded and the pytest cap starts after the slot is held. A waiter gives up after 20 minutes (OPENSWARM_VERIFY_SLOT_WAIT_MS) and runVerify throws, which the tester already turns into the LLM fallback; a waiter that gave up never receives a slot.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
Every attempt that reaches verification runs the target repository's full suite in its own sandbox, and nothing limited how many do so at once. With eight slots that is up to eight full cgf-portal suites together, on a machine other sessions also load. Under that load the pytest run hits its 300 s cap at about 3% progress, the deterministic verdict is lost, and the attempt falls back to the LLM tester (about 16 minutes).
Evidence (daemon log since the 21:01 restart, 2026-10-03)
Deterministic verify unavailable; falling back to LLM tester ... timeout after 300000ms, each at 3% of the suite.1 failed, 5946 passed, 812 skipped in 272.08s): about 100 times slower under load.Change
verifySlots.ts: a FIFO counting semaphore, sized once at startup (OPENSWARM_VERIFY_MAX_CONCURRENT, default 2), with a bounded wait (OPENSWARM_VERIFY_SLOT_WAIT_MS, default 20 min). A waiter that gives up leaves the queue, so a slot is never handed to someone who is gone.runVerifytakes a slot before any sandbox exists, so the 300 s cap counts only the run itself. Its two callers (deterministicTester,qualityHarness) do not nest. On a wait timeoutrunVerifythrows, whichrunTesterWithVerificationalready turns into the LLM fallback: the worst case is today's behaviour.Tests
onWaitreport,runVerifygoes throughverifySlots, default limit 2.runVerify, keeping a gave-up waiter in the queue, and ignoring the limit each fail the matching tests.vitestverify, deterministic tester and pipeline verify suites: 217 of 218 pass. The one failure,runs npm package scripts against head-installed node_modules at base, fails the same way on a checkout without this change (load average 215 while running); it executes real npm.Not in this PR
Selecting tests by changed files, sharing a baseline run across attempts, and a machine-wide lock for other sessions' full runs.
After deploy
At most two verification suites from the daemon at once; the log shows
[Verify] waiting for a verification slot (...)and fewer 300 s timeouts.