Summary
On Claude Code, a background task completing inside an engineer turn kills the implementation loop. The harness delivers the completion notice as a UserPromptSubmit hook event whose prompt body is a synthetic <task-notification> block. markInterruptedTurnIfNeeded() cannot tell it apart from real human input, so it marks the open turn interrupted, sets status: "interrupted" and next_role: "idle", and every later Stop hook bails out with loop-not-running.
The loop then stops silently. In my case it ran for 12 minutes and sat dead for ~11h48m overnight, and nothing in the human-facing epic artifacts said so — state-of-epic.md still read "Active task: Phase 1 task 1".
Environment: Claude Code 2.1.273, platform claude-code, Node v24.18.0, project-local hooks installed by install-hooks.mjs, epic in implementation mode with one bound driver session.
Observed sequence
All times local (UTC+7); the JSONL below is from .epic-loop/epics/<slug>/.runtime/progress-log.jsonl.
| time |
event |
| 01:42:00 |
techlead sets next_role: engineer; turn-start for iteration 3 |
| ~01:44 |
engineer starts the dev stack with a background Bash command (pnpm dev, run_in_background: true) |
| ~01:47 |
engineer kills the dev processes as part of the task; the background command exits |
| 01:47:21 |
harness fires UserPromptSubmit carrying <task-notification>…</task-notification> for that background task → turn-interrupted, reason: user-prompt-interrupted-open-turn |
| 01:50:09 |
second background task exits → second synthetic UserPromptSubmit |
| 01:51:20 |
engineer turn ends, Stop fires → skip, reason: loop-not-running, status: interrupted. Loop dead. |
| 13:39 |
human asks why nothing happened overnight |
{"action":"turn-interrupted","duration_ms":321000,"ended_at":"2026-09-15T18:47:21+00:00","iteration":3,
"reason":"user-prompt-interrupted-open-turn","role":"engineer","slug":"vendure-rebuild",
"started_at":"2026-09-15T18:42:00+00:00"}
{"action":"skip","reason":"loop-not-running","slug":"vendure-rebuild","status":"interrupted",
"timestamp":"2026-09-15T18:51:20+00:00"}
The captured hook payload is kept in .epic-loop/.runtime/hook-events/<session>/20260915T184721Z-userpromptsubmit-no-turn.json; its prompt field starts with:
<task-notification>
<task-id>bqwv481rp</task-id>
<tool-use-id>toolu_…</tool-use-id>
<output-file>/tmp/claude-…/tasks/bqwv481rp.output</output-file>
<status>failed</status>
…
No human typed anything between 01:42 and 13:39.
Root cause
scripts/lib/hooks.mjs:396 calls markInterruptedTurnIfNeeded() for every hook event, and scripts/lib/loop.mjs:317 treats any UserPromptSubmit on the driver session as human interruption:
export function markInterruptedTurnIfNeeded(projectRoot, payload, binding) {
if (payload.hook_event_name !== "UserPromptSubmit") {
return false;
}
…
if (!hasOpenTurn(loop)) {
return false;
}
recordTurnInterrupted(projectRoot, slug, runtime, loop, {
reason: "user-prompt-interrupted-open-turn", …
});
recordTurnInterrupted() (scripts/lib/loop.mjs:510) writes next_role: "idle", status: "interrupted". The Stop continuation gate (scripts/lib/loop.mjs:195) refuses anything whose status is not running, so the loop can never recover on its own.
The assumption "UserPromptSubmit == a human typed something" does not hold on Claude Code. At least these arrive on the same channel with no human involved:
<task-notification> blocks, when a run_in_background Bash command exits — one per background task, at an arbitrary later time
<system-reminder> style injections
So any engineer turn that uses run_in_background (starting a dev server, a long build, a watcher — exactly what a verification-heavy brief asks for) arms a delayed kill switch for the whole loop. The engineer role reference does not warn against background processes, and SKILL.md describes the interrupt path as a user action ("If a bound implementation session receives a new UserPromptSubmit while a turn is still open, treat the open turn as interrupted").
Two secondary problems make it worse:
- Silent death. The stop is recorded only in
.runtime/ logs, which techlead and manager are explicitly forbidden to read in normal flow. state-of-epic.md, tracker.md and implementation-log.md still describe an active task.
- No documented recovery.
SKILL.md says the loop "must not auto-continue until a new implementation start/resume explicitly rebinds or restarts the loop", but there is no resume/restart script — bind-session.mjs --current --mode implementation on an already-bound session does not clear status: "interrupted".
Proposed fix
1. Classify the prompt before treating it as an interrupt.
In markInterruptedTurnIfNeeded(), ignore synthetic prompts. A prompt counts as human input only if, after stripping the wrapper blocks the harness injects, something remains:
const SYNTHETIC_PROMPT_PATTERNS = [
/^\s*<task-notification>[\s\S]*<\/task-notification>\s*$/,
/^\s*<system-reminder>[\s\S]*<\/system-reminder>\s*$/,
/^\s*\[SYSTEM NOTIFICATION - NOT USER INPUT\]/,
];
function isSyntheticPrompt(prompt) {
if (typeof prompt !== "string" || !prompt.trim()) {
return false;
}
const stripped = prompt
.replace(/<system-reminder>[\s\S]*?<\/system-reminder>/g, "")
.replace(/<task-notification>[\s\S]*?<\/task-notification>/g, "")
.trim();
return stripped === "" || SYNTHETIC_PROMPT_PATTERNS.some((re) => re.test(prompt));
}
and bail out early:
if (isSyntheticPrompt(payload.prompt)) {
appendLoopLog(projectRoot, { action: "ignored-synthetic-prompt", reason: "task-notification", … });
return false;
}
The same guard belongs in buildModeReminder() — there is no point injecting the mode marker into a background-task notification.
Conservative variant if you would rather not pattern-match harness internals: keep the interrupt, but only for prompts that are non-empty after stripping known wrapper tags, and treat an unrecognised empty-after-strip prompt as synthetic. Either way the decision should be logged so the behaviour is auditable.
2. Make the stop visible.
When the loop moves to interrupted (or any skip with manual_continue_required), write one line into the human-facing artifacts — state-of-epic.md under Blockers, or an implementation-log.md entry:
## 2026-09-16 - Loop interrupted
Loop stopped at 01:47 (reason: user-prompt-interrupted-open-turn) during Phase 1 task 1.
Resume with: node <skill-dir>/scripts/resume-loop.mjs --slug <slug> --current
Right now the only trace is in files the roles are told not to read, so the next session orients from artifacts that claim work is in progress.
3. Ship an explicit resume path.
scripts/resume-loop.mjs --slug <slug> --current that:
- verifies the caller is the driver session (or rebinds it),
- clears
status: "interrupted",
- sets
next_role back to techlead (never straight to engineer, so the next turn re-verifies state),
- appends a
loop-resumed event.
Without it, recovering from an interrupt means hand-editing .runtime/runtime-state.json, which contradicts "do not hand-edit runtime state".
4. Document the background-process hazard.
Add to references/implementation-engineer-role.md, and ideally to the engineer brief template: long-running processes must stay inside a single Bash invocation (cmd & … kill %1, timeout, until polling) — never run_in_background, because each background task completion is delivered as a UserPromptSubmit and, until fix 1 lands, ends the loop. This is worth saying even after the fix, since the notification still costs a turn boundary.
Impact
Any unattended overnight run that touches a dev server, a watcher, or a long build is a coin flip today: the loop dies at the first background-task completion, with no notification, and stays dead until a human looks. Fix 1 alone removes the failure; fixes 2–4 remove the "silent for 12 hours" part.
Summary
On Claude Code, a background task completing inside an engineer turn kills the implementation loop. The harness delivers the completion notice as a
UserPromptSubmithook event whose prompt body is a synthetic<task-notification>block.markInterruptedTurnIfNeeded()cannot tell it apart from real human input, so it marks the open turn interrupted, setsstatus: "interrupted"andnext_role: "idle", and every laterStophook bails out withloop-not-running.The loop then stops silently. In my case it ran for 12 minutes and sat dead for ~11h48m overnight, and nothing in the human-facing epic artifacts said so —
state-of-epic.mdstill read "Active task: Phase 1 task 1".Environment: Claude Code 2.1.273, platform
claude-code, Node v24.18.0, project-local hooks installed byinstall-hooks.mjs, epic in implementation mode with one bound driver session.Observed sequence
All times local (UTC+7); the JSONL below is from
.epic-loop/epics/<slug>/.runtime/progress-log.jsonl.next_role: engineer;turn-startfor iteration 3pnpm dev,run_in_background: true)UserPromptSubmitcarrying<task-notification>…</task-notification>for that background task →turn-interrupted,reason: user-prompt-interrupted-open-turnUserPromptSubmitStopfires →skip,reason: loop-not-running,status: interrupted. Loop dead.{"action":"turn-interrupted","duration_ms":321000,"ended_at":"2026-09-15T18:47:21+00:00","iteration":3, "reason":"user-prompt-interrupted-open-turn","role":"engineer","slug":"vendure-rebuild", "started_at":"2026-09-15T18:42:00+00:00"} {"action":"skip","reason":"loop-not-running","slug":"vendure-rebuild","status":"interrupted", "timestamp":"2026-09-15T18:51:20+00:00"}The captured hook payload is kept in
.epic-loop/.runtime/hook-events/<session>/20260915T184721Z-userpromptsubmit-no-turn.json; itspromptfield starts with:No human typed anything between 01:42 and 13:39.
Root cause
scripts/lib/hooks.mjs:396callsmarkInterruptedTurnIfNeeded()for every hook event, andscripts/lib/loop.mjs:317treats anyUserPromptSubmiton the driver session as human interruption:recordTurnInterrupted()(scripts/lib/loop.mjs:510) writesnext_role: "idle",status: "interrupted". TheStopcontinuation gate (scripts/lib/loop.mjs:195) refuses anything whose status is notrunning, so the loop can never recover on its own.The assumption "UserPromptSubmit == a human typed something" does not hold on Claude Code. At least these arrive on the same channel with no human involved:
<task-notification>blocks, when arun_in_backgroundBash command exits — one per background task, at an arbitrary later time<system-reminder>style injectionsSo any engineer turn that uses
run_in_background(starting a dev server, a long build, a watcher — exactly what a verification-heavy brief asks for) arms a delayed kill switch for the whole loop. The engineer role reference does not warn against background processes, andSKILL.mddescribes the interrupt path as a user action ("If a bound implementation session receives a newUserPromptSubmitwhile a turn is still open, treat the open turn as interrupted").Two secondary problems make it worse:
.runtime/logs, which techlead and manager are explicitly forbidden to read in normal flow.state-of-epic.md,tracker.mdandimplementation-log.mdstill describe an active task.SKILL.mdsays the loop "must not auto-continue until a new implementation start/resume explicitly rebinds or restarts the loop", but there is noresume/restartscript —bind-session.mjs --current --mode implementationon an already-bound session does not clearstatus: "interrupted".Proposed fix
1. Classify the prompt before treating it as an interrupt.
In
markInterruptedTurnIfNeeded(), ignore synthetic prompts. A prompt counts as human input only if, after stripping the wrapper blocks the harness injects, something remains:and bail out early:
The same guard belongs in
buildModeReminder()— there is no point injecting the mode marker into a background-task notification.Conservative variant if you would rather not pattern-match harness internals: keep the interrupt, but only for prompts that are non-empty after stripping known wrapper tags, and treat an unrecognised empty-after-strip prompt as synthetic. Either way the decision should be logged so the behaviour is auditable.
2. Make the stop visible.
When the loop moves to
interrupted(or anyskipwithmanual_continue_required), write one line into the human-facing artifacts —state-of-epic.mdunder Blockers, or animplementation-log.mdentry:Right now the only trace is in files the roles are told not to read, so the next session orients from artifacts that claim work is in progress.
3. Ship an explicit resume path.
scripts/resume-loop.mjs --slug <slug> --currentthat:status: "interrupted",next_roleback totechlead(never straight toengineer, so the next turn re-verifies state),loop-resumedevent.Without it, recovering from an interrupt means hand-editing
.runtime/runtime-state.json, which contradicts "do not hand-edit runtime state".4. Document the background-process hazard.
Add to
references/implementation-engineer-role.md, and ideally to the engineer brief template: long-running processes must stay inside a single Bash invocation (cmd & … kill %1,timeout,untilpolling) — neverrun_in_background, because each background task completion is delivered as aUserPromptSubmitand, until fix 1 lands, ends the loop. This is worth saying even after the fix, since the notification still costs a turn boundary.Impact
Any unattended overnight run that touches a dev server, a watcher, or a long build is a coin flip today: the loop dies at the first background-task completion, with no notification, and stays dead until a human looks. Fix 1 alone removes the failure; fixes 2–4 remove the "silent for 12 hours" part.