Skip to content

fix: recover interrupted secondary indexes off-thread - #1888

Open
cbeaulieu-gt wants to merge 1 commit into
colbymchenry:mainfrom
cbeaulieu-gt:codex/upstream-secondary-index-recovery
Open

cbeaulieu-gt wants to merge 1 commit into
colbymchenry:mainfrom
cbeaulieu-gt:codex/upstream-secondary-index-recovery

Conversation

@cbeaulieu-gt

Copy link
Copy Markdown
Contributor

Summary

  • rebuild secondary indexes left missing by an interrupted bulk load on a worker-owned SQLite connection
  • make async project opens, MCP retry opens, project-path cache opens, and replacement reopens await recovery before exposing the database
  • coalesce concurrent opens and guard shutdown races so an in-flight recovery cannot resurrect a closed connection
  • add full index-restoration, real watchdog, concurrency, and reopen-lifecycle regression coverage

Why

Crash recovery previously executed every missing CREATE INDEX synchronously inside DatabaseConnection.open(). On a large graph, one index build can block the MCP event loop long enough for the liveness watchdog to kill the process before SQLite commits any visible file progress, so the next restart repeats the same work and can be killed again.

The synchronous API remains available for compatibility. Async callers now use the off-thread recovery path and receive the connection only after every required index has been restored.

Verification

  • npm run build
  • npx vitest run __tests__/foundation.test.ts
  • npx vitest run __tests__/liveness-watchdog.test.ts
  • npx vitest run __tests__/concurrent-locking.test.ts
  • npx vitest run __tests__/mcp-unindexed.test.ts
  • npx vitest run __tests__/db-reopen-on-replace.test.ts

The watchdog regression creates a 600,000-node interrupted database under a scaled timeout. The former synchronous open is killed by the watchdog; the async recovery completes and restores the full index set.

Fixes #1887

🤖 Generated by Codex on behalf of @cbeaulieu-gt

@cbeaulieu-gt

Copy link
Copy Markdown
Contributor Author

I noticed that maintainer draft #1919 carries the #1887 recovery fix as commit 5c41bdf, with my authorship preserved. Since this standalone PR now conflicts with main, I have left it unchanged for the moment to avoid duplicating integration work.

Please let me know whether you would prefer #1888 rebased as the focused integration path, or whether #1919 is intended to carry the fix forward. I’m happy to update this branch if the standalone PR is still useful.

🤖 Generated by Codex on behalf of @cbeaulieu-gt

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Interrupted bulk-load recovery can be repeatedly watchdog-killed while rebuilding indexes

1 participant