Skip to content

perf(search): trigram index for substring name search - #2347

Draft
gtarpenning wants to merge 1 commit into
colbymchenry:mainfrom
gtarpenning:griffin/trigram-substring-prefilter
Draft

gtarpenning wants to merge 1 commit into
colbymchenry:mainfrom
gtarpenning:griffin/trigram-substring-prefilter

Conversation

@gtarpenning

Copy link
Copy Markdown

Summary

name LIKE '%x%' cannot use a B-tree, so every substring lookup scans all nodes. The prompt hook is a fresh process per prompt and runs several of these each time, so the scan cost is paid on every prompt. This adds a trigram FTS5 index (nodes_tri) and uses it to narrow the candidates.

Independent of the cross-process memo PR; either can merge first. #2115 and #2260 touch nearby code in queries.ts.

What changed

  • Migration 12 creates nodes_tri(name, qualified_name), an external-content FTS5 table over nodes with tokenize='trigram', plus insert/delete/update triggers beside the nodes_fts ones. Fresh databases get it from schema.sql.
  • findNodesByNameSubstring and searchNodesLike add rowid IN (SELECT rowid FROM nodes_tri ...) when the token is 3+ characters and has no _ or %. The original LIKE stays in the query, so the index only narrows candidates and result sets are unchanged. Shorter or wildcard tokens take the old path.
  • Bulk loads drop and rebuild the trigram index with the FTS one. The trigger list is now read from schema.sql instead of a hard-coded name list, so the heal after an interrupted bulk load covers both.
  • Without FTS5 ([Bug] v1.5.0 requires FTS5 which is missing in official Node.js builds #1532), migration 12 is skipped, the table is absent, and queries use LIKE. A later bulk load on an FTS5 build creates the table.

Measured

Warm query time on a large TS/Go/Python monorepo index (about 1.5 GB, 350k nodes, 1.2M edges):

Query Before After
Substring lookup (name LIKE) 40-80 ms 1-5 ms
Name or qualified-name lookup 55-85 ms 2-11 ms

Result sets were identical for every token tried (compared by rowid). Tokens under 3 characters stay on the scan path, where the index is no faster.

  • Size: nodes_tri adds about 2.5-3.7% to the database (55 MB on 1.5 GB).
  • Migration: about 1.5 s on that index; about 9 s per 1M synthetic nodes.
  • Write cost: a 300-file sync took 64.9 s with the triggers and 59-63 s without, within run-to-run noise because resolution dominates.

Compatibility

  • Older binaries open a v12 database fine: the triggers are SQLite-level, so their writes keep nodes_tri current, and they never query it.
  • A database migrated on a build without FTS5 has no nodes_tri until the next bulk load on a build that has it; substring search stays correct throughout.

Testing

  • npx tsc --noEmit
  • npx vitest run __tests__/trigram-substring.test.ts __tests__/fts5-fallback.test.ts: results match a LIKE scan; a node removed from nodes_tri alone disappears from long-token results, so the index is consulted (the test fails with the prefilter reverted); a v11 database without FTS5 opens and falls back to LIKE.
  • npx vitest run: 5748 passed, 20 failed. The 20 are in daemon-older-version, daemon-registry and index-command, and fail the same way on unmodified main.

🤖 Generated with Claude Code

Substring lookups (`name LIKE '%x%'`) scanned every node. A trigram FTS5 table
(migration 12) now prefilters them when the token is 3+ chars with no LIKE
wildcard; the original LIKE still decides the result. The table is skipped on
Node builds without FTS5 and healed by the next bulk load.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant