Skip to content

perf(resolution): remove the JS indexing regressions since 1.6.2; the summary time covers the whole run (#2334) - #2362

Merged
colbymchenry merged 1 commit into
mainfrom
perf/2334-destructured-call-scan
Oct 5, 2026
Merged

colbymchenry merged 1 commit into
mainfrom
perf/2334-destructured-call-scan

Conversation

@colbymchenry

Copy link
Copy Markdown
Owner

Summary

Issue #2334 (thanks @bompus): indexing pretix took 2.2x as long since #2163, and the summary line hid it. Two JS/TS resolution checks that first shipped in 1.6.2 redid whole-file work for every reference; both now read a file or a function once. The graph is byte-identical to main. The time codegraph init / index / sync print now covers the whole run.

On pretix: process CPU 160 s → 89 s, wall 76 s → 44 s, peak RSS 2.47 → 1.93 GB (medians of 4 interleaved runs). On CPython's bundled minified d3 (the gap #2344 left): CPU 75 s → 32 s for the 190 files of Lib/profiling + Doc/_static.

Cause

  • matchDestructuredCallResult (fix(resolution): a name destructured from a composable reaches what it returns #2163) ran for every bare JS-family call in a file containing const { / let { / var {. Per call it re-tested that pattern on the whole source, then joined every line above the call, stripped comments, blanked strings, walked the text for braces and matched the binding regex over it. In pretix that is d3.v6.js (566 KB) and pdf.js (501 KB), each with thousands of calls; pdf.worker.js is over the 1 MB limit and never parsed, and files without such text (fabric.js, vue.js, moment) paid only the per-call re-test.
  • jsFunctionLocalScope (fix(js): a name the calling function binds is its local #2226) re-stripped the calling function's lines above every (function, name, line) it was asked about. In a minified script every function's text runs to the end of its one line, so each of those strips was the whole file.
  • Summary time: CodeGraph.indexAll recomputed nodesCreated / edgesCreated after resolution but kept the orchestrator's durationMs, which covers extraction only — the issue's 19.9 s run printed … edges in 1.7s. sync had the same split.

CPU profiles (--cpu-prof, summed over all threads; on a loaded box, so sample times run high), main → this PR:

pretix CPython Lib/profiling + Doc/_static
matchDestructuredCallResult 82.4 s → 1.7 s 35.4 s → 0.37 s
jsFunctionLocalScope 12.3 s → 5.2 s 23.3 s → 1.0 s
all comment stripping (stripCStyle) — 38.3 s → 0.38 s

(What remains in jsFunctionLocalScope on pretix is compiling its two regexes per lookup — linear, and unchanged here.)

Fix

Destructured calls — read the file once. Per file (FIFO of 16 per resolution context, as the other per-file memos): the blanked code, its line starts, each const { … } = f(…) binding (end, callee, keys, the block it sits in and where that block closes), and the spans the blanker read as regex literals. A call looks up its own name; only a call through a destructured name scans the code between that binding and itself for a shadowing declaration — and only after the callee is found to return the key (a const { x } = require(…), for one, needs no scan).

Why the results are identical: blanking the text above a call on its own gives the file's blanked code up to the call — the stripper and blanker look past a character only for a comment opener's second character (a call's name never starts with one) and for the / that closes a regex literal. On that common prefix the binding regex finds exactly the file's matches that end at or before the call (a match never reads past its final (, and one running past the call holds no {…( that could start another), and the brace stack at the call follows from where each block closes.

The one place the prefix reads differently: a call inside what the file's blanking reads as a regex literal — in minified code a division like (a)/b(c)/2 (742 of d3.min.js's 6,248 call sites). There the prefix reads the / as division, so it is the file's code up to the /, then the short tail blanked on its own; a binding statement can only run across that / through its { … } pattern or < … > type arguments, and where one could (the last brace before the / is a pattern's, or the last <>() is a <), the old per-call scan still runs, as it does for an out-of-range line/column or a call preceded by a comment-opening /.

Function-local names — strip a function once. jsFunctionLocalScope strips the function's lines once (LRU of 64 per context) and cuts at the reference's line end — the stripper looks at most one character ahead, and past a line's end that is a newline. Functions that start on the same line share the regex answer by (file, first line, reference line, name): every function of a minified script does (on d3.min.js: 992 regex pairs over 277 MB of text → 320 over 89 MB).

The new memos, and the existing JS_FN_LOCAL_MEMO (which was never cleared), now drop with the other per-file memos in clearNameMatcherMemos, so a long-running server's answer can't outlive an edit.

Summary time. CodeGraph.indexAll and CodeGraph.sync set durationMs to the whole locked run's wall time, as they already did for the totals. Every surface that shows it now reports the whole run: the codegraph init / index summary (and the re-index after opting gitignored child repos in), codegraph sync's line, the MCP server's Auto-synced N file(s) in Xms log, and the telemetry index event's duration bucket (TELEMETRY.md already describes it as the index's duration; buckets will shift up accordingly). On pretix, main printed 32,277 nodes, 80,722 edges in 18.5s for an 84 s run; this branch prints … in 39.0s for a 40 s run.

Verification

  • __tests__/js-resolution-work.test.ts (new, modeled on python-resolution-work.test.ts) counts strips instead of timing: a file with a destructured binding and 40 bare calls is stripped for the check 1× (41× on main); a 40-line function whose every line calls a same-named function is stripped 1× (40× on main). Each count fails when only its own half of the fix is reverted (41/1 and 1/40). The first test also resolves a destructured call written inside a division (the in-literal path).
  • __tests__/integration/full-pipeline.test.ts: durationMs of indexAll and of sync must cover a linking pass delayed 250 ms past what the orchestrator measured — fails on main (expected 218 to be greater than or equal to 468), passes here.
  • Targeted suites (destructured calls, Vue template calls, function-local bindings, function refs, store bindings, CommonJS calls, bare calls, value refs, object-literal methods, Python work, resolution, and main's new Dart / COBOL / Go tests): 18 files / 371 tests pass on the final branch.
  • Differential harness (not committed): old vs new matchDestructuredCallResult on every call position of mealie's 532 JS/TS/Vue files (10,330 positions, 306 resolved through a destructuring) and of CPython's profiling assets (9,170 positions, 742 inside misread regex literals), and on 16,000 random token-soup sources rich in destructuring, slashes, quotes, braces, generics and comments (1.33M positions; ~9.5K decisions resolved inside misread regex literals, ~9.3K through the fallback scan) — 0 differences; old vs new jsFunctionLocalScope on ~500K (function range, line, name) triples — 0 differences. Mutating the new code (scope checks, the crossing guard, the tail blanking, the literal lookup, the line cut, the memo key, …) makes it report differences; the two mutations it can't see are provably equivalent (a call position always starts with an identifier character).
  • npx tsc --noEmit clean; npm run build ok. Full suite on a shared Windows box at 100% load (8998697 + this change): 5,731 passed, 41 failed, 324 skipped — every failure a timeout (36 at vitest's 5 s; the rest at 10–60 s or waiting for a daemon to attach), a status-freshness race or a Windows EBUSY on temp-dir cleanup, in 23 sync / git / MCP-daemon / watcher / explore files. Rerun one file at a time on the final branch, 21 of the 23 pass; sync.test.ts (four 5 s timeouts) and mcp-daemon.test.ts (one 20 s timeout, a different test than in the full run) then pass alone too (58/58, 21/21 + 1 skipped).

Validation (before = origin/main 8998697, after = 8998697 + this change)

Nodes, edges and unresolved_refs (and files) dumped with node:sqlite keyed by file|kind|qualified_name|start_line — node ids replaced by those keys in edges and refs, no row ids — sorted: byte-identical on all 24 runs (8 per repo, both arms): pretix 32,277 nodes / 80,722 edges / 158,546 refs; mealie 19,436 / 42,892 / 40,155 (with #2357's new destructuring and Vue template refs); CPython assets 12,992 / 32,646 / 31,874. Main moved on while this was measured (#2358 COBOL, #2360 Dart, #2361 Go — no JS path); the branch is one commit on deac771, and one more before/after round against 9f69e47 and against deac771 is byte-identical too. (Against the earlier main 26e8488, 18 runs were byte-identical as well.)

Timing — 4 interleaved rounds, node --liftoff-only dist/bin/codegraph.js init -y run in-process so process.cpuUsage() covers every worker; a shared 16-thread Windows box that sat at 100% load, so CPU is the steady number and wall is noisy:

repo wall median (fastest) process CPU median peak RSS median
pretix (1,461 files) main 75.9 s (64.3) 160.2 s 2,465 MB
this PR 43.5 s (39.0) 88.7 s 1,930 MB
CPython Lib/profiling + Doc/_static (190 files) main 37.6 s (29.3) 75.1 s 1,226 MB
this PR 20.4 s (14.6) 32.1 s 809 MB
mealie (1,229 files; Nuxt + Python, no bundles) main 26.7 s (22.8) 39.2 s 768 MB
this PR 42.3 s (20.5) 38.6 s 758 MB

mealie has no bundled library, so it is the control: its CPU is flat; its wall times ranged 22.8–55.7 s on main and 20.5–56.2 s here.

Related

🤖 Generated with Claude Code

… summary time covers the whole run (#2334)

Indexing pretix took 2.2x as long since #2163 (issue #2334), and the
bundled minified d3 was what #2344 left of CPython's 1.6.2 slowdown. Two
JS/TS checks that first shipped in 1.6.2 redid whole-file work for every
reference:

- matchDestructuredCallResult (#2163) ran for every bare call in a file
  holding `const {` / `let {` / `var {`: per call it re-tested that
  pattern on the source, then stripped, blanked and brace-walked every
  line above the call and matched the binding regex over it.
- jsFunctionLocalScope (#2226) re-stripped the calling function's lines
  above each reference. In a minified script every function's text runs
  to the end of its one line, so each strip was the whole file.

CPU profiles, main -> this change: matchDestructuredCallResult 82.4 s
-> 1.7 s and jsFunctionLocalScope 12.3 s -> 5.2 s on pretix; 35.4 s ->
0.37 s and 23.3 s -> 1.0 s on CPython's Lib/profiling + Doc/_static.

Now:
- A file's `const { … } = f(…)` bindings are read once per resolution:
  the blanked code, each binding's end, callee and keys, the block it
  sits in and where that block closes, and the spans the blanker read as
  regex literals. A call looks up its own name; only a call through a
  destructured name scans the code between that binding and itself for
  shadowing, and only once the callee is found to return the key (a
  `require(…)` destructuring, for one, needs no scan).
- Blanking the text above a call on its own gives the file's blanked code
  up to the call: the stripper and blanker look past a character only for
  a comment opener's second character and for a regex literal's closing
  `/`. So the prefix's bindings are the file's that end at or before the
  call (a match never reads past its `(`, and one running past the call
  holds no `{`…`(` to start another), and its brace stack follows from
  where each block closes. For a call inside a span the file reads as a
  regex literal (a division, in minified code) the prefix is the code up
  to the `/` plus the short tail blanked on its own; where a binding
  statement could run across that `/` (inside its `{ … }` or `< … >`), and
  for out-of-range positions, the old per-call scan still runs.
- jsFunctionLocalScope strips a function's lines once and cuts at the
  reference's line end; functions that start on the same line share the
  regex answer by (file, first line, reference line, name).
- The new memos, and JS_FN_LOCAL_MEMO (never cleared before), drop with
  the other per-file memos in clearNameMatcherMemos.

The summary line printed orchestrator.indexAll's durationMs, which covers
extraction only, beside node and edge totals recomputed after resolution:
the issue's 19.9 s run printed "in 1.7s". CodeGraph.indexAll and sync now
report the whole run's wall time, which `codegraph init`, `index` and
`sync`, the MCP auto-sync log and the telemetry duration bucket read.

Verification: nodes, edges and unresolved_refs dumps keyed by
file|kind|qualified_name|start_line are byte-identical to origin/main
8998697 (with #2357's new destructuring and Vue template refs) on all
24 runs: pretix, mealie and CPython's Lib/profiling + Doc/_static, 8
each; and to deac771 (#2358, #2360, #2361) on one more round of each. A differential harness ran old and new matchDestructuredCallResult
on every call position of mealie (10,330) and the CPython assets (9,170,
742 inside misread regex literals) and on 16,000 random token-soup
sources (1.33M positions), and old and new jsFunctionLocalScope on ~500K
(function, line, name) triples: no difference; it reports one for each
deliberate mutation of the new code that is not provably equivalent.

Medians of 4 interleaved runs on a loaded 16-thread Windows box, process
CPU / wall / peak RSS: pretix 160.2 s / 75.9 s / 2.47 GB -> 88.7 s /
43.5 s / 1.93 GB; CPython assets 75.1 s / 37.6 s / 1.23 GB -> 32.1 s /
20.4 s / 0.81 GB; mealie (no bundles, the control) 39.2 s -> 38.6 s CPU.

__tests__/js-resolution-work.test.ts counts strips: a file with a
destructured binding and 40 bare calls is stripped for the check once
(41 on main), a 40-line function once (40 on main); each count fails when
only its own half of the fix is reverted. full-pipeline.test.ts requires
indexAll's and sync's durationMs to cover a linking pass delayed 250 ms
past the orchestrator's time (on main: 218 < 468).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant