Repository navigation
perf(resolution): remove the JS indexing regressions since 1.6.2; the summary time covers the whole run (#2334) - #2362
Merged
Conversation
… summary time covers the whole run (#2334) Indexing pretix took 2.2x as long since #2163 (issue #2334), and the bundled minified d3 was what #2344 left of CPython's 1.6.2 slowdown. Two JS/TS checks that first shipped in 1.6.2 redid whole-file work for every reference: - matchDestructuredCallResult (#2163) ran for every bare call in a file holding `const {` / `let {` / `var {`: per call it re-tested that pattern on the source, then stripped, blanked and brace-walked every line above the call and matched the binding regex over it. - jsFunctionLocalScope (#2226) re-stripped the calling function's lines above each reference. In a minified script every function's text runs to the end of its one line, so each strip was the whole file. CPU profiles, main -> this change: matchDestructuredCallResult 82.4 s -> 1.7 s and jsFunctionLocalScope 12.3 s -> 5.2 s on pretix; 35.4 s -> 0.37 s and 23.3 s -> 1.0 s on CPython's Lib/profiling + Doc/_static. Now: - A file's `const { … } = f(…)` bindings are read once per resolution: the blanked code, each binding's end, callee and keys, the block it sits in and where that block closes, and the spans the blanker read as regex literals. A call looks up its own name; only a call through a destructured name scans the code between that binding and itself for shadowing, and only once the callee is found to return the key (a `require(…)` destructuring, for one, needs no scan). - Blanking the text above a call on its own gives the file's blanked code up to the call: the stripper and blanker look past a character only for a comment opener's second character and for a regex literal's closing `/`. So the prefix's bindings are the file's that end at or before the call (a match never reads past its `(`, and one running past the call holds no `{`…`(` to start another), and its brace stack follows from where each block closes. For a call inside a span the file reads as a regex literal (a division, in minified code) the prefix is the code up to the `/` plus the short tail blanked on its own; where a binding statement could run across that `/` (inside its `{ … }` or `< … >`), and for out-of-range positions, the old per-call scan still runs. - jsFunctionLocalScope strips a function's lines once and cuts at the reference's line end; functions that start on the same line share the regex answer by (file, first line, reference line, name). - The new memos, and JS_FN_LOCAL_MEMO (never cleared before), drop with the other per-file memos in clearNameMatcherMemos. The summary line printed orchestrator.indexAll's durationMs, which covers extraction only, beside node and edge totals recomputed after resolution: the issue's 19.9 s run printed "in 1.7s". CodeGraph.indexAll and sync now report the whole run's wall time, which `codegraph init`, `index` and `sync`, the MCP auto-sync log and the telemetry duration bucket read. Verification: nodes, edges and unresolved_refs dumps keyed by file|kind|qualified_name|start_line are byte-identical to origin/main 8998697 (with #2357's new destructuring and Vue template refs) on all 24 runs: pretix, mealie and CPython's Lib/profiling + Doc/_static, 8 each; and to deac771 (#2358, #2360, #2361) on one more round of each. A differential harness ran old and new matchDestructuredCallResult on every call position of mealie (10,330) and the CPython assets (9,170, 742 inside misread regex literals) and on 16,000 random token-soup sources (1.33M positions), and old and new jsFunctionLocalScope on ~500K (function, line, name) triples: no difference; it reports one for each deliberate mutation of the new code that is not provably equivalent. Medians of 4 interleaved runs on a loaded 16-thread Windows box, process CPU / wall / peak RSS: pretix 160.2 s / 75.9 s / 2.47 GB -> 88.7 s / 43.5 s / 1.93 GB; CPython assets 75.1 s / 37.6 s / 1.23 GB -> 32.1 s / 20.4 s / 0.81 GB; mealie (no bundles, the control) 39.2 s -> 38.6 s CPU. __tests__/js-resolution-work.test.ts counts strips: a file with a destructured binding and 40 bare calls is stripped for the check once (41 on main), a 40-line function once (40 on main); each count fails when only its own half of the fix is reverted. full-pipeline.test.ts requires indexAll's and sync's durationMs to cover a linking pass delayed 250 ms past the orchestrator's time (on main: 218 < 468). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Issue #2334 (thanks @bompus): indexing pretix took 2.2x as long since #2163, and the summary line hid it. Two JS/TS resolution checks that first shipped in 1.6.2 redid whole-file work for every reference; both now read a file or a function once. The graph is byte-identical to main. The time
codegraph init/index/syncprint now covers the whole run.On pretix: process CPU 160 s → 89 s, wall 76 s → 44 s, peak RSS 2.47 → 1.93 GB (medians of 4 interleaved runs). On CPython's bundled minified d3 (the gap #2344 left): CPU 75 s → 32 s for the 190 files of
Lib/profiling+Doc/_static.Cause
matchDestructuredCallResult(fix(resolution): a name destructured from a composable reaches what it returns #2163) ran for every bare JS-family call in a file containingconst {/let {/var {. Per call it re-tested that pattern on the whole source, then joined every line above the call, stripped comments, blanked strings, walked the text for braces and matched the binding regex over it. In pretix that isd3.v6.js(566 KB) andpdf.js(501 KB), each with thousands of calls;pdf.worker.jsis over the 1 MB limit and never parsed, and files without such text (fabric.js,vue.js,moment) paid only the per-call re-test.jsFunctionLocalScope(fix(js): a name the calling function binds is its local #2226) re-stripped the calling function's lines above every (function, name, line) it was asked about. In a minified script every function's text runs to the end of its one line, so each of those strips was the whole file.CodeGraph.indexAllrecomputednodesCreated/edgesCreatedafter resolution but kept the orchestrator'sdurationMs, which covers extraction only — the issue's 19.9 s run printed… edges in 1.7s.synchad the same split.CPU profiles (
--cpu-prof, summed over all threads; on a loaded box, so sample times run high), main → this PR:Lib/profiling+Doc/_staticmatchDestructuredCallResultjsFunctionLocalScopestripCStyle)(What remains in
jsFunctionLocalScopeon pretix is compiling its two regexes per lookup — linear, and unchanged here.)Fix
Destructured calls — read the file once. Per file (FIFO of 16 per resolution context, as the other per-file memos): the blanked code, its line starts, each
const { … } = f(…)binding (end, callee, keys, the block it sits in and where that block closes), and the spans the blanker read as regex literals. A call looks up its own name; only a call through a destructured name scans the code between that binding and itself for a shadowing declaration — and only after the callee is found to return the key (aconst { x } = require(…), for one, needs no scan).Why the results are identical: blanking the text above a call on its own gives the file's blanked code up to the call — the stripper and blanker look past a character only for a comment opener's second character (a call's name never starts with one) and for the
/that closes a regex literal. On that common prefix the binding regex finds exactly the file's matches that end at or before the call (a match never reads past its final(, and one running past the call holds no{…(that could start another), and the brace stack at the call follows from where each block closes.The one place the prefix reads differently: a call inside what the file's blanking reads as a regex literal — in minified code a division like
(a)/b(c)/2(742 ofd3.min.js's 6,248 call sites). There the prefix reads the/as division, so it is the file's code up to the/, then the short tail blanked on its own; a binding statement can only run across that/through its{ … }pattern or< … >type arguments, and where one could (the last brace before the/is a pattern's, or the last<>()is a<), the old per-call scan still runs, as it does for an out-of-range line/column or a call preceded by a comment-opening/.Function-local names — strip a function once.
jsFunctionLocalScopestrips the function's lines once (LRU of 64 per context) and cuts at the reference's line end — the stripper looks at most one character ahead, and past a line's end that is a newline. Functions that start on the same line share the regex answer by (file, first line, reference line, name): every function of a minified script does (ond3.min.js: 992 regex pairs over 277 MB of text → 320 over 89 MB).The new memos, and the existing
JS_FN_LOCAL_MEMO(which was never cleared), now drop with the other per-file memos inclearNameMatcherMemos, so a long-running server's answer can't outlive an edit.Summary time.
CodeGraph.indexAllandCodeGraph.syncsetdurationMsto the whole locked run's wall time, as they already did for the totals. Every surface that shows it now reports the whole run: thecodegraph init/indexsummary (and the re-index after opting gitignored child repos in),codegraph sync's line, the MCP server'sAuto-synced N file(s) in Xmslog, and the telemetryindexevent's duration bucket (TELEMETRY.md already describes it as the index's duration; buckets will shift up accordingly). On pretix, main printed32,277 nodes, 80,722 edges in 18.5sfor an 84 s run; this branch prints… in 39.0sfor a 40 s run.Verification
__tests__/js-resolution-work.test.ts(new, modeled onpython-resolution-work.test.ts) counts strips instead of timing: a file with a destructured binding and 40 bare calls is stripped for the check 1× (41× on main); a 40-line function whose every line calls a same-named function is stripped 1× (40× on main). Each count fails when only its own half of the fix is reverted (41/1 and 1/40). The first test also resolves a destructured call written inside a division (the in-literal path).__tests__/integration/full-pipeline.test.ts:durationMsofindexAlland ofsyncmust cover a linking pass delayed 250 ms past what the orchestrator measured — fails on main (expected 218 to be greater than or equal to 468), passes here.matchDestructuredCallResulton every call position of mealie's 532 JS/TS/Vue files (10,330 positions, 306 resolved through a destructuring) and of CPython's profiling assets (9,170 positions, 742 inside misread regex literals), and on 16,000 random token-soup sources rich in destructuring, slashes, quotes, braces, generics and comments (1.33M positions; ~9.5K decisions resolved inside misread regex literals, ~9.3K through the fallback scan) — 0 differences; old vs newjsFunctionLocalScopeon ~500K (function range, line, name) triples — 0 differences. Mutating the new code (scope checks, the crossing guard, the tail blanking, the literal lookup, the line cut, the memo key, …) makes it report differences; the two mutations it can't see are provably equivalent (a call position always starts with an identifier character).npx tsc --noEmitclean;npm run buildok. Full suite on a shared Windows box at 100% load (8998697 + this change): 5,731 passed, 41 failed, 324 skipped — every failure a timeout (36 at vitest's 5 s; the rest at 10–60 s or waiting for a daemon to attach), a status-freshness race or a WindowsEBUSYon temp-dir cleanup, in 23 sync / git / MCP-daemon / watcher / explore files. Rerun one file at a time on the final branch, 21 of the 23 pass;sync.test.ts(four 5 s timeouts) andmcp-daemon.test.ts(one 20 s timeout, a different test than in the full run) then pass alone too (58/58, 21/21 + 1 skipped).Validation (before = origin/main 8998697, after = 8998697 + this change)
Nodes, edges and unresolved_refs (and files) dumped with node:sqlite keyed by
file|kind|qualified_name|start_line— node ids replaced by those keys in edges and refs, no row ids — sorted: byte-identical on all 24 runs (8 per repo, both arms): pretix 32,277 nodes / 80,722 edges / 158,546 refs; mealie 19,436 / 42,892 / 40,155 (with #2357's new destructuring and Vue template refs); CPython assets 12,992 / 32,646 / 31,874. Main moved on while this was measured (#2358 COBOL, #2360 Dart, #2361 Go — no JS path); the branch is one commit on deac771, and one more before/after round against 9f69e47 and against deac771 is byte-identical too. (Against the earlier main 26e8488, 18 runs were byte-identical as well.)Timing — 4 interleaved rounds,
node --liftoff-only dist/bin/codegraph.js init -yrun in-process soprocess.cpuUsage()covers every worker; a shared 16-thread Windows box that sat at 100% load, so CPU is the steady number and wall is noisy:Lib/profiling+Doc/_static(190 files)mealie has no bundled library, so it is the control: its CPU is flat; its wall times ranged 22.8–55.7 s on main and 20.5–56.2 s here.
Related
🤖 Generated with Claude Code