Skip to content

Indexing slowed 2.2x on pretix since #2163: matchDestructuredCallResult rescans the file for every call #2334

Description

@bompus

Indexing pretix (1,466 files) takes 19.9 s on main at 6560052a (1.6.2), up from 8.3 s. A git bisect over 290e03f7..6560052a points at cf449fda (#2163):

Commit codegraph init on pretix
d03e4dbf (parent) 8.34 s
cf449fda (#2163) 18.65 s
6560052a (1.6.2) 19.91 s

Each step was one run on Node 24 on a 16-vCPU WSL2 host, with CODEGRAPH_NO_DAEMON=1 and telemetry off. Across the nine runs the fast builds took 7.95 to 8.34 s and the slow ones 17.70 to 19.91 s.

Cause. matchDestructuredCallResult runs for every bare JS-family call in a file that contains const {, let { or var {. For each call it joins every line above the call, runs stripCommentsForRegex and blankStringContents over that text, walks it a character at a time in stackAt, and matches the binding regex across it. The work for a file grows with the square of its size. pretix ships bundled libraries under src/pretix/static/: pdfjs/pdf.worker.js (1.95 MB), d3/d3.v6.js (560 KB) and pdfjs/pdf.js (497 KB).

A node --cpu-prof run at cf449fda shows each of the six resolver workers busy for 15.7 s. Self time summed over all threads:

  • stripCStyle: 12.7 s
  • blankStringContents: 11.2 s
  • the const { test and the binding regex: 3.6 s
  • matchDestructuredCallResult: 2.5 s
  • stackAt: 0.6 s

Possible fix. Most calls in such a file don't go through a destructured name. Strip the file once, collect the local names its const { … } = f(…) bindings introduce, cache that set per file, and return early when ref.referenceName isn't in it. Only calls through a destructured name would then pay for the prefix scan, and the links found stay the same. My fork's port of this resolver does this.

The summary line leaves out this time. For the 19.9 s run, codegraph init printed 32,322 nodes, 81,229 edges in 1.7s. In indexAll (src/index.ts), nodesCreated and edgesCreated are recomputed after resolution, but durationMs keeps the value from orchestrator.indexAll, which covers extraction only. The printed time leaves out resolution and linking, which is where this slowdown happens, so the CLI output didn't show it.

Activity

  1. colbymchenry commented on Oct 5, 2026

    @colbymchenry
    Owner

    Thanks @bompus for the bisect and the profile. They pointed straight at the cause.

    Fixed on main in #2362, and both parts of the report are covered:

    • The slowdown. Two checks that shipped in 1.6.2 redid whole-file work for every reference. matchDestructuredCallResult (fix(resolution): a name destructured from a composable reaches what it returns #2163) stripped and scanned all the text above each bare call, as you found. jsFunctionLocalScope (fix(js): a name the calling function binds is its local #2226) also re-stripped the calling function for every lookup, and in a minified bundle that function's text is the whole file. Each file's destructured bindings are now read once and looked up by name, and a function is stripped once and cut at the reference's line. A call inside what the blanker reads as a regex literal, which in minified code is often a division, is handled exactly. The old per-call scan is only a fallback for rare edge cases.
    • The summary time. indexAll and sync now report the whole run's wall time, resolution and linking included, so the init/index/sync summaries, the MCP auto-sync log and the telemetry duration bucket all describe the run they sit beside.

    The graph is unchanged. Node, edge and unresolved_refs dumps were byte-identical in every before and after run on pretix, mealie and CPython's bundled assets. A differential harness comparing the old and new code over about 1.3M fuzzed call positions found no difference. On pretix, process CPU went from 160 s to 89 s and wall time from 76 s to 44 s (medians of 4 interleaved runs on a heavily loaded box), and peak memory went from 2.47 GB to 1.93 GB. The summary now prints the run's real time: in 39.0s for a 40 s run, where main printed in 18.5s for an 84 s run. pdf.worker.js is over the 1 MB file limit, so it was never parsed. The pretix cost came from d3.v6.js and pdf.js.

    This ships in the next release, and you're credited in the changelog. Thanks also for flagging it in #2332. With #2344 that covers both halves of the slowdown since 1.6.2.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions