Skip to content

perf(resolution): read a bare-recorded call's receiver on its own line first — graph byte-identical - #2431

Merged
colbymchenry merged 1 commit into
mainfrom
claude/sweet-shannon-cb4994
Oct 7, 2026
Merged

colbymchenry merged 1 commit into
mainfrom
claude/sweet-shannon-cb4994

Conversation

@colbymchenry

@colbymchenry colbymchenry commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Summary

bareCallReceiver (src/resolution/name-matcher.ts) reads back what a call the extractor recorded by its bare name was written on, such as this.render() or window.App.init(). For every such call it joined the ref's line with the seven lines after it, then cut the text at the ref's column:

const text = lines.slice(ref.line - 1, ref.line + 7).join('\n').slice(Math.max(0, ref.column));

On a minified bundle that line is the whole file, so every call copied it. go-ethereum's graphql/internal/graphiql/graphiql.min.js is 3 lines (69 / 976,444 / 40 chars), so each call on its long line joined about 1 MB, and more than once, because several resolver checks call bareCallReceiver. Only a line followed by another line pays: V8 returns the element itself for a one-element join, and here the sourcemap comment is the next line. This shipped in 1.6.2 (#2223).

It now searches the ref's own line from the column first, and joins the lines only when the line can't settle it:

const span = lines.slice(ref.line - 1, ref.line + 7);
const from = Math.max(0, ref.column);
let text = span[0]?.slice(from) ?? '';
let at = call.exec(text);
if (!at || text.indexOf(ref.referenceName) < at.index) {
  text = span.join('\n').slice(from);
  at = call.exec(text);
}

span[0].slice(from) is a sliced string (no copy), and call is the same regex as before.

Why the result is the same for every input

  1. The own-line text is a prefix of the old joined text, at the same start offset. That also holds for NaN, fractional and negative columns. A column at or past the line's end gives '', which falls back.
  2. A match on the line at i is therefore a match of the joined text at i. The regex has no end anchor or lookahead, its lookbehind sees the same preceding character, and the path it took reads only the line's characters.
  3. Every match starts with the literal name. A match of the joined text before i would start with the name before i, which is still on the line, and indexOf(name) < at.index rules that out. So the joined text's first match is exactly i, before = text.slice(0, i) is the same string, and everything after it depends only on before.
  4. Otherwise (no match on the line, or the name earlier on it), the joined text decides, exactly as before.

The joined text can match earlier than the line alone, so the indexOf guard is load-bearing. The optional <[^<>()]*> and \[…\] groups match newlines, so an earlier name<… or name[… can run past the line break and enclose the line's own match. In window.App.foo[foo(1),\n2](3), the line alone matches foo( at 15, while the joined text matches foo[…]( at 11. The receiver is window.App; without the guard it would be null. No real code in the 12 repos below does this, but the fuzz hit 886 cases, and the new pin test fails with the guard removed.

Validation

Differential check, old vs new compiled function. Each arm's built name-matcher.js was loaded with bareCallReceiver exported, and both were run on the same inputs with deep-equal results asserted:

inputs refs mismatches
every unresolved ref of 12 repos, from DB snapshots taken just before resolution (go-ethereum, excalidraw, qwik, bootstrap4, legend-state, mantis, takenote, gin, django, requests, play-samples, vscode) 3,111,262 0
synthetic calls refs: every identifier of every line (sampled on long lines) at its own column, the start of its member chain, column 0, an earlier column, the line's end and past it, and under the next identifier's name 3,192,761 0
random token soups (call-shaped pieces, newlines, <>[]), random line/column/name, incl. out-of-range lines and NaN/fractional/negative/past-the-end columns 5,000,000 0

Characters joined, real refs: go-ethereum 9,056.9 M → 2.1 M (68,667 joins → 370), vscode 189.4 M → 0.7 M, play-samples 74.3 M → 0.0 M.

Graph identity. Full codegraph init, then scripts/dump-graph.mjs; dumps are byte-identical:

base repo runtime result
ed199e6 (this PR's base) go-ethereum Node 22 and Node 24 (bundled) IDENTICAL on both (38,778 nodes / 156,211 edges / 125,976 refs)
ed199e6 excalidraw, qwik, bootstrap4 Node 22 IDENTICAL
d8a7f86 + #2424 go-ethereum Node 22 IDENTICAL (38,706 / 155,820 / 125,942)
d8a7f86 go-ethereum Node 24 IDENTICAL
d8a7f86 excalidraw, qwik, bootstrap4, legend-state, mantis, takenote Node 22 IDENTICAL

d8a7f86 alone never finished go-ethereum on Node 22 (the stall #2424 fixed), so on that base both Node 22 arms carry #2424's source change. ed199e6 already includes #2424.

CPU, go-ethereum. The box is shared and at 100% load, so these are process CPU and profiles; wall times swing by tens of seconds.

  • Full codegraph init, 4 interleaved rounds per arm with alternating order. Medians; every dump identical:

    runtime arms user+sys sys
    Node 22 main + fix(resolution): a large minified JS bundle no longer stalls resolving on Node 22 #2424 → + this 89.4 → 83.9 s (−6%) 14.8 → 12.4 s
    Node 24 main → + this 81.5 → 74.7 s (−8%) 15.3 → 11.1 s

    A second Node 22 set with phase timings gave 89.5 → 85.3 s, and every run with the change was below every run without it.

  • --cpu-prof on Node 22, all 18 threads: bareCallReceiver self time 6.29 → 1.04 s (2.61 → 0.25 s in the busiest thread); GC 5.48 → 3.00 s.

  • One resolver replaying graphiql.min.js's 18,394 refs, interleaved ×3, with the same 8,863 results: CPU 9.7 → 5.4 s on Node 22 (+ fix(resolution): a large minified JS bundle no longer stalls resolving on Node 22 #2424) and 8.1 → 4.1 s on Node 24.

  • Whole-index wall time doesn't move measurably: Node 22 phase-timed medians were 66.1 vs 65.8 s, while the parse loop (identical code in both arms) alone varied 21–31 s between runs. The pool spreads the saved work across its workers.

Tests.

  • A new counted test in __tests__/js-resolution-work.test.ts puts 300 window.AppN.mN(N) calls on one minified line. Every receiver is still read back. On main the calls join 2,077,500 characters (300 × the 6,925-char window) against a bound of one line's length (6,886); with this change they join 0.
  • A second test pins the multi-line cases: a call that continues on the next lines, the foo[foo(1),\n2](3) case above, and a column past its line's end. It passes on main and fails with the indexOf guard removed.
  • 12 related test files (392 tests) pass on this base, including kernel-tsjs-parity with a kernel built from it.
  • Full suite: the parallel run on this saturated box (100% CPU, 100+ node processes) hit load failures in 191 files (612 test and 40 hook timeouts, EBUSY, timing asserts). A serial rerun of those 191 files with 120 s timeouts passed 2,334 of 2,356 tests; the 6 that failed were timeouts and an EBUSY again. Their 6 files then passed alone (159 tests), on main's code and on this branch alike.

CHANGELOG

No entry, deliberately. On go-ethereum the whole index uses 6–8% less CPU, but its wall time doesn't move measurably, so users won't notice. The per-file effect is larger (graphiql.min.js resolves in about half the CPU).

Overlap

#2424 (merged as 8f0766c, in this PR's base) also edited name-matcher.ts, in jsCodeBindsName, far from this hunk. It fixed a separate Node 22 stall on the same file; this was the per-ref copy left over from that investigation. No open PR touches bareCallReceiver.

🤖 Generated with Claude Code

…e first — graph byte-identical

bareCallReceiver joined the ref's line with the seven after it for every
bare-recorded call, then cut the text at the ref's column. On a minified
bundle that line is the whole file, so every call copied it: about 1 MB
per call on go-ethereum's graphiql.min.js, 9 GB over that repo's call
refs.

It now searches the ref's own line from the column first. Every match of
the call pattern starts with the name, so a match there with no earlier
occurrence of the name on the line is where the joined lines match first
too. With no match on the line (a call that continues on the next one),
or with the name earlier on it (an earlier `name[…` or `name<…` can run
past the line break and enclose the match), the joined lines decide as
before. The result is the same for every input.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant