Read this first every session. Single source of truth for what is done, in progress, and next.
- Current phase: 7 — Backend parsers & federation (v2 horizon). Phase 6 complete (Gate 6, M6). 0.5.0 published to npm.
- Next step: publish 0.6.1 to npm (needs the user's creds —
npm publish). 0.6.0 is already published; 0.6.1 is a metadata-only patch that carries the npm/GitHub SEO work (description, 20 keywords incl. the MCP/AI-agent cluster,repository/homepage/bugs, LICENSE, README rewrite) — registry metadata is immutable per version, so it only takes effect on publish. Then the tester validates the wider data-source coverage. After that, the backend-parser steps (7.4 Python, 7.5 Go, 7.6 federation) — larger, new-package work, still sketch-level; best held until that feedback confirms priority. 6F.6 detection half still blocked on a real failing test file from the tester. - Done: 0.1–0.4, 1.1–1.6, 2.1–2.5, 3.1–3.6, 4.1–4.6, 5.1–5.7, 6F.1–6F.5, 6F.7–6F.10 + 6F.6 (defensive half), 6.1–6.5 (Phase 6 complete), 7.1 (GraphQL adapter), 7.2 (Next.js server data), 7.3 (push channels) · 0.5.0 released & published to npm (tag v0.5.0); 0.6.0 prepared (7.1–7.3; all packages + generator/CLI/MCP version strings bumped, README changelog;
npm packverified; publish pending user's npm creds). Self-validated on Grafana frontend (6,461 files, 15,334 nodes / 18,367 edges in 72 s): 55 RTK-query data sources, 32 routes, 1,009 coverage edges, gibberish declines — all previously 0/1 in the field run. - Gates passed: Gate 0 (CI + red-path, #5/#6) · Gate 1 (precision 1.000, recall 0.895, zero poison) · Gate 2 (C1 instance attribution 1.000 · B1 4-level handler chains · C6 store writers↔readers · A9 portals — scorecard 137/0/0, precision & recall 1.000) · Gate 3 (B3 action effects · B4 routers · B6 cyclic journeys terminate · B7/B8 form & non-JSX events · G5 flag/role conditions — precision & recall 1.000) · Gate 4 (A4 rarity · A10 fuzzy/OCR · A1 structural · A6 subtree · E3 vision annotations · E2 aliases · G4 corrections — high-conf correct 1.000, ambiguity honesty 1.000, poison rate 0.000) · Gate 5 (F1 context bundle · F2 blast radius · F3 test coverage · F4 response schema · F5 git history · MCP server over stdio — scorecard 265/0/0, all honesty metrics 1.000; M5 reached — ticket in → budgeted context bundle out, over MCP)
A context-provider node in a multi-agent development pipeline. Input: a Jira ticket (text + screenshots + links). Output: a context bundle — matched component instances, their data lineage (APIs, state, events), the relevant user-journey slice, blast radius, tests, and git history — sized to a token budget, with evidence and confidence on every claim.
Three core requirements:
- Match — UI snapshot / ticket text → component instance(s)
- Journeys — every user action path, n levels deep, lazily expanded
- Attribution — which APIs feed which UI, per instance (a shared
DataTableon the Users page is fed by/api/users; the same component on the Invoices page by/api/invoices)
Reference docs:
- docs/failure-modes.md — the catalog (IDs A1–G8) every step maps to
- docs/testing-strategy.md — eval harness, metrics, phase gates
- One step = one branch = one PR to
main. Branch:build/phase-{N}/step-{N}.{M}-{slug}. - Update this file (step status + Status block) in the same PR.
- A step is done only when its acceptance criteria all pass and
pnpm evalis green. - A phase is done only when its gate (see testing-strategy.md) passes in CI.
- Every step that addresses a failure mode adds/extends a fixture under
eval/fixtures/<id>-*.
Status legend: [ ] todo · [~] in progress · [x] done
Everything else depends on getting the schema and the measurement loop right first. The v0.1 schema is definition-only, which is wrong for C1 — fix before building on it.
Failure modes: C1 (schema half), D2 (schema half), A5, B5 (edge conditions)
Build: rework packages/core/src/types.ts:
InstanceNode—{ id, kind: "instance", definitionId, parentInstanceId | null, loc, staticProps: Record<string, string> }. One per JSX call site of a project component. Id:instance:<file>:<line>#<DefName>.- Definition nodes keep
component/hookkinds.rendersedges connect instances;instance-ofconnects instance → definition. Evidence—{ kind: "text-match" | "structure" | "edge-chain" | "alias" | "correction", detail: string, loc?: SourceLocation }.Confidence—"high" | "medium" | "low"plus numericscore: number(0–1). Every query result type carriesevidence: Evidence[]andconfidence.EdgeCondition—{ kind: "flag" | "role" | "branch" | "response", expression: string }, optional on every edge.- Result envelope:
QueryResult<T> = { status: "ok" | "ambiguous" | "declined", candidates: Array<{ value: T, confidence, evidence }>, disambiguation?: string }. No query API ever returns a bare value. Accept: - Unit tests for id construction, envelope invariants.
parser-reactupdated: emits an instance node per project-component JSX usage;renders/fetches-from/reads-state/handlesedges originate from instances where applicable, definitions otherwise.- Demo-app scan shows
UserCardwith 1 definition + 1 instance (parentUserList). matchComponentsByText/traceLineagereturnQueryResultenvelopes.
Failure modes: infrastructure for all; first fixtures C1, A4
Build: eval/run.ts (a workspace package @coderadar/eval):
- Discovers
eval/fixtures/*/, scans eachapp/dir, diffs the graph + query results againstgolden.json(format in testing-strategy.md, includingforbiddenentries). - Emits
eval/scorecard.json: per-fixture pass/fail, per-metric values, grouped by failure-mode category; appends toeval/history.jsonl. pnpm evalruns it; non-zero exit if any threshold ineval/thresholds.jsonis violated.- Fixtures:
c1-shared-datatable(DataTable rendered by Users + Invoices pages, different APIs — must attribute per instance,forbiddencatches definition-level attribution),a4-generic-text(three components sharing "Save"), plusdemo-appmoved under fixtures. Accept:pnpm evalruns green locally on the non-C1 assertions; C1 attribution assertions may be red (prop-flow lands in 2.2) but the fixture and its golden file exist and the runner reports them asexpected-fail: phase-2. Expected-fail support is part of the runner.
Failure modes: G2, G3 (foundation), D1 (foundation) Build:
GraphMeta—{ commitSha, dirty: boolean, generatedAt, generator, scanRoot }embedded inLineageGraph.saveGraph(graph, path)/loadGraph(path)in core with schema-version check (refuse to load a newer major version).- JSON Schema for the whole graph exported to
schemas/lineage-graph.schema.json(repo root —dist/is gitignored; generated from the TS types; committed; drift-gated by a test that regenerates and diffs). - CLI:
scanrecords commit SHA (viagit rev-parse,dirtyfromgit status). Accept: round-trip test (scan → save → load → deep-equal); schema drift test;coderadar scanoutput includes SHA.
Failure modes: D1, process
Build: GitHub Actions workflow: install (pnpm cache) → build → typecheck → vitest unit tests → pnpm eval → upload scorecard as artifact. Threshold ratchet rule documented in the workflow file header (raising OK; lowering needs PR-body justification).
Accept: CI green on a PR; a deliberately broken fixture in a test branch turns it red. Gate 0 passes.
Make the parser survive real code instead of demo code. All work is within-file or follow-one-reference; the cross-file instance/prop-flow machinery is Phase 2.
Failure modes: C2 (constants half), C3
Build: in parser-react:
- Resolve identifier arguments to fetch/axios through imports to their declarations (constant folding: string literals,
constobject members likeENDPOINTS.USERS, simple concatenation). - Template literals → route patterns:
`/api/users/${id}`→endpoint: "/api/users/:id",resolved: "partial", unresolved segments listed. - Strip configured base-URL prefixes (scan option
baseUrls: string[]). DataSourceNodegains{ pattern: string, resolved: "full" | "partial" | "none", raw: string }. Accept: fixturesc2-endpoint-constants,c3-dynamic-endpointsgreen; lineage precision holds ≥ 0.90 on all existing fixtures.
Failure modes: C2 (wrapper half) Build:
- Detection heuristic: a function/method whose body reaches
fetch/axiosand takes a path-like parameter → classified as an API wrapper; its call sites become data sources with the path argument resolved per 1.1. - Manual override in scan config:
apiWrappers: ["apiClient.get", "http.post", ...]. - Wrapper chains up to depth 3 (
useApi → apiClient.get → fetch). Accept: fixturec2-api-wrapper(three-layer wrapper) green; wrapper detection has unit tests for both heuristic and config paths.
Failure modes: C5
Build: when useQuery/useMutation/useSWR's fn argument is a reference, resolve to its declaration (same file or import) and extract the endpoint from its body via 1.1/1.2. Data source records both the query key and the resolved endpoint; mutations get method from the inner call.
Accept: fixture c5-queryfn-indirection green (queryFn in a separate api/users.ts).
Failure modes: A2 Build:
- Scan option
i18n: { localeGlobs: string[], defaultLocale: string }; parse JSON/YAML locale files into a key → string-per-locale table (nested keys flattened,{{var}}placeholders preserved). t("key")/<Trans i18nKey="key">call sites resolve intorenderedTextentries{ text, locale, source: "i18n", key }for all locales.renderedTextbecomes structured:{ text: string, source: "jsx" | "attribute" | "i18n", branch?: string, locale?: string }[](schema addition — update JSON Schema + goldens). Accept: fixturea2-i18n-keysgreen: searching "Team Members" and "Équipe" both find the component.
Failure modes: A7 (extraction half), A8 Build:
- Template text:
`${count} items`→"* items"with atemplate: trueflag. - Branch tagging: text inside a conditional (ternary /
&&/ early return) getsbranch: <condition source text>. - Normalization utility in core (shared with the Phase 4 matcher): lowercase, collapse whitespace, strip punctuation, naive singular/plural folding.
Accept: fixtures
a7-transformed-text,a8-conditional-textgreen (error-state text findable and tagged with its branch).
Failure modes: D4 Build:
- Class components:
render()treated as the body;this.state/setState→ state nodes; lifecycle fetches detected. - HOC unwrapping for
connect(...),withRouter(...)etc. (unwrap to the inner component; record the HOC name). - Anything unparseable yields a node with
flags: ["incomplete"]rather than silence; scan summary counts incomplete nodes. Accept: fixtured4-class-componentsgreen; incomplete count surfaced inscanoutput. Gate 1 passes.
The heart of the project. C1 and B1 live here.
Failure modes: C1 (graph half), A5
Build: cross-file pass in parser-react:
- Resolve every JSX tag to its component definition through imports (including
export { X as Y }, barrel files, default exports). - Emit
InstanceNodeper call site withparentInstanceIdforming the render tree; static (literal) props recorded instaticProps. - Design-system components (imported from
node_modulesor a configureddesignSystemPackageslist): instances are still created, flaggedexternal-definition— the instance is ours even when the definition isn't. Accept: fixturea5-design-systemgreen (match resolves to the usage site); instance counts asserted in c1 golden; barrel-file resolution unit-tested.
Failure modes: C1 (the headline) Build:
- For each instance, connect prop values to their origins in the parent scope: identifier props trace back through variable declarations to hook results / fetch results / store reads within the parent.
- New edge
provides-data(parent's data source → instance,via: <propName>). traceLineage(instanceId)followsprovides-data+ own edges;traceLineage(definitionId)returns the per-instance breakdown, never a merged blob.- Depth limit (default 5 hops) with
flags: ["depth-limited"]when hit. Accept: c1-shared-datatable attribution assertions flip from expected-fail to green — DataTable@Users →/api/users, DataTable@Invoices →/api/invoices, zeroforbiddenhits. Instance attribution accuracy = 1.0 on C1 fixtures.
Failure modes: B1 Build: same machinery, callback direction:
onClick={props.onSave}→ resolveonSaveat each parent instance → the passed function → its body (which Phase 3 mines for effects). Chain recorded as evidence (edge-chain).- Handles: inline arrows wrapping prop calls, renamed props, default props, destructured-with-rename.
- Unresolvable after 4 levels →
handler: null, flags: ["unresolved-prop-handler"]— visible, not silent. Accept: fixtureb1-prop-drilled-handler(4 levels) ≥ 0.85 resolution; unresolved cases flagged in graph.
Failure modes: C6, B2 (redux/zustand half) Build:
- Redux/RTK: slice detection (
createSlice),useSelector(s => s.users)→ StateNode per slice path; dispatch sites of thunks/actions that carry API results →writes-stateedges from the data source to the slice. - Zustand:
create()stores, actions that fetch → samewrites-statelinkage. traceLineagethrough a StateNode now continues to its writers: reader component → slice → populating API. Accept: fixturec6-store-decoupledgreen — component with no fetch of its own correctly attributed to the API called on a different page at login.
Failure modes: A9
Build: createPortal detection + adapter list for common modal/toast libs (react-modal, radix Dialog, react-hot-toast): the triggering instance gets a triggers-render edge to the portal-content instance.
Accept: fixture a9-modal-portal green: matching the modal's text surfaces both the modal component and its trigger site. Gate 2 passes.
Failure modes: B4
Build: RouteNode (path, layout, guards) in core. Adapters: React Router (createBrowserRouter / <Route> trees, nested + lazy) and Next.js file-based (app + pages router). routes-to edges route → page-component instance tree.
Accept: fixtures b4-react-router, b4-nextjs-approuter green: every golden route maps to its page component.
Failure modes: B3, B2 (dispatch half)
Build: mine resolved handler bodies (from 2.3) for effects, each a typed triggers edge from the EventNode:
navigate(...)/router.push(...)→navigates-towith route pattern (computed strings →:paramform, B3)- fetch/axios/wrapper calls → existing data-source edges
dispatch(action)→ through the 2.4 adapter to state writessetState/setters →writes-stateon local state Accept: fixtureb3-programmatic-navgreen; every effect kind covered by a unit test.
Failure modes: B5, B6
Build: journeys(graph, startInstanceId | routePath, { depth }) in core:
- BFS over event → effect → route → page-instance → its events…, expanding paths at query time; per-path visited-set for cycle handling (a node may repeat across paths, not within one); cap +
truncatedflag. - Steps carry
EdgeConditions (branch/flag/role) so a journey reads: Users page → [role=admin] Delete click → DELETE /api/users/:id → confirm toast. - Returns
QueryResult<JourneyPath[]>. Accept: fixtureb6-cyclic-journeys(list ↔ detail loop): 3-level golden paths exact, terminates < 1 s; depth-n request on a cyclic graph never hangs.
Failure modes: B7, B8
Build: react-hook-form / Formik adapters (handleSubmit(onSubmit) → real handler); addEventListener in useEffect → EventNode (source: "effect"); adapter list for hotkey libs. Unknown patterns → file-level flags: ["unscanned-events"].
Accept: fixtures b8-react-hook-form, b7-effect-listeners green.
Failure modes: G5, B5
Build: feature-flag detection (configurable call names: useFlag, useFeature, isEnabled) and role checks in render branches → EdgeCondition{kind:"flag"|"role"} on the enclosed renders/handles edges.
Accept: fixture g5-feature-flag green: flag-gated UI's journey step carries the flag name. Gate 3 passes.
Failure modes: B9
Build: ExternalNode + exits-app / enters-at edges. Navigate/window.open/window.location.assign/<a href>/<Link to> to an absolute URL or mailto:/tel: scheme → exits-app (event or component → external, deduped by host). Deep-link/OAuth-callback route paths → enters-at from an inbound external node. Journeys reaching an external end with an exit step.
Accept: fixture b9-cross-app-hops green (OAuth redirect, Stripe window.open, mailto link, /auth/callback entry).
Failure modes: A4, A7 (matching half), A10
Build: replace matchComponentsByText with a scorer over instances:
- Normalization (from 1.5) both sides; fuzzy token match (edit distance ≤ 1 per token, token-set overlap) for OCR noise.
- Rarity weighting: term weight = inverse frequency across the graph ("Save" ≈ 0, "Reconciliation" ≈ high).
- Combination bonus: co-occurrence of multiple terms in one instance subtree outweighs scattered singles.
Accept:
a4-generic-textgreen ("Save" alone →ambiguous; "Save" + "invoice details" → correct top-1); noisy-term fixturea10-ocr-noise(misspelled terms) top-3 correct.
Failure modes: A1, A3, A12 — fixtures: a1-no-static-text, a3-api-text, a12-non-text (structure-only matches capped at medium confidence — honest graceful degradation)
Build: structural signature per instance subtree (child element kinds/counts: table with N columns, form with M inputs, card grid). Query side accepts a structure descriptor (from vision output: "a table with columns Name, Email, Actions") and scores against signatures. Text and structure scores combine into one ranking.
Accept: fixture a1-no-static-text (dashboard, zero literals) top-3 correct via structure alone.
Failure modes: A6
Build: when matches nest (Page > Section > Card all match), return the deepest instance covering the matched term/structure set; ancestors listed as context, not competing candidates.
Accept: fixture a6-composed-page green: full-page term set → page node; card-specific terms → card instance with page as context.
Failure modes: A10, E3, A13
Build: @coderadar/vision package:
VisionAdapterinterface:extract(image) → { terms: string[], structure: StructureDescriptor, annotations: Region[] }. Ships with a Claude-vision implementation; OCR-only fallback stub for tests.- Annotation-region priority (E3): terms inside detected circles/arrows weighted 3×.
- Non-app detection: extraction result includes
looksLikeApp: boolean(Figma frames, marketing pages → decline path). - Ephemeral only (G7): images processed in memory, never written to the graph or disk; document in the package README.
Accept: interface unit-tested against recorded extraction outputs (no live API in CI); annotation weighting verified on
e3-annotated-screenshotfixture (pre-extracted terms + regions checked in).
Failure modes: D2, D6, G1 Build:
- Score → confidence mapping calibrated on the eval set (thresholds chosen so measured precision at
high≥ 0.95 — recorded ineval/calibration.json, regenerated by an eval subcommand). ambiguousresults includedisambiguation: a concrete question generated from the differences between top candidates ("Which page is the table on — Users or Invoices?").declinedresults carry a machine-readable reason (out-of-scope,not-our-app,no-signal). Accept: poison rate ≤ 0.05 and ambiguity honesty ≥ 0.90 on the full matching eval set; everyhigh-confidence answer in the eval set is correct.
Failure modes: E2, G4
Build: aliases.yaml (checked into the target repo): "invoice widget" → BillingSummaryCard, route titles, feature names. Corrections API: recordCorrection(terms, confirmedInstanceId) appends to corrections.jsonl; both feed the matcher as first-class evidence (kind: "alias" | "correction", high weight). Eval subcommand folds corrections into new ticket-eval cases.
Accept: fixture e2-business-vocab green via alias; a recorded correction changes the next identical query's top-1 (integration test). Gate 4 passes.
Failure modes: E1, E5, E6, F6(decline), E4(decline)
Build: @coderadar/agent-sdk package. resolveContext(ticket: { text, screenshots?, links? }):
- Entry-point classification: visual (screenshot present) / textual (UI terms in prose) / behavioral ("clicking X does nothing" → match on event/handler/journey vocabulary, E5) / out-of-domain (backend/infra/perf →
declined, E6) / unsupported-input (video → structured decline, E4). Classification is rule-based + keyword lexicons; no LLM call inside the node (determinism, G8). - Runs the matching engine with all available signals; merges rankings.
Accept:
eval/tickets/suite (≥ 15 hand-written tickets: 5 visual, 4 textual, 3 behavioral, 3 out-of-domain) — OOD rejection ≥ 0.95, entry-point classification accuracy ≥ 0.90.
Failure modes: F1
Build: ContextBundle type + JSON Schema (committed, drift-gated): sections match (instances + evidence), lineage (per-instance sources with patterns/methods), journeys (slice around the match, depth 2 default), blastRadius, tests, history, warnings (staleness, incomplete flags). Budgeter: budgetTokens param; sections trimmed in fixed priority order (match > lineage > blastRadius > journeys > tests > history), each trim recorded in warnings.
Accept: golden bundles for 5 ticket evals; every bundle ≤ budget under a tokenizer test at budgets 2k/4k/8k; trim order unit-tested.
Failure modes: F2
Build: reverse adjacency in core: blastRadius(nodeId) → all instances rendering this definition, all consumers of this data source/endpoint, all readers of this state slice, journeys passing through — each with distance. CLI coderadar impact <node>.
Accept: fixture golden: changing the c1 DataTable definition lists both page instances; changing /api/users lists every consumer across fixtures.
Done: blastRadius(graph, target) in core — a dependency-direction-aware reverse BFS (dependencyOf maps each edge kind to resource↔dependent, so a change propagates the right way regardless of edge direction; journey edges don't propagate impact). Resolves a target by node id, component name, endpoint, state name, or route path; returns ImpactNode[] (node, relation, distance) nearest-first, always high-confidence. Wired into the context bundle's blastRadius section (compact Name@file:line labels) and a CLI impact <node> command (-d depth cap). c1 golden blast block asserts both instances (d1) + both pages (d2) for the shared DataTable, and every /api/users consumer with an over-reach guard forbidding the invoices side. New GoldenBlast type + blast check kind in the eval harness. 5 core unit tests; eval 258/0/0, gate OK.
Failure modes: F3
Build: scan *.test.* / *.spec.* / __tests__: imports + rendered components (render(<UserList/>), testing-library queries) → TestNode + covered-by edges. Bundle's tests section lists test files for matched instances and their lineage.
Accept: fixture with co-located tests: bundle names the right test files; components without tests get warnings: ["untested"].
Done: TestNode (kind test, framework vitest/jest/unknown) + covered-by edge in core (schema regenerated, drift gate green). New detectTests parser pass: test files are excluded from the component/instance scan (isTestFile) so they never emit spurious nodes, then swept — every component a test renders (JSX tag) or imports resolves to a covered-by edge (imports resolved to their source file for precise attribution, name fallback otherwise). Bundle populates tests from the matched component's render subtree (componentSubtree) and pushes an untested warning when the matched component has no coverage; blastRadius counts a test as a dependent of the component it covers. New f3-test-coverage fixture + GoldenCoverage/coverage check kind (covers UserList, Sidebar untested). 6 parser unit tests (incl. two bundle-level); eval 262/0/0, gate OK.
Failure modes: F4
Build: data sources link to response types: generic argument (useQuery<User[]>), annotated variable types, or an OpenAPI spec (scan option openapi: path) matched by endpoint pattern. Bundle lineage entries carry responseType: { name, fields } (one level of fields, not deep).
Accept: fixture f4-typed-responses green for all three sources (generic, annotation, OpenAPI).
Done: ResponseType { name, fields: {name,type}[], source } on DataSourceNode in core (schema regenerated, drift gate green). New response.ts parser module: responseFromCall recovers the type from a call's generic argument (axios.get<User[]>, useQuery<T>) or, failing that, the annotation on the nearest enclosing typed variable whose initializer holds the call (const data: Invoice[] = await fetch(…).then(r => r.json())), stopping at function boundaries; only property signatures are read (one level, methods skipped). loadOpenApi/linkOpenApiResponses is a post-pass that fills untyped sources from an OpenAPI 3 JSON spec (openapi scan option / CLI --openapi), matching ${METHOD} ${endpoint} with {id}→:id normalization and $ref resolution. Bundle lineage dataSources carry responseType; trace prints it. New f4-typed-responses fixture + GoldenResponse/responses check kind (all three sources). 5 parser unit tests incl. bundle-level; eval 265/0/0, gate OK.
Failure modes: F5
Build: history section: last N commits touching matched files (git log --follow), PR numbers parsed from merge/squash subjects. Pure git subprocess, no network.
Accept: integration test on this repo's own history; graceful empty section outside a git repo.
Done: gitHistory(root, files, limit) in agent-sdk — per-file git log --follow (via execFileSync, no network), commits deduped across files and sorted newest-first by committer time, parsePrNumber pulls #NN from merge/squash subjects. Any failure (not a repo, git missing, untracked path) yields [] — the bundle degrades gracefully. Bundle history section (default 5) is populated from the matched component's files + render subtree; BundleCommit gains optional pr (context-bundle schema regenerated, drift gate green). Integration tests run against this repo's own history; 5 agent-sdk unit tests + a bundle-wiring test. eval 265/0/0, gate OK.
Failure modes: G1 (surface), pipeline integration
Build: @coderadar/mcp exposing tools: resolve_context(ticket), find_component(terms), trace_lineage(id), journeys(id, depth), blast_radius(id) — thin wrappers over agent-sdk against a pre-built graph (path from env/config). Tool descriptions written for agent consumption (when to use which, what ambiguous means, that disambiguation should be relayed to a human).
Accept: MCP integration test via stdio client: scan fixture → each tool returns schema-valid envelopes; ambiguous/declined statuses round-trip. Gate 5 passes. CodeRadar is now pluggable into the multi-agent system.
Done: new @coderadar/mcp package (MCP SDK 1.29). createServer(graph) registers all five tools over @modelcontextprotocol/sdk's McpServer with zod input schemas and agent-facing descriptions (resolve_context returns the full budgeted ContextBundle; each tool returns its QueryResult envelope as JSON text; descriptions state that ambiguous/disambiguation goes to a human). resolve_context = buildBundle; find_component/trace_lineage/journeys/blast_radius wrap core. Bin coderadar-mcp loads a graph from CODERADAR_GRAPH (or argv) and serves stdio; also shipped in the published ui-lineage bundle as the ui-lineage-mcp bin (MCP SDK + zod kept external). Integration test via a real stdio Client: lists all five tools and round-trips ok/ambiguous/declined over a4-generic-text (6 tests). Gate 5 passed. M5 reached — ticket in → budgeted context bundle out, over MCP.
Failure modes: D1, G2
Build: per-file content hashes in GraphMeta; scan --update re-parses only changed files + dependents (import graph), rebuilds affected cross-file passes (instances/prop-flow are the tricky part — dependents include all parents of changed components). --watch mode for dev.
Accept: correctness: incremental result deep-equals full re-scan on 20 randomized single-file edits of the bench repo (property test); 10-file change < 15 s.
Done: scanReact split into createScanProject (build + load the ts-morph
project) and scanProject (the full analysis). New IncrementalScanner keeps
one project alive; update() calls refreshFromFileSystemSync on each source
file — ts-morph re-parses only files whose bytes changed — picks up added files
and drops deleted ones, then re-runs scanProject. Correctness by
construction: every node and cross-file edge is re-derived from the current
ASTs each update, so an update's graph is byte-identical to a fresh full scan —
the "dependents" problem (a changed component's parents' instance/prop-flow
attribution, store/route/journey wiring) is handled for free because the global
passes always re-run; the incremental win is parsing (unchanged files keep cached
ASTs). Property test proves it: an interconnected app (pages → components → atoms
- hook, cross-file imports/fetches) stays byte-identical to a full re-scan across
20 randomized single-file edits (rendered text / endpoint / re-pointed
cross-file import), plus add-file and delete-file cases with exact
changedreporting.GraphMeta.fileHashes(relative path → sha256, schema regenerated, drift gate green) records provenance;projectFileHashescomputes it. CLI:scan --updateshort-circuits when every file hashes identically to the prior graph (else reports the changed set and full-rescans, correct at ~2 s), andscan --watchre-emits the graph on each debounced file change (ignoring the output file,node_modules,.coderadar). Verified end-to-end via the CLI (--update no-change/changed, --watch live edit). eval 317/0/0/0, determinism 1.000; full suite (parser-react 151), typecheck, lint green.
Gate 6 — PASSED. Incremental (6.1) · scale bench < 5 min / < 4 GB, budget- gated nightly (6.2) · deterministic two-run byte-identity, eval-gated (6.3) · version-skew/rename tracking (6.4) · generated/vendored classification + PII policy enforced (6.5). M6 reached: production-grade — incremental, fast, deterministic, versioned.
Failure modes: D3
Build: eval/bench/ generator (2,000+ file synthetic app with realistic import depth); profile; apply: lazy ts-morph project loading, file-batch parallelism (worker threads), tree-sitter fast path for the text-extraction pass if ts-morph remains the bottleneck. Perf budget asserted in nightly CI.
Accept: full scan < 5 min, peak RSS < 4 GB on the bench repo.
Done: new eval/src/bench.ts deterministically generates a synthetic app
(default ~2,016 files: N features × page → 8 sections → 8 atoms + a hook, with
real import depth and a direct-fetch per section so the endpoint pass is
exercised — 1,904 components / 896 data sources / 2,688 instances / 8,176 edges),
scans it, and gates a perf budget (--files / --budget-seconds /
--budget-rss-mb), measuring wall time and peak RSS via
process.resourceUsage().maxRSS. Measured: ~2,016 files in ~2 s, ~400 MB RSS
— vastly under the 300 s / 4 GB budget (the synthetic app has no tsconfig, so no
type-resolution cost; the field Grafana run was 6,461 files in 72 s / 1.5 GB).
Because ts-morph is nowhere near the bottleneck at target scale, the optional
worker-thread / tree-sitter fast paths were deliberately not applied — they'd
add complexity for no measured win; revisit if a future budget run regresses.
Root pnpm bench script; nightly .github/workflows/perf.yml (schedule +
manual dispatch) runs the budget-gated bench without slowing per-PR CI; the
generated tree (eval/bench/) is gitignored and self-cleaned after each run.
eval 317/0/0/0, determinism 1.000; typecheck + lint clean. Gate 6 fully
passes once 6.1 lands (incremental re-scan uses this bench for its property
test).
Failure modes: G8
Build: stable ordering everywhere (nodes, edges, candidates — explicit sort keys, no map-iteration order leaks); vision/OCR results cached by image hash; generatedAt excluded from equality. Determinism check in the eval runner (two runs, byte-diff).
Accept: double-run byte-identical on all fixtures + bench repo.
Done: scanReact now iterates source files in a stable path order (new
sortedSourceFiles — ts-morph returns glob-enumeration order, which varies by
platform) across all three file passes, and returns nodes sorted by id + edges
sorted by a total-order key over every identifying field (edgeSortKey).
resolveHookEdges re-sorts and dedups after rewriting unresolved-hook:
placeholders (a rewrite can collide two edges onto one hook id). Candidate
ranking gains a final id tiebreak so same-named components in different files
rank deterministically. The eval runner scans every fixture twice and
byte-compares (dropping only generatedAt): new determinismStablePct metric,
printed each run and a hard gate — any non-deterministic fixture fails the
run outright. 4 new parser-react unit tests (two-run byte-identical · nodes in id
order · edges in key order · no duplicate edges). Full suite green; eval
304/0/0/0, determinism 1.000, all metrics 1.000. One fragile unit assertion in
queryfn.test.ts (relied on node insertion order to find the react-query
/api/stats source ahead of the raw fetch it wraps) rewritten to assert the
react-query resolution exists, order-independently. Bench-repo half of the
accept criterion lands with 6.2. Vision/OCR image-hash caching deferred — the
eval path exercises no live OCR, so it isn't a determinism risk today; revisit if
6.4/vision work reintroduces it.
Failure modes: G3, A11
Build: graph store keyed by SHA (.coderadar/graphs/<sha>.json + latest pointer); resolveContext accepts graphVersion; cross-version diff maps renamed/moved definitions (same structure+text signature, different name/path) → bundle warning: "matched InvoiceCard in prod graph; renamed BillingCard on main".
Accept: fixture pair (pre/post rename): query against old graph + current code yields the rename warning with the new name.
Done: new core diffRenames(fromGraph, toGraph) pairs a definition that is
gone by identity (name+file) in the newer graph with a same-body-signature
definition that is new there — signature = structural fingerprint + normalized
rendered text + props + rendered-children, all stable across a rename or file
move. Confident 1:1 only: a signature unique among both the gone and the arrived
definitions; generic empty bodies and ambiguous many-to-many sets are skipped, so
no false renames. SHA-keyed store in core storage (saveGraphToStore /
loadGraphFromStore / graphStoreDir): .coderadar/graphs/<sha>.json (or
working outside git) + a latest pointer. buildBundle gains a currentGraph
option — when the matched definition was renamed/moved there it emits
version skew — matched \InvoiceCard` (…); renamed `BillingCard` (…) in the
current graph. CLI: scan --storewrites the store;bundle --against <sha|latest> loads the current graph from it (resolveContext/version selection surfaced at the store boundary). Fixture pair g3-version-skew/{old,new}(InvoiceCard.tsx → BillingCard.tsx, renamed **and** moved, identical body); its golden validates the post-rename graph directly, the two-graph diff is covered by unit tests. Verified end-to-end via the CLI: an old-graph ticket against--against latestwarns with the new name+file. 7 core + 4 store + 2 parser-react (real scan) + 3 agent-sdk tests..coderadar/` gitignored. eval 317/0/0/0, determinism 1.000, all metrics
1.000; typecheck + lint clean.
Failure modes: D5, G7
Build: classify generated code (headers like @generated, codegen paths, sourcemap-less minified files): excluded from matching, retained as API metadata. PII policy doc (docs/security.md): screenshots ephemeral, never persisted/embedded/logged; corrections store holds terms only, never images; enforced by a lint test grepping vision-package writes.
Accept: generated-code fixture excluded from match candidates but present in lineage; security doc + lint test in CI. Gate 6 passes.
Done: new isGeneratedFile in the scanner classifies a file as machine-
generated by (a) a __generated__/ or generated/ directory segment or a
.generated./.gen. filename infix, (b) an @generated / DO NOT EDIT /
AUTO-GENERATED banner in the file head, or (c) a sourcemap-less minified line
(≥ 3000 chars). A post-pass tags component and hook nodes in those files with the
generated flag — uniform across function components, class components, and
hooks. matchComponents drops generated components from the candidate pool
(and therefore the IDF corpus), so a screenshot or ticket never resolves to
codegen output, while their nodes/edges — including data sources — stay in the
graph for tracing. New docs/security.md states the G7 PII policy (screenshots
ephemeral, graph/corrections hold no image data) and is enforced by
packages/vision/src/policy.test.ts, which greps the vision source for
filesystem-write and image-persistence APIs and asserts the package never
imports fs. New fixture d5-generated-code (path-classified GlyphCatalog +
banner/filename-classified SchemaViewer + hand-written RevenuePanel): both
generated components trace to their endpoints yet decline on their rendered text,
RevenuePanel matches and isn't poisoned by generated text mixed into the query.
6 parser-react unit tests + 7 vision policy tests. eval 314/0/0/0, determinism
1.000, all metrics 1.000; typecheck + lint clean. Gate 6 criteria met
(6.1/6.2/6.4 still to land for the full lifecycle phase).
Source: external validation of ui-lineage@0.3.0 (MCP id coderadar) against a production
React 18 + Redux Toolkit + RTK Query + MUI + React Router app (664 components, 1,982 instances,
4,843 edges), 2026-07-15. Verdict: trace_lineage, determinism, out-of-domain decline, token
budgeting all solid — but instance resolution hit 28%, RTK Query and object-config routes
produced zero nodes, and a matcher bug made find_component non-discriminating, all while
every gate scores 1.000. The eval suite has a fixture-vs-reality blind spot; this phase fixes
the field defects and closes that blind spot. Prioritized ahead of 6.1–6.5.
Failure modes: A14 (new — added to the catalog in this step), A4, D2, D6
Build: rendered-text targets that normalize to empty (|, /, :, -) currently match
every query — find_component(["zzqwxnomatch12345"]) "matches rendered text", and every call
returns the same ~20 punctuation-rendering components ranked only by the query's own character
rarity. Fix in the matcher: discard rendered-text targets that normalize to empty; require a
minimum alphanumeric token length before a rendered-text target can score. Add a no-signal
decline: when the only surviving matches are empty/punctuation targets, return declined
(reason no-signal) instead of ambiguous with a full candidate list.
Accept: new fixture a14-punctuation-wildcard (components rendering bare | / / / -):
gibberish query → declined: no-signal; a real query ranks the true component top-1 with
punctuation components scoring 0; resolveContext no longer reports high confidence against
punctuation-only matches.
Done: root cause was textMatches (core text.ts): an empty haystack — rendered text like
| normalizes to "" — satisfies needle.includes(haystack) for every query, and a pure-*
template collapses to a match-everything regex the same way. New hasMatchSignal(normalized)
guard (≥ 2 alphanumeric chars, wildcards/spaces excluded) short-circuits textMatches, so
punctuation-only and pure-wildcard targets never match anything; with no surviving targets,
gibberish falls through to the existing declined("no-signal") path (no separate decline
branch needed), and resolveContext inherits the fix through matchComponents. A14 added to
docs/failure-modes.md. New fixture a14-punctuation-wildcard (Divider |, Breadcrumb / -,
CalendarPanel): gibberish declines, a real term ranks CalendarPanel top-1 with no punctuation
noise, gibberish mixed into a real query doesn't poison it. 8 new core unit tests; eval
271/0/0, gate OK.
Failure modes: D2 (the eval blind spot itself)
Build: new fixture eval/fixtures/field-patterns mirroring the shapes the field app used
and current fixtures miss: multi-hop aliased barrel chains (index.ts re-export → tsconfig
paths alias → export { X as Y }), Loadable(lazy(() => import())) pages, an RTK Query
store (createApi + injectEndpoints + builder.query|mutation split across files under
store/api/), object-config createBrowserRouter([...]) assembled from separate route files,
and tests using a custom render wrapper. Goldens assert target behavior (resolution rate,
data-source/route/coverage counts); checks owned by 6F.3–6F.6 land in an explicit skip list and
are enabled by their steps, so pnpm eval stays green throughout.
Accept: fixture scans clean; 6F.1 checks green; skip list names the step that must enable
each remaining check.
Done: new fixture eval/fixtures/field-patterns (11 app files): two-hop barrel + rename
(components/index.ts → ui/index.ts → definition), @ui/@store/* tsconfig-path aliases,
Loadable(lazy(() => import())) page elements, route arrays declared in routes.tsx,
spread-composed, and passed to createBrowserRouter as an imported identifier, an RTK Query
store (createApi base + two injectEndpoints slices with string-form and object-form
query plus a builder.mutation), and a test rendering through a custom
renderWithProviders wrapper. The harness's native xfail is the skip list: 9 target checks
carry expectedFail naming the enabling step (6F.3: DataGrid instance count + blast radius ·
6F.4: two endpoint attributions · 6F.5: three routes, one effect, one journey), and
unexpected-pass gating removes stale marks the moment a capability lands. checks.ts change:
xfail-marked attributions stay out of the lineage precision/recall tallies while marked, so a
known-missing capability doesn't depress global metrics. Scanning the fixture confirmed the
field diagnosis — the @ui alias import yields NO instance node for UsersPage's <Grid/>,
and 0 route/data-source/hook nodes exist. Already working in this shape (kept as passing
checks): relative two-hop barrel rename (InvoicesPage's Grid), covered-by through the custom
wrapper — so 6F.6 must find the actual field-breaking coverage variant and extend the fixture.
eval 280 pass / 0 fail / 9 xfail / 0 unexpected-pass, gate OK, all metrics 1.000.
Failure modes: A5, C1, F2, D4
Build: the field run linked only 563/1,982 instances (28%) — barrel resolution passes unit
tests but real chains break it. Extend import resolution: tsconfig paths aliases, multi-hop
re-export chains, export * from, default-export renames, and
Loadable(lazy(() => import())) / React.lazy wrappers. Unresolved instances keep the
incomplete flag with the failing import path recorded as evidence.
Accept: field-patterns fixture instance→definition resolution ≥ 95%; blastRadius on a
shared component returns its true dependents (field regression: was 0); render graph connects
route roots to leaf instances; enable the 6F.2 checks.
Done: two scanner changes. (1) scanReact now honors the scanned app's own
tsconfig.json (passed as tsConfigFilePath, file discovery still ours) so baseUrl/paths
make alias imports resolvable — @ui → two-hop rename barrel → definition now links, where
before the usage produced no instance node at all. (2) New lazyDefinition resolution in
materializeInstances: a tag naming a module-scope
const X = Loadable(lazy(() => import("./Page"))) (or bare React.lazy) unwraps to the
dynamic import and resolves the target module's default export — exercised by a new
PreviewPane fixture component whose variable name differs from the definition, so only the
unwrap can link it. Fixture golden: 6F.3 marks removed (DataGrid instances = 2, blast radius
DataGrid → UsersPage + InvoicesPage), InvoicesPage now has 1 instance (PreviewPane's lazy
usage). 100% of the fixture's component usages resolve. 2 new parser unit tests (121 total);
eval 283 pass / 0 fail / 7 xfail (all owned by 6F.4–6F.5) / 0 unexpected-pass, gate OK,
metrics 1.000.
Failure modes: B2, C5, C6
Build: createApi / injectEndpoints / builder.query|mutation are unrecognized today —
0 data-source nodes across ~40 store files in the field run, so the entire server-cache
dimension is invisible. Recognize API-slice endpoint definitions as DataSourceNodes (URL from
the query builder, method, tags), link generated hooks (useGetUsersQuery) at component call
sites → per-instance fetches-from edges, and model server-cache selectors
(useSelector on the api reducer path) as StateNode reads.
Accept: field-patterns fixture: every endpoint under store/api/** emits a data-source
node; traceLineage from a component using a generated hook reaches its endpoint; enable the
6F.2 checks.
Done: new rtk-query DataSourceKind in core (schema regenerated, drift gate green). In
detectStores, createApi / api.injectEndpoints declarations now extract endpoints:
rtkBaseCall resolves the (possibly imported) receiver back to its createApi through
chained injectEndpoints hops, rtkBaseUrl reads the fetchBaseQuery baseUrl, and every
builder.query|mutation becomes a DataSourceNode — string-form and object-form ({ url, method }) query configs, template params normalized to :param (e.g.
/api/invoices/:id/pay, POST), baseUrl joined. Generated hooks (useXQuery, useXMutation,
useLazyXQuery) register in a new rtkHooks map, and the component body walk wires
fetches-from edges at call sites. Fixture: 3 data sources, both attribution xfails removed
and passing (UsersPage → /api/users, InvoicesPage → /api/invoices); the unused mutation stays
a consumer-less source. Server-cache useSelector reads remain future work (not needed by
the accept criteria). 4 new parser tests (125 total); eval 285 pass / 0 fail / 5 xfail (all
6F.5) / 0 unexpected-pass, gate OK, metrics 1.000.
Failure modes: B4, B3
Build: step 3.1 covers createBrowserRouter on fixtures, but the field app produced 0
route nodes — diagnose and close the gap: route-object arrays defined in separate
files/variables (not inline at the call site), spread/composed arrays, children nesting,
the lazy: route property, and Loadable(lazy()) page elements. routes-to edges must
survive the lazy wrapper (depends on 6F.3).
Accept: field-patterns fixture: all golden routes emit RouteNodes; journeys("/route")
returns a real multi-step path (field regression: was declined: not-found); enable the 6F.2
checks.
Done: three changes in routes.ts. (1) New routeArrayElements: a router config resolves
whether it's an array literal, an imported identifier (via go-to-definition), or
spread-composed from separately-declared arrays — applied to both the createBrowserRouter
argument and every children property, as/satisfies unwrapped, hop-bounded. (2)
resolveLazyVariable now unwraps wrapper helpers around the lazy call —
Loadable(lazy(() => import())) resolves to the imported page's default export, so
routes-to lands on the real component. (3) No change needed for navigates-to/journeys:
they lit up as soon as route nodes existed. Field-patterns golden: all 5 remaining xfails
removed — 3 routes, the onClick→/users effect, and the "/"→"/users" terminal journey all pass
unmarked, so the fixture is fully green with zero marks and Gate 6F's extractor criteria
are met. 4 new parser tests (129 total); eval 290 pass / 0 fail / 0 xfail / 0
unexpected-pass, gate OK, metrics 1.000.
Failure modes: F3
Build: the field run found 1 covered-by edge app-wide, making untested warnings
near-universal noise. Handle custom render wrappers (renderWithProviders(<X/>) — resolve
through to the JSX argument), test files importing through the same alias/barrel chains as
6F.3, and suites living outside __tests__. When coverage mapping is still near-empty
(< 5% of components covered), replace per-bundle untested warnings with a single graph-level
coverage-unmapped note.
Accept: field-patterns fixture: wrapped renders produce covered-by edges; a near-empty
coverage graph emits coverage-unmapped instead of per-component untested; enable the 6F.2
checks.
Done (defensive half): the coverage-unmapped downgrade shipped — buildBundle now, when
test files exist but < 5% of components carry a covered-by edge, emits one graph-level
coverage-unmapped — only N/M components have mapped test coverage note instead of a
near-universal false untested. A genuinely test-free repo (no test nodes) keeps the accurate
per-component untested. 3 agent-sdk unit tests (near-empty → downgrade · healthy → keep ·
no-tests → keep); verified on the real Grafana graph (34% coverage → per-component untested
preserved). Detection half still open (custom-wrapper/alias/outside-__tests__ resolution):
could NOT reproduce the field failure — the fixture's renderWithProviders shape already maps,
and Grafana produced 1,009 covered-by edges, so detection works on real code. Blocked on a
real failing test file from the tester before building the wrong thing.
Failure modes: A4, D2, D6
Build: weight matches by term specificity against the candidate's own identifiers (name,
props, file path), not global character rarity alone — gibberish must never outscore a real
term (field: 6.50 for gibberish vs 6.09 for "calendar"). Expose a top-line score alongside
confidence on every find_component / resolve_context candidate (currently nested and easy
to misread) — core envelope, CLI output, MCP result schema.
Accept: ranking fixture: identifier-specific terms beat rarity-only scores; score present
at candidate top level across CLI + MCP (schemas regenerated, drift gate green).
Done: matchComponents applies IDENTIFIER_AFFINITY (×1.5) to a matched term that also
names the candidate — its name, props, or file basename, camelCase-split and tokenized
(memoized per component, fuzzy-tolerant) — with the evidence line noting "(also names the
component)". Genuinely-tied generic terms stay tied, so ambiguity honesty is unchanged
(1.000). Candidate gains an optional top-level score (raw ranking score, larger = stronger
within one result; confidence stays the calibrated signal) set by matchComponents and
printed by the CLI (score=… confidence=…); MCP inherits it through the envelope JSON.
Candidate isn't part of the generated schemas, so no schema change (drift gate confirms).
2 new core tests (211 total); eval 290/0/0/0, gate OK, metrics 1.000.
Failure modes: — (user-requested feature, 2026-07-15)
Build: new CLI command visualize -g <graph.json> -o <out.html> emitting a single
self-contained HTML file (graph JSON + all JS/CSS inlined, zero network dependencies) that
renders the lineage graph as a canvas force-directed "galaxy": node kinds color-coded
(component / instance / hook / data-source / state / event / route / external), edges typed.
Interactions: kind + edge-type filter toggles, search with fly-to, click → detail panel (name,
file, loc, props) with neighborhood highlight (visual blast radius) and dim-others, zoom/pan,
physics pause. Must stay responsive at field scale (~2.6k nodes / ~4.8k edges).
Accept: generated from the demo-app graph: opens from file:// with no external requests,
all node kinds rendered and filterable; unit tests on generation (embedded JSON round-trips,
output is self-contained); README section with a screenshot.
Done: new visualize.ts in the CLI package — pure renderVisualization(graph) (graph in,
HTML string out, trivially testable) plus a toViewModel that trims nodes to what the client
needs and drops dangling edges. The command ui-lineage visualize -g <graph> -o <html> writes
one self-contained file: graph JSON embedded (with < → < so rendered text containing
</script> can't break out), all CSS/JS inlined, no network requests. Client is a canvas
force-directed sim with grid-approximated repulsion (neighbor-cell only) so it scales to
thousands of nodes; node radius scales with degree; kinds color-coded with a legend.
Interactions: per-kind node + edge filter toggles (with all/none), search-and-fly, click →
detail panel (name, kind, detail line, file:line, connection count, flags) with neighbor
highlight + dim-others (visual blast radius), drag-to-pin, scroll-zoom, physics pause,
always-labels. Verified live in-browser against the demo-app graph: renders, filters, physics
toggle, and search→select all work with zero console errors; file:///http both load with
no external requests. CLI package gained a test script + vitest; 5 new tests (embed
round-trips, dangling-edge drop, </script> breakout guard, self-contained assertions).
README documents the command. eval unaffected (290/0/0/0). Gate 6F extractor criteria met;
only 6F.6 (coverage, data-blocked) remains before the gate fully closes.
Failure modes: A15 (new), A4, D6
Build: self-found while validating 0.4.0 on Grafana's frontend — find "Find silences by matcher" ranked OrderBySection (renders a bare BY) top-1 over SilencesFilter, because a
lone stopword is a rare literal with high IDF; and find "The" returned confident matches on
"the". New isLowSignal(normalized) in core/text.ts (extends the A14 hasMatchSignal guard):
true when a string is empty, punctuation-only, or entirely stopwords (folded through
foldPlural so it compares to normalized text). matchComponents now drops stopword-only query
terms and skips stopword-only rendered-text targets, so the exact-phrase component wins and a
stopword-only query declines no-signal.
Accept: new fixture a15-stopword-noise (a BY component, an exact-phrase component, a
the-rendering component): the phrase query ranks the exact-phrase component top-1; a stopword
query declines; a stopword mixed with a real term ignores the stopword-rendering component.
Done: implemented + 5 unit tests (2 text, 3 matcher) + fixture. Verified on the real
Grafana graph: Find silences by matcher → SilencesFilter top-1 (OrderBySection gone),
The → no-signal, and resolve on a stopword-heavy ticket went from confidently-wrong to an
honest no-signal. eval 297/0/0/0, gate OK, all metrics 1.000. A16 (HTML-entity rendered text)
spun out to 6F.10 — out of scope for a scoring PR.
Failure modes: A16 (new)
Build: self-found on Grafana — rendered text that is an HTML entity ( , ",
>) normalizes to a junk token (nbsp, 34, gt); numeric entities make gibberish
containing those digits match. Decode/strip HTML entities in the parser's text-extraction pass
(parser-react), or treat entity-only rendered text as low-signal. Not a scoring fix — belongs in
extraction, hence a separate step from 6F.9.
Accept: fixture with entity-only rendered text: it produces no match target; gibberish that
shares digits with a numeric entity declines. eval green.
Done: new entities.ts in parser-react — decodeEntities(text) resolves the HTML entities
React decodes at render time (numeric decimal " / hex " generically, named entities
from a curated map: markup, whitespace, punctuation, symbols, currency, accented Latin-1;
unknown names left verbatim, matching React). extractRenderedText decodes JSX text and quoted
attribute values — the two surfaces React HTML-decodes — while JS string/template literals stay
untouched (React renders {" "} literally). Decoded entities become the character React
renders, which the normalizer strips: →space→dropped, >/"/·→
punctuation that normalizes to empty, so an entity-only component yields no discriminating
target (verified: EntitySpacer.renderedText = ["\"", ">", "·", "<", "›"], zero
alphanumeric tokens). New fixture a16-html-entities (entity-only EntitySpacer + real
QuotaNotice): the named-entity token nbsp, the numeric-entity token 34, and a gibberish
query sharing those digits (zzqwxnomatch12345, which pre-fix matched via "→"34") all
decline no-signal, while the real query still lands on QuotaNotice and isn't poisoned by the
digit-sharing gibberish. 7 new tests (4 unit decode + 3 fixture integration), 136 parser-react
total; eval 304/0/0/0, gate OK, all metrics 1.000.
Gate 6F: field-patterns fixture fully green (skip list empty ✅) · instance resolution ≥ 95%
(✅ 100%) · RTK data sources > 0 (✅) · route nodes > 0 (✅) · gibberish queries decline
no-signal (✅) · pnpm eval green end-to-end (✅ 304/0/0/0). Remaining before the gate is
formally stamped: 6F.6 test-coverage hardening (blocked on a real failing sample).
Sketch level for 7.2–7.6 — detail each before starting, after v1 feedback.
Failure modes: C4
Build: recognize GraphQL operations a component runs (Apollo/urql/graphql-request useQuery/useMutation/useSubscription/useLazyQuery with a gql/graphql tagged-template argument, inline or a resolved const) and emit them as data sources: operation NAME as identity, operation TYPE as method.
Accept: fixture with query/mutation/subscription operations (imported const + inline tag) → one graphql data source each with the right method and a fetches-from edge; react-query useQuery/useMutation (non-gql arg) still classify as react-query.
Done: new graphql.ts — parseGraphqlOperation reads the operation type/name (+ root selection fields as an anonymous-op fallback) from a gql document, ignoring # comments and handling shorthand { … }; graphqlOperationFromArg resolves a hook's first argument to an operation (inline tagged template, or an identifier followed via getDefinitionNodes() so an imported/co-located gql const resolves cross-file). detectDataSource gains a GraphQL branch before the react-query branch (Apollo/urql reuse useQuery/useMutation, so it only claims the call when the argument is an actual gql document, else falls through). Emits sourceKind: "graphql", method = query/mutation/subscription, endpoint = operation name (the value federation/attribution join on) — reusing the existing data-source node + fetches-from wiring, no schema change (graphql was already a DataSourceKind). New fixture c4-graphql (imported-const useQuery/useMutation + inline useSubscription); 9 unit tests incl. the react-query disambiguation on the c5 fixture. eval 326/0/0/0, determinism 1.000, all metrics 1.000; typecheck + lint clean. Codegen response types (feeding 5.5) deferred — needs a .graphql/codegen fixture; the operation→data-source spine is what unblocks federation.
Failure modes: C9
Build: attribute server-side data fetching to the page it feeds — async RSC server components (fetch in their own body) and pages-router getServerSideProps/getStaticProps/getStaticPaths (fetch in a separate exported function).
Accept: fixture with an async RSC page, a getServerSideProps page, and a getStaticProps page → each page attributes its server-side endpoint; the data functions produce no component/hook nodes.
Done: RSC async server components already worked — the async component is recognized and its inline await fetch(...) is caught by the body walk (locked with a test). The gap was pages-router data functions: getServerSideProps/getStaticProps/getStaticPaths are top-level exported functions, not components/hooks, so their fetches were never scanned. The main loop now collects those functions per file and, once the file's default-export page component is known, runs the data-source extraction on their bodies attributed to that page (fetches-from). New fixture c9-nextjs-server-data (async RSC + getServerSideProps + getStaticProps); 4 unit tests. eval 341/0/0/0, determinism 1.000, all metrics 1.000; parser-react 168 tests, typecheck + lint clean. Server actions deferred ("use server" mutations bound to <form action> — a mutation/event surface, fuzzier than the data-in path; revisit with feedback).
Failure modes: C8
Build: recognize server-push channels (new WebSocket(url), new EventSource(url)) as data sources — long-lived connections, not request/response fetches — with a fetches-from edge from the component that opens them.
Accept: fixture with a WebSocket and an SSE channel → one websocket / one sse data source with the URL as endpoint and a fetches-from edge each; no stray sources from unrelated constructors.
Done: new sse DataSourceKind (schema regenerated, drift gate green); new pushchannels.ts — detectPushChannel matches a new WebSocket(...) / new EventSource(...) expression (by constructor name, allowing a window./self. qualifier — a very low false-positive set) and returns a data source with the URL resolved through the existing resolveEndpoint (constant-folded, :param-normalized). extractBodyFacts gains a NewExpression pass that emits the node + fetches-from edge, mirroring the fetch path (skipped inside API wrappers). websocket for WebSocket, sse for EventSource, transport in method (WS/SSE). New fixture c8-push-channels (WebSocket wss://… + SSE /api/…/stream); 4 unit tests. eval 332/0/0/0, determinism 1.000, all metrics 1.000; parser-react 164 tests, typecheck + lint clean. socket.io deferred (io() needs import-aware detection to avoid false positives on the common io identifier).
- 7.4 Python backend parser (C10, G6): FastAPI/Django route decorators →
ServesNode{ method, pattern }; tree-sitter based, emits the same LineageGraph JSON. - 7.5 Go backend parser: mux/gin/echo handler registration →
ServesNode. - 7.6 Federation (G6): endpoint-pattern join across graphs (
fetches-from: /api/users↔serves: /api/users); multi-graph loader in agent-sdk; bundle lineage extends to the owning service + handler file. - Gate 7: cross-repo fixture (React app + FastAPI service):
resolveContextreaches the Python handler.
- C11 — field-level attribution (response field → rendered value)
- F6 — CSS/styling lineage
- E4 — video/GIF frame extraction
| Milestone | Meaning | Phases |
|---|---|---|
| M0 | Measurement loop exists — every later claim is testable | 0 |
| M1 | Survives real-world code patterns | 1 |
| M2 | Headline correct: per-instance API attribution (C1) + handler resolution (B1) | 2 |
| M3 | n-level journeys, lazily expanded | 3 |
| M4 | Screenshot/text → ranked, calibrated, honest matches | 4 |
| M5 ✅ | Pluggable node: ticket in → budgeted context bundle out, over MCP | 5 |
| M6F | Field-hardened: v0.3.0 feedback closed — trustworthy matching + real-world extractors + visualizer | 6F |
| M6 ✅ | Production-grade: incremental, fast, deterministic, versioned | 6 |
| M7 | Full-stack lineage: pixel → backend handler | 7 |