Repository navigation
Commit cf0eae2
feat(embeddings): split camelCase identifiers by default; give vectors an identity
Natural-language search missed symbols whose names carry the words being
searched for. `getUserById` reaches the embedder as one rare token, so a
query of "get user by id" had little to match.
Identifier splitting
Measured on a doc->symbol retrieval eval, BGE-small, pure semantic R@1:
camelCase (TypeScript, 377 symbols) 0.355 -> 0.637 +79%
PascalCase (Rust, 200 symbols) 0.520 -> 0.640 +23%
snake_case (Rust, 723 symbols) 0.683 -> 0.692 +1.3%
snake_case (SystemVerilog, 696) 0.510 -> 0.497 -2.5%
So it applies only where words run together. snake_case and kebab-case
already tokenise into the same words; a leading or trailing delimiter
separates nothing, so `_handleClick` splits like `handleClick`. The Rust
snake-vs-Pascal pair is the controlled comparison: same repo, same docs,
only the casing differs.
Exposed as --split-identifiers, codegraph.splitIdentifiers in VS Code and a
checkbox in JetBrains, wired through the CLI, MCP builder, engine config,
daemon, LSP initializationOptions and the cross-client parity check.
Flags that could not be turned off
`--full-body-embedding` was `#[arg(long, default_value = "true")]` on a
bool, which clap parses as a flag that rejects a value: always true, and
`=false` errored. Both it and --split-identifiers now take an optional
value.
Vector identity
Stored vectors had no record of what produced them, so nothing could tell
whether they were comparable to the ones a process was about to make.
Switching --embedding-model was the sharpest case: there is no dimension
check anywhere, and cosine_similarity zips its inputs, so a 768d query
against a stored 384d vector silently scored a dot product over the first
384 dimensions against norms of different lengths. Semantic ranking was
garbage, with no error, until a manual reindex. Present before this change.
Vectors are now stamped with the schema, model, full-body and split
settings that built them, and a mismatched set is never loaded.
- A rebuild takes ownership of the project once, after it knows it has
something to write, in one atomic step that replaces the stored set and
writes its stamp. Saves only add; writes are chunked so a save under
memory pressure does not triple the footprint; checkpoints keep a crashed
rebuild resumable.
- A store that cannot be read is left alone. graph.db is one RocksDB shared
by every project, and a lock held elsewhere reads exactly like an absent
store; treating it as absent claimed the project and cleared a valid set.
An unreadable store now embeds in memory for the session.
- A live --watch daemon owns its project's vectors; sessions never claim
over it, and the daemon carries the same embed-text settings as the
sessions that read from it.
- An auto-spawned engine is passed --full-body-embedding and
--split-identifiers with their values. Both are in the stamp, so an engine
started with defaults would have served its client no vectors at all.
An index written by 0.20.1 or earlier carries no stamp and is re-embedded
the first time this version opens it. A --watch daemon left running across
the upgrade keeps writing unstamped vectors that this build will not load;
restart it after upgrading.
Known trade-off: an LSP client that omits fullBodyEmbedding still defaults
it off, unlike every other client. VS Code and JetBrains always send it;
a bare nvim/emacs/helix client and an IDE client on the same project would
replace each other's vector set. Aligning the default would move every
bare client to ~3x slower indexing, so it is left and documented in place.
Also: the eval harness read CODEGRAPH_SPLIT_IDS with is_ok(), so `=0`
enabled it and an A/B run had two identical arms; StorageBackend gains
scan_prefix_keys with a default body so out-of-tree backends still compile.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017rVbt7rENTwXkdHt3Bpgb51 parent 1e6084d commit cf0eae2
21 files changed
Lines changed: 1171 additions & 175 deletions
File tree
- crates
- codegraph-memory/examples
- codegraph-server/src
- ai_query
- mcp
- codegraph/src/storage
- jetbrains
- scripts
- src/main/kotlin/ai/codegraph/jetbrains
- mcp
- server
- settings
- vscode
- src
- telemetry
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
94 | 94 | | |
95 | 95 | | |
96 | 96 | | |
97 | | - | |
| 97 | + | |
| 98 | + | |
98 | 99 | | |
99 | 100 | | |
100 | 101 | | |
101 | 102 | | |
102 | 103 | | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
| 111 | + | |
| 112 | + | |
| 113 | + | |
| 114 | + | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
103 | 119 | | |
104 | 120 | | |
105 | 121 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
168 | 168 | | |
169 | 169 | | |
170 | 170 | | |
171 | | - | |
| 171 | + | |
| 172 | + | |
| 173 | + | |
| 174 | + | |
| 175 | + | |
| 176 | + | |
| 177 | + | |
| 178 | + | |
| 179 | + | |
172 | 180 | | |
173 | 181 | | |
174 | 182 | | |
| |||
0 commit comments