Skip to content

Commit cf0eae2

Browse files
anvansterclaude
andcommitted
feat(embeddings): split camelCase identifiers by default; give vectors an identity
Natural-language search missed symbols whose names carry the words being searched for. `getUserById` reaches the embedder as one rare token, so a query of "get user by id" had little to match. Identifier splitting Measured on a doc->symbol retrieval eval, BGE-small, pure semantic R@1: camelCase (TypeScript, 377 symbols) 0.355 -> 0.637 +79% PascalCase (Rust, 200 symbols) 0.520 -> 0.640 +23% snake_case (Rust, 723 symbols) 0.683 -> 0.692 +1.3% snake_case (SystemVerilog, 696) 0.510 -> 0.497 -2.5% So it applies only where words run together. snake_case and kebab-case already tokenise into the same words; a leading or trailing delimiter separates nothing, so `_handleClick` splits like `handleClick`. The Rust snake-vs-Pascal pair is the controlled comparison: same repo, same docs, only the casing differs. Exposed as --split-identifiers, codegraph.splitIdentifiers in VS Code and a checkbox in JetBrains, wired through the CLI, MCP builder, engine config, daemon, LSP initializationOptions and the cross-client parity check. Flags that could not be turned off `--full-body-embedding` was `#[arg(long, default_value = "true")]` on a bool, which clap parses as a flag that rejects a value: always true, and `=false` errored. Both it and --split-identifiers now take an optional value. Vector identity Stored vectors had no record of what produced them, so nothing could tell whether they were comparable to the ones a process was about to make. Switching --embedding-model was the sharpest case: there is no dimension check anywhere, and cosine_similarity zips its inputs, so a 768d query against a stored 384d vector silently scored a dot product over the first 384 dimensions against norms of different lengths. Semantic ranking was garbage, with no error, until a manual reindex. Present before this change. Vectors are now stamped with the schema, model, full-body and split settings that built them, and a mismatched set is never loaded. - A rebuild takes ownership of the project once, after it knows it has something to write, in one atomic step that replaces the stored set and writes its stamp. Saves only add; writes are chunked so a save under memory pressure does not triple the footprint; checkpoints keep a crashed rebuild resumable. - A store that cannot be read is left alone. graph.db is one RocksDB shared by every project, and a lock held elsewhere reads exactly like an absent store; treating it as absent claimed the project and cleared a valid set. An unreadable store now embeds in memory for the session. - A live --watch daemon owns its project's vectors; sessions never claim over it, and the daemon carries the same embed-text settings as the sessions that read from it. - An auto-spawned engine is passed --full-body-embedding and --split-identifiers with their values. Both are in the stamp, so an engine started with defaults would have served its client no vectors at all. An index written by 0.20.1 or earlier carries no stamp and is re-embedded the first time this version opens it. A --watch daemon left running across the upgrade keeps writing unstamped vectors that this build will not load; restart it after upgrading. Known trade-off: an LSP client that omits fullBodyEmbedding still defaults it off, unlike every other client. VS Code and JetBrains always send it; a bare nvim/emacs/helix client and an IDE client on the same project would replace each other's vector set. Aligning the default would move every bare client to ~3x slower indexing, so it is left and documented in place. Also: the eval harness read CODEGRAPH_SPLIT_IDS with is_ok(), so `=0` enabled it and an A/B run had two identical arms; StorageBackend gains scan_prefix_keys with a default body so out-of-tree backends still compile. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017rVbt7rENTwXkdHt3Bpgb5
1 parent 1e6084d commit cf0eae2

21 files changed

Lines changed: 1171 additions & 175 deletions

File tree

‎README.md‎

Lines changed: 17 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -94,12 +94,28 @@ one tool and exits without the MCP stdio handshake — ideal for scripting.
9494
| `--workspace <path>` | current dir | Directories to index (repeatable for multi-project) |
9595
| `--exclude <dir>` | — | Directories to skip (repeatable) |
9696
| `--embedding-model <model>` | `bge-small` | `bge-small` (384d, fast), `jina-code-v2` (768d, 6× slower), `granite-97m` (384d, 32K ctx, ~3× slower), or `static` (model2vec, 256d — ~100× faster indexing, no ONNX; needs a local model dir, see below) |
97-
| `--full-body-embedding` | `true` | Embed full function body (~50 lines) for better semantic search and duplicate detection |
97+
| `--full-body-embedding` | `true` | Embed full function body (~50 lines) for better semantic search and duplicate detection. Takes a value: `--full-body-embedding=false` turns it off |
98+
| `--split-identifiers` | `true` | Also embed the word-split form of identifiers whose words are run together, so `getUserById` embeds as "get user by id" too. Names already separated by `_` or `-` are left alone, since they tokenise into the same words; a leading or trailing delimiter separates nothing, so `_handleClick` is split like `handleClick`. Takes a value: `--split-identifiers=false` embeds raw names, as releases up to 0.20.1 did |
9899
| `--max-files <n>` | 5000 | Maximum files to index |
99100
| `--profile <name>` | `all` | Filter the exposed MCP tool surface to a named subset (see below) |
100101
| `--graph-only` | off | Skip embedding generation — build the graph and serve structural tools only. No ONNX model load, 10-50× faster indexing. Semantic search and memory tools unavailable. For CI / one-shot graph queries. |
101102
| `--run-tool <name>` | — | One-shot mode: index, run a single tool, print its result, exit. No MCP handshake. Pair with `--tool-args '<json>'`. |
102103

104+
`--split-identifiers` and `--full-body-embedding` both change the text every
105+
symbol is embedded from, and vectors built from different text cannot be ranked
106+
against each other.
107+
A project's stored vectors are therefore stamped with the settings that built
108+
them - `--embedding-model` included, since models differ in dimension - and are
109+
ignored by any process configured differently.
110+
An index written by 0.20.1 or earlier carries no stamp, so it is re-embedded the
111+
first time this version opens it; changing any of the three re-embeds in the
112+
background rather than requiring a manual reindex.
113+
Run `--watch` with the same flags as the sessions that read the project, so both
114+
sides share one set instead of re-embedding over each other.
115+
A daemon left running across an upgrade keeps writing vectors the new build
116+
cannot match, and sessions attached to it serve without semantic search until it
117+
is restarted - so restart `--watch` after upgrading.
118+
103119
#### `--embedding-model static` — model2vec fast indexing
104120

105121
Static (model2vec) embeddings replace the ONNX transformer with a token→vector

‎crates/codegraph-memory/examples/embed_eval.rs‎

Lines changed: 9 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -168,7 +168,15 @@ fn metrics(ranks: &[usize]) -> Scores {
168168

169169
/// Returns (semantic-only, hybrid 0.4*BM25 + 0.6*cosine) scores.
170170
fn evaluate(engine: &VectorEngine, syms: &[Sym]) -> (Scores, Scores) {
171-
let split = std::env::var("CODEGRAPH_SPLIT_IDS").is_ok();
171+
// Presence is not the question - `CODEGRAPH_SPLIT_IDS=0` must mean off.
172+
// `is_ok()` made setting it to 0 turn the feature ON, which silently turned
173+
// an A/B run into two identical arms. `=1` is the one spelling this harness
174+
// documents, so it is the one spelling it accepts.
175+
//
176+
// Note this forces splitting on every identifier, unlike the engine, which
177+
// applies it only to delimiter-free names: the point of this harness is to
178+
// measure the lever in isolation.
179+
let split = std::env::var("CODEGRAPH_SPLIT_IDS").as_deref() == Ok("1");
172180
let sym_texts: Vec<String> = syms
173181
.iter()
174182
.map(|s| {

0 commit comments

Comments
 (0)