Repository navigation
feat(sessions): return a hit's whole passage on request - #427
Merged
Merged
Conversation
codegraph_sessions and codegraph sessions take full / --full to return each hit's stored passage (up to 4,000 characters) instead of a 24-token snippet. One byte budget of 16,000 is spent in rank order; hits past it keep their snippet and the answer says how many. Snippets stay the default.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A
codegraph_sessionshit shows about 24 tokens of text. For hosts stored in SQLite (OpenCode, Devin, AGY) there is no transcript file to open for the rest, and for JSONL hosts an agent has to read a large file that also holds the tool traffic the index excludes. The index already stores the whole passage (at most 4,000 characters), so this returns it.full: trueon the MCP tool and--fullon the CLI return each hit's stored passage instead of the snippet. The default is unchanged.--jsoncarriestextper hit andfullCut.Measured on this project's index (35,107 passages): median passage 198 characters, 90th percentile 1,885, 99th 3,702. The top 10 hits for four real queries came to 2.7, 11.5, 11.9 (5 hits) and 3.3 KB in full.
Not in this change: stable hit references and neighboring passages (they need a stored ordinal and a migration), and redaction of secrets in expanded text. Expanded text is what an agent could already read from the transcript; the docs say so.
Tests: three new cases in sessions-index (stored passage beside an unchanged snippet and printed as a quoted block, budget spent as a ranked prefix with the cut count, role filter) and a CLI end-to-end case; the related suites pass (77 tests);
tsc --noEmitclean. Run live on this project, it printed whole passages.README rows checked: the CLI command line and the
codegraph_sessionsrow in the tools table (both updated), plus the site CLI and MCP pages, server instructions and CHANGELOG. No number the README quotes moves.Summary by CodeRabbit
--fullin the CLI orfull: truein the MCP tool. Full passages share a 16,000-byte limit; remaining results keep their snippets, with the response indicating how many were limited.