Retrieval navigation on a Judge, and real per-page text (HAL-1371, HAL-1375) - #64
Conversation
…ges, no generative call HAL-1371. TreeWalkStrategy navigates with up to eight chat-completion hops, each carrying the structure and every page read so far. But navigation is two judgements over candidate sets the tree already provides. JudgeNavigator asks one Noul per leaf — would the answer be found in this section — in one request, reads the pages of the best few, and asks one Noul per page — does this page contain the specific facts or figures — batched under the request budget. The pages above threshold are the evidence; when none are, the best two are returned with their real probability rather than nothing. JudgeWalkStrategy is the same over a section tree: sections with page ranges are the leaves, a section's body is chunked into page-sized units for the page ranking, and the result names the sections the evidence came from. It writes no answer: answering is the one generative step and belongs to the caller, over the evidence only. Selectable as strategy=judgewalk in cmd/server and cmd/engine whenever llm.judge is configured; the judge is now built before the strategies. cmd/navbench measures it on FinanceBench: for each question, are the gold evidence pages in the evidence set, how many pages were read, at what cost and time.
…s cross-references, and stops prejudging risk sections A 10-K's Item 8 is 70 pages; the 40-page budget cut it off. A coarse pass over page heads (700 chars, ~150 tokens each) now picks which of up to 120 gathered pages deserve their full text. A page two overlapping leaves both cover is gathered once. Item 3 of a 10-K is one sentence that says "see Note 21". When an evidence page refers to a leaf that was not read, the navigator reads it — code finds the reference, the tree names the leaf, the Judge judges its pages — for one more request. The leaf instruction said a risk-factors section "holds neither" the figures nor the discussion; three Boeing questions answered in Risk Factors ranked it last. The instruction now says what each kind of section holds without ruling any out.
HAL-1375. Page text was reconstructed from sections by keying each section's whole content under its first page. A section that starts on page 58 and runs into 59 put page 59's opening under "page 58"; a page that started no section had no entry at all. AMD_2022_10K had 104 page entries for a 121-page PDF. Detection, resolution, the coverage gate and judgewalk's page ranking all worked on that — the navigation bench showed it as evidence one page before the gold page. ParsedDoc.Pages is now built from the parser's own rows, grouped by the physical page they were read from, with the running-header and boilerplate filtering already applied. Ingest and every bench command take pages from it; the section assembly stays only as the fallback for formats without pages. AMD: 121 entries, contiguous.
…text fix 40 FinanceBench questions: 34/40 with every gold page in the evidence set, 37/40 in the right section, ~4 Jev requests and $0.003 each. The per-page fix also took the corpus to 21/21 trees and 47/47 gold pages.
Reviewer's GuideThis PR adds a non-generative Judge-based retrieval path that hierarchically selects sections and evidence pages, makes it available through both engine binaries, and fixes page text construction to preserve physical-page boundaries; accompanying benchmarks, tests, and evaluation documentation measure the resulting FinanceBench behavior. Sequence diagram for Judge-based retrieval navigationsequenceDiagram
participant Caller
participant JudgeWalkStrategy
participant JudgeNavigator
participant Judge
participant PageLoader
Caller->>JudgeWalkStrategy: SelectWithCost(ctx, tree, query, budget)
JudgeWalkStrategy->>JudgeNavigator: Navigate(ctx, query, leaves, loadPages)
JudgeNavigator->>Judge: RankLeaves(ctx, query, leaves)
Judge-->>JudgeNavigator: Leaf scores
JudgeNavigator->>PageLoader: Load pages for best leaves
PageLoader-->>JudgeNavigator: Candidate pages
opt More pages than MaxPages
JudgeNavigator->>Judge: rankPages(ctx, query, page heads, coarse)
Judge-->>JudgeNavigator: Coarse page scores
end
JudgeNavigator->>Judge: RankPages(ctx, query, full pages)
Judge-->>JudgeNavigator: Evidence page scores
opt Evidence references an unread leaf
JudgeNavigator->>PageLoader: Load referenced leaf pages
PageLoader-->>JudgeNavigator: Referenced pages
JudgeNavigator->>Judge: RankPages(ctx, query, referenced pages)
Judge-->>JudgeNavigator: Additional evidence scores
end
JudgeNavigator-->>JudgeWalkStrategy: NavResult
JudgeWalkStrategy-->>Caller: SelectedIDs and cited pages
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
|
Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (17)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Closes HAL-1371, HAL-1375. Likely closes HAL-1365.
Navigation on the Judge.
JudgeNavigatorranks the tree's leaves in one request, gathers the pages of the best five, ranks page heads to pick 40, ranks those in full, and follows "see Note 21" when the tree has that leaf.JudgeWalkStrategyis the same over a section tree, selectable asstrategy=judgewalkin both binaries whenllm.judgeis configured. No generative call; answering is the caller's one call over the evidence.FinanceBench, 40 questions, no chat model (evaluation):
The bug it found. Page text was assembled from sections, keyed by the page each section started on: a section running 58→59 put page 59's opening under "page 58", and pages that started no section had no entry (AMD: 104 entries for 121 pages).
ParsedDoc.Pagesis now built from the parser's own rows per physical page, and ingest plus every bench read it. Re-running the TOC stage on real pages: 21/21 trees (GENERALMILLS's "one-page parse" was this), 47/47 gold pages inside a leaf, title recall 0.975 against the previous trees.Also: the leaf-ranking prompt no longer rules out risk sections (it had ranked Boeing's Risk Factors last for three questions answered there);
cmd/navbench. Full short suite green locally; CI red is the billing lock (HAL-1354).Summary by Sourcery
Enable Judge-based retrieval navigation and correct document page assembly so retrieval can select evidence using accurate physical-page text.
New Features:
judgewalkas a selectable strategy.Bug Fixes:
Enhancements:
Documentation:
Tests:
Summary by CodeRabbit
New Features
judgewalkretrieval strategy, with automatic fallback to tree-based navigation when no Judge is configured.Bug Fixes