Skip to content

Retrieval navigation on a Judge, and real per-page text (HAL-1371, HAL-1375) - #64

Merged
hallelx2 merged 4 commits into
mainfrom
halleluyaholudele/hal-1371-retrieval-navigation-on-a-judge
Sep 18, 2026
Merged

hallelx2 merged 4 commits into
mainfrom
halleluyaholudele/hal-1371-retrieval-navigation-on-a-judge

Conversation

@hallelx2

@hallelx2 hallelx2 commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Closes HAL-1371, HAL-1375. Likely closes HAL-1365.

Navigation on the Judge. JudgeNavigator ranks the tree's leaves in one request, gathers the pages of the best five, ranks page heads to pick 40, ranks those in full, and follows "see Note 21" when the tree has that leaf. JudgeWalkStrategy is the same over a section tree, selectable as strategy=judgewalk in both binaries when llm.judge is configured. No generative call; answering is the caller's one call over the evidence.

FinanceBench, 40 questions, no chat model (evaluation):

value
every gold evidence page in the evidence set 34 / 40 (0.850)
gold pages inside the chosen section 37 / 40 (0.925)
requests / pages read / cost per question 4.1 / 40.7 / $0.0029

The bug it found. Page text was assembled from sections, keyed by the page each section started on: a section running 58→59 put page 59's opening under "page 58", and pages that started no section had no entry (AMD: 104 entries for 121 pages). ParsedDoc.Pages is now built from the parser's own rows per physical page, and ingest plus every bench read it. Re-running the TOC stage on real pages: 21/21 trees (GENERALMILLS's "one-page parse" was this), 47/47 gold pages inside a leaf, title recall 0.975 against the previous trees.

Also: the leaf-ranking prompt no longer rules out risk sections (it had ranked Boeing's Risk Factors last for three questions answered there); cmd/navbench. Full short suite green locally; CI red is the billing lock (HAL-1354).

Summary by Sourcery

Enable Judge-based retrieval navigation and correct document page assembly so retrieval can select evidence using accurate physical-page text.

New Features:

  • Add Judge-based retrieval navigation that ranks relevant tree leaves and pages, follows applicable cross-references, and exposes judgewalk as a selectable strategy.
  • Add a navigation benchmark for measuring evidence-page recall, section hits, usage, latency, and cost without a generative answering call.

Bug Fixes:

  • Build per-page text from the parser's physical page rows so page boundaries, page numbering, and pages without section starts are preserved correctly.

Enhancements:

  • Use Judge navigation in both the engine and server when configured, with a treewalk fallback when it is unavailable.
  • Expose parser-level page text and update ingestion and benchmarking tools to consume real per-page content.
  • Revise leaf ranking guidance to include risk sections as valid answer sources and support following references such as notes and items.

Documentation:

  • Document the Judge navigation evaluation, page-text bug, benchmark results, remaining misses, and reproduction steps.

Tests:

  • Add coverage for physical-page grouping, parser page exposure, per-page ingestion preference, Judge navigation, batching, cross-reference following, and tree strategy integration.

Summary by CodeRabbit

  • New Features

    • Added Judge-based retrieval navigation for improved evidence selection.
    • Added the judgewalk retrieval strategy, with automatic fallback to tree-based navigation when no Judge is configured.
    • Added per-page PDF text extraction to improve page-level evidence accuracy.
    • Added a navigation benchmark command and FinanceBench evaluation reporting.
  • Bug Fixes

    • Corrected page assembly to preserve physical page boundaries and avoid off-by-one evidence misses.

…ges, no generative call

HAL-1371. TreeWalkStrategy navigates with up to eight chat-completion
hops, each carrying the structure and every page read so far. But
navigation is two judgements over candidate sets the tree already
provides. JudgeNavigator asks one Noul per leaf — would the answer be
found in this section — in one request, reads the pages of the best
few, and asks one Noul per page — does this page contain the specific
facts or figures — batched under the request budget. The pages above
threshold are the evidence; when none are, the best two are returned
with their real probability rather than nothing.

JudgeWalkStrategy is the same over a section tree: sections with page
ranges are the leaves, a section's body is chunked into page-sized
units for the page ranking, and the result names the sections the
evidence came from. It writes no answer: answering is the one
generative step and belongs to the caller, over the evidence only.
Selectable as strategy=judgewalk in cmd/server and cmd/engine whenever
llm.judge is configured; the judge is now built before the strategies.

cmd/navbench measures it on FinanceBench: for each question, are the
gold evidence pages in the evidence set, how many pages were read, at
what cost and time.
…s cross-references, and stops prejudging risk sections

A 10-K's Item 8 is 70 pages; the 40-page budget cut it off. A coarse
pass over page heads (700 chars, ~150 tokens each) now picks which of
up to 120 gathered pages deserve their full text. A page two
overlapping leaves both cover is gathered once.

Item 3 of a 10-K is one sentence that says "see Note 21". When an
evidence page refers to a leaf that was not read, the navigator reads
it — code finds the reference, the tree names the leaf, the Judge
judges its pages — for one more request.

The leaf instruction said a risk-factors section "holds neither" the
figures nor the discussion; three Boeing questions answered in Risk
Factors ranked it last. The instruction now says what each kind of
section holds without ruling any out.
HAL-1375. Page text was reconstructed from sections by keying each
section's whole content under its first page. A section that starts on
page 58 and runs into 59 put page 59's opening under "page 58"; a page
that started no section had no entry at all. AMD_2022_10K had 104 page
entries for a 121-page PDF. Detection, resolution, the coverage gate
and judgewalk's page ranking all worked on that — the navigation bench
showed it as evidence one page before the gold page.

ParsedDoc.Pages is now built from the parser's own rows, grouped by
the physical page they were read from, with the running-header and
boilerplate filtering already applied. Ingest and every bench command
take pages from it; the section assembly stays only as the fallback
for formats without pages. AMD: 121 entries, contiguous.
…text fix

40 FinanceBench questions: 34/40 with every gold page in the evidence
set, 37/40 in the right section, ~4 Jev requests and $0.003 each. The
per-page fix also took the corpus to 21/21 trees and 47/47 gold pages.
Copilot AI lite review requested due to automatic review settings September 18, 2026 18:15

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @hallelx2, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 6 days by commenting @sourcery-ai review. Upgrade to get a review now.

@sourcery-ai

sourcery-ai Bot commented Sep 18, 2026

Copy link
Copy Markdown

Reviewer's Guide

This PR adds a non-generative Judge-based retrieval path that hierarchically selects sections and evidence pages, makes it available through both engine binaries, and fixes page text construction to preserve physical-page boundaries; accompanying benchmarks, tests, and evaluation documentation measure the resulting FinanceBench behavior.

Sequence diagram for Judge-based retrieval navigation

sequenceDiagram
    participant Caller
    participant JudgeWalkStrategy
    participant JudgeNavigator
    participant Judge
    participant PageLoader

    Caller->>JudgeWalkStrategy: SelectWithCost(ctx, tree, query, budget)
    JudgeWalkStrategy->>JudgeNavigator: Navigate(ctx, query, leaves, loadPages)
    JudgeNavigator->>Judge: RankLeaves(ctx, query, leaves)
    Judge-->>JudgeNavigator: Leaf scores
    JudgeNavigator->>PageLoader: Load pages for best leaves
    PageLoader-->>JudgeNavigator: Candidate pages
    opt More pages than MaxPages
        JudgeNavigator->>Judge: rankPages(ctx, query, page heads, coarse)
        Judge-->>JudgeNavigator: Coarse page scores
    end
    JudgeNavigator->>Judge: RankPages(ctx, query, full pages)
    Judge-->>JudgeNavigator: Evidence page scores
    opt Evidence references an unread leaf
        JudgeNavigator->>PageLoader: Load referenced leaf pages
        PageLoader-->>JudgeNavigator: Referenced pages
        JudgeNavigator->>Judge: RankPages(ctx, query, referenced pages)
        Judge-->>JudgeNavigator: Additional evidence scores
    end
    JudgeNavigator-->>JudgeWalkStrategy: NavResult
    JudgeWalkStrategy-->>Caller: SelectedIDs and cited pages
Loading

File-Level Changes

Change Details Files
Adds Judge-driven hierarchical retrieval navigation that ranks TOC leaves, narrows candidate pages with coarse and full-text passes, follows recognized cross-references, and exposes the strategy in both binaries.
  • Introduces JudgeNavigator for batched leaf and page judgments with configurable budgets, thresholds, page limits, and evidence fallback.
  • Adds JudgeWalkStrategy over section trees, including page loading, usage reporting, and strategy=judgewalk selection with treewalk fallback when no Judge is configured.
  • Adds Note/Item-style reference detection and one-hop navigation to unread referenced leaves.
  • Adds FinanceBench navigation benchmark with per-question recall, selected sections, pages read, request, token, latency, and cost metrics.
pkg/retrieval/judgewalk.go
pkg/retrieval/judgewalk_test.go
cmd/engine/main.go
cmd/server/main.go
cmd/navbench/main.go
Reworks parsed document page text to use parser-provided physical pages instead of reconstructing pages from section boundaries.
  • Adds ParsedDoc.Pages and parser Page types populated by grouping filtered PDF rows by physical page.
  • Uses parser pages for TOC ingestion and all benchmark/page-dump/TOC-resolution readers, retaining section assembly only as a fallback for formats without page data.
  • Adds regression coverage for page grouping, parser page exposure, and fallback behavior.
pkg/parser/parser.go
pkg/parser/pdf.go
pkg/parser/pages_test.go
pkg/ingest/ingest.go
pkg/ingest/bench_export.go
pkg/ingest/toc_builder_test.go
cmd/ingestbench/main.go
cmd/jevbench/main.go
cmd/pagedump/main.go
cmd/tocdump/main.go
cmd/tocresolve/main.go
Documents the retrieval evaluation, implementation tradeoffs, and remaining misses, including the impact of correcting page-level text.
  • Records FinanceBench results for leaf selection, evidence recall, request volume, pages read, latency, and cost.
  • Explains the section-based page-numbering bug and validation against real parser pages.
  • Documents coarse page-head ranking, prompt changes, cross-reference following, reproduction commands, and known limitations.
docs/evaluations/2026-09-18-retrieval-navigation-on-a-judge.md

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@coderabbitai

coderabbitai Bot commented Sep 18, 2026

Copy link
Copy Markdown

Review Change StackReview Change Stack

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: e0d4aa98-b48b-41c6-8fad-34c87cca06dc

📥 Commits

Reviewing files that changed from the base of the PR and between 619a997 and b5d94fe.

📒 Files selected for processing (17)
  • cmd/engine/main.go
  • cmd/ingestbench/main.go
  • cmd/jevbench/main.go
  • cmd/navbench/main.go
  • cmd/pagedump/main.go
  • cmd/server/main.go
  • cmd/tocdump/main.go
  • cmd/tocresolve/main.go
  • docs/evaluations/2026-09-18-retrieval-navigation-on-a-judge.md
  • pkg/ingest/bench_export.go
  • pkg/ingest/ingest.go
  • pkg/ingest/toc_builder_test.go
  • pkg/parser/pages_test.go
  • pkg/parser/parser.go
  • pkg/parser/pdf.go
  • pkg/retrieval/judgewalk.go
  • pkg/retrieval/judgewalk_test.go
 __________________________________
< The only good bug is a dead bug. >
 ----------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@hallelx2
hallelx2 merged commit b07a642 into main Sep 18, 2026
1 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants