Repository navigation
Measured claims in the differential comments go stale unnoticed (a reviewer read "751 names" as fact; the corpus holds 1116) #497
Description
Activity
Scope note — the sweep reaches three files this body does not name, and it
found a fourth stale claim of its own.The magnitude sweep
Work on this issue started by measuring the differential run, because several
comments describe the pinned-wheel worker pass as "multi-minute". Measured
2026-09-03, it is sub-second:measured whole run, --baseline 1.4.00.33s wall whole run, --baseline 2.2.0(the default)0.57s worker's share, timed inside main()0.10s / 0.29s all four baselines back to back 1.97s a run with uv's cache emptied, so the wheel is downloadedunder a second The phrase entered on 2026-08-05 in
7767ba2, at ONE site, with no measurement
recorded beside it, and was copied outward until eight sites carried it:
four intools/differential/compare.py, three intests/v2/test_differential.py,
and aDeclined:bullet indocs/design/decisions.md— which gave the false
cost as "the whole of the reason" for declining precise per-name contest
detection.That bullet is the reason this belongs here rather than in a separate issue: it
is this issue's own defect class — an unmeasured number that nothing recomputes
— reaching its most consequential form, a recorded decision resting on it. The
decline itself survives, rewritten to rest on coverage, which is true: a
run-time check sees one ledger per invocation and only when somebody runs the
tool, where the static predicate covers a rule from the moment it is written.What that means for Scope
The Scope section above names
expected_since_*.toml,
tests/v2/test_ledger_guards.pyandtests/v2/_differential_fixtures.py. The
sweep also had to touchtools/differential/compare.py,
tests/v2/test_differential.pyanddocs/design/decisions.md. Read Scope as
naming where the ORIGINAL three instances were found, not as bounding the class
— a measured claim in a comment goes stale the same way wherever it sits, and
the phrase had spread across three files nobody had reason to check.Two figures in this body are also off
Both were used while planning and both are corrected here rather than silently:
- The corpus population is 1120 deduped entries / 1116 distinct names / 1113
compared at baseline 1.4.0 (seven order-bearing shape-4/5 entries are
dropped by the baseline-minimum skip). "1116" is a NAME count; it is not the
entry count, and the two are easy to interchange. - Timing the worker by calling
_run_workerdirectly over every loaded entry
aborts at a 1.4.0 baseline, for the same reason — those seven entries have
to be dropped first. Wrap_run_workerin a timer and callmain()instead.
Recompute:
time (uv run python tools/differential/compare.py --baseline 1.4.0 >/dev/null).- The corpus population is 1120 deduped entries / 1116 distinct names / 1113
- added a commit that references this issue
on Sep 3, 2026 - added 20 commits that reference this issue
on Oct 8, 2026 - added a commit that references this issue
on Oct 8, 2026
Rationale
The differential harness keeps measured facts in prose — how many names a rule
explains, what shape a name diffs, how large a corpus is. Nothing recomputes
them. Three have now been caught stale in a single session, one of them while
actively misleading a reader.
docs/design/AGENTS.mdalready states the convention: "A count in a datedentry is evidence, not a live fact", and "A count that carries no argument is
better deleted than dated." These sites predate or ignore it.
The three, measured 2026-09-02
1.
tests/v2/_differential_fixtures.py— the comment above_UNCLASSIFIED_NAMESsays the old expression "rescanned all 751 names percandidate name rather than per rule".
len(_CORPUS_NAMES)is 1116. During acode review the same day, a reviewer read 751 off this comment and reported it
as the live corpus population in a review of unrelated code. The count carries a
real argument (why a rescan was expensive — a measured ~400x) so the fix is to
keep the argument and drop the digits: "rescanned every corpus name per
candidate rather than per rule".
2.
_CROSS_RULE_WINNERS, the田中さん IIrow — recorded the diff shape as("given", "suffix"); measured against the 1.4.0 wheel it is{family, given, suffix}(first '田中さん'→'',last 'II'→'田中',suffix ''→'さん, II'). Four sites carried the wrong shape — the row, aneighbouring note, the
_MUST_NOT_MATCHcomment, and two more inexpected_since_1.4.0.toml. Fixed in the #382 branch; listed here becauseof what it exposes, below.
3.
expected_since_1.4.0.toml,fix(comma-family) lone post-comma piece routes to suffix/title, not first— its comment says "this rule now claimsexactly one name, 'Andrews, M.D.'". A differential run at baseline 1.4.0
measures eight. Not fixed; it is the reason this issue exists rather than a
fourth commit on #382.
The structural half, which matters more than the three
test_the_recorded_rule_still_wins_each_contested_namecannot catch a wrongdiff shape. It feeds
classify()the recorded shape and asserts therecorded winner — so a guessed shape agrees with itself forever. That is how
田中さん IIsat wrong under a roster whose own docstring says "The diff shapesare measured against the 1.4.0 wheel, not guessed." The roster is a
mechanisms.md#RECORDED-ROSTERSinstance whose recorded half is unverified inone dimension.
A guard that compares each
_CROSS_RULE_WINNERSkey's shape against a realcomparison would close it. That needs the pinned-wheel worker pass, so it
belongs in the differential run rather than the unit suite — which is the same
placement
undeclared_contestsuses in #382.A trap for whoever does this. "How many names does a rule explain" cannot be
answered by driving
classify()over the names its regex reaches at its owndeclared
fields— that asks which rule would win if a name diffed in thewidest shape the rule admits, and returns 278 for the rule in item 3. A rule
explains a name only when that name ACTUALLY diffs and the rule classifies that
real diff, which requires the 1.4.0 worker pass. The cheap proxy is off by a
factor of thirty here, and it is the obvious thing to reach for.
Scope
expected_since_*.toml,tests/v2/test_ledger_guards.pyandtests/v2/_differential_fixtures.pyfor measured claims in prose — counts ofnames, diff shapes, reach figures, corpus sizes.
outlives the digits, or delete it where it carries no argument. Prefer the
rephrasing —
AGENTS.mdgives "the two shares differ by orders of magnitude"outliving "58% vs 0.65%" as the model.
recur silently.
Recorded roster VALUES are ground truth and must not be edited to match a new
measurement — a moved row is a finding. This issue is about the prose around
them, plus the one guard that would make the values self-checking.
Not in scope
#382's own exemption reasons. Every claim in those was measured in the pass that
wrote them, and each names its recompute.