Skip to content

Measured claims in the differential comments go stale unnoticed (a reviewer read "751 names" as fact; the corpus holds 1116) #497

Description

@derek73

Rationale

The differential harness keeps measured facts in prose — how many names a rule
explains, what shape a name diffs, how large a corpus is. Nothing recomputes
them. Three have now been caught stale in a single session, one of them while
actively misleading a reader.

docs/design/AGENTS.md already states the convention: "A count in a dated
entry is evidence, not a live fact"
, and "A count that carries no argument is
better deleted than dated."
These sites predate or ignore it.

The three, measured 2026-09-02

1. tests/v2/_differential_fixtures.py — the comment above
_UNCLASSIFIED_NAMES says the old expression "rescanned all 751 names per
candidate name rather than per rule". len(_CORPUS_NAMES) is 1116. During a
code review the same day, a reviewer read 751 off this comment and reported it
as the live corpus population in a review of unrelated code. The count carries a
real argument (why a rescan was expensive — a measured ~400x) so the fix is to
keep the argument and drop the digits: "rescanned every corpus name per
candidate rather than per rule".

2. _CROSS_RULE_WINNERS, the 田中さん II row — recorded the diff shape as
("given", "suffix"); measured against the 1.4.0 wheel it is
{family, given, suffix} (first '田中さん'→'', last 'II'→'田中',
suffix ''→'さん, II'). Four sites carried the wrong shape — the row, a
neighbouring note, the _MUST_NOT_MATCH comment, and two more in
expected_since_1.4.0.toml. Fixed in the #382 branch; listed here because
of what it exposes, below.

3. expected_since_1.4.0.toml, fix(comma-family) lone post-comma piece routes to suffix/title, not first — its comment says "this rule now claims
exactly one name, 'Andrews, M.D.'". A differential run at baseline 1.4.0
measures eight. Not fixed; it is the reason this issue exists rather than a
fourth commit on #382.

The structural half, which matters more than the three

test_the_recorded_rule_still_wins_each_contested_name cannot catch a wrong
diff shape.
It feeds classify() the recorded shape and asserts the
recorded winner — so a guessed shape agrees with itself forever. That is how
田中さん II sat wrong under a roster whose own docstring says "The diff shapes
are measured against the 1.4.0 wheel, not guessed."
The roster is a
mechanisms.md#RECORDED-ROSTERS instance whose recorded half is unverified in
one dimension.

A guard that compares each _CROSS_RULE_WINNERS key's shape against a real
comparison would close it. That needs the pinned-wheel worker pass, so it
belongs in the differential run rather than the unit suite — which is the same
placement undeclared_contests uses in #382.

A trap for whoever does this. "How many names does a rule explain" cannot be
answered by driving classify() over the names its regex reaches at its own
declared fields — that asks which rule would win if a name diffed in the
widest shape the rule admits, and returns 278 for the rule in item 3. A rule
explains a name only when that name ACTUALLY diffs and the rule classifies that
real diff, which requires the 1.4.0 worker pass. The cheap proxy is off by a
factor of thirty here, and it is the obvious thing to reach for.

Scope

  1. Sweep expected_since_*.toml, tests/v2/test_ledger_guards.py and
    tests/v2/_differential_fixtures.py for measured claims in prose — counts of
    names, diff shapes, reach figures, corpus sizes.
  2. For each: reground it with a stated recompute, or rephrase so the argument
    outlives the digits, or delete it where it carries no argument. Prefer the
    rephrasing — AGENTS.md gives "the two shares differ by orders of magnitude"
    outliving "58% vs 0.65%" as the model.
  3. Add the diff-shape guard to the differential run so item 2's class cannot
    recur silently.

Recorded roster VALUES are ground truth and must not be edited to match a new
measurement — a moved row is a finding. This issue is about the prose around
them, plus the one guard that would make the values self-checking.

Not in scope

#382's own exemption reasons. Every claim in those was measured in the pass that
wrote them, and each names its recompute.

Activity

  1. added this to the v2.3 milestone on Sep 2, 2026
  2. self-assigned this
    on Sep 2, 2026
  3. derek73 commented on Sep 3, 2026

    @derek73
    OwnerAuthor

    Scope note — the sweep reaches three files this body does not name, and it
    found a fourth stale claim of its own.

    The magnitude sweep

    Work on this issue started by measuring the differential run, because several
    comments describe the pinned-wheel worker pass as "multi-minute". Measured
    2026-09-03, it is sub-second:

    measured
    whole run, --baseline 1.4.0 0.33s wall
    whole run, --baseline 2.2.0 (the default) 0.57s
    worker's share, timed inside main() 0.10s / 0.29s
    all four baselines back to back 1.97s
    a run with uv's cache emptied, so the wheel is downloaded under a second

    The phrase entered on 2026-08-05 in 7767ba2, at ONE site, with no measurement
    recorded beside it, and was copied outward until eight sites carried it:
    four in tools/differential/compare.py, three in tests/v2/test_differential.py,
    and a Declined: bullet in docs/design/decisions.md — which gave the false
    cost as "the whole of the reason" for declining precise per-name contest
    detection.

    That bullet is the reason this belongs here rather than in a separate issue: it
    is this issue's own defect class — an unmeasured number that nothing recomputes
    — reaching its most consequential form, a recorded decision resting on it. The
    decline itself survives, rewritten to rest on coverage, which is true: a
    run-time check sees one ledger per invocation and only when somebody runs the
    tool, where the static predicate covers a rule from the moment it is written.

    What that means for Scope

    The Scope section above names expected_since_*.toml,
    tests/v2/test_ledger_guards.py and tests/v2/_differential_fixtures.py. The
    sweep also had to touch tools/differential/compare.py,
    tests/v2/test_differential.py and docs/design/decisions.md. Read Scope as
    naming where the ORIGINAL three instances were found, not as bounding the class
    — a measured claim in a comment goes stale the same way wherever it sits, and
    the phrase had spread across three files nobody had reason to check.

    Two figures in this body are also off

    Both were used while planning and both are corrected here rather than silently:

    • The corpus population is 1120 deduped entries / 1116 distinct names / 1113
      compared at baseline 1.4.0
      (seven order-bearing shape-4/5 entries are
      dropped by the baseline-minimum skip). "1116" is a NAME count; it is not the
      entry count, and the two are easy to interchange.
    • Timing the worker by calling _run_worker directly over every loaded entry
      aborts at a 1.4.0 baseline, for the same reason — those seven entries have
      to be dropped first. Wrap _run_worker in a timer and call main() instead.

    Recompute: time (uv run python tools/differential/compare.py --baseline 1.4.0 >/dev/null).

  4. added 20 commits that reference this issue on Oct 8, 2026
  5. added a commit that references this issue on Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions