Skip to content

Fix two quadratic walks: tail_reading's fixed point and the particle chain's title test - #572

Merged
derek73 merged 5 commits into
masterfrom
fix/issue-558-559-quadratics
Oct 1, 2026
Merged

derek73 merged 5 commits into
masterfrom
fix/issue-558-559-quadratics

Conversation

@derek73

@derek73 derek73 commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Closes #558, closes #559. Both are found by the _pipeline/ sweep for #553, and both are performance fixes that move no field, report or view.

What changed

#559, _group.py chain(). The ambiguous-particle emitter asked all(is_leading_title(...) for x in range(k)) at every chain site, which costs leading titles × sites. merge(k, j) changes only indices from k on, and k only grows, so every piece behind k is final. A forward-only titled cursor therefore answers the same question with one walk per chain.

#558, _pieces.py tail_reading. The S2-peel / H5-chain fixed point re-ran peel_trailing over the whole walk once per title the chain took. A pass now resumes the walk at the splice, over rest as written:

  • The pieces in front of the splice are the ones a fresh walk meets.
  • credential_anchors is a front-to-back pass, so its carried values still hold for those positions.

Two of the walk's tests read position counts of pieces already peeled, and the splice lowers those counts. So a pass resumes only where neither count can move:

  • kept >= 2: every carried member still counts ≥ 3 for the acronym fork's words-to-spare test.
  • behind_count >= 2: the numeral's last-pair test reads the same pair.

Otherwise the pass walks afresh. A pick that the stopping title made is dropped. The splice is never materialized between passes; its runs are joined once at the end, so no C-level copy goes quadratic either.

The first pass is unchanged (end= replaces a slice), so ordinary names pay no new frame and the _CALL_BASELINE band holds.

Evidence that nothing moves

  • Oracle on every tail_reading call. I wrapped every call site to run the old re-peeling loop next to the new one and compared (rest, titled, names, numeral, picks). The inputs were the differential corpora plus 150,000 fuzzed names under four configurations. One configuration is a lexicon putting ma/ba/sa in the titles, which makes a stopping title an ambiguous-acronym pick. Result: 698,009 calls, 108,453 through the resume path, all identical.
  • Negative controls. Weakening each condition fails the oracle:
    • kept >= 1 fails on phd rev. ma jr..
    • behind_count >= 1 fails on John Smith V Prof. VI.
    • Never dropping the stopper pick fails only under the overlap lexicon, because in the default vocabulary no title is an ambiguous acronym.
  • Whole-parse diff against master. repr, ambiguities and initials() were compared over 121,745 names × 3 configurations: byte-identical, with nameparser.__file__ asserted on each side. An independent review fuzz (72,663 names, plus a 36,012-name chain fuzz with 13,279 particle-or-given reports) also found no difference.
  • Differential gate. Exit 0 at all five baselines (1.4.0, 2.0.0, 2.1.0, 2.2.0, 2.3.0). The output is identical to master, so the classified summaries are too.
  • Checks. Suite 10616 passed, plus mypy and ruff.

Guards

_REREAD_SHAPES in tests/v2/test_benchmark.py counts the frames of the one function each defect re-ran, at 8 units against 32:

shape fixed at fc682e3
tail (listed_lean) 3.67x 14.67x
clause (listed_lean) 3.67x 14.67x
chain (is_leading_title) 3.77x 11.00x

Each row fails against a copy reverting only its own half of the fix, and passes against one reverting only the other.

These are frame-counted rather than _PREFIXED_SHAPES clock rows, which is what #558 suggested. The cost is Python-level, so a frame count is deterministic and keeps running under a line tracer, while that table exists for C-level costs a frame can't see (#553). AGENTS.md's perf gotcha now names the new table.

test_a_title_the_fixed_point_splices_out_keeps_no_peel_pick (tests/v2/pipeline/test_pieces.py) covers the stopper-pick drop, the one diff line no other test reached. The drop is reachable only through a caller's lexicon, so the test adds ma to the titles. Its recorded control: without the drop, John Smith Ma. PhD Jr. keeps its fields but reports suffix-or-name on the title Ma..

Release log

Two bullets, measured on Python 3.11 through HumanName, 400 → 1,600 units:

shape 2.0.0 2.2.0 2.3.0 this branch
#558 tail 4.2x 4.1x 12.9x 4.1x
#559 chain 12.5x 12.2x 12.2x 4.1x

The #558 maiden-clause shape was never released, so no bullet claims it.

Left alone

decisions.md's 2026-09-09 fixed-point entry says John Prof. MA Prof. "reads … family MA". Since #289, mixed-case MA takes the credential, so that reading is now one-case only. The entry is a dated record, so I've left it as it was; the live docstring next to it now names the one-case pair.

🤖 Generated with Claude Code

derek73 and others added 4 commits October 1, 2026 13:56
`chain()` asked `all(is_leading_title(...) for x in range(k))` at every
ambiguous-particle chain site, so leading titles x chain sites
`is_leading_title` calls: `"Dr. " * n + "Jan " + "van Berg " * n` cost
12.2x the time for 4x the input at 2.2.0 and 2.3.0 (12.5x at 2.0.0).

merge(k, j) changes only indices from k on and k only grows, so a piece
behind k is final: a forward-only cursor over the leading-title run
answers the same question with one walk per chain. Full parse output
(repr, ambiguities, initials) is byte-identical to master over the
differential corpora plus 120,000 fuzzed names under three
configurations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The S2 peel / H5 chain fixed point re-ran `peel_trailing` over the
whole walk once per title the chain took, re-reading every suffix the
passes before had peeled: `"John Smith " + "MA Prof. " * n` cost 12.9x
for 4x the input since 2.3.0, and the maiden clause, which runs the
fixed point three times, 14.8x on master.

A pass now resumes the walk at the splice, over `rest` as written. The
pieces in front of the splice are the ones a fresh walk meets, and
`credential_anchors` is a front-to-back pass, so its carried values
hold for them. Two of the walk's tests read position COUNTS of pieces
already peeled, which the splice lowers, so a pass resumes only where
neither can move: two or more pieces in front of the splice (the
acronym fork's words to spare) and two or more peeled pieces behind it
(the numeral's last-pair test). Otherwise it walks afresh. A pick the
stopping title made is dropped, a fresh walk never meeting it.

Checked against the re-peeling loop on every tail_reading call over
the corpora plus 150,000 fuzzed names under four configurations,
including a lexicon putting title words in the ambiguous acronym
class: 698,009 calls, 108,453 through the resume path, all identical.
Weakening each of the three conditions fails that check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`_REREAD_SHAPES` counts the one function each defect re-ran, at 8
units against 32: under 3.8x on this tree, 11.0-14.7x at fc682e3.
Each row fails against a copy reverting only its own half of the fix
and passes against one reverting only the other. Frame-counted rather
than a `_PREFIXED_SHAPES` clock row because the cost is Python-level,
which keeps the guard deterministic and running under a line tracer.

Release-log bullets for both fixes, and AGENTS.md's perf gotcha names
the new table as the home for a prefixed Python-level shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
'John Prof. MA Prof.' never reaches the resume path's words-to-spare
condition: one peeled piece behind the splice already forces a fresh
walk, and since #289 the capitals take a mixed-case MA whatever the
count. 'john prof. ma ma prof.' is the input the condition decides
(a copy resuming at one piece in front reads family 'john', suffix
'ma ma'). The transparency paragraph's 'John Prof. MA' pair went
stale the same way with #289 and now names the one-case spelling,
which still reads family 'ma'. Peel's `anchors` note no longer reads
as a promise about every returned Peel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added this to the 2.4 milestone Oct 1, 2026
@derek73 derek73 self-assigned this Oct 1, 2026
@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.87%. Comparing base (fc682e3) to head (7a4fa38).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #572   +/-   ##
=======================================
  Coverage   98.87%   98.87%           
=======================================
  Files          45       45           
  Lines        3990     4017   +27     
=======================================
+ Hits         3945     3972   +27     
  Misses         45       45           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

The one line of the #558 diff no test reached: a resumed pass drops the
pick the previous walk made at the title the chain then took. No
default title is an ambiguous acronym, so the row configures one
('ma'). Without the drop, 'John Smith Ma. PhD Jr.' keeps its fields and
reports 'suffix-or-name' on the title 'Ma.'.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73
derek73 merged commit 30825f8 into master Oct 1, 2026
10 of 11 checks passed
@derek73
derek73 deleted the fix/issue-558-559-quadratics branch October 1, 2026 21:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

1 participant