From 60aa9028ee12006bada838f5f1cc140eb579975c Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Mon, 28 Sep 2026 23:07:44 -0700 Subject: [PATCH 1/4] fix(#553): the family-comma given-part walk asks membership of sets The walk over the given part after a family comma asked, once per piece, whether the piece was among the H5 chain's titles (`m in titled_idx`, a tuple) and, where the chain took nothing, whether it was among the pieces the first pass read as name words (`m not in walkable`, a list). Both scan, so a long given part cost time quadratic in its length: 8.4x for 4x the words on 'Doe, Jane ' + 'Smith ' * 1600 and 10x on a trailing title run, since 2.3.0 (py3.11, measured 2026-09-28; 2.2.0 reads 4.1-4.3x on both). The same tuple keyed `floors` and `anchor_memo`, and a tuple re-hashes its whole length at every lookup where a frozenset caches its hash. Both are frozensets now, the ordered list kept as `candidates` only for the two reads that need order (`trailing_titles` and its slice). No reading moves: every consumer asks membership or keys a memo, and the gate exits 0 at every ledger baseline. The frame guards could not see this -- a C-level `in` emits no frame -- and `_SHAPES` cannot express a prefix, so the new _PREFIXED_SHAPES table times (prefix, unit, probe) rows on the clock at base 1600. At `_BASE` the broken run shape read 5.97x once, under `_MAX_RATIO`; at 1600 the broken tree reads 8.43-10.95x against 3.94-4.32x fixed, and each row fails against a copy reverting only its own half of the fix. Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 2 +- docs/release_log.rst | 2 ++ nameparser/_pipeline/_assign.py | 35 ++++++++++++------- tests/v2/test_benchmark.py | 61 +++++++++++++++++++++++++++++---- 4 files changed, 81 insertions(+), 19 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index c32ec491..ae3f3904 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -382,7 +382,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_ **`_normalize` must reach a fixed point** — storage and match-time share the one fold, and `Lexicon.__setstate__` re-validates, so a value that changes on re-normalization changes under its owner. `strip().strip(".")` alone is not idempotent (`'. a .'` → `' a '` → `'a'`). The loop is the fix; keep any new stripping inside it. **Anything built on `_normalize` must converge too** — `_fold_words` runs `_normalize` per word and DROPS the words that fold away (`_title_key` is that list space-joined, and `_run_addresses_by_given` reads the list itself, so its last-word arm is the last word of the FOLDED key by construction); keeping the empty slot stored `'lt .'` as `'lt '`, a key match-time can never rebuild (so the entry is silently inert) and `__setstate__` rejects on the next round-trip as "not written by this version". -**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). +**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work -- a list scanned by `in`, a tuple re-hashed as a dict key -- emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at `_BASE` on the broken tree, under the bound). **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). **Expected-failure tests use `@pytest.mark.xfail`** — the conftest parametrized fixture breaks `@unittest.expectedFailure`; always use `@pytest.mark.xfail` instead. diff --git a/docs/release_log.rst b/docs/release_log.rst index 04430c78..0078c4e8 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -46,6 +46,8 @@ Release Log - **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546) + - **Fix a long given part after a family comma costing quadratic time.** Since 2.3.0, parsing ``"Doe, Jane " + "Smith " * n`` took time growing with the square of the part's length: at 1,600 words, 4x the words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (``Doe, Jane Smith Prof. Prof. ...``) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553) + **Additions** - **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index d6898f0e..f46e7acb 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -621,8 +621,8 @@ def assign(state: ParseState) -> ParseState: # path below, which reads the whole segment as a # credential run and leaves no piece for the walk to # place. - titled_idx: tuple[int, ...] = () - walkable: list[int] = [] + titled_idx: frozenset[int] = frozenset() + walkable: frozenset[int] = frozenset() #: Where the trailing suffix run starts, for the #531 #: report below: `trailing_floor`'s answer, read once on #: the first member the loop meets and -1 until then (no @@ -633,7 +633,7 @@ def assign(state: ParseState) -> ParseState: #: LESS of the walk than the first one did. run_floor = -1 - def previous_kept(m: int, titled: tuple[int, ...]) -> int: + def previous_kept(m: int, titled: frozenset[int]) -> int: """The piece before `m` that the H5 chain did NOT take. Both readings this segment needs are that one: the piece the lenient tail test measures against, and @@ -678,15 +678,15 @@ def previous_kept(m: int, titled: tuple[int, ...]) -> int: #: between the two passes is the role, through `_set_roles`, #: which is a `copy_with(role=...)` and leaves #: text and tags identical. - floors: dict[tuple[int, ...], tuple[int, bool]] = {} + floors: dict[frozenset[int], tuple[int, bool]] = {} #: #544's anchors, per `titled` value like `floors` and for #: the same reason: `given_slot_anchors` over the pieces the #: chain kept, computed once, the first time a member's own #: writing declines, so a run of members is read in one #: forward pass rather than one look-behind per member. - anchor_memo: dict[tuple[int, ...], list[bool]] = {} + anchor_memo: dict[frozenset[int], list[bool]] = {} - def anchored(m: int, titled: tuple[int, ...]) -> bool: + def anchored(m: int, titled: frozenset[int]) -> bool: memo = anchor_memo.get(titled) if memo is None: # nothing in front the pass could read as an @@ -705,7 +705,7 @@ def anchored(m: int, titled: tuple[int, ...]) -> bool: anchor_memo[titled] = memo return memo[m] - def trailing_floor(m: int, titled: tuple[int, ...]) -> int: + def trailing_floor(m: int, titled: frozenset[int]) -> int: """Where the trailing suffix run starts, walked as far down as `m` needs it: `m >= trailing_floor(m, titled)` is exactly "every kept piece behind `m` reads as a @@ -767,7 +767,7 @@ def trailing_floor(m: int, titled: tuple[int, ...]) -> int: floors[titled] = (low, final) return low - def reads_as_a_suffix(m: int, titled: tuple[int, ...]) -> bool: + def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: """Does this segment's walk read piece `m` as a suffix? Asked twice, and by one predicate rather than by two @@ -955,10 +955,21 @@ def reads_as_a_suffix(m: int, titled: tuple[int, ...]) -> bool: # 'Smith, II Mr. V' a middle 'Mr.' where the title is # (24 inputs of that shape move, of 191,146 generated, # measured 2026-09-09). - walkable = [k for k in range(n, len(pieces)) - if k == n or not reads_as_a_suffix(k, ())] - kept = trailing_titles(walkable, pieces, ptags, tokens) - titled_idx = tuple(walkable[kept:]) + # + # Kept as a list only as long as ORDER is asked of it: + # the walk below asks membership, once per piece, and + # a list answers that by scanning, so the walk was + # quadratic in the part's length (#553). The same goes + # for the chain's pieces, which the walk and every + # memo below it key on -- and a tuple, unlike a + # frozenset, re-hashes its whole length at every + # lookup. + candidates = [k for k in range(n, len(pieces)) + if k == n + or not reads_as_a_suffix(k, frozenset())] + kept = trailing_titles(candidates, pieces, ptags, tokens) + walkable = frozenset(candidates) + titled_idx = frozenset(candidates[kept:]) for k in titled_idx: _set_roles(tokens, pieces[k], Role.TITLE) # v1 walk order: the first non-title piece is ALWAYS the diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index 350675ce..62224f40 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -342,13 +342,14 @@ def _best(text: str, parse_: Callable[[str], object], def _assert_grows_linearly(unit: str, - parse_: Callable[[str], object]) -> None: - small = _best(unit * _BASE, parse_) - large = _best(unit * (_BASE * _FACTOR), parse_) + parse_: Callable[[str], object], + prefix: str = "", base: int = _BASE) -> None: + small = _best(prefix + unit * base, parse_) + large = _best(prefix + unit * (base * _FACTOR), parse_) ratio = large / small assert ratio < _MAX_RATIO, ( - f"{unit!r} x{_BASE} took {small * 1e3:.2f}ms, " - f"x{_BASE * _FACTOR} took {large * 1e3:.2f}ms -- {ratio:.1f}x for " + f"{prefix!r} + {unit!r} x{base} took {small * 1e3:.2f}ms, " + f"x{base * _FACTOR} took {large * 1e3:.2f}ms -- {ratio:.1f}x for " f"{_FACTOR}x the input, which is superlinear (linear is ~{_FACTOR})") @@ -357,6 +358,53 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: _assert_grows_linearly(unit, parse) +# Shapes that need a PREFIX, which `_SHAPES` cannot express: it repeats +# a unit and nothing else, and the walk these guard runs only on the +# given part AFTER a family comma. `"a, "` puts the comma in every unit +# and measures segment count instead; `"Smith "` alone never takes the +# comma path. So each row is (prefix, unit, reachability probe), timed +# on the clock like `_SHAPES` -- the defect is C-level work (a list +# scanned by `in`, a tuple re-hashed as a dict key), which emits no +# frame, so the frame-ratio guards below were blind to it (#553). +# +# Base 1600, not `_BASE`, and the recorded negative control is why. +# Measured 2026-09-28 on py3.11, three runs of each on the tree before +# the fix, ratio for 4x the run: +# +# base given_part_run given_part_titles +# 800 6.71 7.60 5.97 7.83 8.31 8.16 +# 1600 8.67 8.43 8.51 10.09 10.95 10.29 +# +# against 3.94-4.32 for both rows at both bases on this tree. At 800 +# the run row's quadratic read 5.97 once, under `_MAX_RATIO` -- a +# coin-flip guard; at 1600 the bound sits ~1.4x over the worst clean +# run and ~1.4x under the weakest broken one, so it did not move. +_PREFIXED_BASE = 1600 +_PREFIXED_SHAPES: dict[str, tuple[str, str, Callable[[str], bool]]] = { + # every word of the run is a middle name, so the walk asks the + # untitled membership test once per piece + "given_part_run": ( + "Doe, Jane ", "Smith ", + lambda text: parse(text).middle.split() == text.split()[2:], + ), + # every trailing title joins the H5 chain, so the walk asks the + # titled membership test once per piece and keys its memos on the + # whole chain + "given_part_titles": ( + "Doe, Jane Smith ", "Prof. ", + lambda text: parse(text).title.split() == text.split()[3:], + ), +} + + +@pytest.mark.parametrize("prefix,unit,reaches", _PREFIXED_SHAPES.values(), + ids=list(_PREFIXED_SHAPES)) +def test_prefixed_cost_grows_no_worse_than_linearly( + prefix: str, unit: str, reaches: Callable[[str], bool]) -> None: + assert reaches(prefix + unit * 4), "shape no longer reaches the walk" + _assert_grows_linearly(unit, parse, prefix=prefix, base=_PREFIXED_BASE) + + # Shapes that need a NON-DEFAULT POLICY to reach the code they guard. # Every _SHAPES entry runs bare parse(), and a stage gated on an opt-in # Policy field is dead there -- #329's clause loop exits on its first @@ -394,13 +442,14 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: def test_shape_tables_are_not_empty() -> None: # pytest turns an EMPTY parametrize into a SKIP, not a failure, so - # deleting the last entry of either table would retire its guard + # deleting the last entry of any of these tables would retire its guard # into the skip count with nothing going red. _POLICY_SHAPES is # the nearer risk, holding only shapes whose stage a default parse # cannot reach at all -- so it gains an entry only when an opt-in # Policy field turns out to have a scaling cliff behind it. assert _SHAPES assert _POLICY_SHAPES + assert _PREFIXED_SHAPES @pytest.mark.parametrize("unit,parser,reaches", _POLICY_SHAPES.values(), From a8998dbf192a380ee79778f6017699d889ab8764 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Mon, 28 Sep 2026 23:30:50 -0700 Subject: [PATCH 2/4] test(#553): the prefixed guard skips under a tracer; probes pin the comma path Coverage slows every Python line and leaves a C-level scan alone, so under `coverage run` (py3.11, 2026-09-28) the pre-fix tree's run row reads 5.46-5.52x at base 1600 -- under `_MAX_RATIO`, passing the very regression it guards -- and the titles row 6.50-6.54x. Base 6400 separates them again (8.19-10.23x broken, 4.21-4.28x fixed) but costs 15.6s a job under coverage against ~1s, so the rows skip while a tracer is installed (settrace or sys.monitoring's coverage tool). CI's ja-extra job runs tests/v2/ on py3.14 without coverage, reading 9.27-10.95x broken against 3.95-4.03x fixed; its workflow step now says it is this guard's only CI home, and AGENTS.md's perf gotcha records the skip. The probes run at the measured size and assert the family-comma reading too: the titles row read the same title on the comma-less path, and would have gone on timing that instead of the walk. Comments: the titles row names its real mechanism (the membership test per piece, plus `previous_kept` walking down through the chain; the memos are never reached), `reads_as_a_suffix`'s docstring says the first pass hands it an empty set, "walkable pass" is "candidates pass", the new block in `_assign.py` says what each container is for, the release-log bullet states its sizes, and the AGENTS.md figure is dated with its base. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/python-package.yml | 6 +++ AGENTS.md | 2 +- docs/release_log.rst | 2 +- nameparser/_pipeline/_assign.py | 31 +++++++------ tests/v2/test_benchmark.py | 67 +++++++++++++++++++++++----- 5 files changed, 81 insertions(+), 27 deletions(-) diff --git a/.github/workflows/python-package.yml b/.github/workflows/python-package.yml index d7a43a36..3b0c4a7c 100644 --- a/.github/workflows/python-package.yml +++ b/.github/workflows/python-package.yml @@ -95,6 +95,12 @@ jobs: # the step still passes green. This line is what makes that # failure loud -- without it the job proves nothing on the day it # matters. + # + # NO COVERAGE here, on purpose: this is the one job that runs the + # _PREFIXED_SHAPES clock guard in tests/v2/test_benchmark.py, + # which skips under a tracer because coverage dilutes the C-level + # cost it measures below its bound (#553). Adding --cov to this + # step retires that guard from CI. run: | uv run python -c "import namedivider" uv run pytest tests/v2/ -q diff --git a/AGENTS.md b/AGENTS.md index ae3f3904..e5557d92 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -382,7 +382,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_ **`_normalize` must reach a fixed point** — storage and match-time share the one fold, and `Lexicon.__setstate__` re-validates, so a value that changes on re-normalization changes under its owner. `strip().strip(".")` alone is not idempotent (`'. a .'` → `' a '` → `'a'`). The loop is the fix; keep any new stripping inside it. **Anything built on `_normalize` must converge too** — `_fold_words` runs `_normalize` per word and DROPS the words that fold away (`_title_key` is that list space-joined, and `_run_addresses_by_given` reads the list itself, so its last-word arm is the last word of the FOLDED key by construction); keeping the empty slot stored `'lt .'` as `'lt '`, a key match-time can never rebuild (so the entry is silently inert) and `__setstate__` rejects on the next round-trip as "not written by this version". -**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work -- a list scanned by `in`, a tuple re-hashed as a dict key -- emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at `_BASE` on the broken tree, under the bound). **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). +**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a tracer**, because coverage slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run`, same date). So CI's `ja-extra` job, the one step running `tests/v2/` without `--cov`, is the only place it guards; the workflow step says so, and adding coverage there retires the guard. **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). **Expected-failure tests use `@pytest.mark.xfail`** — the conftest parametrized fixture breaks `@unittest.expectedFailure`; always use `@pytest.mark.xfail` instead. diff --git a/docs/release_log.rst b/docs/release_log.rst index 0078c4e8..d87a9eac 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -46,7 +46,7 @@ Release Log - **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546) - - **Fix a long given part after a family comma costing quadratic time.** Since 2.3.0, parsing ``"Doe, Jane " + "Smith " * n`` took time growing with the square of the part's length: at 1,600 words, 4x the words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (``Doe, Jane Smith Prof. Prof. ...``) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553) + - **Fix a long given part after a family comma costing quadratic time.** Since 2.3.0, parsing ``"Doe, Jane " + "Smith " * n`` took time growing with the square of the part's length: going from 1,600 to 6,400 words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (``Doe, Jane Smith Prof. Prof. ...``) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553) **Additions** diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index f46e7acb..afcd0784 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -778,9 +778,10 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: site that places them. `titled` is a PARAMETER because the two passes hand it - different values -- () on the first, the chain's own - pieces on the second, which is what makes that second - reading the one 'as if the titled pieces were absent'. + different values -- an empty set on the first, the + chain's own pieces on the second, which is what makes + that second reading the one 'as if the titled pieces + were absent'. A closure over the caller's local said the same thing, but only by WHEN it was rebound. """ @@ -788,7 +789,7 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: return True # rules.md#S2, the given part's trailing slot (#531). # INLINE, and that is the frame budget talking rather - # than taste: the walkable pass asks this closure once + # than taste: the candidates pass asks this closure once # per piece of every family-comma segment, so a helper # call would cost a frame on every non-member piece and # a generator expression would cost its own on 3.11. @@ -801,7 +802,7 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # only it. A member is asked a second time by # trailing_floor()'s descent above, and twice is the # whole of it: the descent takes each piece once and - # the walkable pass asks each piece once, which is what + # the candidates pass asks each piece once, which is what # makes the run linear. What the slot costs, measured # against cc78c960: 'Smith, John', 'Doe, John Q.', # 'Smith, John V', 'Smith, MA' and 'Berg, Jan vd' are @@ -956,14 +957,14 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # (24 inputs of that shape move, of 191,146 generated, # measured 2026-09-09). # - # Kept as a list only as long as ORDER is asked of it: - # the walk below asks membership, once per piece, and - # a list answers that by scanning, so the walk was - # quadratic in the part's length (#553). The same goes - # for the chain's pieces, which the walk and every - # memo below it key on -- and a tuple, unlike a - # frozenset, re-hashes its whole length at every - # lookup. + # `candidates` is the ORDER, which `trailing_titles` and + # its slice read; everything after asks MEMBERSHIP, + # once per piece, and a list or tuple answers that by + # scanning, so the walk was quadratic in the part's + # length (#553). Hence the two sets. A tuple keying + # `floors` or `anchor_memo` above would also re-hash + # its whole length at every lookup, where a frozenset + # caches its hash. candidates = [k for k in range(n, len(pieces)) if k == n or not reads_as_a_suffix(k, frozenset())] @@ -1046,7 +1047,9 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # the frame-ratio test in tests/v2/test_benchmark.py # was structurally blind to it, and the clock-based # shapes beside it repeat ONE unit where this cost - # needs a name holding two runs. What keeps it gone + # needs a name holding two runs (`_PREFIXED_SHAPES` + # adds a prefix but still repeats one run, so it + # does not reach this either). What keeps it gone # is the structure: one walk, read by both callers, # so a second would have to be written on purpose. # diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index 62224f40..3a24fb4d 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -18,6 +18,9 @@ because the stage is gated on an opt-in Policy field that bare parse() leaves empty. A shape guards nothing if the default policy cannot reach the code under it. + +_PREFIXED_SHAPES exists for a shape that needs a prefix before its +repeated run, which _SHAPES cannot express (#553). """ import sys import time @@ -369,16 +372,37 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: # # Base 1600, not `_BASE`, and the recorded negative control is why. # Measured 2026-09-28 on py3.11, three runs of each on the tree before -# the fix, ratio for 4x the run: +# the fix (003b4962), ratio for 4x the run: +# +# base given_part_run given_part_titles +# 800 (_BASE) 6.71 7.60 5.97 7.83 8.31 8.16 +# 1600 8.67 8.43 8.51 10.09 10.95 10.29 +# +# against 3.94-4.32 for both rows at both bases on the fixed tree +# (60aa9028). At 800 the run row's quadratic read 5.97 once, under +# `_MAX_RATIO` -- a coin-flip guard; at 1600 the bound sits ~1.4x over +# the worst clean run and ~1.4x under the weakest broken one, so +# `_MAX_RATIO` did not move. Each row also fails, three runs of three, +# against a copy reverting only ITS half of the fix, and passes against +# a copy reverting only the other half. # -# base given_part_run given_part_titles -# 800 6.71 7.60 5.97 7.83 8.31 8.16 -# 1600 8.67 8.43 8.51 10.09 10.95 10.29 +# SKIPPED UNDER A TRACER, which is every CI build job but `ja-extra`: +# coverage slows every Python line and leaves a C-level scan alone, so +# the quadratic becomes a smaller share of the parse. Measured the same +# day under `coverage run` on py3.11, three runs each: # -# against 3.94-4.32 for both rows at both bases on this tree. At 800 -# the run row's quadratic read 5.97 once, under `_MAX_RATIO` -- a -# coin-flip guard; at 1600 the bound sits ~1.4x over the worst clean -# run and ~1.4x under the weakest broken one, so it did not move. +# base broken: run broken: titles fixed: run / titles +# 1600 5.46 - 5.52 6.50 - 6.54 -- +# 3200 6.61 - 6.70 8.19 - 8.23 4.16 / 4.17 +# 6400 8.19 - 8.29 10.03 - 10.23 4.21-4.28 / 4.22-4.25 +# +# so at this base the run row PASSES a broken tree under coverage. +# 6400 separates the populations again and costs 15.6s a job under +# coverage (4.9s without), against ~1s here; declined for the CI time +# AGENTS.md's grid rules guard. `ja-extra` runs `tests/v2/` on py3.14 +# with no coverage, where this base reads 9.27-10.95 broken against +# 3.95-4.03 fixed -- that job is this guard, and its workflow step +# says so. _PREFIXED_BASE = 1600 _PREFIXED_SHAPES: dict[str, tuple[str, str, Callable[[str], bool]]] = { # every word of the run is a middle name, so the walk asks the @@ -388,8 +412,9 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: lambda text: parse(text).middle.split() == text.split()[2:], ), # every trailing title joins the H5 chain, so the walk asks the - # titled membership test once per piece and keys its memos on the - # whole chain + # titled membership test once per piece, and the name word in + # front of the chain asks it again at every step of walking down + # through it (`previous_kept`) "given_part_titles": ( "Doe, Jane Smith ", "Prof. ", lambda text: parse(text).title.split() == text.split()[3:], @@ -397,11 +422,31 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: } +def _tracer_installed() -> bool: + """A coverage tracer, by either of the two hooks coverage uses: + `sys.settrace` (the C tracer) or, on 3.12+, `sys.monitoring`.""" + if sys.gettrace() is not None: + return True + monitoring = getattr(sys, "monitoring", None) + return (monitoring is not None + and monitoring.get_tool(monitoring.COVERAGE_ID) is not None) + + @pytest.mark.parametrize("prefix,unit,reaches", _PREFIXED_SHAPES.values(), ids=list(_PREFIXED_SHAPES)) def test_prefixed_cost_grows_no_worse_than_linearly( prefix: str, unit: str, reaches: Callable[[str], bool]) -> None: - assert reaches(prefix + unit * 4), "shape no longer reaches the walk" + # at the measured size, and with the comma's own reading asserted: + # without it the titles row reads the same title on the comma-less + # path, and would go on timing that instead of the walk + text = prefix + unit * _PREFIXED_BASE + assert reaches(text), "shape no longer reaches the walk" + name = parse(text) + assert (name.given, name.family) == ("Jane", "Doe"), ( + "shape no longer takes the family-comma path") + if _tracer_installed(): + pytest.skip("a tracer dilutes a C-level cost below the bound; " + "CI's ja-extra job runs this without one") _assert_grows_linearly(unit, parse, prefix=prefix, base=_PREFIXED_BASE) From 73326cea01c7cfa41b3c99d188ab027909d944aa Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Mon, 28 Sep 2026 23:35:23 -0700 Subject: [PATCH 3/4] test(#553): skip only under a line tracer; ja-extra fails rather than skips The skip fired under coverage's sys.monitoring core as well, on figures measured only under its settrace core. Measured 2026-09-28 on GIL py3.14, where coverage.py 7.15 uses sys.monitoring: the pre-fix tree reads 8.87-10.91x at base 1600 under `coverage run` against 3.99-4.08x fixed -- no dilution. So the rows now skip only while `sys.gettrace()` is set, which in CI is the 3.11-3.13 build jobs, and run in the 3.14+ jobs and in ja-extra. And the skip no longer rests on a YAML comment: ja-extra sets NAMEPARSER_REQUIRE_CLOCK_GUARDS, under which a line tracer FAILS these rows instead of skipping them, so adding coverage to that job (or to pytest's addopts) cannot retire the guard in silence. Both modes checked on py3.11 under --cov: two skips without the variable, two failures naming the fix with it. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/python-package.yml | 13 +++++---- AGENTS.md | 2 +- tests/v2/test_benchmark.py | 40 ++++++++++++++-------------- 3 files changed, 29 insertions(+), 26 deletions(-) diff --git a/.github/workflows/python-package.yml b/.github/workflows/python-package.yml index 3b0c4a7c..e16c1497 100644 --- a/.github/workflows/python-package.yml +++ b/.github/workflows/python-package.yml @@ -96,11 +96,14 @@ jobs: # failure loud -- without it the job proves nothing on the day it # matters. # - # NO COVERAGE here, on purpose: this is the one job that runs the - # _PREFIXED_SHAPES clock guard in tests/v2/test_benchmark.py, - # which skips under a tracer because coverage dilutes the C-level - # cost it measures below its bound (#553). Adding --cov to this - # step retires that guard from CI. + # NO COVERAGE here, on purpose, and the env var below enforces it: + # the _PREFIXED_SHAPES clock guard in tests/v2/test_benchmark.py + # skips under a line tracer (coverage's core below py3.14), which + # dilutes the C-level cost it measures below its bound (#553). + # With NAMEPARSER_REQUIRE_CLOCK_GUARDS set, a tracer here fails + # those rows instead of skipping them. + env: + NAMEPARSER_REQUIRE_CLOCK_GUARDS: "1" run: | uv run python -c "import namedivider" uv run pytest tests/v2/ -q diff --git a/AGENTS.md b/AGENTS.md index e5557d92..ee18456d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -382,7 +382,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_ **`_normalize` must reach a fixed point** — storage and match-time share the one fold, and `Lexicon.__setstate__` re-validates, so a value that changes on re-normalization changes under its owner. `strip().strip(".")` alone is not idempotent (`'. a .'` → `' a '` → `'a'`). The loop is the fix; keep any new stripping inside it. **Anything built on `_normalize` must converge too** — `_fold_words` runs `_normalize` per word and DROPS the words that fold away (`_title_key` is that list space-joined, and `_run_addresses_by_given` reads the list itself, so its last-word arm is the last word of the FOLDED key by construction); keeping the empty slot stored `'lt .'` as `'lt '`, a key match-time can never rebuild (so the entry is silently inert) and `__setstate__` rejects on the next round-trip as "not written by this version". -**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a tracer**, because coverage slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run`, same date). So CI's `ja-extra` job, the one step running `tests/v2/` without `--cov`, is the only place it guards; the workflow step says so, and adding coverage there retires the guard. **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). +**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), so the table runs in CI's 3.14+ jobs and in `ja-extra`, which sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a tracer there fails the rows instead of skipping them. **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). **Expected-failure tests use `@pytest.mark.xfail`** — the conftest parametrized fixture breaks `@unittest.expectedFailure`; always use `@pytest.mark.xfail` instead. diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index 3a24fb4d..0fd9d437 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -22,6 +22,7 @@ _PREFIXED_SHAPES exists for a shape that needs a prefix before its repeated run, which _SHAPES cannot express (#553). """ +import os import sys import time from collections.abc import Callable @@ -386,8 +387,9 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: # against a copy reverting only ITS half of the fix, and passes against # a copy reverting only the other half. # -# SKIPPED UNDER A TRACER, which is every CI build job but `ja-extra`: -# coverage slows every Python line and leaves a C-level scan alone, so +# SKIPPED UNDER A LINE TRACER (`sys.settrace`), the core coverage.py +# 7.15 uses below py3.14 -- so CI's 3.11-3.13 build jobs. A line +# tracer slows every Python line and leaves a C-level scan alone, so # the quadratic becomes a smaller share of the parse. Measured the same # day under `coverage run` on py3.11, three runs each: # @@ -396,13 +398,17 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: # 3200 6.61 - 6.70 8.19 - 8.23 4.16 / 4.17 # 6400 8.19 - 8.29 10.03 - 10.23 4.21-4.28 / 4.22-4.25 # -# so at this base the run row PASSES a broken tree under coverage. +# so at this base the run row PASSES a broken tree under that tracer. # 6400 separates the populations again and costs 15.6s a job under # coverage (4.9s without), against ~1s here; declined for the CI time -# AGENTS.md's grid rules guard. `ja-extra` runs `tests/v2/` on py3.14 -# with no coverage, where this base reads 9.27-10.95 broken against -# 3.95-4.03 fixed -- that job is this guard, and its workflow step -# says so. +# AGENTS.md's grid rules guard. The `sys.monitoring` core coverage +# uses from 3.14 does NOT dilute it -- 8.87-10.91 broken against +# 3.99-4.08 fixed at this base, GIL py3.14, same day -- so the rows +# run there, and in `ja-extra`, which runs `tests/v2/` on py3.14 with +# no coverage (9.27-10.95 broken against 3.95-4.03 fixed). That job +# sets NAMEPARSER_REQUIRE_CLOCK_GUARDS, under which a line tracer +# FAILS these rows instead of skipping them: whatever else changes, +# one job cannot retire this guard in silence. _PREFIXED_BASE = 1600 _PREFIXED_SHAPES: dict[str, tuple[str, str, Callable[[str], bool]]] = { # every word of the run is a middle name, so the walk asks the @@ -422,16 +428,6 @@ def test_parse_cost_grows_no_worse_than_linearly(unit: str) -> None: } -def _tracer_installed() -> bool: - """A coverage tracer, by either of the two hooks coverage uses: - `sys.settrace` (the C tracer) or, on 3.12+, `sys.monitoring`.""" - if sys.gettrace() is not None: - return True - monitoring = getattr(sys, "monitoring", None) - return (monitoring is not None - and monitoring.get_tool(monitoring.COVERAGE_ID) is not None) - - @pytest.mark.parametrize("prefix,unit,reaches", _PREFIXED_SHAPES.values(), ids=list(_PREFIXED_SHAPES)) def test_prefixed_cost_grows_no_worse_than_linearly( @@ -444,9 +440,13 @@ def test_prefixed_cost_grows_no_worse_than_linearly( name = parse(text) assert (name.given, name.family) == ("Jane", "Doe"), ( "shape no longer takes the family-comma path") - if _tracer_installed(): - pytest.skip("a tracer dilutes a C-level cost below the bound; " - "CI's ja-extra job runs this without one") + if sys.gettrace() is not None: + reason = ("a line tracer dilutes a C-level cost below the bound " + "(#553)") + if os.environ.get("NAMEPARSER_REQUIRE_CLOCK_GUARDS"): + pytest.fail(f"{reason}, and this run is the one that must " + f"measure it: drop the tracer from this job") + pytest.skip(reason) _assert_grows_linearly(unit, parse, prefix=prefix, base=_PREFIXED_BASE) From fe38249bb98aab73ce6f94a26ce61301da84e02c Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Mon, 28 Sep 2026 23:37:14 -0700 Subject: [PATCH 4/4] docs(#553): the ja-extra env var catches a line tracer, not coverage as such On py3.14 coverage runs its sys.monitoring core, which leaves sys.gettrace() unset and does not dilute the guarded cost, so the workflow comment's "no coverage here, enforced" claimed more than the variable does. It names what the variable catches -- a line tracer, by a Python downgrade, COVERAGE_CORE=ctrace or a debugger -- and AGENTS.md says "line tracer" to match. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/python-package.yml | 14 ++++++++------ AGENTS.md | 2 +- 2 files changed, 9 insertions(+), 7 deletions(-) diff --git a/.github/workflows/python-package.yml b/.github/workflows/python-package.yml index e16c1497..72b2a423 100644 --- a/.github/workflows/python-package.yml +++ b/.github/workflows/python-package.yml @@ -96,12 +96,14 @@ jobs: # failure loud -- without it the job proves nothing on the day it # matters. # - # NO COVERAGE here, on purpose, and the env var below enforces it: - # the _PREFIXED_SHAPES clock guard in tests/v2/test_benchmark.py - # skips under a line tracer (coverage's core below py3.14), which - # dilutes the C-level cost it measures below its bound (#553). - # With NAMEPARSER_REQUIRE_CLOCK_GUARDS set, a tracer here fails - # those rows instead of skipping them. + # The _PREFIXED_SHAPES clock guard in tests/v2/test_benchmark.py + # skips under a LINE tracer (sys.settrace -- coverage's core below + # py3.14), which dilutes the C-level cost it measures below its + # bound (#553). This job must measure it, so the env var below + # makes a line tracer here fail those rows rather than skip them: + # a Python downgrade under coverage, COVERAGE_CORE=ctrace, or a + # debugger. Coverage's sys.monitoring core, its 3.14 default, does + # not dilute the cost and is not caught, nor needs to be. env: NAMEPARSER_REQUIRE_CLOCK_GUARDS: "1" run: | diff --git a/AGENTS.md b/AGENTS.md index ee18456d..dde7b3d9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -382,7 +382,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_ **`_normalize` must reach a fixed point** — storage and match-time share the one fold, and `Lexicon.__setstate__` re-validates, so a value that changes on re-normalization changes under its owner. `strip().strip(".")` alone is not idempotent (`'. a .'` → `' a '` → `'a'`). The loop is the fix; keep any new stripping inside it. **Anything built on `_normalize` must converge too** — `_fold_words` runs `_normalize` per word and DROPS the words that fold away (`_title_key` is that list space-joined, and `_run_addresses_by_given` reads the list itself, so its last-word arm is the last word of the FOLDED key by construction); keeping the empty slot stored `'lt .'` as `'lt '`, a key match-time can never rebuild (so the entry is silently inert) and `__setstate__` rejects on the next round-trip as "not written by this version". -**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), so the table runs in CI's 3.14+ jobs and in `ja-extra`, which sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a tracer there fails the rows instead of skipping them. **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). +**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), so the table runs in CI's 3.14+ jobs and in `ja-extra`, which sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a line tracer there fails the rows instead of skipping them. **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently). **Expected-failure tests use `@pytest.mark.xfail`** — the conftest parametrized fixture breaks `@unittest.expectedFailure`; always use `@pytest.mark.xfail` instead.