Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -383,7 +383,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_

**`_normalize` must reach a fixed point** — storage and match-time share the one fold, and `Lexicon.__setstate__` re-validates, so a value that changes on re-normalization changes under its owner. `strip().strip(".")` alone is not idempotent (`'. a .'` → `' a '` → `'a'`). The loop is the fix; keep any new stripping inside it. **Anything built on `_normalize` must converge too** — `_fold_words` runs `_normalize` per word and DROPS the words that fold away (`_title_key` is that list space-joined, and `_run_addresses_by_given` reads the list itself, so its last-word arm is the last word of the FOLDED key by construction); keeping the empty slot stored `'lt .'` as `'lt '`, a key match-time can never rebuild (so the entry is silently inert) and `__setstate__` rejects on the next round-trip as "not written by this version".

**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), and CI collects coverage in its 3.14 build job alone, so every CI job runs the table; `ja-extra` also sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a line tracer there fails the rows instead of skipping them. **Coverage is collected on ONE interpreter, the one that uploads it** (`COVERAGE_ON` in `.github/workflows/python-package.yml`): the package has no version-conditional code, so one report stands for all of them, and under a line tracer the property grids in `tests/v2/test_properties.py` had cost the 3.12 job most of its ~12 minutes (measured 2026-09-29: the three #397 grids take 139s under coverage's settrace core on 3.12, 24s under its sys.monitoring core, 17s with none). **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently).
**Perf regressions are caught by the scaling test, not the absolute-time ones** — `tests/v2/test_benchmark.py::test_parse_cost_grows_no_worse_than_linearly` times a repeated unit at n vs 4n over a table of shapes (one per pipeline inner loop) and bounds the ratio; the `_thousand_names` tests use constant-size, delimiter-free input and are structurally blind to a complexity regression. Two rules when touching it: calibrate `_MAX_RATIO` against the WEAKEST quadratic's signal (a mixed quadratic surfaces far below the textbook 16×, so the operating point `_BASE` matters more than the bound), and confirm a planted regression fails it across REPEATED runs — one failure is a coin-flip on a timing test. The shapes cover different dimensions (segment count only via `commas`, intra-piece accumulation only via `particles`/`conjunctions`, non-ASCII input only via `honorifics` — every other unit is pure ASCII, so `script_segment` returns at its bail and the CJK stages go unmeasured, M2's clause view only via `maiden_clause`, whose unit has to END on a class member: `MA nee ` holds the same two words, the peel stops at the trailing marker, and the shape reaches nothing — and connective RUN LENGTH only via `link_run`, whose unit has to be MIXED CASE and hold a connective of the generational class: `i und ` reads the letter as an initial and reaches nothing where `i Und ` reaches everything, and S2's anchor pass only via `credential_run`, whose unit has to hold a Title-case member behind an unambiguous credential: `PhD MA ` reads the same fields and never asks the pass, the capitals deciding each member first); measure before pruning one. **A shape whose input needs a PREFIX cannot be a `_SHAPES` row at all**, since that table repeats a unit and nothing else — a maiden clause needs a name word and a marker before the run it is about, and `"nee i Und " * n` reaches the clause rule not at all, measuring the identical ratio on a broken tree and a fixed one. That one is `test_a_clause_link_run_does_not_cost_quadratically`, which builds its own input and counts FRAMES. A prefixed shape whose cost is Python-level joins `_REREAD_SHAPES` beside it — (input builder, the one function the defect re-runs, reachability probe) rows counted in that function's frames at 8 units against 32 — which holds #558's trailing fixed point, its maiden-clause twin, and #559's particle chain, a shape that grows at both ends and so has no unit to repeat. A prefixed shape whose cost is C-level work — a list scanned by `in`, a tuple re-hashed as a dict key — emits no frame for that count to see, so it goes in `_PREFIXED_SHAPES` instead: (prefix, unit, reachability probe) rows timed on the clock like `_SHAPES`, at a base of their own recorded beside the table (#553, whose run shape read 5.97× once at base 800 on the pre-fix tree, under the bound; py3.11, 2026-09-28). **That table skips under a line tracer** (`sys.settrace`, coverage's core below py3.14), which slows every Python line and not the C-level scan, diluting the quadratic below the bound (5.46–5.52× broken at base 1600 under `coverage run` on py3.11, same date). The `sys.monitoring` core coverage uses from 3.14 does not dilute it (8.87–10.91× broken), and CI collects coverage in its 3.14 build job alone, so every CI job runs the table; `ja-extra` also sets `NAMEPARSER_REQUIRE_CLOCK_GUARDS` so that a line tracer there fails the rows instead of skipping them. **Coverage is collected on ONE interpreter, the one that uploads it** (`COVERAGE_ON` in `.github/workflows/python-package.yml`): the package has no version-conditional code, so one report stands for all of them, and under a line tracer the property grids in `tests/v2/test_properties.py` had cost the 3.12 job most of its ~12 minutes (measured 2026-09-29: the three #397 grids take 139s under coverage's settrace core on 3.12, 24s under its sys.monitoring core, 17s with none). **A shape the CLOCK cannot reach needs a FRAME-count guard instead**, which is the second scaling test in that file (`test_a_trailing_credential_run_does_not_cost_exponentially`, #531): where the defect is an exponential rather than a quadratic, the input length that separates the curves on a timing test does not finish, so the guard counts frames over 8 units against 16 and bounds THAT ratio. One pair does not see every curve, and the fix round for #531 measured why: at 2× the input the per-member LINEAR work swamps a quadratic (2.08× for a genuine one against 1.73× clean), so that pair guards the exponential alone and a second, longer pair — 16 against 64, where the same quadratic reads 7.42× against 3.53× clean — is what can see one. Assert them in that order: an exponential never returns from the longer run, so the cheap pair has to have failed first. Frame counts do not move under load, so this shape needs no repeated-run calibration — but it does need the same reachability assertion `_POLICY_SHAPES` rows carry, since the walk under measurement runs only while every unit still reads as a credential. A stage gated on an opt-in `Policy` field needs a `_POLICY_SHAPES` entry instead, since bare `parse()` never enters it — and that table's rows carry a **reachability probe** run before the measurement, because a precedence change can quietly stop the shape reaching the stage and leave a green test measuring a no-op (`_POLICY_SHAPES` is also asserted non-empty: an empty `parametrize` is a skip, not a failure, so deleting its last row would retire the guard silently).

**Expected-failure tests use `@pytest.mark.xfail`** — the conftest parametrized fixture breaks `@unittest.expectedFailure`; always use `@pytest.mark.xfail` instead.

Expand Down
4 changes: 4 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,10 @@ Release Log

- **Fix a long given part after a family comma costing quadratic time.** Since 2.3.0, parsing ``"Doe, Jane " + "Smith " * n`` took time growing with the square of the part's length: going from 1,600 to 6,400 words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (``Doe, Jane Smith Prof. Prof. ...``) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553)

- **Fix a name ending in alternating credentials and titles costing quadratic time.** Since 2.3.0, parsing ``"John Smith " + "MA Prof. " * n`` took time growing with the square of ``n``: going from 400 to 1,600 pairs cost 12.9x the time, where 2.2.0 cost 4x. It costs 4x again (Python 3.11, ``HumanName``, measured 2026-10-01). No field moves (closes #558)

- **Fix many leading titles plus many surname particles costing quadratic time.** Parsing ``"Dr. " * n + "Jan " + "van Berg " * n`` took time growing with the product of the two counts in every 2.x release: going from 400 to 1,600 of each cost 12.2x the time at 2.2.0 and 2.3.0, and 12.5x at 2.0.0. It costs 4x now (Python 3.11, ``HumanName``, measured 2026-10-01). No field moves (closes #559)

**Additions**

- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)
Expand Down
33 changes: 20 additions & 13 deletions nameparser/_pipeline/_group.py
Original file line number Diff line number Diff line change
Expand Up @@ -1724,6 +1724,12 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(),
tail = len(pieces) - trailing_start(name_start, pieces, ptags,
tokens, one_case=one_case)
def chain(tail: int) -> None:
# `pieces[:titled]` are known to be leading titles. merge(k,
# j) changes only indices from k on, and k only grows, so a
# piece behind k is final and its answer can be kept: the
# cursor makes the all-titles-ahead test below one walk per
# chain rather than one per chain site (#559).
titled = 0
k = 0
while k < len(pieces):
if k == leading or not prefix(k):
Expand All @@ -1743,8 +1749,8 @@ def chain(tail: int) -> None:
# A fork whose two sides are decided in different stages
# needs an emitter in each.
#
# Narrow, and #367 is why. `all(is_leading_title(...))`
# says every piece ahead of this one is a title, and the
# Narrow, and #367 is why. `titled == k` says every
# piece ahead of this one is a title, and the
# loop skipped k == leading, so `leading` is STRICTLY
# before k -- and being before k it is one of those titles,
# while being `leading` it satisfies `not title or prefix`.
Expand Down Expand Up @@ -1784,17 +1790,18 @@ def chain(tail: int) -> None:
# an ambiguous particle, while title() is a call per piece.)
if (j > k + 1
and "vocab:particle-ambiguous"
in tokens[pieces[k][0]].tags
and all(is_leading_title(pieces[x], ptags[x],
tokens)
for x in range(k))):
i = pieces[k][0]
ambiguities.append(PendingAmbiguity(
AmbiguityKind.PARTICLE_OR_GIVEN,
f"{tokens[i].text!r} was chained onto the following "
f"name piece; it is also a given name in other "
f"names",
(i,)))
in tokens[pieces[k][0]].tags):
while titled < k and is_leading_title(
pieces[titled], ptags[titled], tokens):
titled += 1
if titled == k:
i = pieces[k][0]
ambiguities.append(PendingAmbiguity(
AmbiguityKind.PARTICLE_OR_GIVEN,
f"{tokens[i].text!r} was chained onto the "
f"following name piece; it is also a given "
f"name in other names",
(i,)))
# rules.md#S2: "A BARE ambiguous acronym is consumed
# only when the name has words to spare — as the second
# of two words it stays the family name — and at the
Expand Down
Loading
Loading