From 83e0e9631e9134e707c1725a0984ee7f870fe06d Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 11:28:34 -0700 Subject: [PATCH 1/4] fix(#562): a particle chain unsettles a comma credential run C1's run test left a run whose every member is listed and written in capitals in a mixed-case name to the family-comma path, on the promise that path reads it wholly as credentials. Two particles side by side break that promise: group chains them into one particle run, assign reads the part as name text, and P6 attaches the chain to the family. 'John Smith, PhD DO DO' read given 'PhD', family 'DO DO John Smith'; 'John Smith, DO DO DO' given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a chained member. The settled check now gives way when two particles stand next to each other in the part (classify's own particle test, asked only while the run is still settled), so the count reads the run: two name words before the comma flip it to the credential run, reported. That covers a non-member particle too ('John Smith, PhD vd DO'). 1.4.0 read every one of these names this way. rules.md#C1 states the exception with examples; decisions.md records the choice of #562's fix 1 over fix 2, the measurement and the two neighbours left open. The four 2.x ledgers gain a fix(#562) rule, the settled-grid pin and its two negative controls are re-measured, and the ten #562 exceptions come out of it. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 1 + docs/design/rules.md | 12 ++++-- docs/release_log.rst | 2 + nameparser/_pipeline/_segment.py | 14 ++++++ tests/v2/cases.py | 29 +++++++++++++ tests/v2/test_ledger_guards.py | 42 +++++++++++++++--- tests/v2/test_properties.py | 45 ++++++++------------ tools/differential/corpus_rules.jsonl | 2 + tools/differential/corpus_shapes.jsonl | 2 + tools/differential/expected_since_2.0.0.toml | 22 ++++++++++ tools/differential/expected_since_2.1.0.toml | 22 ++++++++++ tools/differential/expected_since_2.2.0.toml | 22 ++++++++++ tools/differential/expected_since_2.3.0.toml | 22 ++++++++++ 13 files changed, 201 insertions(+), 36 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index a4140b52..4a38d730 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -875,6 +875,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-06 #511 — a suffix value handed to `revise()` derives its entries from its own commas, by the rule a whole name uses. `Parser.revise` sub-parsed each value and forced every harvested token to the named role AFTER the sub-parse had run R1's entry pass, which keys on Role.SUFFIX; a bare 'MD PhD' reads there as a title and a family name, so the pass joined nothing and the field rendered 'MD, PhD'. The fix runs the same pass again: `revise` now sub-parses to a ParseState, forces the role on every non-dropped token, and calls `suffix_entries` — the pass, lifted unedited out of post_rules' tail into a function the two callers share (mechanisms.md#ONE-PREDICATE-PER-QUESTION) — over the forced state, then assembles and harvests as before. So a comma in the value parts two credentials and a space joins them. TWO SPELLINGS of the lifted pass, and the reason is the call budget: a first draft had post_rules call the state-in/state-out wrapper, and the second ParseState build cost three more calls per parse against the band tests/v2/test_benchmark.py holds (py3.11, 2026-09-06, tools/perf/call_count.py: 450 calls/name before the move, 451 with the in-place worker `_mark_suffix_entries` that post_rules now calls, 454 with the draft; the facade band tops at 455.9), so the worker writes in place as every other post rule does and the wrapper exists for the one caller with no token list of its own. NO `_run` HELPER for the same reason: a draft routed `parse()` through a private state builder shared with `revise`, and that cost one frame per parse on the hot path (py3.11, 2026-09-06: 415 calls/name against 414 without it, in a 402-418 band), so the four-line construction is spelled twice and `parse()` is untouched. MEASURED 2026-09-06 over the 1117 distinct names in `tools/differential/corpus*.jsonl`, comparing `revise(p, suffix=p.suffix).suffix` against `p.suffix` under the default parser for every name with a non-empty suffix (368 of the 1117): 38 differed before and 1 after; the differential gate is byte-identical at all four baselines, `revise` not being on the compare path. RECOMPUTE: parse each corpus name, skip an empty suffix, revise the parse with its own suffix, count the names whose suffix moved (the script is in the #511 issue body). THE ONE LEFT is '김민준씨, J.씨', and it is not entry structure: the whole-name parse keeps 'J.씨' one glued suffix token, suffix '씨, J.씨', while the sub-parse of the bare value '씨, J.씨' peels the honorific off the initial, so the revised field renders '씨, J. 씨' — right entries, the spurious comma of before ('씨, J., 씨') gone, one word read differently by the value's own parse than by the whole name's. That is the "classified ON ITS OWN" limit `revise`'s docstring has always recorded, and CJK honorific peeling is a W-rule question this change does not move; pinned in `test_revise_reads_a_glued_honorific_on_its_own`. SUPERSEDES the phd-merge acceptance of 'Ph., D.' on this path (its bullet says how). STALE TAGS, tried and backed out: a draft cleared the sub-parse's own "joined" with the forcing and re-derived it, because the pass only ADDS the tag and `revise(n, family="Jones MD PhD")` carries the sub-parse's between-piece suffix mark on 'PhD' onto a FAMILY token. Measured 2026-09-06 against `330ee55`, the clear also destroyed every WITHIN-piece mark on a non-suffix value — 'D.' of `revise(n, family="John Ph. D. Smith")` lost the merge mark the #436 bullet's DECLINED list calls role-blind and correct for every role — while on a ParsedName only the suffix string view reads "joined" (`_text_for`'s suffix_join gate) — the facade's `_list_for` heals it for every role, but no path puts a revised name into a HumanName, the v1 setters going through `replace()`, so wiring those setters onto `revise` is the change that would show the stale mark, as `last_list == ['Jones', 'MD PhD']` — `initials()` and every field string being identical at both trees for every shape measured. So the tags are kept minus FOLDED_TAG as before; the between-piece mark on a forced non-suffix role is a tag-only oddity that predates this change and stays, and for a suffix value nothing depends on the sub-parse's marks, every pair it joined sharing a bucket with no parting token and the pass setting it again; pinned in `test_revise_keeps_the_sub_parses_within_piece_mark`. Dropped tokens keep their role and tags through the forcing, because the pass filters them by index and assemble omits them. ONE MORE LIMIT, pinned in `test_revise_leaves_a_policy_delimiter_unparted_without_a_tail_segment`, and it is the "classified ON ITS OWN" limit again rather than a rule of revise's: a delimiter the policy names through `extra_suffix_delimiters` is dropped, and so parts entries, only on a segment after a comma that the reading of the words makes a tail, and a value with no comma of its own has none — under `Policy(extra_suffix_delimiters=frozenset({" - "}))` the whole name 'Doe, John, MD PhD - FACS' renders 'MD PhD, FACS' while `revise(n, suffix="MD PhD - FACS")` renders 'MD PhD - FACS', the dash surviving as a token and the forced role making it a suffix word; measured 2026-09-06, a value whose own words read with a tail segment does part ('John Doe, MD - FACS' revises to 'John Doe, MD, FACS') and one after a suffix comma does not ('MD, PhD - FACS' stays), which is how the value's words read as a name deciding it. The round-trip is unaffected, the whole-name view having already rendered that boundary as a comma. DECLINED: a list-valued `revise(p, suffix=["MD PhD", "FACS"])`, which widens the API and leaves the string path where it was; documenting the limit with pins alone, the docstring already recording it and #511's measurement showing it reachable on every space-joined run #436 produced; and a text-only read of the delimiter cores inside `revise`, a second rule for a case no round-trip reaches. Out of scope and left as is: `ParsedName.replace(suffix="MD, PhD")` renders 'MD,, PhD', `replace` whitespace-splitting by contract and the facade setters riding on it for v1 parity. - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. +- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index 077c8546..afefb244 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1702,7 +1702,9 @@ C1. Rationale: a credential run after the comma means the name is in credential (S2), and the part reads as the credential run on that evidence rather than on the count; a word of the class by shape alone carries no such lean, so a part holding one is read by the - count. A decision either way at this comma + count, and so is a part holding two particles side by side, + which S2 joins into one particle run rather than leaving either + to its capitals. A decision either way at this comma is reported; for a run of words the decision is the flip to the credential run, reported once over the whole part. A flip in which no listed word of this class takes part is the exception and is @@ -1718,8 +1720,9 @@ C1. Rationale: a credential run after the comma means the name is in in it: a word of this class read as the credential because a credential in front speaks for it reports, and one its own capitals made the credential does not, so a run whose every such - word is written in capitals reads whole in silence - ('John Smith, PhD MA', 'Smith, PhD MA'), as does a part read as + word is written in capitals, no two particles side by side, + reads whole in silence ('John Smith, PhD MA', 'Smith, PhD MA'), as does a part + read as titles before a lone given name ('Smith, Ms MD Ma'). It is one of TWO places the comma's own decision is reported, the other being the word trailing the given part after it (S2), which is a second decision about a second word and never @@ -1793,6 +1796,9 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, Ed Ma" → suffix="Ed Ma" "Jane Doe, MS LAc" → suffix="MS LAc" "Smith, PhD MEng" → family="Smith" · boundary + "John Smith, PhD DO DO" → suffix="PhD DO DO" + "John Smith, PhD DO DO" → ambiguities=("suffix-or-name",) + "John Smith, PhD vd DO" → suffix="PhD vd DO" "John Smith, X.Y.Z." → suffix="X.Y.Z." "John Smith, X.Y.Z." → ambiguities=() "John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z." diff --git a/docs/release_log.rst b/docs/release_log.rst index b00c54ee..ea51980f 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -16,6 +16,8 @@ Release Log - **Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.** ``HumanName("John Smith, Ed Ma")`` gives first ``John``, last ``Smith``, suffix ``Ed Ma``, where 2.0 through 2.3 gave first ``Ed``, middle ``Ma``, last ``John Smith`` -- 1.4.0's reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (``John Smith, MA``), and ``parse()`` reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report, the capitals having decided it (``John Smith, PhD MA``, ``John Smith, MS MA``). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so ``john smith, v phd ma`` gives first ``john``, last ``smith``, suffix ``v phd ma``, where 2.3 gave first ``v``, middle ``ma``, last ``john smith``. ``john smith, md ma`` and ``John Smith, Ms Ma`` move the same way, where 2.0 through 2.3 gave title ``md``/``Ms`` -- the second is the accepted cost, ``Ms`` read as the suffix word it also is, as ``John Smith, Ms`` alone already reads it -- while one name word before the comma keeps the listing form (``Smith, Ms Ma`` gives title ``Ms``, first ``Ma``). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential's company whatever its case: ``John Smith PhD MEng`` and ``Doe, Jane PhD MEng`` give suffix ``PhD MEng``, the fields every release gave (``PhD, MEng`` through 2.2), now reported, and ``doe, jane v phd do`` gives suffix ``v phd do`` where 2.3.0 gave last ``do doe`` -- a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: ``Wang Ma PhD`` keeps last ``Ma``. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: ``Smith, PhD Ma`` gives last ``Smith``, suffix ``PhD Ma``, where 2.3 gave first ``PhD``, middle ``Ma``. A title that is also a credential (``MD``, ``Ms``) opening that part stays a title and nothing in the part speaks, so ``Smith, MD PhD Ma`` keeps title ``MD``, first ``PhD``, middle ``Ma`` and ``Smith, Ms MD Ma`` title ``Ms MD``, first ``Ma``, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so ``Wang M.Eng.`` gives suffix ``M.Eng.``, as ``Wang M.A.`` does and as 2.0 through 2.3 did. See the #544 entry under ``S2`` in ``docs/design/decisions.md`` (closes #544) + - **Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run.** ``HumanName("John Smith, PhD DO DO")`` gives first ``John``, last ``Smith``, suffix ``PhD DO DO``, where 2.2 and 2.3 gave first ``PhD``, last ``DO DO John Smith`` and 2.0 and 2.1 gave title ``PhD``, first ``DO DO`` -- 1.4.0's reading, restored. ``DO`` is the one acronym that is also a surname particle, and two particles next to each other join into one particle run whatever their capitals, so the capitals no longer settle such a run silently: it is read by the count of name words before the comma, and ``parse()`` reports the call. The same holds when the other particle is a suffix word such as ``vd`` (``John Smith, PhD vd DO`` gives suffix ``PhD vd DO``), and runs like ``John Smith, MD DO DO`` that already read correctly now report too. A single ``DO`` is still left to its capitals (``John Smith, PhD DO`` gives suffix ``PhD DO`` with no report), and one name word before the comma keeps the listing form (``Smith, PhD DO DO`` gives first ``PhD``, last ``DO DO Smith``, as 2.3 did). See the #562 entry under ``C1`` in ``docs/design/decisions.md`` (closes #562) + - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516) diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index e47675e9..eb0efd69 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -306,6 +306,7 @@ def class_run(seg: tuple[int, ...]) -> bool: shaped = 0 pairs = 0 lexicon = state.lexicon + prev_particle = False for i in groups[1]: text = state.tokens[i].text fold = run_word_fold(text, lexicon, state.policy) @@ -337,6 +338,19 @@ def class_run(seg: tuple[int, ...]) -> bool: break else: rest.append(text) + # #562, rules.md#C1: "and so is a part holding two + # particles side by side, which S2 joins into one particle + # run rather than leaving either to its capitals" -- group + # chains the pair ('PhD DO DO', 'PhD vd DO', 'MA vd vd'), + # the family-comma path then reads the chain as name text, + # so the run is the count's to read, whatever its case. + # Asked only while the run is still settled -- nothing + # re-settles it -- with classify's own particle test, since + # group's chain is what the capitals lose to. + if settled: + particle = _normalize(text) in lexicon.particles + settled = not (particle and prev_particle) + prev_particle = particle else: # Every word is a member or left to the suffix predicate. # A run whose every member the WRITING already settles as a diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 9444a053..f69b27e5 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -835,6 +835,35 @@ def _check_cjk_shape_purity(self) -> None: "make. Pinned because the reading now rests on the " "family-comma path alone. 1.4.0 read the same; 2.3.0 " "read given 'PhD', middle 'MA'"), + Case("comma_run_with_a_particle_chain_takes_the_count", + "John Smith, PhD DO DO", + {"given": "John", "family": "Smith", "suffix": "PhD DO DO"}, + ambiguities=("suffix-or-name",), + notes="#562: the two 'DO's stand side by side, so S2 joins " + "them into one particle run rather than leaving either " + "to its capitals, and the row above's stand-down does " + "not apply: the run is the count's, flipped and " + "reported. 1.4.0 read the same; 2.0.0 read title " + "'PhD', given 'DO DO', and 2.3.0 given 'PhD', family " + "'DO DO John Smith'", + shape=3), + Case("comma_run_chained_behind_a_suffix_particle_takes_the_count", + "John Smith, PhD vd DO", + {"given": "John", "family": "Smith", "suffix": "PhD vd DO"}, + ambiguities=("suffix-or-name",), + notes="#562: the particles need not be members of the class " + "-- 'vd' is particle and unambiguous suffix vocabulary, " + "and 'DO' beside it is chained all the same. 1.4.0 read the same; 2.3.0 read given 'PhD', " + "family 'vd DO John Smith'", + shape=3), + Case("comma_run_with_a_lone_trailing_particle_member_stays_settled", + "John Smith, PhD DO", + {"given": "John", "family": "Smith", "suffix": "PhD DO"}, + notes="#562's boundary: one 'DO' beside no other particle " + "is left to its capitals (S2), so the run stays settled " + "and silent, as 'John Smith, PhD MA' is. Untagged: every " + "2.x release misread it, and its fix is #289's (#530), " + "not #562's"), Case("comma_run_by_shape_member_takes_the_count", "John Smith, X.Y.Z. MA", {"given": "John", "family": "Smith", "suffix": "X.Y.Z. MA"}, diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 22603096..909d6496 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1415,6 +1415,11 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: "García Márquez, G.J.R.", "Smith Jr., A.B.", "Smith, J. R. R.", "John Smith X.Y.", "García Márquez, Ms G.J.", "García Márquez, Ed G.J."), + # #562: the shapes the particle chain does not reach. One particle + # beside no other is left to its capitals and stays settled; one + # name word before the comma keeps the listing form, chain or not. + "fix(#562) a comma credential run holding a particle chain reads by the name-word count": + ("John Smith, PhD DO", "Smith, PhD DO DO", "John Smith, PhD MA"), # Policy.unlisted_caps_suffixes is OFF by default, so its whole # population is a probe here: the corpora run at the default, and # a rule of this arc reaching one of these names would mean the @@ -2476,6 +2481,10 @@ class _LatinCopy(NamedTuple): r"García Márquez, G\.J\.", r"García Márquez, PhD G\.J\.", r"John Smith, A\.B\.", r"John Smith, A\.B\. Ph\.D\.", r"John Smith, PhD X\.Y\.", r"John Smith, X\.Y\. P\.Q\."}), + # #562's particle-chain rule, literal-anchored to its two movers: + # what selects them is two particles side by side in a comma + # credential run, which no wordlist spells. + frozenset({"John Smith, PhD DO DO", "John Smith, PhD vd DO"}), frozenset({"Doe, MA Smith", r"J\.A\. K\.D\.", r"Jack X\.Y\.Z\.", r"John Smith J\.u\.n\.i\.o\.r\.", r"John Smith R\.A\.I\.", "Royce, Ed", r"Smith Jr\., A\.B\.", r"Smith, A\.B\.", @@ -3639,8 +3648,11 @@ def _claim(rule: dict) -> _Claim: # added. Reach, not explanation: all three are the contest # fix(#274) is now declared to outrank. # 2026-09-27, #544: 26 -> 27; gains 'doe, jane v phd do'. + # 2026-10-01, #562: 27 -> 29, 'John Smith, PhD DO DO' + # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples + # and case rows. Reach, verified name by name. "fix(#379) a tussenvoegsel after a family comma attaches to the family": - _Claim(27, ('family', 'middle'), "48aafe72402e", None), + _Claim(29, ('family', 'middle'), "36879da9fc51", None), "fix(#380) a trailing vd after a family comma is the tussenvoegsel, not a post-nominal": _Claim(2, ('family', 'suffix'), "ec0d45289dc1", None), # 279 -> 280 with #371, and the growth is corpus, not behavior: @@ -3746,8 +3758,11 @@ def _claim(rule: dict) -> _Claim: # 2026-09-30, #563: 396 -> 408, the twelve names #563's rules.md#C1 # examples and case rows add, every one written with a comma. # Reach, verified name by name. + # 2026-10-01, #562: 408 -> 410, 'John Smith, PhD DO DO' + # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples + # and case rows. Reach, verified name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(408, ('given', 'suffix', 'title'), "6f26a85296fb", None), + _Claim(410, ('given', 'suffix', 'title'), "2b67e34c920c", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3839,8 +3854,11 @@ def _claim(rule: dict) -> _Claim: # 2026-09-30, #563: 396 -> 408, the twelve names #563's rules.md#C1 # examples and case rows add, every one written with a comma. # Reach, verified name by name. + # 2026-10-01, #562: 408 -> 410, 'John Smith, PhD DO DO' + # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples + # and case rows. Reach, verified name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(408, ('family', 'given'), "6f26a85296fb", None), + _Claim(410, ('family', 'given'), "2b67e34c920c", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4627,7 +4645,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-19, #533: 21 -> 24, the same three new corpus # names as the 1.4.0 copy -- the do pair this change added # after a family comma. - _Claim(27, ('_ambiguities', 'family', 'middle'), "48aafe72402e", None), + # 2026-10-01, #562: 27 -> 29, 'John Smith, PhD DO DO' + # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples + # and case rows. Reach, verified name by name. + _Claim(29, ('_ambiguities', 'family', 'middle'), "36879da9fc51", None), # 2026-09-18: 126 -> 131. Five corpus names arrived with # #289/#516's own case rows -- 'J.씨', 'John Smith 田.中.', # '毛泽东, MA', '田中 太郎, MA', '마틴 킹, MA' -- all of them @@ -4932,6 +4953,8 @@ def _claim(rule: dict) -> _Claim: # four role movers and five names whose only diff is the report. "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "3550ea3b5804", ('DEFAULT',)), + "fix(#562) a comma credential run holding a particle chain reads by the name-word count": + _Claim(2, ('_ambiguities', 'family', 'given', 'suffix', 'title'), "29c98be5c6c7", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -5425,6 +5448,8 @@ def _claim(rule: dict) -> _Claim: # four role movers and five names whose only diff is the report. "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "3550ea3b5804", ('DEFAULT',)), + "fix(#562) a comma credential run holding a particle chain reads by the name-word count": + _Claim(2, ('_ambiguities', 'family', 'given', 'suffix'), "29c98be5c6c7", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -5795,7 +5820,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-19, #533: 21 -> 24, the same three new corpus # names as the 1.4.0 copy -- the do pair this change added # after a family comma. - _Claim(27, ('_ambiguities', 'family', 'middle'), "48aafe72402e", None), + # 2026-10-01, #562: 27 -> 29, 'John Smith, PhD DO DO' + # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples + # and case rows. Reach, verified name by name. + _Claim(29, ('_ambiguities', 'family', 'middle'), "36879da9fc51", None), "fix(#424) an unlisted abbreviation is as transparent as a listed title to the leading particle": _Claim(1, ('_ambiguities', 'family', 'given'), "ca7b37af6cf8", None), "fix(#367) a title no longer displaces a leading particle out of the leading position": @@ -6053,6 +6081,8 @@ def _claim(rule: dict) -> _Claim: # four role movers and five names whose only diff is the report. "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "3550ea3b5804", ('DEFAULT',)), + "fix(#562) a comma credential run holding a particle chain reads by the name-word count": + _Claim(2, ('_ambiguities', 'family', 'given', 'suffix', 'title'), "29c98be5c6c7", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -6412,6 +6442,8 @@ def _claim(rule: dict) -> _Claim: # four role movers and five names whose only diff is the report. "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "3550ea3b5804", ('DEFAULT',)), + "fix(#562) a comma credential run holding a particle chain reads by the name-word count": + _Claim(2, ('_ambiguities', 'family', 'given', 'suffix'), "29c98be5c6c7", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index bffcf305..601abc07 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -930,18 +930,7 @@ def test_a_title_first_word_counts_as_a_word() -> None: _SETTLED_MEMBERS = ("MA", "BA", "ED", "DO", "JD", "MENG", "LAC", "X.Y.", *_SETTLED_TITLE_CASE) _SETTLED_WORDS = ("PhD", "MD", "MS", "Jr", "Esq.", "Sr", "III", "Ms") -#: A KNOWN DEFECT (#562): a run ending in two particle members. Group -#: chains the pair into one particle run (P2), which S2 says the -#: capitals no longer decide, so C1's shortcut, which assumes the capitals made every -#: member a credential, settles a run the family-comma path does not -#: read as one -- assign already reads the part as name text. 'John -#: Smith, PhD DO DO' reads given 'PhD', family 'DO DO John Smith' (P6 -#: moving the pair), and 'John Smith, DO DO DO' given 'DO', middle 'DO -#: DO', where the declined flip reads each as suffix. Pinned by count so -#: that the fix fails here and takes the pin out with it. -_SETTLED_EXCEPTION_TAIL = " DO DO" -_SETTLED_EXCEPTIONS = 10 -_SETTLED_COUNT = 3024 +_SETTLED_COUNT = 2994 def _comma_state(text: str) -> ParseState: @@ -971,27 +960,32 @@ def test_a_run_c1_leaves_as_settled_is_read_wholly_as_credentials( two Title-case members and the unlisted dotted 'X.Y.' are there for the controls below: each is a member the mirror must NOT settle, and only such members can show it over-promising. Measured - 2026-09-29: 5,580 texts, 3,024 of them on the settled path (the - pin below), two parses each, 1.1s on 3.11 (`--durations`). + 2026-09-29: 5,580 texts, 3,024 of them on the settled path, two + parses each, 1.1s on 3.11 (`--durations`); 2,994 since #562 (the + pin below, 2026-10-01), whose particle chains ('PhD DO DO') left + the settled path for the count and so are no longer exceptions + here. RECORDED NEGATIVE CONTROL: the two halves of the mirror cover for each other, so removing either alone fails nothing -- forcing the lean to "credential" leaves Title-case members behind `isupper()`, and dropping `isupper()` leaves them to a lean that answers "name". - With both removed it fails on 1,239 texts beyond the exceptions - ('John Smith, MA Ma' reading family 'John Smith', given 'MA', - middle 'Ma'), and the exceptions grow from 10 to 12. RECORDED - NEGATIVE CONTROL for the listed-set half (a member admitted only - by shape has no lean, S2): dropped, it fails on 751 texts ('John - Smith, MA X.Y.' reading family 'John Smith', given 'MA', suffix - 'X.Y.'), with 11 exceptions. 0 and 10 here (measured 2026-09-29). + With both removed ("credential" for any lean at all) it fails on + 1,142 texts ('John Smith, MA Ma' reading family 'John Smith', given + 'MA', middle 'Ma'). RECORDED NEGATIVE CONTROL for the listed-set + half (a member admitted only by shape has no lean, S2): dropped, it + fails on 215 texts ('John Smith, PhD X.Y.' reading 'PhD' as name + text). 0 here (measured 2026-10-01, #562). Over the 2026-09-29 + tree, before #563 and #562, the two read 1,239 and 751 beyond ten + pinned #562 exceptions; #563 moved the second to 215 and #562 the + first to 1,142. """ texts = [f"John Smith, {' '.join(words)}" for n in (2, 3) for words in itertools.product( _SETTLED_MEMBERS + _SETTLED_WORDS, repeat=n) if any(w in _SETTLED_MEMBERS for w in words)] - settled, exceptions, failures = 0, [], [] + settled, failures = 0, [] titled: list[str] = [] for text in texts: with monkeypatch.context() as m: @@ -1008,16 +1002,11 @@ def test_a_run_c1_leaves_as_settled_is_read_wholly_as_credentials( named = [t.text for t in state.tokens if t.span.start > comma and t.role not in (Role.SUFFIX, Role.TITLE)] - if not named: - continue - if text.endswith(_SETTLED_EXCEPTION_TAIL): - exceptions.append(text) - else: + if named: failures.append(f"{text!r}: {named} read as name text") assert not failures, ( f"{len(failures)} run(s) C1 left as settled were not read as " f"credentials: {failures[:5]}") - assert len(exceptions) == _SETTLED_EXCEPTIONS, exceptions # a Title-case member is never settled: the writing leans it a name. # The count pin below would catch one too, short of a compensating # move; this assert is the message that names it. diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 475c5850..895bb3e1 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -215,8 +215,10 @@ "John Smith, Mr." "John Smith, Mr. Jr." "John Smith, PhD" +"John Smith, PhD DO DO" "John Smith, PhD MEng" "John Smith, PhD X.Y." +"John Smith, PhD vd DO" "John Smith, V." "John Smith, X.Y. P.Q." "John Smith, X.Y.Z." diff --git a/tools/differential/corpus_shapes.jsonl b/tools/differential/corpus_shapes.jsonl index dc9c8aec..a7801994 100644 --- a/tools/differential/corpus_shapes.jsonl +++ b/tools/differential/corpus_shapes.jsonl @@ -247,9 +247,11 @@ {"name": "John Smith, MEng PhD", "shape": 3} {"name": "John Smith, Ms Ma", "shape": 3} {"name": "John Smith, PhD", "shape": 3} +{"name": "John Smith, PhD DO DO", "shape": 3} {"name": "John Smith, PhD MEng", "shape": 3} {"name": "John Smith, PhD Ma", "shape": 3} {"name": "John Smith, PhD X.Y.", "shape": 3} +{"name": "John Smith, PhD vd DO", "shape": 3} {"name": "John Smith, X.Y. P.Q.", "shape": 3} {"name": "John Smith, X.Y.Z.", "shape": 3} {"name": "John Smith, X.Y.Z. MA", "shape": 3} diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index c760aa7b..47440953 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2574,6 +2574,28 @@ name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\ fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#562) a comma credential run holding a particle chain reads by the name-word count" +# rules.md#C1: a part whose every member is a listed word in capitals +# in a mixed-case name reads as the credential run on that writing +# rather than on the count, "and so is a part holding two particles +# side by side, which S2 joins into one particle run rather than +# leaving either to its capitals" -- read by the count, that is, +# which two name words before the comma flip to the credential run, +# reported. The family-comma path had read the part as name text: +# this baseline read title 'PhD', given 'DO DO', family 'John +# Smith'. 'John Smith, PhD vd DO' is the same chain with a particle +# that is not a member ('vd', particle and unambiguous suffix +# vocabulary). 1.4.0 read both as suffixes, so that ledger has no +# rule. +# +# Literal. Probes: 'John Smith, PhD DO' (one particle beside no other +# stays settled and silent) and 'Smith, PhD DO DO' (one name word +# before the comma keeps the listing form) are _MUST_NOT_MATCH. +name_regex = "^(?:John Smith, PhD DO DO|John Smith, PhD vd DO)$" +fields = ["title", "given", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] + [[change]] issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" # rules.md#C1: "Paired initials are the exception to the count" -- diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index b4e5381e..ec88d86c 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2461,6 +2461,28 @@ name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\ fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#562) a comma credential run holding a particle chain reads by the name-word count" +# rules.md#C1: a part whose every member is a listed word in capitals +# in a mixed-case name reads as the credential run on that writing +# rather than on the count, "and so is a part holding two particles +# side by side, which S2 joins into one particle run rather than +# leaving either to its capitals" -- read by the count, that is, +# which two name words before the comma flip to the credential run, +# reported. The family-comma path had read the part as name text: +# this baseline read title 'PhD', given 'DO DO', family 'John +# Smith'. 'John Smith, PhD vd DO' is the same chain with a particle +# that is not a member ('vd', particle and unambiguous suffix +# vocabulary). 1.4.0 read both as suffixes, so that ledger has no +# rule. +# +# Literal. Probes: 'John Smith, PhD DO' (one particle beside no other +# stays settled and silent) and 'Smith, PhD DO DO' (one name word +# before the comma keeps the listing form) are _MUST_NOT_MATCH. +name_regex = "^(?:John Smith, PhD DO DO|John Smith, PhD vd DO)$" +fields = ["title", "given", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] + [[change]] issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" # rules.md#C1: "Paired initials are the exception to the count" -- diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index bc8254b1..b181dd05 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1050,6 +1050,28 @@ name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\ fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#562) a comma credential run holding a particle chain reads by the name-word count" +# rules.md#C1: a part whose every member is a listed word in capitals +# in a mixed-case name reads as the credential run on that writing +# rather than on the count, "and so is a part holding two particles +# side by side, which S2 joins into one particle run rather than +# leaving either to its capitals" -- read by the count, that is, +# which two name words before the comma flip to the credential run, +# reported. The family-comma path had read the part as name text: +# this baseline read given 'PhD', family 'DO DO John Smith', P6 +# attaching the chain. 'John Smith, PhD vd DO' is the same chain +# with a particle that is not a member ('vd', particle and +# unambiguous suffix vocabulary). 1.4.0 read both as suffixes, so +# that ledger has no rule. +# +# Literal. Probes: 'John Smith, PhD DO' (one particle beside no other +# stays settled and silent) and 'Smith, PhD DO DO' (one name word +# before the comma keeps the listing form) are _MUST_NOT_MATCH. +name_regex = "^(?:John Smith, PhD DO DO|John Smith, PhD vd DO)$" +fields = ["given", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] + [[change]] issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" # rules.md#C1: "Paired initials are the exception to the count" -- diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 9f8670fa..a202b3e5 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -338,6 +338,28 @@ name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\ fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#562) a comma credential run holding a particle chain reads by the name-word count" +# rules.md#C1: a part whose every member is a listed word in capitals +# in a mixed-case name reads as the credential run on that writing +# rather than on the count, "and so is a part holding two particles +# side by side, which S2 joins into one particle run rather than +# leaving either to its capitals" -- read by the count, that is, +# which two name words before the comma flip to the credential run, +# reported. The family-comma path had read the part as name text: +# this baseline read given 'PhD', family 'DO DO John Smith', P6 +# attaching the chain. 'John Smith, PhD vd DO' is the same chain +# with a particle that is not a member ('vd', particle and +# unambiguous suffix vocabulary). 1.4.0 read both as suffixes, so +# that ledger has no rule. +# +# Literal. Probes: 'John Smith, PhD DO' (one particle beside no other +# stays settled and silent) and 'Smith, PhD DO DO' (one name word +# before the comma keeps the listing form) are _MUST_NOT_MATCH. +name_regex = "^(?:John Smith, PhD DO DO|John Smith, PhD vd DO)$" +fields = ["given", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] + [[change]] issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" # rules.md#C1: "Paired initials are the exception to the count" -- From 41141fb7e2a1c6c6e04b6614f54148cb1aee1054 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 11:38:14 -0700 Subject: [PATCH 2/4] fix(#562): review round -- P2 does the join, and the release note's claims The docs review found the C1 clause crediting S2 with a join that is P2's (P2 joins C1's interacts:), and the release bullet wrong twice: 'John Smith, MD DO DO' read title 'MD', first 'DO', middle 'DO' at 2.3, not correctly, and DO is not the only credential that is also a particle -- MC and VD are too. The 'no two particles side by side' insert in C1's silence sentence is reverted: that sentence is about runs the count leaves in the listing form, which a particle pair no longer is. The code review found the check reaching runs the family-comma path did read whole ('John Smith, vd DO', 'VD DO', 'MD DO DO'): those flip to the same fields and gain the flip's report. Accepted as C1's own "a decision either way at this comma is reported", recorded in the decisions entry and pinned by a case row; the segment comment no longer claims every chain was misread. The S2 #544 entry's boundary (5) points forward to the narrowing, and the settled test's docstring states the old controls' exception counts exactly. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 ++-- docs/design/rules.md | 11 +++++------ docs/release_log.rst | 2 +- nameparser/_pipeline/_segment.py | 12 +++++++----- tests/v2/cases.py | 10 ++++++++++ tests/v2/test_properties.py | 8 +++++--- tools/differential/expected_since_2.0.0.toml | 4 ++-- tools/differential/expected_since_2.1.0.toml | 4 ++-- tools/differential/expected_since_2.2.0.toml | 4 ++-- tools/differential/expected_since_2.3.0.toml | 4 ++-- 10 files changed, 38 insertions(+), 25 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 4a38d730..d190f785 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -670,7 +670,7 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f - 2026-09-18 (Derek), #531 — THE COMMA REPORT'S REACH NOW INCLUDES THE GIVEN SEGMENT'S TRAILING SLOT, AND THE OPEN FOLLOW-UP IS CLOSED. The 2026-09-18 bullet above headed THE COMMA REPORT'S REACH IS THE FIRST POST-COMMA PIECE recorded this slot as an open maintainer decision and named the two questions it turned on — whether a middle initial's neighbourhood should start reporting (a noise judgement, rules.md#A1) and whether the silence was a 1.4 parity gap. Both are answered here, and that bullet stands as it landed. The slot now reads and reports: `Doe, John MA` gives suffix `MA` and `Doe, John Ma` keeps middle `Ma`, each saying which way it went. Derek chose to restore the ROLE and report both ways — one rule for both spellings — over a report-only change and over a capitals-only one, because the issue exists in the first place because two spellings of one name disagree. The noise question was settled by MEASURING THE DISAGREEMENT rather than by argument: over a generated sweep of 78 pairs — five listed members and two by-shape tokens in three cased spellings each, plus five controls in one spelling apiece, so 26 words against three name shapes — 48 pairs disagreed about whether the word was a credential or a name, and every one of the 48 disagreed in the same direction, the comma form declining what the comma-less form took. After this change 3 disagree and all three are the lower-case `do` rows P6 owns. The sweep and its allowlist are `tests/v2/test_properties.py`; a slot that answers differently from the same name written without a comma is not a quiet slot, it is an inconsistent one. Recorded under `3-0-reevaluations`' standing rule because v1 parity is LOAD-BEARING for one half of this and explicitly NOT for the other: `Doe, John MA` reads suffix `MA` on the 1.4.0 wheel, so the bare-acronym half RESTORES v1's role and the report is all that is new there — its 1.4.0 ledger rule neither retired nor narrowed to `_ambiguities` (nothing below baseline 2.0 can diff on that pseudo-field) but was RE-POINTED to the four names where the writing declines the credential, which keep the 2.0-era middle name against v1. `Doe, John X.Y.Z.` reads middle `X.Y.Z.` at 1.4.0 too, so the dotted half LEAVES v1 and carries a 1.4.0 ledger rule of its own; it moves to match the comma-less `John Doe X.Y.Z.` and rules.md#S3's shape rule, not to restore anything. BLAST RADIUS, stated the way this log's own rule asks: twenty corpus names move, and only TWO of them were in any corpus before this branch (`Doe, John MA` and `Doe, John X.Y.Z.`, both admitted by #530's own arc) — the other eighteen are this change's own case rows, so what the differential measures on pre-existing data is two names, and the population the rule reaches is a SHAPE (every family-comma listing whose given part ends in a member of this class) that the corpora barely sample. ACCEPTED COSTS, all measured on the differential corpora: `SMITH, JOHN DO` keeps family `DO SMITH` where 1.4.0 read suffix `DO`, paired with the Nascimento record in the bullet above; one-case and caseless names take the credential with no case evidence at all, so `DOE, MARY JO MA`, `doe, john ma`, `田中, 太郎 MA` and `김, 민준 MA` all read a suffix, which is what 1.4.0 read for each of the four — "caseless is inert" holds for the LEAN and not for the outcome, since `is_one_case` answers True for a script with no case and the positional reading then decides; `Doe, John van MA` loses its middle to the family, reading family `van Doe`, suffix `MA` with two reports where it read middle `van MA` in silence, which is `Berg, Jan van Jr.`'s reading arriving through a shape it could not reach before (Derek accepted it 2026-09-18 as P6 working correctly rather than as a cascade to carve out); `Doe, John Prof. MA` gains a TITLE role `Prof.` never had, H5's transparency reaching it once `MA` leaves the walk, landing it on the same answer as `Doe, John MA Prof.`; `Smith, LEED AP` moves under the default-off caps switch alone, to given `LEED`, family `Smith`, suffix `AP` with two reports, so no default reading is at stake; and a genuine middle name that is also a class member is now a credential wherever it ends the given part and is not Title-cased — `DOE, JOHN ED` reads suffix `ED` — which is the same cost the comma-less form has carried since 2.0. NOT REPAIRED HERE and left open: the maiden walk claims `Smith MA` whole in `Doe, Jane nee Smith MA` before this slot exists, so the rule cannot reach it and the name stays silent with maiden `Smith MA` — the same gap #530's close-out saw from the other side with `John Smith nee Jones R.A.I.`, and it is out of scope for #531. - 2026-09-19 (Derek), #533 — THE CLASS'S LAST SILENT TRAILING POSITION IS CLOSED, AND #531'S READING NOW LIVES IN ONE PLACE. The slot list in rules.md#S2 gains the trailing slot of a maiden marker's clause, and #S3's enumeration gains it with the by-shape spellings; the rule that governs what the clause does with those words is M2's, and the entry under `### M2` above is where its reasoning lives. The gap the bullet above left open as out of scope for #531 is the one this closes, from both sides: `Doe, Jane nee Smith MA` now gives maiden 'Smith' with suffix 'MA' and reports, and `John Smith nee Jones R.A.I.` gives suffix 'R.A.I.' again — a RESTORATION rather than a change, since 2.3.0 read it that way (measured on the wheel) and this unreleased cycle moved it into the maiden name when #516 took dotted tokens out of the certain-suffix class, where no corpus file held the name and no gate could see it. Two things belong here rather than under M2. FIRST, `credential_at_the_given_slot` in `_pipeline/_pieces.py` is now where #531's reading of a class member ending the given part is written, with two callers — assign's walk over that part, and the maiden walk's second check over the name a take would leave. Spelling it twice is precisely the "condition written to match it" mechanisms.md#ONE-PREDICATE-PER-QUESTION names, and the drift would have been silent, since each site's own tests would have gone on passing. It is a text-and-tags question, which is what puts it in `_pieces` rather than beside either caller — the destination follows the LAYER, not the topic. It costs one frame PER MEMBER asked at that slot: measured 2026-09-19 against 2f57ff21 per `Parser.parse`, `Doe, John MA` goes 310 → 311 and `Doe, John MA Ma MA`, which asks four times, 439 → 443, while a name with no member there never reaches it and pays nothing. The reference band does not move (412/449 on `uv run python tools/perf/call_count.py`), no test pins 310, and Derek took the trade rather than keep two conditions in step across two stages with no test that asks them both. The maiden path got one frame CHEAPER in the same pass, the numeral reading calling the peel pair it had wrapped rather than the wrapper: `Jane Doe nee Smith` 249 → 248. SECOND, THE ONE-CASE HEAD IS AN ACCEPTED EXCEPTION AND IT IS THE M2 INSTANCE OF #492'S DEFERRED QUESTION about whether a cased suffix token counts as case evidence. `DOE, JANE nee Smith Ma` reads suffix 'Ma' where the clause-less `DOE, JANE Ma` keeps middle 'Ma', because the own-words span stops at the marker — rules.md#P3 puts "a maiden marker's run and every word after it" outside the name's own words from the moment the marker is tagged — so the member's own Title-casing is not in the span `one_case` is computed over and the clause HIDES the contrast. That is THE PREDICATE KEEPS THE JUDGED TOKEN IN THE SPAN, the 2026-09-14 #289/#516 entry at the head of this section, failing structurally rather than by oversight: at this slot the judged token is never in the span, and that entry's own answer — include it — cannot be had here. Widening the span for this one question would change `one_case` for the whole name, and three sites read it, so the exception is accepted instead. Measured 2026-09-19: 114 of 2016 generated pairs — every listed member and both by-shape spellings, in three cased spellings, against sixteen heads and six clause bodies of NAME WORDS ONLY — against 186 allowlisted and 984 disagreeing outside the class before the change, and 0 outside it after. The PR review widened that grid to three markers (`nee`, `née`, `geb.`) and three policies (the default and each 2.4 switch), and the class is INDIFFERENT to both: 1,026 of 18,144, exactly 9x the original in both columns, and still 0 disagreeing outside it. `tests/v2/test_properties.py` carries the sweep with the class defined STRUCTURALLY (`one_case` true of the clause form and false of the clause-less one) rather than as a name list. The recorded control beside it is now the SET and not only its size, as a digest over the members: a count cannot notice a swap, one pair leaving and another arriving, which is the same blindness the count was added to close one level up. - 2026-09-23 #492 — THE DEFERRED QUESTION the 2026-09-14 predicate paragraph and the 2026-09-19 #533 bullet above name — whether a cased suffix token counts as case evidence — IS ANSWERED FOR RENDER ONLY. R5's gate now leaves the suffixes out (decisions.md#R5, 2026-09-23); the parse-time fact this section rests on keeps counting the judged token exactly as written above, so `Jack MA` still leans credential and `DOE, JANE nee Smith Ma` still reads suffix `Ma`. The split is decided and recorded under R5, with the shapes it makes visible there (`jack MA` at this slot, `john e jones III` at P3's). -- 2026-09-27 (Derek), #544 — A CREDENTIAL RUN KEEPS ITS AMBIGUOUS MEMBERS: C1'S NAME-WORD COUNT READS A RUN AS IT READS ONE WORD, AN UNAMBIGUOUS CREDENTIAL IN FRONT OF A MEMBER SPEAKS FOR IT AT EVERY TRAILING SLOT, AND A LISTED MEMBER'S CHUNKED DOTTED SPELLING PASSES THE PERIOD GATE. Three readers never looked past the member itself: C1 flipped a part to the suffix comma by the name-word count only when it was ONE token, and the trailing readers — the no-comma peel, the comma part read wholly as credentials, the given part's trailing slot — met the member first, took its name lean and stopped, so the credential in front of it was never reached. #540 marked `meng` and `lac`, whose conventional spellings lean name, which turned the one-case-row cost this section's 2026-09-15 item (ii) accepted into the ordinary spelling of two degrees. DECIDED, ten forks, (9) and (10) on 2026-09-28. (1) THE WHOLE RUN AT C1: two or more words after the comma, every one suffix vocabulary or a class candidate, at least one a candidate, behind two or more name words, is the credential run — all-member runs (`John Smith, Ed Ma`) and a leading member (`John Smith, MEng PhD`) included; offered and declined were an anchored-only rule and the issue's own "last word ambiguous" wording, which contradicted its `Ed Ma` consequence. (2) ALL TRAILING PATHS, ANCHORED IN FRONT: a listed member with an unambiguous credential in front of it, through other members, reads as the credential on the no-comma peel, in the comma part read wholly as credentials and at the given slot; a credential BEHIND it anchors nothing (`Wang Ma PhD` keeps family 'Ma'). (3) THE CHUNKED DOTTED GATE: a listed member written in two or more period-closed letter chunks (`M.Eng.`, `L.Ac.`) passes as `M.A.` does; one trailing period (`Ma.`, `Ed.`) still does not. (4) SINGLE-LETTER ROMAN NUMERALS neither anchor nor count toward a C1 run, S3's single-character retirement being the precedent; they peel as before, and the anchor's half of the exclusion is a matter of their shape, boundary (d) below. (5) NO NEW REPORT WHERE THE WRITING SETTLED IT: a run whose every member is LISTED and leans credential — S2's lean, capitals in a mixed-case name — is not flipped, the family-comma path already reading it as the credential run with the fields and the report it always had (`John Smith, PhD MA`); the alternative, flipping such a run and suppressing only the flip's report, removes reports the family-comma path makes today, `John Smith, MA Jr` losing the one it makes for 'MA'. Because the lean is the listed set's, a by-shape member in capitals never settles a run: `John Smith, X.Y.Z. MA` flips to suffix 'X.Y.Z. MA' and reports. (6) THE ANCHOR DOES NOT REACH ACROSS A MAIDEN CLAUSE: `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' where `Jane Doe Jr. Ma` reads suffix 'Jr. Ma' — rules.md#M2's Accepted boundary, and the M2 bullet of this date. (7) AN ANCHOR OVERRIDES P6 AT THE GIVEN SLOT: `doe, jane v phd do` reads suffix 'v phd do' where it read family 'do doe'; the bullet below amends #531's pairing. (8) THE DUAL EXCLUSION IS SCOPED TO THE GIVEN PART'S LEADING TITLE RUN, NOT THE C1 RUN: after a full name the dual counts as the suffix vocabulary it is, as the legacy disjunct already counted it (`John Smith, MS MA`, `John Smith, MD`), and that is what reads the issue's headline `Jane Doe, MS LAc` as suffix 'MS LAc'. (9) A MEMBER THE COMPANY DECIDES REPORTS WHERE IT WAS DECIDED, in a family comma's part read wholly as credentials as at the peel: each member whose own writing declined and which the credential in front made the credential reports `suffix-or-name` at that reading. `Smith, PhD Ma`, `Smith, PhD MEng` and `Smith, Ph. D. MEng` report the member; `Smith, Dr. PhD LAc` and `John Quincy Smith, Prof. PhD LAc` report 'LAc'; `john smith, v phd ma` reports 'ma'; `John Smith, PhD Ma Prof.`, which C1 does not flip for its trailing title, reports 'Ma' where its twin `John Smith, PhD Ma` reports C1's flip once over the part. A member its own capitals made the credential stays silent (`Smith, PhD MA`), and the part's first piece, which nothing stands in front of, is the first-slot report's word and never this one's. (10) A DUAL IN THE GIVEN PART'S LEADING TITLE RUN SILENCES THE WHOLE PART: refinement (b) below reads such a dual as a title that speaks for nothing, and in a part whose leading title run holds one no credential behind the dual speaks for a member of the ambiguous class either, so the company never reads such a part wholly as credentials: the titles give it its reading, the dual a title and the next word the given name, and behind that given name the given slot reads as it reads any (`Smith, MD PhD Jr Ma` reads suffix 'Jr Ma', as `Smith, MD Jane Jr Ma` does). A part with no member to speak for, or only members whose capitals decide them, reads as it did: `Smith, MD PhD` suffix 'MD PhD', `Smith, Ms MD MA` suffix 'Ms MD MA'. `Smith, Ms MD Ma`, `Smith, MD MS Ma` and `SMITH, MD MS BA` read title 'Ms MD' / 'MD MS', given 'Ma' / 'BA'; `Smith, MD PhD Ma` reads title 'MD', given 'PhD', middle 'Ma', the given slot reporting 'Ma'; e10e83b4 read each of them so. The run is the walk's leading title run, so a dual behind a plain title stands in it too (`Smith, Prof. MD Ma` and `smith, prof. md ma` read title 'Prof. MD', given 'Ma'; `Smith, Dr. MD PhD Ma` keeps given 'PhD'), while a plain title alone silences nothing (`Smith, Dr. PhD LAc` reads title 'Dr.', suffix 'PhD LAc'). Behind a real given name the company reads as ever (`Doe, Jane MD Ma` and `Doe, Jane Ms Ma` read suffix 'MD Ma' and 'Ms Ma'), and past the title run a dual speaks like any suffix word (`Smith, PhD Ms Ma` reads suffix 'PhD Ms Ma'); `Smith, Ms Ma`, `Nguyen, Sr Ba` and `Smith, MD Do` are unchanged. +- 2026-09-27 (Derek), #544 — A CREDENTIAL RUN KEEPS ITS AMBIGUOUS MEMBERS: C1'S NAME-WORD COUNT READS A RUN AS IT READS ONE WORD, AN UNAMBIGUOUS CREDENTIAL IN FRONT OF A MEMBER SPEAKS FOR IT AT EVERY TRAILING SLOT, AND A LISTED MEMBER'S CHUNKED DOTTED SPELLING PASSES THE PERIOD GATE. Three readers never looked past the member itself: C1 flipped a part to the suffix comma by the name-word count only when it was ONE token, and the trailing readers — the no-comma peel, the comma part read wholly as credentials, the given part's trailing slot — met the member first, took its name lean and stopped, so the credential in front of it was never reached. #540 marked `meng` and `lac`, whose conventional spellings lean name, which turned the one-case-row cost this section's 2026-09-15 item (ii) accepted into the ordinary spelling of two degrees. DECIDED, ten forks, (9) and (10) on 2026-09-28. (1) THE WHOLE RUN AT C1: two or more words after the comma, every one suffix vocabulary or a class candidate, at least one a candidate, behind two or more name words, is the credential run — all-member runs (`John Smith, Ed Ma`) and a leading member (`John Smith, MEng PhD`) included; offered and declined were an anchored-only rule and the issue's own "last word ambiguous" wording, which contradicted its `Ed Ma` consequence. (2) ALL TRAILING PATHS, ANCHORED IN FRONT: a listed member with an unambiguous credential in front of it, through other members, reads as the credential on the no-comma peel, in the comma part read wholly as credentials and at the given slot; a credential BEHIND it anchors nothing (`Wang Ma PhD` keeps family 'Ma'). (3) THE CHUNKED DOTTED GATE: a listed member written in two or more period-closed letter chunks (`M.Eng.`, `L.Ac.`) passes as `M.A.` does; one trailing period (`Ma.`, `Ed.`) still does not. (4) SINGLE-LETTER ROMAN NUMERALS neither anchor nor count toward a C1 run, S3's single-character retirement being the precedent; they peel as before, and the anchor's half of the exclusion is a matter of their shape, boundary (d) below. (5) NO NEW REPORT WHERE THE WRITING SETTLED IT: a run whose every member is LISTED and leans credential — S2's lean, capitals in a mixed-case name — is not flipped, the family-comma path already reading it as the credential run with the fields and the report it always had (`John Smith, PhD MA`) — narrowed 2026-10-01 by the #562 entry under C1: a part holding two particles side by side is read by the count instead, and reports; the alternative, flipping such a run and suppressing only the flip's report, removes reports the family-comma path makes today, `John Smith, MA Jr` losing the one it makes for 'MA'. Because the lean is the listed set's, a by-shape member in capitals never settles a run: `John Smith, X.Y.Z. MA` flips to suffix 'X.Y.Z. MA' and reports. (6) THE ANCHOR DOES NOT REACH ACROSS A MAIDEN CLAUSE: `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' where `Jane Doe Jr. Ma` reads suffix 'Jr. Ma' — rules.md#M2's Accepted boundary, and the M2 bullet of this date. (7) AN ANCHOR OVERRIDES P6 AT THE GIVEN SLOT: `doe, jane v phd do` reads suffix 'v phd do' where it read family 'do doe'; the bullet below amends #531's pairing. (8) THE DUAL EXCLUSION IS SCOPED TO THE GIVEN PART'S LEADING TITLE RUN, NOT THE C1 RUN: after a full name the dual counts as the suffix vocabulary it is, as the legacy disjunct already counted it (`John Smith, MS MA`, `John Smith, MD`), and that is what reads the issue's headline `Jane Doe, MS LAc` as suffix 'MS LAc'. (9) A MEMBER THE COMPANY DECIDES REPORTS WHERE IT WAS DECIDED, in a family comma's part read wholly as credentials as at the peel: each member whose own writing declined and which the credential in front made the credential reports `suffix-or-name` at that reading. `Smith, PhD Ma`, `Smith, PhD MEng` and `Smith, Ph. D. MEng` report the member; `Smith, Dr. PhD LAc` and `John Quincy Smith, Prof. PhD LAc` report 'LAc'; `john smith, v phd ma` reports 'ma'; `John Smith, PhD Ma Prof.`, which C1 does not flip for its trailing title, reports 'Ma' where its twin `John Smith, PhD Ma` reports C1's flip once over the part. A member its own capitals made the credential stays silent (`Smith, PhD MA`), and the part's first piece, which nothing stands in front of, is the first-slot report's word and never this one's. (10) A DUAL IN THE GIVEN PART'S LEADING TITLE RUN SILENCES THE WHOLE PART: refinement (b) below reads such a dual as a title that speaks for nothing, and in a part whose leading title run holds one no credential behind the dual speaks for a member of the ambiguous class either, so the company never reads such a part wholly as credentials: the titles give it its reading, the dual a title and the next word the given name, and behind that given name the given slot reads as it reads any (`Smith, MD PhD Jr Ma` reads suffix 'Jr Ma', as `Smith, MD Jane Jr Ma` does). A part with no member to speak for, or only members whose capitals decide them, reads as it did: `Smith, MD PhD` suffix 'MD PhD', `Smith, Ms MD MA` suffix 'Ms MD MA'. `Smith, Ms MD Ma`, `Smith, MD MS Ma` and `SMITH, MD MS BA` read title 'Ms MD' / 'MD MS', given 'Ma' / 'BA'; `Smith, MD PhD Ma` reads title 'MD', given 'PhD', middle 'Ma', the given slot reporting 'Ma'; e10e83b4 read each of them so. The run is the walk's leading title run, so a dual behind a plain title stands in it too (`Smith, Prof. MD Ma` and `smith, prof. md ma` read title 'Prof. MD', given 'Ma'; `Smith, Dr. MD PhD Ma` keeps given 'PhD'), while a plain title alone silences nothing (`Smith, Dr. PhD LAc` reads title 'Dr.', suffix 'PhD LAc'). Behind a real given name the company reads as ever (`Doe, Jane MD Ma` and `Doe, Jane Ms Ma` read suffix 'MD Ma' and 'Ms Ma'), and past the title run a dual speaks like any suffix word (`Smith, PhD Ms Ma` reads suffix 'PhD Ms Ma'); `Smith, Ms Ma`, `Nguyen, Sr Ba` and `Smith, MD Do` are unchanged. WHAT MAY ANCHOR, four boundaries of the company clause, each a rule. (a) A CONNECTIVE (P3) anchors nothing and ends the run: the generational `i` is also Catalan's conjunction, and `Jane Doe nee Puig i Ma` keeps its link, maiden 'Puig i Ma'. (b) A TITLE/SUFFIX DUAL IN THE GIVEN PART'S LEADING TITLE RUN after a one-word family comma is a title there and anchors nothing — and, by (10), no credential behind it speaks for a member in that part — so the given slot's pass starts past the leading title run: `Smith, Ms Ma`, `Nguyen, Sr Ba` and `Smith, MD Do` keep their given names, and `Smith, MD MA Ma` keeps middle 'Ma', no unambiguous credential standing in its run once the title is set aside. Everywhere else a dual anchors like any suffix word (`John Smith MD MEng`, and (8)). (c) THE WALK'S OWN LEADING PIECE NEVER ANCHORS. In a name with no comma before the run, the leading piece is the name word H4's carve-out keeps whatever its vocabulary, and letting it anchor spends that name: in `Om Ma` and `PhD Ma` 'Ma' stays a name word under every name order (family 'Ma' in the default order, given 'Ma' under both family-first orders), where an anchoring `PhD` would leave given 'PhD', suffix 'Ma' and no family at all. One name word in front frees the same word — `John Om Ma` reads suffix 'Om Ma'. A family comma's given part has no such position, its first piece being read by the comma part's own credential reading, so `Smith, PhD MEng` reads family 'Smith', suffix 'PhD MEng'. (d) A SINGLE-LETTER ROMAN NUMERAL, IN ANY CASE (`V`, `v`, `I.`), anchors nothing for its SHAPE — one letter is how a middle initial is written, and S3 retired single-character vocabulary matches for the same reason (`John Smith PhD V Ma` keeps family 'Ma') — not because a generation is no credential. The rule is stated by that shape and the code tests exactly it (`_vocab.is_single_letter_numeral`); "initial-shaped" was the wrong word for it, `is_initial_shaped` answering False for a lower-case `v` that the exclusion covers too. And the exclusion is not about generations: a multi-letter numeral or generational word anchors as any suffix word does (`John Smith PhD III Ma` reads suffix 'PhD III Ma', `abdul Smith Jr Ma` suffix 'Jr Ma'). ACCEPTED COSTS, each a reading this design chooses. `John Smith, Ms Ma` → suffix 'Ms Ma', reported, as `John Smith, Ms` alone already reads suffix (8). `Om Jr Ma` → given 'Om', suffix 'Jr Ma' and no family name: 'Jr' is not the leading piece, so it anchors 'Ma', and this section's standing Accepted clause already consumes an unambiguous suffix that leaves no family (`Om Jr` → given 'Om', suffix 'Jr'); every release from 2.0.0 made that same role assignment. `doe, jane v phd do` reports `suffix-or-name` where it reported `particle-or-given`, the fork reported being the credential pick rather than P6's attachment (7). And a one-word family comma whose given part OPENS with a capitalised `MA` in a mixed-case name, with an unambiguous credential behind it, reads the whole part as a credential run, family 'Smith' and no given name, where the parent kept given 'MA': `Smith, MA PhD Ma` read given 'MA', middle 'Ma', suffix 'PhD' and reads suffix 'MA PhD Ma'. 'MA' leans credential on its capitals, so it is read as the credential it is written as, and what follows it is then the company's (or, for `M.Eng.`, the chunked gate's). Measured 2026-09-28 on the grid below: 138 parses (46 texts, three orders each) go from given 'MA' to no given name, 72 of them with no `M.Eng.` in the run; 12 of the 138 are the parses the attribution below leaves unattributed, the rest falling in the anchor and `M.Eng.` classes. Each reports 'MA' at the first slot, and a member behind the credential reports as the company's pick (9). ACCEPTED LIMITS, each keeping the member a name word, rules.md#S2's Accepted block and case rows: the merged `Ph. D.` is outside the no-comma walk (`John Smith Ph. D. MEng` family 'MEng'; the comma and given-slot spellings do anchor it, `Smith, Ph. D. MEng` and `Doe, Jane Ph. D. MEng` reading suffix 'Ph. D. MEng'), a title between the credential and the member breaks the run (`John Smith PhD Prof. Ma`), and a particle member P2 has chained before the peel is out of reach (`John Smith PhD Do Do` family 'Do Do'; `Smith, PhD Do Ma` given 'PhD', middle 'Do Ma'). @@ -875,7 +875,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-06 #511 — a suffix value handed to `revise()` derives its entries from its own commas, by the rule a whole name uses. `Parser.revise` sub-parsed each value and forced every harvested token to the named role AFTER the sub-parse had run R1's entry pass, which keys on Role.SUFFIX; a bare 'MD PhD' reads there as a title and a family name, so the pass joined nothing and the field rendered 'MD, PhD'. The fix runs the same pass again: `revise` now sub-parses to a ParseState, forces the role on every non-dropped token, and calls `suffix_entries` — the pass, lifted unedited out of post_rules' tail into a function the two callers share (mechanisms.md#ONE-PREDICATE-PER-QUESTION) — over the forced state, then assembles and harvests as before. So a comma in the value parts two credentials and a space joins them. TWO SPELLINGS of the lifted pass, and the reason is the call budget: a first draft had post_rules call the state-in/state-out wrapper, and the second ParseState build cost three more calls per parse against the band tests/v2/test_benchmark.py holds (py3.11, 2026-09-06, tools/perf/call_count.py: 450 calls/name before the move, 451 with the in-place worker `_mark_suffix_entries` that post_rules now calls, 454 with the draft; the facade band tops at 455.9), so the worker writes in place as every other post rule does and the wrapper exists for the one caller with no token list of its own. NO `_run` HELPER for the same reason: a draft routed `parse()` through a private state builder shared with `revise`, and that cost one frame per parse on the hot path (py3.11, 2026-09-06: 415 calls/name against 414 without it, in a 402-418 band), so the four-line construction is spelled twice and `parse()` is untouched. MEASURED 2026-09-06 over the 1117 distinct names in `tools/differential/corpus*.jsonl`, comparing `revise(p, suffix=p.suffix).suffix` against `p.suffix` under the default parser for every name with a non-empty suffix (368 of the 1117): 38 differed before and 1 after; the differential gate is byte-identical at all four baselines, `revise` not being on the compare path. RECOMPUTE: parse each corpus name, skip an empty suffix, revise the parse with its own suffix, count the names whose suffix moved (the script is in the #511 issue body). THE ONE LEFT is '김민준씨, J.씨', and it is not entry structure: the whole-name parse keeps 'J.씨' one glued suffix token, suffix '씨, J.씨', while the sub-parse of the bare value '씨, J.씨' peels the honorific off the initial, so the revised field renders '씨, J. 씨' — right entries, the spurious comma of before ('씨, J., 씨') gone, one word read differently by the value's own parse than by the whole name's. That is the "classified ON ITS OWN" limit `revise`'s docstring has always recorded, and CJK honorific peeling is a W-rule question this change does not move; pinned in `test_revise_reads_a_glued_honorific_on_its_own`. SUPERSEDES the phd-merge acceptance of 'Ph., D.' on this path (its bullet says how). STALE TAGS, tried and backed out: a draft cleared the sub-parse's own "joined" with the forcing and re-derived it, because the pass only ADDS the tag and `revise(n, family="Jones MD PhD")` carries the sub-parse's between-piece suffix mark on 'PhD' onto a FAMILY token. Measured 2026-09-06 against `330ee55`, the clear also destroyed every WITHIN-piece mark on a non-suffix value — 'D.' of `revise(n, family="John Ph. D. Smith")` lost the merge mark the #436 bullet's DECLINED list calls role-blind and correct for every role — while on a ParsedName only the suffix string view reads "joined" (`_text_for`'s suffix_join gate) — the facade's `_list_for` heals it for every role, but no path puts a revised name into a HumanName, the v1 setters going through `replace()`, so wiring those setters onto `revise` is the change that would show the stale mark, as `last_list == ['Jones', 'MD PhD']` — `initials()` and every field string being identical at both trees for every shape measured. So the tags are kept minus FOLDED_TAG as before; the between-piece mark on a forced non-suffix role is a tag-only oddity that predates this change and stays, and for a suffix value nothing depends on the sub-parse's marks, every pair it joined sharing a bucket with no parting token and the pass setting it again; pinned in `test_revise_keeps_the_sub_parses_within_piece_mark`. Dropped tokens keep their role and tags through the forcing, because the pass filters them by index and assemble omits them. ONE MORE LIMIT, pinned in `test_revise_leaves_a_policy_delimiter_unparted_without_a_tail_segment`, and it is the "classified ON ITS OWN" limit again rather than a rule of revise's: a delimiter the policy names through `extra_suffix_delimiters` is dropped, and so parts entries, only on a segment after a comma that the reading of the words makes a tail, and a value with no comma of its own has none — under `Policy(extra_suffix_delimiters=frozenset({" - "}))` the whole name 'Doe, John, MD PhD - FACS' renders 'MD PhD, FACS' while `revise(n, suffix="MD PhD - FACS")` renders 'MD PhD - FACS', the dash surviving as a token and the forced role making it a suffix word; measured 2026-09-06, a value whose own words read with a tail segment does part ('John Doe, MD - FACS' revises to 'John Doe, MD, FACS') and one after a suffix comma does not ('MD, PhD - FACS' stays), which is how the value's words read as a name deciding it. The round-trip is unaffected, the whole-name view having already rendered that boundary as a comma. DECLINED: a list-valued `revise(p, suffix=["MD PhD", "FACS"])`, which widens the API and leaves the string path where it was; documenting the limit with pins alone, the docstring already recording it and #511's measurement showing it reachable on every space-joined run #436 produced; and a text-only read of the delimiter cores inside `revise`, a second rule for a case no round-trip reaches. Out of scope and left as is: `ParsedName.replace(suffix="MD, PhD")` renders 'MD,, PhD', `replace` whitespace-splitting by contract and the facade setters riding on it for v1 parity. - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. -- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. +- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), or a credential that is also a title in front (`John Smith, MD DO DO`) — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED: with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index afefb244..f0b1de96 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1703,8 +1703,8 @@ C1. Rationale: a credential run after the comma means the name is in evidence rather than on the count; a word of the class by shape alone carries no such lean, so a part holding one is read by the count, and so is a part holding two particles side by side, - which S2 joins into one particle run rather than leaving either - to its capitals. A decision either way at this comma + which P2 joins into one particle run that S2 leaves to no word's + capitals. A decision either way at this comma is reported; for a run of words the decision is the flip to the credential run, reported once over the whole part. A flip in which no listed word of this class takes part is the exception and is @@ -1720,9 +1720,8 @@ C1. Rationale: a credential run after the comma means the name is in in it: a word of this class read as the credential because a credential in front speaks for it reports, and one its own capitals made the credential does not, so a run whose every such - word is written in capitals, no two particles side by side, - reads whole in silence ('John Smith, PhD MA', 'Smith, PhD MA'), as does a part - read as + word is written in capitals reads whole in silence + ('John Smith, PhD MA', 'Smith, PhD MA'), as does a part read as titles before a lone given name ('Smith, Ms MD Ma'). It is one of TWO places the comma's own decision is reported, the other being the word trailing the given part after it (S2), which is a second decision about a second word and never @@ -1855,7 +1854,7 @@ C1. Rationale: a credential run after the comma means the name is in V` reads the suffix and `Smith, John PhD I.` continues the run, while adding a suffix comma after either turns that same letter into the middle initial. - history: decisions.md#C1 · interacts: H2, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py + history: decisions.md#C1 · interacts: H2, P2, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py C2. Rationale: text beyond the recognized comma parts should be taken in without silent guessing. diff --git a/docs/release_log.rst b/docs/release_log.rst index ea51980f..f3df48df 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -16,7 +16,7 @@ Release Log - **Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.** ``HumanName("John Smith, Ed Ma")`` gives first ``John``, last ``Smith``, suffix ``Ed Ma``, where 2.0 through 2.3 gave first ``Ed``, middle ``Ma``, last ``John Smith`` -- 1.4.0's reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (``John Smith, MA``), and ``parse()`` reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report, the capitals having decided it (``John Smith, PhD MA``, ``John Smith, MS MA``). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so ``john smith, v phd ma`` gives first ``john``, last ``smith``, suffix ``v phd ma``, where 2.3 gave first ``v``, middle ``ma``, last ``john smith``. ``john smith, md ma`` and ``John Smith, Ms Ma`` move the same way, where 2.0 through 2.3 gave title ``md``/``Ms`` -- the second is the accepted cost, ``Ms`` read as the suffix word it also is, as ``John Smith, Ms`` alone already reads it -- while one name word before the comma keeps the listing form (``Smith, Ms Ma`` gives title ``Ms``, first ``Ma``). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential's company whatever its case: ``John Smith PhD MEng`` and ``Doe, Jane PhD MEng`` give suffix ``PhD MEng``, the fields every release gave (``PhD, MEng`` through 2.2), now reported, and ``doe, jane v phd do`` gives suffix ``v phd do`` where 2.3.0 gave last ``do doe`` -- a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: ``Wang Ma PhD`` keeps last ``Ma``. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: ``Smith, PhD Ma`` gives last ``Smith``, suffix ``PhD Ma``, where 2.3 gave first ``PhD``, middle ``Ma``. A title that is also a credential (``MD``, ``Ms``) opening that part stays a title and nothing in the part speaks, so ``Smith, MD PhD Ma`` keeps title ``MD``, first ``PhD``, middle ``Ma`` and ``Smith, Ms MD Ma`` title ``Ms MD``, first ``Ma``, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so ``Wang M.Eng.`` gives suffix ``M.Eng.``, as ``Wang M.A.`` does and as 2.0 through 2.3 did. See the #544 entry under ``S2`` in ``docs/design/decisions.md`` (closes #544) - - **Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run.** ``HumanName("John Smith, PhD DO DO")`` gives first ``John``, last ``Smith``, suffix ``PhD DO DO``, where 2.2 and 2.3 gave first ``PhD``, last ``DO DO John Smith`` and 2.0 and 2.1 gave title ``PhD``, first ``DO DO`` -- 1.4.0's reading, restored. ``DO`` is the one acronym that is also a surname particle, and two particles next to each other join into one particle run whatever their capitals, so the capitals no longer settle such a run silently: it is read by the count of name words before the comma, and ``parse()`` reports the call. The same holds when the other particle is a suffix word such as ``vd`` (``John Smith, PhD vd DO`` gives suffix ``PhD vd DO``), and runs like ``John Smith, MD DO DO`` that already read correctly now report too. A single ``DO`` is still left to its capitals (``John Smith, PhD DO`` gives suffix ``PhD DO`` with no report), and one name word before the comma keeps the listing form (``Smith, PhD DO DO`` gives first ``PhD``, last ``DO DO Smith``, as 2.3 did). See the #562 entry under ``C1`` in ``docs/design/decisions.md`` (closes #562) + - **Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run.** ``HumanName("John Smith, PhD DO DO")`` gives first ``John``, last ``Smith``, suffix ``PhD DO DO``, where 2.2 and 2.3 gave first ``PhD``, last ``DO DO John Smith`` and 2.0 and 2.1 gave title ``PhD``, first ``DO DO`` -- 1.4.0's reading, restored. ``DO``, ``MC`` and ``VD`` are credentials and surname particles at once, and two particles next to each other join into one particle run whatever their capitals, so the capitals no longer settle such a run silently: it is read by the count of name words before the comma, and ``parse()`` reports the call. ``John Smith, PhD vd DO`` gives suffix ``PhD vd DO`` the same way, and ``John Smith, MD DO DO`` gives suffix ``MD DO DO`` where 2.3 gave title ``MD``, first ``DO``, middle ``DO``. A single ``DO`` is still left to its capitals (``John Smith, PhD DO`` gives suffix ``PhD DO`` with no report), and one name word before the comma keeps the listing form (``Smith, PhD DO DO`` gives first ``PhD``, last ``DO DO Smith``, as 2.3 did). See the #562 entry under ``C1`` in ``docs/design/decisions.md`` (closes #562) - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index eb0efd69..2c2e3776 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -339,11 +339,13 @@ def class_run(seg: tuple[int, ...]) -> bool: else: rest.append(text) # #562, rules.md#C1: "and so is a part holding two - # particles side by side, which S2 joins into one particle - # run rather than leaving either to its capitals" -- group - # chains the pair ('PhD DO DO', 'PhD vd DO', 'MA vd vd'), - # the family-comma path then reads the chain as name text, - # so the run is the count's to read, whatever its case. + # particles side by side, which P2 joins into one particle + # run that S2 leaves to no word's capitals" -- group chains + # the pair, and the family-comma path can no longer promise + # to read the part whole: it read 'PhD DO DO', 'PhD vd DO' + # and 'MA vd vd' as name text. Where it did read the part + # whole ('VD DO', 'MD DO DO'), the count flips it to the + # same fields and reports the call, as any flip does. # Asked only while the run is still settled -- nothing # re-settles it -- with classify's own particle test, since # group's chain is what the capitals lose to. diff --git a/tests/v2/cases.py b/tests/v2/cases.py index f69b27e5..d767d031 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -856,6 +856,16 @@ def _check_cjk_shape_purity(self) -> None: "and 'DO' beside it is chained all the same. 1.4.0 read the same; 2.3.0 read given 'PhD', " "family 'vd DO John Smith'", shape=3), + Case("comma_run_a_particle_pair_opens_is_flipped_and_reported", + "John Smith, vd DO", + {"given": "John", "family": "Smith", "suffix": "vd DO"}, + ambiguities=("suffix-or-name",), + notes="#562's accepted cost: the family-comma path already " + "read this part whole, but the pair is a particle run " + "the capitals do not settle, so the count flips it to " + "the same fields and reports the call, as every flip " + "at this comma does. 2.3.0 read family 'John Smith vd " + "DO'"), Case("comma_run_with_a_lone_trailing_particle_member_stays_settled", "John Smith, PhD DO", {"given": "John", "family": "Smith", "suffix": "PhD DO"}, diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index 601abc07..4a4a5272 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -976,9 +976,11 @@ def test_a_run_c1_leaves_as_settled_is_read_wholly_as_credentials( half (a member admitted only by shape has no lean, S2): dropped, it fails on 215 texts ('John Smith, PhD X.Y.' reading 'PhD' as name text). 0 here (measured 2026-10-01, #562). Over the 2026-09-29 - tree, before #563 and #562, the two read 1,239 and 751 beyond ten - pinned #562 exceptions; #563 moved the second to 215 and #562 the - first to 1,142. + tree, before #563 and #562, the two read 1,239 and 751 failures + beyond the pinned #562 exceptions, which grew from 10 to 12 and to + 11 under them; #563 moved the second to 215 and #562 the first to + 1,142. Without #562's particle check the test fails on exactly the + ten texts it used to pin as exceptions. """ texts = [f"John Smith, {' '.join(words)}" for n in (2, 3) diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 47440953..3af76509 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2579,8 +2579,8 @@ issue = "fix(#562) a comma credential run holding a particle chain reads by the # rules.md#C1: a part whose every member is a listed word in capitals # in a mixed-case name reads as the credential run on that writing # rather than on the count, "and so is a part holding two particles -# side by side, which S2 joins into one particle run rather than -# leaving either to its capitals" -- read by the count, that is, +# side by side, which P2 joins into one particle run that S2 leaves +# to no word's capitals" -- read by the count, that is, # which two name words before the comma flip to the credential run, # reported. The family-comma path had read the part as name text: # this baseline read title 'PhD', given 'DO DO', family 'John diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index ec88d86c..0d9b6c1a 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2466,8 +2466,8 @@ issue = "fix(#562) a comma credential run holding a particle chain reads by the # rules.md#C1: a part whose every member is a listed word in capitals # in a mixed-case name reads as the credential run on that writing # rather than on the count, "and so is a part holding two particles -# side by side, which S2 joins into one particle run rather than -# leaving either to its capitals" -- read by the count, that is, +# side by side, which P2 joins into one particle run that S2 leaves +# to no word's capitals" -- read by the count, that is, # which two name words before the comma flip to the credential run, # reported. The family-comma path had read the part as name text: # this baseline read title 'PhD', given 'DO DO', family 'John diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index b181dd05..f75d0b52 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1055,8 +1055,8 @@ issue = "fix(#562) a comma credential run holding a particle chain reads by the # rules.md#C1: a part whose every member is a listed word in capitals # in a mixed-case name reads as the credential run on that writing # rather than on the count, "and so is a part holding two particles -# side by side, which S2 joins into one particle run rather than -# leaving either to its capitals" -- read by the count, that is, +# side by side, which P2 joins into one particle run that S2 leaves +# to no word's capitals" -- read by the count, that is, # which two name words before the comma flip to the credential run, # reported. The family-comma path had read the part as name text: # this baseline read given 'PhD', family 'DO DO John Smith', P6 diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index a202b3e5..89f5b803 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -343,8 +343,8 @@ issue = "fix(#562) a comma credential run holding a particle chain reads by the # rules.md#C1: a part whose every member is a listed word in capitals # in a mixed-case name reads as the credential run on that writing # rather than on the count, "and so is a part holding two particles -# side by side, which S2 joins into one particle run rather than -# leaving either to its capitals" -- read by the count, that is, +# side by side, which P2 joins into one particle run that S2 leaves +# to no word's capitals" -- read by the count, that is, # which two name words before the comma flip to the credential run, # reported. The family-comma path had read the part as name text: # this baseline read given 'PhD', family 'DO DO John Smith', P6 From d8a87278cacd7d540e5a29c49e0f65e2caad9f45 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 11:40:16 -0700 Subject: [PATCH 3/4] docs(#562): the report-only movers are listed by example, not as a census The second review found 'Esq. DO DO', 'Jr. DO DO' and 'PhD. DO DO' keeping their fields and gaining the report too -- a credential closed by a period in front, which neither named kind covers. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index d190f785..b004514c 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -875,7 +875,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-06 #511 — a suffix value handed to `revise()` derives its entries from its own commas, by the rule a whole name uses. `Parser.revise` sub-parsed each value and forced every harvested token to the named role AFTER the sub-parse had run R1's entry pass, which keys on Role.SUFFIX; a bare 'MD PhD' reads there as a title and a family name, so the pass joined nothing and the field rendered 'MD, PhD'. The fix runs the same pass again: `revise` now sub-parses to a ParseState, forces the role on every non-dropped token, and calls `suffix_entries` — the pass, lifted unedited out of post_rules' tail into a function the two callers share (mechanisms.md#ONE-PREDICATE-PER-QUESTION) — over the forced state, then assembles and harvests as before. So a comma in the value parts two credentials and a space joins them. TWO SPELLINGS of the lifted pass, and the reason is the call budget: a first draft had post_rules call the state-in/state-out wrapper, and the second ParseState build cost three more calls per parse against the band tests/v2/test_benchmark.py holds (py3.11, 2026-09-06, tools/perf/call_count.py: 450 calls/name before the move, 451 with the in-place worker `_mark_suffix_entries` that post_rules now calls, 454 with the draft; the facade band tops at 455.9), so the worker writes in place as every other post rule does and the wrapper exists for the one caller with no token list of its own. NO `_run` HELPER for the same reason: a draft routed `parse()` through a private state builder shared with `revise`, and that cost one frame per parse on the hot path (py3.11, 2026-09-06: 415 calls/name against 414 without it, in a 402-418 band), so the four-line construction is spelled twice and `parse()` is untouched. MEASURED 2026-09-06 over the 1117 distinct names in `tools/differential/corpus*.jsonl`, comparing `revise(p, suffix=p.suffix).suffix` against `p.suffix` under the default parser for every name with a non-empty suffix (368 of the 1117): 38 differed before and 1 after; the differential gate is byte-identical at all four baselines, `revise` not being on the compare path. RECOMPUTE: parse each corpus name, skip an empty suffix, revise the parse with its own suffix, count the names whose suffix moved (the script is in the #511 issue body). THE ONE LEFT is '김민준씨, J.씨', and it is not entry structure: the whole-name parse keeps 'J.씨' one glued suffix token, suffix '씨, J.씨', while the sub-parse of the bare value '씨, J.씨' peels the honorific off the initial, so the revised field renders '씨, J. 씨' — right entries, the spurious comma of before ('씨, J., 씨') gone, one word read differently by the value's own parse than by the whole name's. That is the "classified ON ITS OWN" limit `revise`'s docstring has always recorded, and CJK honorific peeling is a W-rule question this change does not move; pinned in `test_revise_reads_a_glued_honorific_on_its_own`. SUPERSEDES the phd-merge acceptance of 'Ph., D.' on this path (its bullet says how). STALE TAGS, tried and backed out: a draft cleared the sub-parse's own "joined" with the forcing and re-derived it, because the pass only ADDS the tag and `revise(n, family="Jones MD PhD")` carries the sub-parse's between-piece suffix mark on 'PhD' onto a FAMILY token. Measured 2026-09-06 against `330ee55`, the clear also destroyed every WITHIN-piece mark on a non-suffix value — 'D.' of `revise(n, family="John Ph. D. Smith")` lost the merge mark the #436 bullet's DECLINED list calls role-blind and correct for every role — while on a ParsedName only the suffix string view reads "joined" (`_text_for`'s suffix_join gate) — the facade's `_list_for` heals it for every role, but no path puts a revised name into a HumanName, the v1 setters going through `replace()`, so wiring those setters onto `revise` is the change that would show the stale mark, as `last_list == ['Jones', 'MD PhD']` — `initials()` and every field string being identical at both trees for every shape measured. So the tags are kept minus FOLDED_TAG as before; the between-piece mark on a forced non-suffix role is a tag-only oddity that predates this change and stays, and for a suffix value nothing depends on the sub-parse's marks, every pair it joined sharing a bucket with no parting token and the pass setting it again; pinned in `test_revise_keeps_the_sub_parses_within_piece_mark`. Dropped tokens keep their role and tags through the forcing, because the pass filters them by index and assemble omits them. ONE MORE LIMIT, pinned in `test_revise_leaves_a_policy_delimiter_unparted_without_a_tail_segment`, and it is the "classified ON ITS OWN" limit again rather than a rule of revise's: a delimiter the policy names through `extra_suffix_delimiters` is dropped, and so parts entries, only on a segment after a comma that the reading of the words makes a tail, and a value with no comma of its own has none — under `Policy(extra_suffix_delimiters=frozenset({" - "}))` the whole name 'Doe, John, MD PhD - FACS' renders 'MD PhD, FACS' while `revise(n, suffix="MD PhD - FACS")` renders 'MD PhD - FACS', the dash surviving as a token and the forced role making it a suffix word; measured 2026-09-06, a value whose own words read with a tail segment does part ('John Doe, MD - FACS' revises to 'John Doe, MD, FACS') and one after a suffix comma does not ('MD, PhD - FACS' stays), which is how the value's words read as a name deciding it. The round-trip is unaffected, the whole-name view having already rendered that boundary as a comma. DECLINED: a list-valued `revise(p, suffix=["MD PhD", "FACS"])`, which widens the API and leaves the string path where it was; documenting the limit with pins alone, the docstring already recording it and #511's measurement showing it reachable on every space-joined run #436 produced; and a text-only read of the delimiter cores inside `revise`, a second rule for a case no round-trip reaches. Out of scope and left as is: `ParsedName.replace(suffix="MD, PhD")` renders 'MD,, PhD', `replace` whitespace-splitting by contract and the facade setters riding on it for v1 parity. - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. -- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), or a credential that is also a title in front (`John Smith, MD DO DO`) — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED: with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. +- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED: with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. ### T1 — separators, not joiners From 82f2388f400e0132a7310eb44cfd2e50effdd30b Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 11:56:52 -0700 Subject: [PATCH 4/4] docs(#562): the report-only movers are accepted by Derek Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index b004514c..d0c84c37 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -875,7 +875,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-06 #511 — a suffix value handed to `revise()` derives its entries from its own commas, by the rule a whole name uses. `Parser.revise` sub-parsed each value and forced every harvested token to the named role AFTER the sub-parse had run R1's entry pass, which keys on Role.SUFFIX; a bare 'MD PhD' reads there as a title and a family name, so the pass joined nothing and the field rendered 'MD, PhD'. The fix runs the same pass again: `revise` now sub-parses to a ParseState, forces the role on every non-dropped token, and calls `suffix_entries` — the pass, lifted unedited out of post_rules' tail into a function the two callers share (mechanisms.md#ONE-PREDICATE-PER-QUESTION) — over the forced state, then assembles and harvests as before. So a comma in the value parts two credentials and a space joins them. TWO SPELLINGS of the lifted pass, and the reason is the call budget: a first draft had post_rules call the state-in/state-out wrapper, and the second ParseState build cost three more calls per parse against the band tests/v2/test_benchmark.py holds (py3.11, 2026-09-06, tools/perf/call_count.py: 450 calls/name before the move, 451 with the in-place worker `_mark_suffix_entries` that post_rules now calls, 454 with the draft; the facade band tops at 455.9), so the worker writes in place as every other post rule does and the wrapper exists for the one caller with no token list of its own. NO `_run` HELPER for the same reason: a draft routed `parse()` through a private state builder shared with `revise`, and that cost one frame per parse on the hot path (py3.11, 2026-09-06: 415 calls/name against 414 without it, in a 402-418 band), so the four-line construction is spelled twice and `parse()` is untouched. MEASURED 2026-09-06 over the 1117 distinct names in `tools/differential/corpus*.jsonl`, comparing `revise(p, suffix=p.suffix).suffix` against `p.suffix` under the default parser for every name with a non-empty suffix (368 of the 1117): 38 differed before and 1 after; the differential gate is byte-identical at all four baselines, `revise` not being on the compare path. RECOMPUTE: parse each corpus name, skip an empty suffix, revise the parse with its own suffix, count the names whose suffix moved (the script is in the #511 issue body). THE ONE LEFT is '김민준씨, J.씨', and it is not entry structure: the whole-name parse keeps 'J.씨' one glued suffix token, suffix '씨, J.씨', while the sub-parse of the bare value '씨, J.씨' peels the honorific off the initial, so the revised field renders '씨, J. 씨' — right entries, the spurious comma of before ('씨, J., 씨') gone, one word read differently by the value's own parse than by the whole name's. That is the "classified ON ITS OWN" limit `revise`'s docstring has always recorded, and CJK honorific peeling is a W-rule question this change does not move; pinned in `test_revise_reads_a_glued_honorific_on_its_own`. SUPERSEDES the phd-merge acceptance of 'Ph., D.' on this path (its bullet says how). STALE TAGS, tried and backed out: a draft cleared the sub-parse's own "joined" with the forcing and re-derived it, because the pass only ADDS the tag and `revise(n, family="Jones MD PhD")` carries the sub-parse's between-piece suffix mark on 'PhD' onto a FAMILY token. Measured 2026-09-06 against `330ee55`, the clear also destroyed every WITHIN-piece mark on a non-suffix value — 'D.' of `revise(n, family="John Ph. D. Smith")` lost the merge mark the #436 bullet's DECLINED list calls role-blind and correct for every role — while on a ParsedName only the suffix string view reads "joined" (`_text_for`'s suffix_join gate) — the facade's `_list_for` heals it for every role, but no path puts a revised name into a HumanName, the v1 setters going through `replace()`, so wiring those setters onto `revise` is the change that would show the stale mark, as `last_list == ['Jones', 'MD PhD']` — `initials()` and every field string being identical at both trees for every shape measured. So the tags are kept minus FOLDED_TAG as before; the between-piece mark on a forced non-suffix role is a tag-only oddity that predates this change and stays, and for a suffix value nothing depends on the sub-parse's marks, every pair it joined sharing a bucket with no parting token and the pass setting it again; pinned in `test_revise_keeps_the_sub_parses_within_piece_mark`. Dropped tokens keep their role and tags through the forcing, because the pass filters them by index and assemble omits them. ONE MORE LIMIT, pinned in `test_revise_leaves_a_policy_delimiter_unparted_without_a_tail_segment`, and it is the "classified ON ITS OWN" limit again rather than a rule of revise's: a delimiter the policy names through `extra_suffix_delimiters` is dropped, and so parts entries, only on a segment after a comma that the reading of the words makes a tail, and a value with no comma of its own has none — under `Policy(extra_suffix_delimiters=frozenset({" - "}))` the whole name 'Doe, John, MD PhD - FACS' renders 'MD PhD, FACS' while `revise(n, suffix="MD PhD - FACS")` renders 'MD PhD - FACS', the dash surviving as a token and the forced role making it a suffix word; measured 2026-09-06, a value whose own words read with a tail segment does part ('John Doe, MD - FACS' revises to 'John Doe, MD, FACS') and one after a suffix comma does not ('MD, PhD - FACS' stays), which is how the value's words read as a name deciding it. The round-trip is unaffected, the whole-name view having already rendered that boundary as a comma. DECLINED: a list-valued `revise(p, suffix=["MD PhD", "FACS"])`, which widens the API and leaves the string path where it was; documenting the limit with pins alone, the docstring already recording it and #511's measurement showing it reachable on every space-joined run #436 produced; and a text-only read of the delimiter cores inside `revise`, a second rule for a case no round-trip reaches. Out of scope and left as is: `ParsedName.replace(suffix="MD, PhD")` renders 'MD,, PhD', `replace` whitespace-splitting by contract and the facade setters riding on it for v1 parity. - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. -- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED: with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. +- 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. ### T1 — separators, not joiners