From 496c179e907c3ae4fb53472539219d6b70223bed Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 20:15:45 -0700 Subject: [PATCH 1/5] fix(C1): a particle surname before a comma is one name word (#575) rules.md#C1 counts the words before a comma three times -- v1's "more than one word" for an unambiguous credential, the ambiguous class's name-word count (#289, #544), and assign's two-name-word test for the positional read -- and all three counted tokens or pieces. So 'De La Cruz, Ed' lost its given name and 'van der Berg, MA' split the chain (an unreleased regression against 2.3.0), and every release split 'van der Berg, PhD'. A particle and the name word it attaches to now count as one word, the reach P1's fold uses, so 'de Mesnil Jean' stays two words under a family-first order. The comma settles P1's fork for a leading 'van': the listing form puts a surname before it. Bound given-name pairs and connective joins are not counted as one (P5 builds a given name; P3 leaves 'Ortega y Gasset' three words), stated at _vocab.SURNAME_UNIT_TAGS. The unit walk moves from post_rules into _vocab.unit_ends, shared by post_rules' fold (full chain) and the C1 counts (fold reach). Segment runs before classify, so it builds its facts from the vocabulary (surname_unit_tags), held to classify's tags by a sweep test with a recorded negative control. 'De La Cruz, M.J. K.L.' now reads given 'M.J.' (Derek's call); 'Van Johnson, Dr.' reads family 'Van Johnson'. No pre-existing corpus name moves at any baseline; the gate exits 0 at all five. Cost: 'John Smith, PhD' 209 -> 217 frames, other measured names unchanged. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 8 + docs/design/mechanisms.md | 2 +- docs/design/rules.md | 32 +++- docs/release_log.rst | 2 + nameparser/_pipeline/_assign.py | 19 ++- nameparser/_pipeline/_post_rules.py | 39 +---- nameparser/_pipeline/_segment.py | 9 +- nameparser/_pipeline/_vocab.py | 164 ++++++++++++++++++- tests/v2/cases.py | 98 ++++++++++- tests/v2/pipeline/test_assign.py | 20 ++- tests/v2/pipeline/test_classify.py | 35 +++- tests/v2/test_ledger_guards.py | 89 ++++++++-- tools/differential/compare.py | 8 + tools/differential/corpus_rules.jsonl | 7 + tools/differential/corpus_shapes.jsonl | 10 +- tools/differential/expected_since_1.4.0.toml | 31 ++++ tools/differential/expected_since_2.0.0.toml | 25 ++- tools/differential/expected_since_2.1.0.toml | 25 ++- tools/differential/expected_since_2.2.0.toml | 25 ++- tools/differential/expected_since_2.3.0.toml | 25 ++- 20 files changed, 586 insertions(+), 87 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 8418c3b8..73a23a7f 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -880,6 +880,14 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had brought back a split of `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd'. That split is not only an intermediate state of #552: 2.2.0 and 2.3.0 shipped it (2.0.0 and 2.1.0 read family 'Smith vd Ma'), #530 removed it in this cycle, and #552 kept it removed. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them. The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. It is also 1.4.0's reading, verified against the released wheel the same day: 1.4.0 read suffix 'vd Ma' and 'VD MA', and 2.0.0 through 2.3.0 read the whole name as the family. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` before a comma and after a family comma, and the silent mixed-case `Doe, Jane PhD vd Ma`. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. +- 2026-10-01 (Derek), #575 — THE COUNT BEFORE THE COMMA COUNTS A PARTICLE SURNAME ONCE. Every count rules.md#C1 makes of the part before the comma — v1's "more than one word" for an unambiguous credential, the ambiguous class's name-word count (#289, #544), and assign's two-name-word test for the positional read after a comma followed by no name word (#296/#325) — counted tokens or pieces, so `De La Cruz` was three words and `van der Berg` two (the ambiguous `van` stands as its own piece, P1's fork). This cycle's count had therefore read `De La Cruz, Ed` as a credential comma with no given name and split `van der Berg, MA` into given 'van', family 'der Berg' — an unreleased regression against 2.3.0, which read both as the listing form — and every release split `van der Berg, PhD`. DECIDED: a particle and the name word it attaches to are one word in all three counts, so a particle surname reads as a one-word surname does. + THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. + THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. Under the default order the positional read P1 applies makes the same part all surname anyway, so the narrower reach costs nothing there. `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. + ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set, and both exclusions are stated where the count lives (`_vocab.SURNAME_UNIT_TAGS`). A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. A connective join is P3's to decide, and P3 declines the commonest connective surname outright — a single-letter connective in a three-word name stays a name word — so `Ortega y Gasset` is three words even standing alone; reproducing P3's conditions before classify would copy P3. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). + `De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name. + SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`) while assign intersects classify's tags with the same set. `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored. Negative control: with that mirror removed, the test reports exactly those two. + BLAST RADIUS, measured 2026-10-01 with the gate at all five baselines: no pre-existing corpus name moves; the movers are the change's own rules.md and case-row names. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname names diff there only by this cycle's #289 count and report, which is the regression this fixes having never shipped. + COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. ### T1 — separators, not joiners diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 1c8637fa..82096abc 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -29,7 +29,7 @@ Problem shape. A rule needs to know how words were JOINED (chained titles, parti ## UNIT-PARTITION — count the units the joining rules built -Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all. How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_post_rules.py (`_unit_end`, `_units`); rules.md P1 is the counting rule, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. +Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all. How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma (`_vocab.name_word_count`, `surname_unit_count`, and assign's positional-read test), which run partly before classify and so build their facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name and P3 declines the commonest connective surname, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. ## MARK-DONT-STRIP — record the decision, keep the fact diff --git a/docs/design/rules.md b/docs/design/rules.md index d55988ff..2f3a8714 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1669,6 +1669,18 @@ C1. Rationale: a credential run after the comma means the name is in than one word precedes the comma; otherwise it reads as the listing form, the part before the comma being the family name. Only the part after the first comma decides. + Wherever this rule counts the words before the comma, a particle + and the name word it attaches to are one word, as P2 joins them: + the listing form puts a surname before the comma, and a particle + surname is one surname ('De La Cruz, Ed' reads as 'Royce, Ed' + does, 'van der Berg, PhD' as 'Berg, PhD'). That settles P1's fork + for a leading particle that could be a given name, the comma + being the evidence that the part is a surname. The particle + reaches the one word it attaches to and no further, so a word + after that one is a second word. A connective join and a bound + given name are not one word here: a single-letter connective in + a three-word name stays a name word (P3), and a bound given name + builds a given name rather than a surname (P5). For the ambiguous credential class — a bare acronym the vocabulary marks as also an ordinary name, and a word admitted to the class by shape, which S3 defines and bounds — the @@ -1694,9 +1706,10 @@ C1. Rationale: a credential run after the comma means the name is in company has it, and where nothing but words of both the suffix and the title vocabulary stands in front of them, those are the titles of the given part the initials open, as S2 reads them - there. Paired initials that are not the only shape word in the - part are a credential however little else speaks for them, since - no one writes a person's initials as two dotted groups. The same count reads a part of two or + there. Behind two or more name words, paired initials that are + not the only shape word in the part are a credential however + little else speaks for them, since no one writes a person's + initials as two dotted groups. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any @@ -1833,7 +1846,16 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, A.B." → given="A.B." "John Smith, A.B." → ambiguities=("suffix-or-name",) "John Smith, X.Y. P.Q." → suffix="X.Y. P.Q." - "De La Cruz, M.J. K.L." → ambiguities=("suffix-or-name",) + "De La Cruz, M.J. K.L." → given="M.J." + "De La Cruz, M.J. K.L." → ambiguities=("suffix-or-name", "suffix-or-name") + "De La Cruz, Ed" → given="Ed" + "van der Berg, MA" → family="van der Berg" + "van der Berg, MA" → suffix="MA" + "Van Buren, Ed" → family="Van Buren" + "van der Berg, PhD" → family="van der Berg" + "John van Buren, Ed" → suffix="Ed" · boundary + "Ortega y Gasset, Ed" → suffix="Ed" · boundary + "abdul Salam, Ed" → suffix="Ed" · boundary "John Smith, PhD X.Y." → suffix="PhD X.Y." "García Márquez, G.J" → given="G.J" "De La Cruz, M.J. PhD" → given="M.J." · boundary @@ -1880,7 +1902,7 @@ C1. Rationale: a credential run after the comma means the name is in V` reads the suffix and `Smith, John PhD I.` continues the run, while adding a suffix comma after either turns that same letter into the middle initial. - history: decisions.md#C1 · interacts: H2, P2, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py + history: decisions.md#C1 · interacts: H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py C2. Rationale: text beyond the recognized comma parts should be taken in without silent guessing. diff --git a/docs/release_log.rst b/docs/release_log.rst index 9631c654..75cc3c01 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -6,6 +6,8 @@ Release Log **Behavior Changes** + - **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where every release gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. A given name in front still makes two words, so ``John van Buren, Ed`` keeps suffix ``Ed``. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575) + - **Fix a one-letter connective joining a name that gives no sign it is a connective.** ``HumanName("jose e maria santos")`` gives first ``jose``, middle ``e maria``, last ``santos``, where 1.4.0 through 2.3.0 gave first ``jose e maria``; and ``JUAN GARCIA Y LOPEZ`` gives last ``GARCIA Y LOPEZ``, where every release since 1.4.0 read the bare capital as an initial and gave middle ``GARCIA Y``. A single letter is an initial where the writing says so -- a bare Latin capital in a name that is not written wholly in one case -- and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there: ``e`` reads as an initial and ``y`` joins. Mixed-case input is untouched in both directions: ``Jose e Maria Santos`` still gives first ``Jose e Maria`` and ``Jose E Maria Santos`` still gives middle ``E Maria``. Short names move in the derived views rather than the fields, P3's three-word carve-out being unchanged: ``parse("john e smith").initials()`` is ``j. e. s.`` where 2.3.0 gave ``j. s.``, and ``HumanName("john e smith").capitalize()`` gives ``John E Smith`` where 2.3.0 gave ``John e Smith``; ``JUAN Y GARCIA`` moves its ``capitalize()`` the same way in reverse, giving ``Juan y Garcia``, while its initials do not move at all: ``parse(...).initials()`` is ``J. Y. G.``, what 2.3.0 gave and what 1.4.0's own view gave, the ``Y`` holding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461). ``HumanName.initials()`` agrees with the core on both -- see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (``Хосе И Мария Сантос`` still gives first ``Хосе И Мария``), and Arabic ``و`` never enters the rule, having no case to be written against. A ``Lexicon`` knob decides which letters are marked, so the reading is configurable rather than fixed. See the ``P3`` entry of ``docs/design/decisions.md`` (closes #383, closes #479) - **Fix HumanName.initials() reading a one-letter connective by vocabulary and written shape instead of by the parse.** ``HumanName("john e smith").initials()`` gives ``j. e. s.``, where every release from 1.4.0 through 2.3.0 gave ``j. s.``; ``JUAN GARCIA Y LOPEZ`` gives ``J. G. L.`` where 2.3.0 gave ``J. G. Y. L.``; ``JUAN Y GARCIA`` does not move at all this cycle, giving ``J. Y. G.`` on both surfaces as 2.3.0 and 1.4.0 did, since the connective-initials fix further down this list (closes #461) gives its ``Y`` an initial again. That first name read 1.4.0's way at 2.3.0 and only there: 2.0.0 through 2.2.0 already gave today's answer, by the unrelated bug the 2.3.0 note below records as fixed (the facade dropping a bare capital that is also a one-letter conjunction, #462), so against those three releases it does not move at all. The v1 facade decided whether a word was the connective by looking the word up and checking its shape, while ``parse(...).initials()`` read the tag the parse recorded -- so the change above, which reads a single letter in a one-case name from the vocabulary rather than from its case, moved one view and not the other. Both views of a parse now give the same answer. Mixed-case names are untouched on both, the writing having decided the letter: ``John E Smith`` is still ``J. E. S.`` and ``Scott E. Werner`` still ``S. E. W.``. So is a one-case name whose letter is outside the marked set -- ``maria y lopez`` is still ``m. l.``, ``y`` having joined before this release and after it. Two costs, and both match what ``capitalize()`` has always done: editing ``C.conjunctions`` after a name is parsed no longer changes its initials until ``full_name`` is assigned again, and a name restored from a pickle, copied with ``copy.copy``/``copy.deepcopy`` (the same state hooks), or built from keyword fields (``HumanName(first=..., middle=..., last=...)``) carries no tags, so its initials come from the vocabulary and can differ from a fresh parse of the same string. One private break, stated because a v1 subclass can hit it: an override of ``_process_initial`` written to v1's ``(name_part, firstname=False)`` signature now raises ``TypeError`` the first time ``initials()`` runs, since ``initials()`` passes the part's tokens. Such an override has to accept a ``tokens`` keyword *and pass it on* -- ``return super()._process_initial(name_part, firstname, tokens=tokens)`` -- to receive this fix. Widening the signature without forwarding still works, but on the pre-#528 STRING path: the token call hands the override the group's own text as ``name_part`` rather than an empty placeholder, so ``john e smith`` initials ``j. s.`` under such an override, not the ``j. e. s.`` above. A subclass overriding one of the public ``first_list``, ``middle_list`` or ``last_list`` properties keeps working too: that member takes the pre-2.4 vocabulary reading instead of the change above, while an un-overridden member still moves. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #528) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index 7785fe88..5b36eed5 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -71,7 +71,8 @@ from nameparser._lexicon import Lexicon from nameparser._pipeline._vocab import ( - effective_script, is_suffix_lenient, resolve_script_set, + SURNAME_UNIT_TAGS, effective_script, is_suffix_lenient, + resolve_script_set, unit_ends, ) from nameparser._pipeline._pieces import ( anchor_in_reach, credential_at_the_given_slot, given_slot_anchors, @@ -1079,9 +1080,19 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: f"read as " f"{'a credential' if suffix_here else 'a name'}", (i2,))) - if reading is not None and sum( - 1 for k, piece in enumerate(fam_pieces) - if not is_suffix_piece(piece, fam_tags[k], tokens)) > 1: + # The two name words are counted as UNITS (#575, + # mechanisms.md#UNIT-PARTITION): group leaves an ambiguous + # leading particle a piece of its own (P1's fork), so 'van der + # Berg, PhD' holds two name pieces and one surname, and read + # positionally lost 'van' to the given name. The comma is the + # evidence that settles that fork: the listing form puts a + # surname before it. The same facts as segment's count + # (`_vocab.SURNAME_UNIT_TAGS`), so the two agree. + if reading is not None and len(unit_ends([ + tokens[i].tags & SURNAME_UNIT_TAGS + for k, piece in enumerate(fam_pieces) + if not is_suffix_piece(piece, fam_tags[k], tokens) + for i in piece], chain=False)) > 1: order = _assign_main(0, state, tokens, ambiguities) else: for k, piece in enumerate(fam_pieces): diff --git a/nameparser/_pipeline/_post_rules.py b/nameparser/_pipeline/_post_rules.py index ec030523..00f0cd99 100644 --- a/nameparser/_pipeline/_post_rules.py +++ b/nameparser/_pipeline/_post_rules.py @@ -28,7 +28,7 @@ AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure, WorkToken, _NEVER_FLIPPED, comma_bucket, copy_with, ) -from nameparser._pipeline._vocab import delimiter_cores +from nameparser._pipeline._vocab import delimiter_cores, unit_ends from nameparser._policy import PatronymicRule from nameparser._types import ( FOLDED_TAG, UNJOINED_CONJUNCTION_TAG, UNJOINED_TAG, AmbiguityKind, Role, @@ -256,40 +256,6 @@ def _retag(tokens: list[WorkToken], i: int, role: Role) -> None: # rule counts them" # rules.md#P5: "a recognized bound given-name word joins the word # after it into one given name" -def _unit_end(tokens: list[WorkToken], idx: list[int], i: int) -> int: - """One past the end of the unit starting at `idx[i]`. - - RECURSIVE, and that is the whole point: what a conjunction or a - bound given-name word joins is the next UNIT, not the next word. - Absorbing a single index instead strands a particle at the end of - the unit, severed from the words it chains -- "de la Vega y la - Vega" cut between `la` and `Vega`, reporting family - "de la Vega y la", which is the same defect as a bare particle - opening the given name, mirrored.""" - if "particle" in tokens[idx[i]].tags: - j = i - while j + 1 < len(idx) and "particle" in tokens[idx[j + 1]].tags: - j += 1 - # ... then the words it joins, stopping where the next - # particle starts a group of its own, at a suffix word (the - # stop _group's chain uses), or at a conjunction, which the - # shared loop below joins to the whole unit after it rather - # than to the one word after it. - while (j + 1 < len(idx) - and "particle" not in tokens[idx[j + 1]].tags - and "conjunction" not in tokens[idx[j + 1]].tags - and "vocab:suffix" not in tokens[idx[j + 1]].tags): - j += 1 - end = j + 1 - else: - end = i + 1 - if "vocab:bound-given" in tokens[idx[i]].tags and end < len(idx): - end = _unit_end(tokens, idx, end) - while end + 1 < len(idx) and "conjunction" in tokens[idx[end]].tags: - end = _unit_end(tokens, idx, end + 1) - return end - - def _units(tokens: list[WorkToken], idx: list[int]) -> list[list[int]]: """`idx` split into the units other rules COUNT: one name word each, except where another rule has already made several words one @@ -316,8 +282,7 @@ def _units(tokens: list[WorkToken], idx: list[int]) -> list[list[int]]: out separately.""" units: list[list[int]] = [] i = 0 - while i < len(idx): - end = _unit_end(tokens, idx, i) + for end in unit_ends([tokens[k].tags for k in idx]): units.append(list(idx[i:end])) i = end return units diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 2c2e3776..3fff4d02 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -56,6 +56,7 @@ ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, caps_shape_candidate, is_one_case, is_paired_initials, is_single_letter_numeral, is_wholly_suffix, name_word_count, + surname_unit_count, run_word_fold, ) from nameparser._types import AmbiguityKind @@ -391,9 +392,15 @@ def class_run(seg: tuple[int, ...]) -> bool: case_class() pre_comma_names = name_word_count(texts(groups[0]), state.lexicon, state.policy) + # rules.md#C1's "more than one word precedes the comma" counts a + # particle chain and a connective join as one word (#575): v1's + # token count split 'van der Berg, PhD' into given 'van', family + # 'der Berg'. The token count stays first, so a one-token part + # never builds the units. structure = ( Structure.SUFFIX_COMMA - if ((suffixy(groups[1]) and len(groups[0]) > 1) + if ((suffixy(groups[1]) and len(groups[0]) > 1 + and surname_unit_count(texts(groups[0]), state.lexicon) > 1) or (pre_comma_names is not None and pre_comma_names >= 2)) else Structure.FAMILY_COMMA) ambiguities = list(state.ambiguities) diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 84fd6dcf..ced988a4 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -43,7 +43,7 @@ import functools import re import unicodedata -from collections.abc import Callable, Iterable, Sequence +from collections.abc import Callable, Iterable, Sequence, Set from typing import Literal from nameparser._lexicon import ( @@ -775,6 +775,134 @@ def ambiguous_class_candidate(text: str, lexicon: Lexicon, return False +# mechanisms.md#UNIT-PARTITION: "A rule that counts name words counts +# those units, and takes each whole or not at all." The walk is shared +# by two readers that hold the facts in different forms: post_rules +# reads classify's TAGS, and segment, which runs before classify has +# tagged anything, builds the same tag names from the vocabulary +# (`surname_unit_tags`). One walk over tag sets, so the two cannot +# disagree about where a unit ends (mechanisms.md#ONE-PREDICATE-PER- +# QUESTION) -- only, at worst, about a token's facts, which +# test_vocab's agreement test pins. +def unit_ends(tags: Sequence[Set[str]], chain: bool = True) -> list[int]: + """The END (one past the last index) of each unit of `tags`, in + order, partitioning `range(len(tags))`: one name word each, except + where another rule has already made several words one name -- a + particle and the words it chains (P2), a conjunction-joined run + (P3), and a bound given-name word with the word it completes (P5). + Each element is the tag set of one token, read for "particle", + "conjunction", "vocab:bound-given" and "vocab:suffix". + + `chain` picks how far a particle reaches. True is P2's chain: the + particle run and every word after it up to the next stop ("van der + Berg Smith" is one unit). False is P1's fold reach: the particle + run and the ONE unit it attaches to, so "de Mesnil Jean" is two -- + the count rules.md#C1 needs before a comma, where under a + family-first order that part is family 'de Mesnil', given 'Jean' + (#575). + + RECURSIVE, and that is the whole point: what a conjunction or a + bound given-name word joins is the next UNIT, not the next word. + Absorbing a single index instead strands a particle at the end of + the unit, severed from the words it chains -- "de la Vega y la + Vega" cut between `la` and `Vega`, reporting family + "de la Vega y la", which is the same defect as a bare particle + opening the given name, mirrored.""" + n = len(tags) + + def end_of(i: int) -> int: + if "particle" in tags[i] and not chain: + j = i + 1 + while j < n and "particle" in tags[j]: + j += 1 + end = (end_of(j) if j < n and "conjunction" not in tags[j] + and "vocab:suffix" not in tags[j] else j) + elif "particle" in tags[i]: + j = i + while j + 1 < n and "particle" in tags[j + 1]: + j += 1 + # ... then the words it joins, stopping where the next + # particle starts a group of its own, at a suffix word (the + # stop _group's chain uses), or at a conjunction, which the + # shared loop below joins to the whole unit after it rather + # than to the one word after it. + while (j + 1 < n + and "particle" not in tags[j + 1] + and "conjunction" not in tags[j + 1] + and "vocab:suffix" not in tags[j + 1]): + j += 1 + end = j + 1 + else: + end = i + 1 + if "vocab:bound-given" in tags[i] and end < n: + end = end_of(end) + while end + 1 < n and "conjunction" in tags[end]: + end = end_of(end + 1) + return end + + ends: list[int] = [] + i = 0 + while i < n: + i = end_of(i) + ends.append(i) + return ends + + +#: The facts rules.md#C1's count before a comma reads (#575): a +#: particle chain is one surname, and a suffix word stops it. Neither +#: of the other two joins `unit_ends` knows is read there. A bound +#: given-name pair builds a GIVEN name, and P5 gives up a family word +#: where the name has no other ('abdul Salam' alone is given 'abdul', +#: family 'Salam'), so before a comma it is not one surname. A +#: connective join is P3's to decide, and P3 declines the commonest +#: connective surname outright -- a single-letter connective in a +#: three-word name stays a name word, so 'Ortega y Gasset' is three -- +#: on conditions (that exception, the case fork, generational +#: connectives) segment could only rebuild by copying P3. Segment +#: builds these from the vocabulary (`surname_unit_tags`); assign +#: intersects classify's tags with this set, so the two counts read +#: one set of facts. +SURNAME_UNIT_TAGS = frozenset({"particle", "vocab:suffix"}) +_PARTICLE_ONLY = frozenset({"particle"}) +_SUFFIX_ONLY = frozenset({"vocab:suffix"}) +_NO_TAGS: frozenset[str] = frozenset() + + +def surname_unit_tags(text: str, lexicon: Lexicon) -> frozenset[str]: + """`SURNAME_UNIT_TAGS` for one token from the vocabulary alone -- + segment's view of classify's tags, built before classify runs, and + classify's own tests: particle membership, `suffix_as_written`, and + the period-joined derivation. Kept from drifting by + test_classify.test_surname_unit_tags_agree_with_classify.""" + n = _normalize(text) + # classify's whole-token test, then its period-joined derivation + # for a token no whole-token title or suffix claimed ('JD.CPA'), + # asked only of a token with a period, which it needs + suffix = suffix_as_written(n, text, lexicon) or ( + "." in text and n not in lexicon.titles + and period_joined_vocab(text, lexicon) == "suffix") + if n in lexicon.particles: + return SURNAME_UNIT_TAGS if suffix else _PARTICLE_ONLY + return _SUFFIX_ONLY if suffix else _NO_TAGS + + +def surname_unit_count(texts: Sequence[str], lexicon: Lexicon) -> int: + """How many units (`unit_ends`) the part before a comma holds, a + particle chain (P2) and a connective join (P3) each counting once: + rules.md#C1's count before the comma, for the credential reading + of a part that is wholly suffix words (#575), with a particle + reaching as P1's fold does (`unit_ends`'s `chain=False`). 'van der + Berg' is one surname, so 'van der Berg, PhD' reads as 'Berg, PhD' + does.""" + # With no particle, every token is a unit of its own, so the count + # is the token count: the suffix tests and the walk are skipped for + # the commonest comma names ('John Smith, PhD'). + if not any(_normalize(t) in lexicon.particles for t in texts): + return len(texts) + return len(unit_ends([surname_unit_tags(t, lexicon) for t in texts], + chain=False)) + + def name_word_count(texts: Sequence[str], lexicon: Lexicon, policy: Policy) -> int: """How many of these texts are NAME words -- not suffix @@ -792,16 +920,38 @@ def name_word_count(texts: Sequence[str], lexicon: Lexicon, leading peel reads it as a title, and that divergence is recorded once, at `is_title_shaped` itself (mechanisms.md#ONE-PREDICATE-PER-QUESTION). + + It counts UNITS holding a name word, not tokens (#575, + mechanisms.md#UNIT-PARTITION): a particle chain (P2) builds one + family name, so 'De La Cruz' is one name word and 'De La Cruz, Ed' + reads as 'Royce, Ed' does. Why the other two joins are not read + here: `SURNAME_UNIT_TAGS`. """ predicate = (is_suffix_lenient if policy.lenient_comma_suffixes else is_suffix_strict) - n = 0 + + # One pass: each token's name-word verdict, and whether any token + # is a particle. With none, every token is a unit of its own and + # the count is a sum, so the commonest comma names never build the + # units -- and pay no frame the token loop did not already pay. + names: list[bool] = [] + has_particle = False for text in texts: - if (predicate(text, lexicon) or _normalize(text) in lexicon.titles - or is_title_shaped(text)): - continue - n += 1 - return n + n = _normalize(text) + if n in lexicon.particles: + has_particle = True + names.append(not (predicate(text, lexicon) or n in lexicon.titles + or is_title_shaped(text))) + if not has_particle: + return sum(names) + count = 0 + start = 0 + for end in unit_ends([surname_unit_tags(t, lexicon) for t in texts], + chain=False): + if any(names[start:end]): + count += 1 + start = end + return count def is_wholly_suffix(texts: Sequence[str], lexicon: Lexicon, diff --git a/tests/v2/cases.py b/tests/v2/cases.py index d767d031..44ab2500 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -992,6 +992,85 @@ def _check_cjk_shape_purity(self) -> None: "family around a suffix 'vd'. 2.3.0 read family 'Smith " "Ma', suffix 'vd', and 1.4.0 last 'Smith', suffix 'vd, " "Ma'"), + # #575: rules.md#C1's count before the comma counts a particle + # chain (P2) as ONE name word -- it builds one family name -- so a + # particle surname reads as a one-word surname does ('Royce, Ed', + # 'Berg, MA', 'Berg, PhD'). + Case("a_particle_surname_before_a_comma_is_one_name_word", + "De La Cruz, Ed", + {"given": "Ed", "family": "De La Cruz"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="the count read 'De La Cruz' as three name words and " + "flipped to a credential comma, leaving no given name. " + "2.3.0 read given 'Ed' too; 1.4.0 read suffix 'Ed'", + shape=2), + Case("a_lowercase_particle_surname_before_a_comma_is_one_name_word", + "de la Cruz, Ma", + {"given": "Ma", "family": "de la Cruz"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="as 'Cruz, Ma' reads. 2.3.0 read given 'Ma' too; 1.4.0 " + "read suffix 'Ma'", + shape=2), + Case("an_ambiguous_particle_chain_before_a_comma_stays_whole", + "van der Berg, MA", + {"family": "van der Berg", "suffix": "MA"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="one name word before the comma, so C1 reads the case " + "and capitals in a mixed-case name make 'MA' the " + "credential, as 'Berg, MA' reads. The count had split " + "the chain: given 'van', family 'der Berg'. 2.3.0 read " + "given 'MA', family 'van der Berg'", + shape=2), + Case("a_leading_particle_that_could_be_a_given_name_is_settled_by_the_comma", + "Van Buren, Ed", + {"given": "Ed", "family": "Van Buren"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="standing alone 'Van Buren' is P1's fork (given 'Van'); " + "the listing form puts a surname before the comma, so " + "the comma settles it. The count had read two name " + "words: given 'Van', family 'Buren', suffix 'Ed'. " + "2.3.0 read given 'Ed', family 'Van Buren'", + shape=2), + Case("a_connective_surname_is_not_one_name_word_before_a_comma", + "Ortega y Gasset, Ed", + {"given": "Ortega", "middle": "y", "family": "Gasset", + "suffix": "Ed"}, + ambiguities=("suffix-or-name",), + notes="#575 boundary: only a particle chain counts as one name " + "word. P3 leaves a single-letter connective in a " + "three-word name a name word, so 'Ortega y Gasset' is " + "three even standing alone (given 'Ortega', middle 'y'), " + "and the count follows it. 2.3.0 read given 'Ed'", + shape=3), + Case("a_particle_surname_before_an_unambiguous_credential_stays_whole", + "van der Berg, PhD", + {"family": "van der Berg", "suffix": "PhD"}, + classification="fix(#575)", + notes="as 'Berg, PhD' reads. Every release split the chain on " + "v1's token count: given 'van', family 'der Berg', " + "suffix 'PhD'", + shape=2), + Case("a_particle_surname_after_a_given_name_is_still_a_full_name", + "John van Buren, Ed", + {"given": "John", "family": "van Buren", "suffix": "Ed"}, + ambiguities=("suffix-or-name",), + notes="#575 boundary: a given name and a particle surname are " + "two name words, so the part after the comma is still " + "the credential run", + shape=3), + Case("a_bound_given_pair_before_a_comma_counts_as_two_name_words", + "abdul Salam, Ed", + {"given": "abdul", "family": "Salam", "suffix": "Ed"}, + ambiguities=("suffix-or-name",), + notes="#575 boundary: a bound pair (P5) builds a GIVEN name and " + "gives up a family word when the name has no other " + "('abdul Salam' alone is given 'abdul', family 'Salam'), " + "so it is not one surname before a comma", + shape=3), Case("the_anchor_does_not_reach_across_a_maiden_clause", "Jane Doe Jr. nee Smith Ma", {"given": "Jane", "family": "Doe", "suffix": "Jr.", @@ -2347,16 +2426,17 @@ def _check_cjk_shape_purity(self) -> None: "takes for a name. 1.4.0 read given 'X.Y.', middle " "'P.Q.'", shape=3), - Case("two_pairs_behind_a_two_word_surname_report_the_flip", + Case("two_pairs_behind_a_particle_surname_open_the_given_part", "De La Cruz, M.J. K.L.", - {"family": "De La Cruz", "suffix": "M.J. K.L."}, - classification="fix(#563)", - ambiguities=("suffix-or-name",), - notes="#563 second review round: the cost of the two-pair " - "line is a surname with no given name, so the flip it " - "makes is reported rather than silent (rules.md#C1). " - "1.4.0 read given 'M.J.', middle 'K.L.'", - shape=3), + {"given": "M.J.", "family": "De La Cruz", "suffix": "K.L."}, + classification="fix(#575)", + ambiguities=("suffix-or-name", "suffix-or-name"), + notes="'De La Cruz' is one name word (#575), so the count no " + "longer reaches the part and it reads as 'Cruz, M.J. " + "K.L.' does (Derek, 2026-10-01). #563 had read it as a " + "flipped credential run with no given name. 1.4.0 read " + "given 'M.J.', middle 'K.L.'", + shape=2), Case("a_class_member_in_front_of_paired_initials_does_not_speak", "García Márquez, Ed G.J.", {"given": "Ed", "family": "García Márquez", "suffix": "G.J."}, diff --git a/tests/v2/pipeline/test_assign.py b/tests/v2/pipeline/test_assign.py index 50af5ebd..0c0733d6 100644 --- a/tests/v2/pipeline/test_assign.py +++ b/tests/v2/pipeline/test_assign.py @@ -648,14 +648,28 @@ def test_non_title_post_comma_segment_is_untouched() -> None: def test_positional_segment_zero_reports_the_particle_fork() -> None: # The comma no longer fixed the family, so the leading ambiguous # particle IS a live fork again -- emitted at the site that decides - # it, per the ambiguity doctrine. - out = _assigned("Van Johnson, Dr.") + # it, per the ambiguity doctrine. Three words, since the read needs + # two name words counted as units (#575): the particle and the word + # it attaches to are one. + out = _assigned("Van Johnson Smith, Dr.") assert _by_role(out, Role.GIVEN) == "Van" - assert _by_role(out, Role.FAMILY) == "Johnson" + assert _by_role(out, Role.FAMILY) == "Smith" assert [a.kind for a in out.ambiguities] == \ [AmbiguityKind.PARTICLE_OR_GIVEN] +def test_a_lone_particle_surname_before_a_title_only_comma_is_the_family() -> None: + # #575: 'Van Johnson' is ONE name word before the comma -- a + # particle and the word it attaches to -- so there are not two to + # read positionally, and the listing form's surname reading + # settles P1's fork. This test pinned given 'Van', family + # 'Johnson' and a particle-or-given report until then. + out = _assigned("Van Johnson, Dr.") + assert _by_role(out, Role.FAMILY) == "Van Johnson" + assert _by_role(out, Role.GIVEN) == "" + assert out.ambiguities == () + + def test_a_credential_run_after_a_family_comma_reads_as_suffixes() -> None: # 'Smith, Jr.' -- the peel's whole-segment exception claimed this # even with 'jr' out of TITLES, because is_leading_title also diff --git a/tests/v2/pipeline/test_classify.py b/tests/v2/pipeline/test_classify.py index ed2666d8..58366edf 100644 --- a/tests/v2/pipeline/test_classify.py +++ b/tests/v2/pipeline/test_classify.py @@ -14,7 +14,8 @@ ) from nameparser._pipeline._tokenize import tokenize from nameparser._pipeline._vocab import ( - ambiguous_class_candidate, ambiguous_class_member, caps_shape_candidate, + SURNAME_UNIT_TAGS, ambiguous_class_candidate, ambiguous_class_member, + caps_shape_candidate, surname_unit_tags, ) from nameparser._policy import Policy from nameparser._types import AmbiguityKind, Role @@ -731,3 +732,35 @@ def test_the_caps_shape_is_a_testable_predicate() -> None: assert (_tags_by_text("John Smith X.Y", policy=on)["X.Y"] == _tags_by_text("John Smith X.Y")["X.Y"]) + + +def _surname_unit_sweep() -> list[str]: + lex = Lexicon.default() + words: set[str] = {"Smith", "Ma", "M.D.", "Ph.D.", "JD.CPA", "Msc.Ed.", + "Lt.Gov.", "X.Y.Z.", "V.", "I", "v", "Jr.", "de."} + for field in ("particles", "suffix_acronyms", "suffix_words", + "conjunctions", "bound_given_names", "titles"): + for w in getattr(lex, field): + if " " not in w: + words |= {w, w.title(), w.upper()} + return sorted(words) + + +def test_surname_unit_tags_agree_with_classify() -> None: + # #575: rules.md#C1's count before the comma reads two facts per + # token -- particle, suffix -- and `segment` must build them from + # the vocabulary (`_vocab.surname_unit_tags`), since it runs before + # classify has tagged anything; assign reads classify's own tags + # for the same count. Swept over every single-word entry of the + # vocabularies a token can be tagged from, in three casings, plus + # the period shapes classify derives a suffix from ('JD.CPA', + # 'Msc.Ed.' -- the two this sweep caught on the first draft). + lex = Lexicon.default() + disagree = [] + for word in _surname_unit_sweep(): + tags = _tags_by_text(f"Smith {word}, John", lexicon=lex).get(word) + if tags is None: # tokenize split it; nothing to compare + continue + if surname_unit_tags(word, lex) != tags & SURNAME_UNIT_TAGS: + disagree.append(word) + assert disagree == [] diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 94b647e9..f6a55830 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1381,6 +1381,10 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: "fix(#289) a written case contrast decides a bare ambiguous acronym": ("JOHN SMITH MA", "ANH DO", "anh van do", "jack ma", "Jack Ma", "Jack Ma.", "毛泽东, MA", "Smith, A.B.C.", "Jack M.A."), + # #575: two name words before the comma, a given name and a + # particle surname, keep the credential reading. + "fix(#575) a particle surname before a comma is one name word": + ("John van Buren, Ed", "John van der Berg, PhD"), # #516's dotted rule must never reach the shapes the vocabulary # SETTLES (a whole-token match, or a surviving chunk claim), the # leading dotted runs rules.md#S2 excludes by position, the single @@ -2458,13 +2462,19 @@ class _LatinCopy(NamedTuple): "John de Ma", "John van der Berg Ma", r"Smith Jr\., MA", r"Smith Jr\., Ma", "Smith, MA", "abdul Smith Berg Ma", "abdul Smith Ma", "john smith, ma"}), - frozenset({"Davis Royce, Ed", r"Doe, Dr\. MA", "Doe, MA", "Doe, MA PhD", - r"Doe, Mr\. MA PhD", "Freiherr von Berg MA", "JOHN SMITH, MA", - "Jack MA", r"Jack MA\.", "Jack Wei Ma", r"John Prof\. MA", + # 2026-10-01, #575: the 2.x copy gains seven C1 examples, comma + # names whose post-comma word is a listed member -- no set copied. + frozenset({"Davis Royce, Ed", "De La Cruz, Ed", r"Doe, Dr\. MA", + "Doe, MA", "Doe, MA PhD", r"Doe, Mr\. MA PhD", + "Freiherr von Berg MA", "JOHN SMITH, MA", "Jack MA", + r"Jack MA\.", "Jack Wei Ma", r"John Prof\. MA", "John Smith Ma", "John Smith, Ed", "John Smith, MA", - "John Smith, Ma", "John de Ma", "John van der Berg Ma", + "John Smith, Ma", "John de Ma", "John van Buren, Ed", + "John van der Berg Ma", "Ortega y Gasset, Ed", r"Smith Jr\., MA", r"Smith Jr\., Ma", "Smith, MA", - "abdul Smith Berg Ma", "abdul Smith Ma", "john smith, ma"}), + "Van Buren, Ed", "abdul Salam, Ed", "abdul Smith Berg Ma", + "abdul Smith Ma", "de la Cruz, Ma", "john smith, ma", + "van der Berg, MA"}), frozenset({r"García Márquez, G\.J\.R\.", r"Jack X\.Y\.I\.", r"John Smith B\.Tech\.", r"John Smith C\.H\.A\.", r"John Smith E\.S\.Q\.", r"John Smith Q\.W\.E\.R\.T\.", @@ -3067,6 +3077,9 @@ class _LatinCopy(NamedTuple): frozenset({"Doe, Jane nee Smith PhD MEng", "Jane Doe nee Smith PhD MA", "Jane Doe nee Smith PhD MEng"}), frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), + # 2026-10-01, #575: rules.md#C1's particle-surname examples at + # 1.4.0, literal names that copy no set. + frozenset({"De La Cruz, Ed", "Van Buren, Ed", "de la Cruz, Ma"}), # 2026-10-01, #554: the run rule gains rules.md#C1's particle-and- # suffix examples, a vocabulary word in front of a member, which # copies no set either. @@ -3775,8 +3788,11 @@ def _claim(rule: dict) -> _Claim: # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, # Jr do', #554's rules.md#C1 examples. Reach, verified name by # name. + # 2026-10-01, #575: 414 -> 422, its eight rules.md#C1 and + # case-row names, every one a comma name. Reach, verified + # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(414, ('given', 'suffix', 'title'), "46663c4aa90a", None), + _Claim(422, ('given', 'suffix', 'title'), "de00a0ee1e57", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3875,8 +3891,15 @@ def _claim(rule: dict) -> _Claim: # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, # Jr do', #554's rules.md#C1 examples. Reach, verified name by # name. + # 2026-10-01, #575: 414 -> 422, its eight rules.md#C1 and + # case-row names, every one a comma name. Reach, verified + # name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(414, ('family', 'given'), "46663c4aa90a", None), + _Claim(422, ('family', 'given'), "de00a0ee1e57", None), + # 2026-10-01, #575: new, 3; 'De La Cruz, Ed', 'Van Buren, Ed', + # 'de la Cruz, Ma'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(3, ('family', 'given', 'suffix'), "a1d6626eb35f", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4137,16 +4160,18 @@ def _claim(rule: dict) -> _Claim: # pinning that a trailing link does not count itself for an # unrelated 'y', which is in the corpus because its row # carries a shape tag. Reach again, verified name by name. + # 2026-10-01, #575: 104 -> 105, 'Ortega y Gasset, Ed'. Reach. "fix(initials-per-word) a connective run initials each word (facade, since 2.0.0)": - _Claim(104, ('_initials',), "bc7dfeba1da5", ('DEFAULT',)), + _Claim(105, ('_initials',), "231bbda6768a", ('DEFAULT',)), # 2026-09-19, #533: 41 -> 43. Two new corpus names opening # with a bound-given word, 'Berg, abdul MA' and 'Berg, abdul # nee Jones MA' -- the P5 pair this change added to record # that the clause form now agrees with the bare one. # 2026-09-26, #535: 43 -> 44, 'Berg, abdul nee Smith V' -- # another opening bound-given word. Reach again. + # 2026-10-01, #575: 44 -> 45, 'abdul Salam, Ed'. Reach. "fix(initials-per-word) a bound-given run initials each word (facade, since 2.0.0)": - _Claim(44, ('_initials',), "4d6f497bebf7", ('DEFAULT',)), + _Claim(45, ('_initials',), "690d9625f120", ('DEFAULT',)), # 2026-09-18: 109 -> 110. One new corpus name, # 'john van der berg ma' -- rules.md#P2's one-case contrast, # and a particle chain like every other member. @@ -4159,8 +4184,12 @@ def _claim(rule: dict) -> _Claim: # 2026-09-30, #563: 115 -> 117, 'De La Cruz, M.J. PhD' and # 'De La Cruz, M.J. K.L.' -- a particle chain, 'De La Cruz'. # Reach, verified name by name. + # 2026-10-01, #575: 117 -> 123, its six particle-surname + # names ('De La Cruz, Ed', 'de la Cruz, Ma', 'van der Berg, + # MA', 'Van Buren, Ed', 'van der Berg, PhD', 'John van Buren, + # Ed'). Reach, verified name by name. "fix(initials-per-word) a particle chain inside a name part initials each word (facade, since 2.0.0)": - _Claim(117, ('_initials',), '0820bd80678d', ('DEFAULT',)), + _Claim(123, ('_initials',), 'ca2598e60110', ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": @@ -4960,7 +4989,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -5457,7 +5491,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6097,7 +6136,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6460,7 +6504,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -7280,6 +7329,14 @@ def test_every_rule_claims_the_recorded_share_of_the_corpus() -> None: "MD, DO, DDS": "fix(#296) a dropped prenominal takes the name position it " "occupies", + # #575 (2026-10-01): the routing rule describes a lone + # post-comma piece LEAVING `first`; these two move a piece INTO + # `given`, because the particle surname before the comma is one + # name word (rules.md#C1). The rule that says so wins. + "De La Cruz, Ed": + "fix(#575) a particle surname before a comma is one name word", + "de la Cruz, Ma": + "fix(#575) a particle surname before a comma is one name word", }, # The two 2.x ledgers had NO section here until #452, and the # coverage assertion below was `<=`, so their absence read as "no @@ -8429,6 +8486,10 @@ def test_a_rule_reaching_no_corpus_name_says_why_it_is_kept() -> None: "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 2), ("fix(#296) a credential-only comma string reads a name and its postnominal", "fix(comma-precomma-family) pre-comma run reads as family, not given", 2), + # 2026-10-01, #575: declared on the #575 rule; the routing + # rule's half of its contest is a pinned overlap, not this. + ("fix(#575) a particle surname before a comma is one name word", + "fix(comma-precomma-family) pre-comma run reads as family, not given", 3), ("fix(#296) a lone post-comma credential is a suffix", "fix(suffix-routing) a two-token name ending in the suffix word jr keeps it in `suffix`", 2), ("fix(#400/#274) bound-given join and maiden consumption in one name", diff --git a/tools/differential/compare.py b/tools/differential/compare.py index 814fa18b..80e484b7 100644 --- a/tools/differential/compare.py +++ b/tools/differential/compare.py @@ -1881,6 +1881,14 @@ class _ShapeMismatch(NamedTuple): # nesting, so neither is the narrower and file order is the # whole decision. The shape is this run's, not guessed. "Smith, MA": ("given", "suffix"), + # #575 (2026-10-01): 1.4.0 read suffix 'Ed'/'Ma' with no given + # name, and the particle surname is now one name word, so the + # word after the comma is the given name. The lone-post-comma + # routing rule's Latin comma regex reaches both and `fields` + # overlap on {given, suffix} without nesting, so the #575 rule + # stands ahead of it on purpose. The shape is this run's. + "De La Cruz, Ed": ("given", "suffix"), + "de la Cruz, Ma": ("given", "suffix"), }, # #501's six, moved here from _WATCHED_DIFFS with their shapes # unchanged. The four CJK rows sit at 2.0.0 alone: the honorific diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index d816b00b..6ef80d64 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -37,6 +37,7 @@ "Carod i Rovira, Josep" "Carod y de Rovira i" "Davis Royce, Ed" +"De La Cruz, Ed" "De La Cruz, M.J. K.L." "De La Cruz, M.J. PhD" "Del Toro" @@ -231,6 +232,7 @@ "John née Jones Smith MA" "John née Jones Smith Ma" "John née Jones Smith V" +"John van Buren, Ed" "John van Mc" "John van der Berg" "John van der Berg Ma" @@ -295,6 +297,7 @@ "Nguyen, Van Le" "Nguyễn Thị Minh Khai" "Nguyễn, Thị Vân" +"Ortega y Gasset, Ed" "Ph. D. Van Johnson" "Ph. D., John" "Prince of Wales Jr" @@ -363,6 +366,7 @@ "Smith. John" "Steven Hardman, MD, DO, DDS" "The Right Hon. the President of the Queen's Bench Division" +"Van Buren, Ed" "Van Johnson" "Vega, Juan de la" "Vincent van Gogh van Beethoven" @@ -379,6 +383,7 @@ "abdul Jr Smith Berg" "abdul Ma Smith" "abdul Ph. D. Smith Berg" +"abdul Salam, Ed" "abdul Sir Smith Berg" "abdul Smith Berg Ma" "abdul Smith Jr Ma" @@ -446,6 +451,8 @@ "tran lac" "van Berg Jan de" "van Gogh" +"van der Berg, MA" +"van der Berg, PhD" "van der Berg, abdul née Jones" "wang meng" "Иван Петрович Абрамович" diff --git a/tools/differential/corpus_shapes.jsonl b/tools/differential/corpus_shapes.jsonl index a7801994..7cbaf51b 100644 --- a/tools/differential/corpus_shapes.jsonl +++ b/tools/differential/corpus_shapes.jsonl @@ -139,6 +139,8 @@ {"name": "Carod i Rovira, Josep", "shape": 2} {"name": "DOE, JOHN MA", "shape": 2} {"name": "DOE, MARY JO MA", "shape": 2} +{"name": "De La Cruz, Ed", "shape": 2} +{"name": "De La Cruz, M.J. K.L.", "shape": 2} {"name": "Doe, Dr. John MA", "shape": 2} {"name": "Doe, Dr. MA", "shape": 2} {"name": "Doe, Dr. MA Smith", "shape": 2} @@ -222,11 +224,14 @@ {"name": "Smith, PhD MEng", "shape": 2} {"name": "Smith, PhD Ma", "shape": 2} {"name": "Smith, meng", "shape": 2} +{"name": "Van Buren, Ed", "shape": 2} +{"name": "de la Cruz, Ma", "shape": 2} {"name": "de la Vega, Juan", "shape": 2} {"name": "doe, jane v phd do", "shape": 2} {"name": "doe, john ma", "shape": 2} +{"name": "van der Berg, MA", "shape": 2} +{"name": "van der Berg, PhD", "shape": 2} {"name": "Davis Royce, Ed", "shape": 3} -{"name": "De La Cruz, M.J. K.L.", "shape": 3} {"name": "De La Cruz, M.J. PhD", "shape": 3} {"name": "Dr. John P. Doe-Ray, CLU, CFP, LUTC", "shape": 3} {"name": "García Márquez, Ed G.J.", "shape": 3} @@ -255,7 +260,10 @@ {"name": "John Smith, X.Y. P.Q.", "shape": 3} {"name": "John Smith, X.Y.Z.", "shape": 3} {"name": "John Smith, X.Y.Z. MA", "shape": 3} +{"name": "John van Buren, Ed", "shape": 3} +{"name": "Ortega y Gasset, Ed", "shape": 3} {"name": "Steven Hardman, MD, DO, DDS", "shape": 3} +{"name": "abdul Salam, Ed", "shape": 3} {"name": "john smith, ma", "shape": 3} {"name": "john smith, md ma", "shape": 3} {"name": "john smith, phd meng", "shape": 3} diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 34268a49..44ae2016 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -766,6 +766,37 @@ issue = "fix(#325) a credential run across a second comma reads as suffixes" name_regex = "(?i)^smith,\\s*jr\\.?,\\s*phd$" fields = ["title", "suffix"] +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle and the name word it attaches to are one +# word, as P2 joins them" -- so the part before the comma is one +# surname and the word after it is the given name, as 'Royce, Ed' +# reads. 1.4.0 counted tokens: 'De La Cruz, Ed' and 'de la Cruz, Ma' +# read suffix 'Ed'/'Ma' with no given name, and 'Van Buren, Ed' read +# given 'Van', family 'Buren', suffix 'Ed'. +# +# Ahead of the routing rule below on purpose: that rule's alternation +# reaches the two 'De La Cruz' spellings, and it describes the +# opposite move (a lone post-comma piece leaving `first`), so letting +# it absorb a piece ARRIVING in `given` would explain the diff with a +# rule that contradicts it. +# +# Literal; the probe 'John van Buren, Ed' (two name words: the +# credential reading stays) is _MUST_NOT_MATCH. +name_regex = "^(?:De La Cruz, Ed|Van Buren, Ed|de la Cruz, Ma)$" +fields = ["given", "family", "suffix"] + +[[change.precedes_narrower]] +issue = "fix(comma-precomma-family) pre-comma run reads as family, not given" +why = """ +The later rule describes 1.4.0 handing a pre-comma word to `given` +that 2.x reads as family, which is half of these moves ('Van' leaving +`given` in 'Van Buren, Ed'). It says nothing about the word AFTER the +comma arriving in `given`, the other half, nor why the particle +surname is one name: that is rules.md#C1's count (#575), which this +rule names. +""" + [[change]] issue = "fix(comma-family) lone post-comma piece routes to suffix/title, not first" # 'Smith, Dr.' / 'Andrews, M.D.': v1 put the lone strict-suffix-or-title diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 209325e5..389c137b 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2524,7 +2524,16 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: seven names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -3878,3 +3887,17 @@ issue = "fix(#296/#531) phd is not a prenominal, so behind a dual title it is th name_regex = "^Smith, MD PhD Ma$" fields = ["title", "given", "middle", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle and the name word it attaches to are one +# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index d96f4faf..65bab4ab 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2411,7 +2411,16 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: seven names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -3789,3 +3798,17 @@ issue = "fix(#296/#531) phd is not a prenominal, so behind a dual title it is th name_regex = "^Smith, MD PhD Ma$" fields = ["title", "given", "middle", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle and the name word it attaches to are one +# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index c99de14a..d47efb3b 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1000,7 +1000,16 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: seven names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -2172,3 +2181,17 @@ issue = "fix(#544) a credential in front anchors a member after a one-word famil name_regex = "^Smith, PhD Ma$" fields = ["given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle and the name word it attaches to are one +# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index c6103486..f203f09f 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -288,7 +288,16 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: seven names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -1433,3 +1442,17 @@ issue = "fix(#544) a credential in front anchors a member after a one-word famil name_regex = "^Smith, PhD Ma$" fields = ["given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle and the name word it attaches to are one +# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] From 7d9da21ca8572b93616153c7ebc081b7d34e19df Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 20:37:40 -0700 Subject: [PATCH 2/5] fix(C1): #575 review round -- title/bound particles, suffix stop, docs Code (both reviews): - A word that is also title vocabulary or a bound given-name head ('Freiherr', 'St', 'Abu') is no particle to the count before a comma: the first draft folded 'Freiherr von Berg, PhD' into the family and moved 'St John, PhD' and 'Abu Bakar, Ed' against every release. - Assign's count walks every token, so a suffix word still stops a particle: 'van Jr. Berg, Mr.' had read family 'van Berg'. Segment and assign now derive their facts through one function each from one set (surname_unit_tags / surname_unit_facts). - The agreement test's negative control is stored and asserted with the period-joined mirror patched out; stale comments corrected. Docs (design review): - rules.md#C1: the reach is P1's fold, not P2's chain; title and bound particles excluded; the P3 reason restated (P3's join depends on the whole name); a title in front is out of scope; 'Freiherr von Berg, Ed' recorded as an accepted consequence (2.3.0's reading too). P3's statement and mechanisms.md#UNIT-PARTITION's Contract name C1's count as their exception. - decisions.md#C1: corrected blast radius (one corpus name moves, of a population of two), the real cost of the one-word reach, and the version range. - release_log: precise version range, the Van Johnson move, and the #563 bullet's stale 'De La Cruz, M.J. K.L.' example replaced; the #563 ledger comments updated to match. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 13 ++-- docs/design/mechanisms.md | 2 +- docs/design/rules.md | 40 ++++++++---- docs/release_log.rst | 4 +- nameparser/_pipeline/_assign.py | 29 ++++++--- nameparser/_pipeline/_segment.py | 2 +- nameparser/_pipeline/_vocab.py | 64 +++++++++++++------- tests/v2/cases.py | 35 +++++++++++ tests/v2/pipeline/test_classify.py | 49 +++++++++++---- tests/v2/test_ledger_guards.py | 61 +++++++++---------- tools/differential/compare.py | 8 --- tools/differential/corpus_rules.jsonl | 1 + tools/differential/expected_since_1.4.0.toml | 27 +++++++-- tools/differential/expected_since_2.0.0.toml | 14 +++-- tools/differential/expected_since_2.1.0.toml | 14 +++-- tools/differential/expected_since_2.2.0.toml | 14 +++-- tools/differential/expected_since_2.3.0.toml | 14 +++-- 17 files changed, 266 insertions(+), 125 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 73a23a7f..eca4865e 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -880,13 +880,14 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had brought back a split of `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd'. That split is not only an intermediate state of #552: 2.2.0 and 2.3.0 shipped it (2.0.0 and 2.1.0 read family 'Smith vd Ma'), #530 removed it in this cycle, and #552 kept it removed. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them. The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. It is also 1.4.0's reading, verified against the released wheel the same day: 1.4.0 read suffix 'vd Ma' and 'VD MA', and 2.0.0 through 2.3.0 read the whole name as the family. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` before a comma and after a family comma, and the silent mixed-case `Doe, Jane PhD vd Ma`. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. -- 2026-10-01 (Derek), #575 — THE COUNT BEFORE THE COMMA COUNTS A PARTICLE SURNAME ONCE. Every count rules.md#C1 makes of the part before the comma — v1's "more than one word" for an unambiguous credential, the ambiguous class's name-word count (#289, #544), and assign's two-name-word test for the positional read after a comma followed by no name word (#296/#325) — counted tokens or pieces, so `De La Cruz` was three words and `van der Berg` two (the ambiguous `van` stands as its own piece, P1's fork). This cycle's count had therefore read `De La Cruz, Ed` as a credential comma with no given name and split `van der Berg, MA` into given 'van', family 'der Berg' — an unreleased regression against 2.3.0, which read both as the listing form — and every release split `van der Berg, PhD`. DECIDED: a particle and the name word it attaches to are one word in all three counts, so a particle surname reads as a one-word surname does. - THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. - THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. Under the default order the positional read P1 applies makes the same part all surname anyway, so the narrower reach costs nothing there. `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. - ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set, and both exclusions are stated where the count lives (`_vocab.SURNAME_UNIT_TAGS`). A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. A connective join is P3's to decide, and P3 declines the commonest connective surname outright — a single-letter connective in a three-word name stays a name word — so `Ortega y Gasset` is three words even standing alone; reproducing P3's conditions before classify would copy P3. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). +- 2026-10-01 (Derek), #575 — THE COUNT BEFORE THE COMMA COUNTS A PARTICLE SURNAME ONCE. Every count rules.md#C1 makes of the part before the comma — v1's "more than one word" for an unambiguous credential, the ambiguous class's name-word count (#289, #544), and assign's two-name-word test for the positional read after a comma followed by no name word (#296/#325) — counted tokens or pieces, so `De La Cruz` was three words and `van der Berg` two (the ambiguous `van` stands as its own piece, P1's fork). This cycle's count had therefore read `De La Cruz, Ed` as a credential comma with no given name and split `van der Berg, MA` into given 'van', family 'der Berg' — an unreleased regression against 2.3.0, which read both as the listing form — and 1.4.0 through 2.3.0 all split `van der Berg, PhD`. DECIDED: a particle run and the one name word it attaches to are one word in all three counts, so a particle surname reads as a one-word surname does. + THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. + THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. + ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). + A WORD THAT IS ALSO A TITLE OR A BOUND GIVEN NAME IS NO PARTICLE TO THE COUNT. Found by the docs review of the first draft, which counted every particle: `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`), and that draft read `Freiherr von Berg, PhD` as family 'Freiherr von Berg', `St John, PhD` as family 'St John' and `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', each against every release. In front of a name such a word reads as the title or the given-name join it also is, so all three keep master's reading. ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.2.0's and 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. `De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name. - SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`) while assign intersects classify's tags with the same set. `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored. Negative control: with that mirror removed, the test reports exactly those two. - BLAST RADIUS, measured 2026-10-01 with the gate at all five baselines: no pre-existing corpus name moves; the movers are the change's own rules.md and case-row names. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname names diff there only by this cycle's #289 count and report, which is the regression this fixes having never shipped. + SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out). + BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the names already in the corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, followed by a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: every corpus name whose part before its first comma has `_vocab.surname_unit_count` 1 and more than one token. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. The move outside the corpora against 2.3.0 is `Van Johnson, Dr.`, above; `Freiherr von Berg, Ed` moves only against this cycle's master. COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. ### T1 — separators, not joiners diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 82096abc..e818d42d 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -29,7 +29,7 @@ Problem shape. A rule needs to know how words were JOINED (chained titles, parti ## UNIT-PARTITION — count the units the joining rules built -Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all. How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma (`_vocab.name_word_count`, `surname_unit_count`, and assign's positional-read test), which run partly before classify and so build their facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name and P3 declines the commonest connective surname, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. +Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma (`_vocab.name_word_count`, `surname_unit_count`, and assign's positional-read test), which run partly before classify and so build their facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. ## MARK-DONT-STRIP — record the decision, keep the fact diff --git a/docs/design/rules.md b/docs/design/rules.md index 2f3a8714..9303fab2 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -518,7 +518,8 @@ P3. Rationale: connective words ("y", "of the") bind name words into have gone the other way, and is reported. The joined part is ONE name word wherever another rule counts them, so a rule taking "one name word" takes the whole join and - never half of it. + never half of it — except C1's count of the words before a comma, + which counts particle surnames alone and states why. A connective counts as a name word wherever this rule counts them, whatever else the vocabulary says the word is, where it is placed to join. A word can be a connective and a generation at once — the @@ -1670,17 +1671,24 @@ C1. Rationale: a credential run after the comma means the name is in listing form, the part before the comma being the family name. Only the part after the first comma decides. Wherever this rule counts the words before the comma, a particle - and the name word it attaches to are one word, as P2 joins them: - the listing form puts a surname before the comma, and a particle - surname is one surname ('De La Cruz, Ed' reads as 'Royce, Ed' - does, 'van der Berg, PhD' as 'Berg, PhD'). That settles P1's fork - for a leading particle that could be a given name, the comma - being the evidence that the part is a surname. The particle - reaches the one word it attaches to and no further, so a word - after that one is a second word. A connective join and a bound - given name are not one word here: a single-letter connective in - a three-word name stays a name word (P3), and a bound given name - builds a given name rather than a surname (P5). + run and the one name word it attaches to are one word, the reach + P1's fold counts rather than the whole of P2's chain: the listing + form puts a surname before the comma, and a particle surname is + one surname ('De La Cruz, Ed' reads as 'Royce, Ed' does, 'van der + Berg, PhD' as 'Berg, PhD'). Where that surname is the whole part, + the comma settles P1's fork for a leading particle that could be + a given name, being the evidence that the part is a surname; a + word after the one the particle attaches to is a second word, and + a suffix word ends the particle's reach. A word that is also + title vocabulary or a bound given name is not a particle to this + count, reading in front of a name as the title (H1) or the + given-name join (P5) it also is. A connective join and a bound + given name are not one word here: whether P3 joins a connective + depends on the words of the whole name, which this count is part + of deciding, and a bound given name builds a given name rather + than a surname (P5). A title in front is a word to the count an + unambiguous credential takes, so 'Dr. van der Berg, PhD' reads as + its part does alone. For the ambiguous credential class — a bare acronym the vocabulary marks as also an ordinary name, and a word admitted to the class by shape, which S3 defines and bounds — the @@ -1862,6 +1870,14 @@ C1. Rationale: a credential run after the comma means the name is in "García Márquez, Ms G.J." → given="G.J." · boundary "García Márquez, Ed G.J." → given="Ed" · boundary "John Smith, A.B. Ph.D." → given="A.B." · boundary + Accepted: a title in front of a particle surname, before the + credential class's word, is read into the family. The surname + counts once and the title is no name word, so the count reads the + listing form, and the listing form keeps a title before the comma + in the family, as it already did for any one-word surname ('Prof. + Cruz, Ed', 'Freiherr Berg, Ed'). The count had read 'von' as a + second name word and kept the title apart. + "Freiherr von Berg, Ed" → family="Freiherr von Berg" Accepted: three or more initials run together with periods behind a surname of two words read as a credential. Initials are conventionally written apart ('J. R. R.'), and a token of three or diff --git a/docs/release_log.rst b/docs/release_log.rst index 75cc3c01..9616b7bc 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -6,7 +6,7 @@ Release Log **Behavior Changes** - - **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where every release gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. A given name in front still makes two words, so ``John van Buren, Ed`` keeps suffix ``Ed``. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575) + - **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where 1.4.0 through 2.3.0 gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name: ``Van Johnson, Dr.`` gives last ``Van Johnson``, where 2.2 and 2.3 gave first ``Van``. A given name in front still makes two words, so ``John van Buren, Ed`` reads suffix ``Ed`` as ``John Smith, Ed`` does. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575) - **Fix a one-letter connective joining a name that gives no sign it is a connective.** ``HumanName("jose e maria santos")`` gives first ``jose``, middle ``e maria``, last ``santos``, where 1.4.0 through 2.3.0 gave first ``jose e maria``; and ``JUAN GARCIA Y LOPEZ`` gives last ``GARCIA Y LOPEZ``, where every release since 1.4.0 read the bare capital as an initial and gave middle ``GARCIA Y``. A single letter is an initial where the writing says so -- a bare Latin capital in a name that is not written wholly in one case -- and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there: ``e`` reads as an initial and ``y`` joins. Mixed-case input is untouched in both directions: ``Jose e Maria Santos`` still gives first ``Jose e Maria`` and ``Jose E Maria Santos`` still gives middle ``E Maria``. Short names move in the derived views rather than the fields, P3's three-word carve-out being unchanged: ``parse("john e smith").initials()`` is ``j. e. s.`` where 2.3.0 gave ``j. s.``, and ``HumanName("john e smith").capitalize()`` gives ``John E Smith`` where 2.3.0 gave ``John e Smith``; ``JUAN Y GARCIA`` moves its ``capitalize()`` the same way in reverse, giving ``Juan y Garcia``, while its initials do not move at all: ``parse(...).initials()`` is ``J. Y. G.``, what 2.3.0 gave and what 1.4.0's own view gave, the ``Y`` holding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461). ``HumanName.initials()`` agrees with the core on both -- see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (``Хосе И Мария Сантос`` still gives first ``Хосе И Мария``), and Arabic ``و`` never enters the rule, having no case to be written against. A ``Lexicon`` knob decides which letters are marked, so the reading is configurable rather than fixed. See the ``P3`` entry of ``docs/design/decisions.md`` (closes #383, closes #479) @@ -24,7 +24,7 @@ Release Log - **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516) - - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``De La Cruz, M.J. K.L.`` gives last ``De La Cruz``, suffix ``M.J. K.L.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` + - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` - **Fix a credential ending the given part of a family-comma listing being read as a middle name in silence.** ``HumanName("Doe, John MA")`` gives first ``John``, last ``Doe``, suffix ``MA``, where 2.0 through 2.3 gave middle ``MA`` -- and 1.4.0 gave the suffix, so this restores v1's reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone: ``Doe, John Ma`` keeps middle ``Ma``, written the way a name is written, and ``Doe, John Ed`` keeps middle ``Ed``. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields -- a declined name re-rendered without its comma, ``John Ma Doe``, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent -- ``Doe, John MA Smith`` gives middle ``MA Smith`` and reports nothing -- while a credential run or a trailing title is transparent to it: ``Doe, John MA PhD`` gives suffix ``MA PhD`` and ``Doe, John MA Prof.`` gives title ``Prof.`` with suffix ``MA``. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further: ``Doe, John Prof. MA`` now gives title ``Prof.`` where it gave middle ``Prof. MA``, the trailing-title chain reaching a word the credential used to hide; and ``Doe, John van MA`` gives last ``van Doe`` with suffix ``MA`` where it gave middle ``van MA``, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read: ``DOE, MARY JO MA``, ``doe, john ma``, ``田中, 太郎 MA`` and ``김, 민준 MA`` all give a suffix. The unlisted dotted spelling moves with them without the parity claim -- ``Doe, John X.Y.Z.`` gives suffix ``X.Y.Z.`` where 1.4.0 and 2.3.0 both gave a middle name -- to match the comma-less ``John Doe X.Y.Z.``. One word is carved out: ``do`` is the only member of this class that is also a surname particle, so capitals decide it and, with nothing in front of it, the particle reading keeps every other spelling (a degree in front is the other exception, the next-but-one entry). ``Doe, John DO`` gives suffix ``DO``, while ``Doe, John do``, ``Doe, John Do``, ``DOE, JOHN DO`` and ``doe, john do`` are unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, so ``SMITH, JOHN DO`` keeps last ``DO SMITH`` as ``NASCIMENTO, EDSON ARANTES DO`` does -- right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See the ``S2`` and ``P6`` entries of ``docs/design/decisions.md`` (closes #531) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index 5b36eed5..2ae225c7 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -71,8 +71,8 @@ from nameparser._lexicon import Lexicon from nameparser._pipeline._vocab import ( - SURNAME_UNIT_TAGS, effective_script, is_suffix_lenient, - resolve_script_set, unit_ends, + effective_script, is_suffix_lenient, resolve_script_set, + surname_unit_facts, unit_ends, ) from nameparser._pipeline._pieces import ( anchor_in_reach, credential_at_the_given_slot, given_slot_anchors, @@ -1087,12 +1087,25 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # positionally lost 'van' to the given name. The comma is the # evidence that settles that fork: the listing form puts a # surname before it. The same facts as segment's count - # (`_vocab.SURNAME_UNIT_TAGS`), so the two agree. - if reading is not None and len(unit_ends([ - tokens[i].tags & SURNAME_UNIT_TAGS - for k, piece in enumerate(fam_pieces) - if not is_suffix_piece(piece, fam_tags[k], tokens) - for i in piece], chain=False)) > 1: + # (`_vocab.surname_unit_facts`), over EVERY token, so a suffix + # word still stops a particle ('van Jr. Berg, Mr.' is two name + # words); a unit counts when a non-suffix piece holds a token + # of it. + positional = False + if reading is not None: + idx = [i for piece in fam_pieces for i in piece] + named = {i for k, piece in enumerate(fam_pieces) + if not is_suffix_piece(piece, fam_tags[k], tokens) + for i in piece} + names = 0 + start = 0 + for end in unit_ends([surname_unit_facts(tokens[i].tags) + for i in idx], chain=False): + if not named.isdisjoint(idx[start:end]): + names += 1 + start = end + positional = names > 1 + if positional: order = _assign_main(0, state, tokens, ambiguities) else: for k, piece in enumerate(fam_pieces): diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 3fff4d02..d926d225 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -393,7 +393,7 @@ def class_run(seg: tuple[int, ...]) -> bool: pre_comma_names = name_word_count(texts(groups[0]), state.lexicon, state.policy) # rules.md#C1's "more than one word precedes the comma" counts a - # particle chain and a connective join as one word (#575): v1's + # particle and the word it attaches to as one word (#575): v1's # token count split 'van der Berg, PhD' into given 'van', family # 'der Berg'. The token count stays first, so a one-token part # never builds the units. diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index ced988a4..297320cc 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -776,14 +776,15 @@ def ambiguous_class_candidate(text: str, lexicon: Lexicon, # mechanisms.md#UNIT-PARTITION: "A rule that counts name words counts -# those units, and takes each whole or not at all." The walk is shared +# those units, and takes each whole or not at all" -- C1's count before +# a comma being the stated exception (`SURNAME_UNIT_TAGS`). The walk is shared # by two readers that hold the facts in different forms: post_rules # reads classify's TAGS, and segment, which runs before classify has # tagged anything, builds the same tag names from the vocabulary # (`surname_unit_tags`). One walk over tag sets, so the two cannot # disagree about where a unit ends (mechanisms.md#ONE-PREDICATE-PER- # QUESTION) -- only, at worst, about a token's facts, which -# test_vocab's agreement test pins. +# test_classify's agreement test pins. def unit_ends(tags: Sequence[Set[str]], chain: bool = True) -> list[int]: """The END (one past the last index) of each unit of `tags`, in order, partitioning `range(len(tags))`: one name word each, except @@ -849,30 +850,46 @@ def end_of(i: int) -> int: #: The facts rules.md#C1's count before a comma reads (#575): a -#: particle chain is one surname, and a suffix word stops it. Neither -#: of the other two joins `unit_ends` knows is read there. A bound +#: particle starts a surname unit, and a suffix word stops it. A word +#: that is ALSO title vocabulary ('Freiherr', 'St') or a bound +#: given-name head ('Abu') is not a particle here: in front of a name +#: it reads as the title or the given-name join it also is (rules.md +#: H1, P5), so 'Freiherr von Berg, PhD' keeps its title. Neither of +#: the other two joins `unit_ends` knows is read either. A bound #: given-name pair builds a GIVEN name, and P5 gives up a family word #: where the name has no other ('abdul Salam' alone is given 'abdul', -#: family 'Salam'), so before a comma it is not one surname. A -#: connective join is P3's to decide, and P3 declines the commonest -#: connective surname outright -- a single-letter connective in a -#: three-word name stays a name word, so 'Ortega y Gasset' is three -- -#: on conditions (that exception, the case fork, generational -#: connectives) segment could only rebuild by copying P3. Segment -#: builds these from the vocabulary (`surname_unit_tags`); assign -#: intersects classify's tags with this set, so the two counts read -#: one set of facts. +#: family 'Salam'), so before a comma it is not one surname. Whether +#: P3 joins a single-letter connective depends on the words of the +#: whole name -- 'Carod i Rovira, Josep' joins as four words where +#: 'Ortega y Gasset' alone, three words, does not -- and the count is +#: part of deciding what the whole name is, so segment could reach +#: P3's answer only by copying P3. Segment builds these facts from the +#: vocabulary (`surname_unit_tags`); assign derives them from +#: classify's tags (`surname_unit_facts`), so the two counts read one +#: set of facts. SURNAME_UNIT_TAGS = frozenset({"particle", "vocab:suffix"}) _PARTICLE_ONLY = frozenset({"particle"}) _SUFFIX_ONLY = frozenset({"vocab:suffix"}) _NO_TAGS: frozenset[str] = frozenset() +_NOT_A_SURNAME_PARTICLE = frozenset({"vocab:title", "vocab:bound-given"}) + + +def surname_unit_facts(tags: Set[str]) -> frozenset[str]: + """`SURNAME_UNIT_TAGS` for one token, from classify's tags: the + reading assign's count before a comma takes.""" + particle = ("particle" in tags + and tags.isdisjoint(_NOT_A_SURNAME_PARTICLE)) + if "vocab:suffix" in tags: + return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY + return _PARTICLE_ONLY if particle else _NO_TAGS def surname_unit_tags(text: str, lexicon: Lexicon) -> frozenset[str]: - """`SURNAME_UNIT_TAGS` for one token from the vocabulary alone -- - segment's view of classify's tags, built before classify runs, and - classify's own tests: particle membership, `suffix_as_written`, and - the period-joined derivation. Kept from drifting by + """`surname_unit_facts` for one token from the vocabulary alone -- + segment's view of classify's tags, built before classify runs, with + classify's own tests: particle, title and bound given-name + membership, `suffix_as_written`, and the period-joined derivation. + Kept from drifting by test_classify.test_surname_unit_tags_agree_with_classify.""" n = _normalize(text) # classify's whole-token test, then its period-joined derivation @@ -881,9 +898,11 @@ def surname_unit_tags(text: str, lexicon: Lexicon) -> frozenset[str]: suffix = suffix_as_written(n, text, lexicon) or ( "." in text and n not in lexicon.titles and period_joined_vocab(text, lexicon) == "suffix") - if n in lexicon.particles: - return SURNAME_UNIT_TAGS if suffix else _PARTICLE_ONLY - return _SUFFIX_ONLY if suffix else _NO_TAGS + particle = (n in lexicon.particles and n not in lexicon.titles + and n not in lexicon.bound_given_names) + if suffix: + return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY + return _PARTICLE_ONLY if particle else _NO_TAGS def surname_unit_count(texts: Sequence[str], lexicon: Lexicon) -> int: @@ -893,10 +912,11 @@ def surname_unit_count(texts: Sequence[str], lexicon: Lexicon) -> int: of a part that is wholly suffix words (#575), with a particle reaching as P1's fold does (`unit_ends`'s `chain=False`). 'van der Berg' is one surname, so 'van der Berg, PhD' reads as 'Berg, PhD' - does.""" + does. A connective join is not one unit here: `SURNAME_UNIT_TAGS`.""" # With no particle, every token is a unit of its own, so the count # is the token count: the suffix tests and the walk are skipped for - # the commonest comma names ('John Smith, PhD'). + # the commonest comma names ('John Smith, PhD'). A superset test -- + # a title-particle passes it -- so it only ever skips work. if not any(_normalize(t) in lexicon.particles for t in texts): return len(texts) return len(unit_ends([surname_unit_tags(t, lexicon) for t in texts], diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 44ab2500..4ddb9e0d 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -1071,6 +1071,41 @@ def _check_cjk_shape_purity(self) -> None: "('abdul Salam' alone is given 'abdul', family 'Salam'), " "so it is not one surname before a comma", shape=3), + Case("a_title_particle_before_a_comma_is_still_a_title", + "Freiherr von Berg, PhD", + {"title": "Freiherr", "family": "von Berg", "suffix": "PhD"}, + ambiguities=("particle-or-given",), + notes="#575 boundary: a word of both the title and the particle " + "vocabulary is not a particle to the count before the " + "comma, so it stays a title in front of the surname. The " + "first draft of #575 counted 'Freiherr von Berg' as one " + "surname and folded the title into the family"), + Case("a_bound_given_particle_before_a_comma_is_still_a_given_name", + "Abu Bakar, Ed", + {"given": "Abu", "family": "Bakar", "suffix": "Ed"}, + ambiguities=("suffix-or-name", "particle-or-given"), + notes="#575 boundary: 'abu' is a bound given-name head as well " + "as a particle, and in front of a name it reads as the " + "given-name join (P5), so it is not one surname with the " + "word after it"), + Case("a_suffix_inside_the_part_stops_the_particle", + "van Jr. Berg, Mr.", + {"title": "Mr.", "given": "van", "middle": "Jr.", + "family": "Berg"}, + ambiguities=("particle-or-given",), + notes="#575 boundary: the particle reaches no further than a " + "suffix word, so 'van' and 'Berg' are two name words and " + "the part keeps its positional read. A first draft " + "walked past the suffix and read family 'van Berg'"), + Case("a_title_in_front_keeps_the_particle_fork", + "Dr. van der Berg, PhD", + {"title": "Dr.", "given": "van", "family": "der Berg", + "suffix": "PhD"}, + ambiguities=("particle-or-given",), + notes="#575 boundary: the count for an unambiguous credential " + "counts the title as a word, so the comma reads as a " + "credential comma and the part reads as it does alone, " + "P1's fork and all. Unchanged"), Case("the_anchor_does_not_reach_across_a_maiden_clause", "Jane Doe Jr. nee Smith Ma", {"given": "Jane", "family": "Doe", "suffix": "Jr.", diff --git a/tests/v2/pipeline/test_classify.py b/tests/v2/pipeline/test_classify.py index 58366edf..81c5ad89 100644 --- a/tests/v2/pipeline/test_classify.py +++ b/tests/v2/pipeline/test_classify.py @@ -6,6 +6,7 @@ from nameparser._lexicon import Lexicon, _normalize from nameparser._pipeline import STAGES from nameparser._pipeline import _classify as _classify_module +from nameparser._pipeline import _vocab from nameparser._pipeline._classify import classify from nameparser._pipeline._extract import extract_delimited from nameparser._pipeline._segment import segment @@ -14,8 +15,8 @@ ) from nameparser._pipeline._tokenize import tokenize from nameparser._pipeline._vocab import ( - SURNAME_UNIT_TAGS, ambiguous_class_candidate, ambiguous_class_member, - caps_shape_candidate, surname_unit_tags, + ambiguous_class_candidate, ambiguous_class_member, + caps_shape_candidate, surname_unit_facts, surname_unit_tags, ) from nameparser._policy import Policy from nameparser._types import AmbiguityKind, Role @@ -746,21 +747,43 @@ def _surname_unit_sweep() -> list[str]: return sorted(words) -def test_surname_unit_tags_agree_with_classify() -> None: - # #575: rules.md#C1's count before the comma reads two facts per - # token -- particle, suffix -- and `segment` must build them from - # the vocabulary (`_vocab.surname_unit_tags`), since it runs before - # classify has tagged anything; assign reads classify's own tags - # for the same count. Swept over every single-word entry of the - # vocabularies a token can be tagged from, in three casings, plus - # the period shapes classify derives a suffix from ('JD.CPA', - # 'Msc.Ed.' -- the two this sweep caught on the first draft). +def _surname_unit_disagreements() -> list[str]: lex = Lexicon.default() disagree = [] for word in _surname_unit_sweep(): tags = _tags_by_text(f"Smith {word}, John", lexicon=lex).get(word) if tags is None: # tokenize split it; nothing to compare continue - if surname_unit_tags(word, lex) != tags & SURNAME_UNIT_TAGS: + if surname_unit_tags(word, lex) != surname_unit_facts(tags): disagree.append(word) - assert disagree == [] + return disagree + + +#: The agreement test's recorded negative control: what it reports +#: with `surname_unit_tags`' period-joined mirror switched off -- the +#: two words its first run caught, which classify tags as suffixes +#: through that derivation alone. +_SURNAME_UNIT_CONTROL = ["JD.CPA", "Msc.Ed."] + + +def test_surname_unit_tags_agree_with_classify() -> None: + # #575: rules.md#C1's count before the comma reads two facts per + # token -- particle, suffix -- and `segment` must build them from + # the vocabulary (`_vocab.surname_unit_tags`), since it runs before + # classify has tagged anything, while assign derives them from + # classify's tags (`_vocab.surname_unit_facts`). Swept over every + # single-word entry of the vocabularies a token can be tagged + # from, in three casings, plus the period shapes classify derives + # a suffix from; a title-particle ('Freiherr') and a bound + # given-name particle ('Abu') are in the sweep through those lists. + assert _surname_unit_disagreements() == [] + + +def test_the_surname_unit_agreement_test_can_fail( + monkeypatch: pytest.MonkeyPatch) -> None: + # The guard above, with the mirror it exists to hold switched off: + # `_vocab.surname_unit_tags` reads `period_joined_vocab` from its + # own module, so patching it there leaves classify untouched. + monkeypatch.setattr(_vocab, "period_joined_vocab", + lambda text, lexicon: None) + assert _surname_unit_disagreements() == _SURNAME_UNIT_CONTROL diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index f6a55830..80bf60fb 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -2462,11 +2462,12 @@ class _LatinCopy(NamedTuple): "John de Ma", "John van der Berg Ma", r"Smith Jr\., MA", r"Smith Jr\., Ma", "Smith, MA", "abdul Smith Berg Ma", "abdul Smith Ma", "john smith, ma"}), - # 2026-10-01, #575: the 2.x copy gains seven C1 examples, comma + # 2026-10-01, #575: the 2.x copy gains eight C1 examples, comma # names whose post-comma word is a listed member -- no set copied. frozenset({"Davis Royce, Ed", "De La Cruz, Ed", r"Doe, Dr\. MA", "Doe, MA", "Doe, MA PhD", r"Doe, Mr\. MA PhD", - "Freiherr von Berg MA", "JOHN SMITH, MA", "Jack MA", + "Freiherr von Berg MA", "Freiherr von Berg, Ed", + "JOHN SMITH, MA", "Jack MA", r"Jack MA\.", "Jack Wei Ma", r"John Prof\. MA", "John Smith Ma", "John Smith, Ed", "John Smith, MA", "John Smith, Ma", "John de Ma", "John van Buren, Ed", @@ -3079,7 +3080,8 @@ class _LatinCopy(NamedTuple): frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), # 2026-10-01, #575: rules.md#C1's particle-surname examples at # 1.4.0, literal names that copy no set. - frozenset({"De La Cruz, Ed", "Van Buren, Ed", "de la Cruz, Ma"}), + frozenset({"De La Cruz, Ed", "Freiherr von Berg, Ed", "Van Buren, Ed", + "de la Cruz, Ma"}), # 2026-10-01, #554: the run rule gains rules.md#C1's particle-and- # suffix examples, a vocabulary word in front of a member, which # copies no set either. @@ -3788,11 +3790,11 @@ def _claim(rule: dict) -> _Claim: # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, # Jr do', #554's rules.md#C1 examples. Reach, verified name by # name. - # 2026-10-01, #575: 414 -> 422, its eight rules.md#C1 and + # 2026-10-01, #575: 414 -> 423, its nine rules.md#C1 and # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(422, ('given', 'suffix', 'title'), "de00a0ee1e57", None), + _Claim(423, ('given', 'suffix', 'title'), "f5edb96b96cb", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3891,15 +3893,15 @@ def _claim(rule: dict) -> _Claim: # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, # Jr do', #554's rules.md#C1 examples. Reach, verified name by # name. - # 2026-10-01, #575: 414 -> 422, its eight rules.md#C1 and + # 2026-10-01, #575: 414 -> 423, its nine rules.md#C1 and # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(422, ('family', 'given'), "de00a0ee1e57", None), - # 2026-10-01, #575: new, 3; 'De La Cruz, Ed', 'Van Buren, Ed', - # 'de la Cruz, Ma'. + _Claim(423, ('family', 'given'), "f5edb96b96cb", None), + # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von + # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": - _Claim(3, ('family', 'given', 'suffix'), "a1d6626eb35f", None), + _Claim(4, ('family', 'given', 'suffix', 'title'), "30163564e03d", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4184,12 +4186,12 @@ def _claim(rule: dict) -> _Claim: # 2026-09-30, #563: 115 -> 117, 'De La Cruz, M.J. PhD' and # 'De La Cruz, M.J. K.L.' -- a particle chain, 'De La Cruz'. # Reach, verified name by name. - # 2026-10-01, #575: 117 -> 123, its six particle-surname + # 2026-10-01, #575: 117 -> 124, its seven particle-surname # names ('De La Cruz, Ed', 'de la Cruz, Ma', 'van der Berg, # MA', 'Van Buren, Ed', 'van der Berg, PhD', 'John van Buren, - # Ed'). Reach, verified name by name. + # Ed', 'Freiherr von Berg, Ed'). Reach, verified name by name. "fix(initials-per-word) a particle chain inside a name part initials each word (facade, since 2.0.0)": - _Claim(123, ('_initials',), 'ca2598e60110', ('DEFAULT',)), + _Claim(124, ('_initials',), 'bf190fecb582', ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": @@ -4989,9 +4991,9 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the # alternation gained. Reach, verified name by name. - _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), @@ -5491,9 +5493,9 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the # alternation gained. Reach, verified name by name. - _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), @@ -6136,9 +6138,9 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the # alternation gained. Reach, verified name by name. - _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), @@ -6504,9 +6506,9 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - # 2026-10-01, #575: 23 -> 30, the seven C1 examples the + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the # alternation gained. Reach, verified name by name. - _Claim(30, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "dda7ec083d1e", ('DEFAULT',)), + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), @@ -7329,14 +7331,6 @@ def test_every_rule_claims_the_recorded_share_of_the_corpus() -> None: "MD, DO, DDS": "fix(#296) a dropped prenominal takes the name position it " "occupies", - # #575 (2026-10-01): the routing rule describes a lone - # post-comma piece LEAVING `first`; these two move a piece INTO - # `given`, because the particle surname before the comma is one - # name word (rules.md#C1). The rule that says so wins. - "De La Cruz, Ed": - "fix(#575) a particle surname before a comma is one name word", - "de la Cruz, Ma": - "fix(#575) a particle surname before a comma is one name word", }, # The two 2.x ledgers had NO section here until #452, and the # coverage assertion below was `<=`, so their absence read as "no @@ -8486,10 +8480,13 @@ def test_a_rule_reaching_no_corpus_name_says_why_it_is_kept() -> None: "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 2), ("fix(#296) a credential-only comma string reads a name and its postnominal", "fix(comma-precomma-family) pre-comma run reads as family, not given", 2), - # 2026-10-01, #575: declared on the #575 rule; the routing - # rule's half of its contest is a pinned overlap, not this. + # 2026-10-01, #575: both declared on the #575 rule, over its + # four names; it stands ahead of the two comma rules because + # each describes the opposite or the other half of the move. ("fix(#575) a particle surname before a comma is one name word", - "fix(comma-precomma-family) pre-comma run reads as family, not given", 3), + "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 4), + ("fix(#575) a particle surname before a comma is one name word", + "fix(comma-precomma-family) pre-comma run reads as family, not given", 4), ("fix(#296) a lone post-comma credential is a suffix", "fix(suffix-routing) a two-token name ending in the suffix word jr keeps it in `suffix`", 2), ("fix(#400/#274) bound-given join and maiden consumption in one name", diff --git a/tools/differential/compare.py b/tools/differential/compare.py index 80e484b7..814fa18b 100644 --- a/tools/differential/compare.py +++ b/tools/differential/compare.py @@ -1881,14 +1881,6 @@ class _ShapeMismatch(NamedTuple): # nesting, so neither is the narrower and file order is the # whole decision. The shape is this run's, not guessed. "Smith, MA": ("given", "suffix"), - # #575 (2026-10-01): 1.4.0 read suffix 'Ed'/'Ma' with no given - # name, and the particle surname is now one name word, so the - # word after the comma is the given name. The lone-post-comma - # routing rule's Latin comma regex reaches both and `fields` - # overlap on {given, suffix} without nesting, so the #575 rule - # stands ahead of it on purpose. The shape is this run's. - "De La Cruz, Ed": ("given", "suffix"), - "de la Cruz, Ma": ("given", "suffix"), }, # #501's six, moved here from _WATCHED_DIFFS with their shapes # unchanged. The four CJK rows sit at 2.0.0 alone: the honorific diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 6ef80d64..7151dfac 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -85,6 +85,7 @@ "Duke of Edinburgh" "Esq. Smith" "Freiherr von Berg MA" +"Freiherr von Berg, Ed" "Freiherr von Richthofen V" "Gal·la Serra" "Garcia" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 44ae2016..ba52c291 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -783,8 +783,23 @@ issue = "fix(#575) a particle surname before a comma is one name word" # # Literal; the probe 'John van Buren, Ed' (two name words: the # credential reading stays) is _MUST_NOT_MATCH. -name_regex = "^(?:De La Cruz, Ed|Van Buren, Ed|de la Cruz, Ma)$" -fields = ["given", "family", "suffix"] +# 'Freiherr von Berg, Ed' joins as rules.md#C1's accepted +# consequence: the surname counts once and 'Freiherr' is no name word, +# so the listing form reads it and keeps the title in the family, as +# it does for 'Prof. Cruz, Ed'. 1.4.0 read title 'Freiherr', first +# 'von Berg', suffix 'Ed'. +name_regex = "^(?:De La Cruz, Ed|Freiherr von Berg, Ed|Van Buren, Ed|de la Cruz, Ma)$" +fields = ["title", "given", "family", "suffix"] + +[[change.precedes_narrower]] +issue = "fix(comma-family) lone post-comma piece routes to suffix/title, not first" +why = """ +The later rule describes a lone post-comma piece LEAVING `first` for +`suffix` or `title`. These names move the other way: the word after +the comma ARRIVES in `given`, because the particle surname before it +is one name word (rules.md#C1, #575), and that is what this rule +names. +""" [[change.precedes_narrower]] issue = "fix(comma-precomma-family) pre-comma run reads as family, not given" @@ -4051,8 +4066,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 389c137b..42383138 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2524,7 +2524,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -# 2026-10-01, #575: seven names join, the C1 examples of that change. +# 2026-10-01, #575: eight names join, the C1 examples of that change. # Against these baselines none of them is #575's doing: each moves # on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y # Gasset, Ed' and 'abdul Salam, Ed' are two name words before the @@ -2533,7 +2533,9 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" +# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -2619,8 +2621,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 65bab4ab..38363584 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2411,7 +2411,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -# 2026-10-01, #575: seven names join, the C1 examples of that change. +# 2026-10-01, #575: eight names join, the C1 examples of that change. # Against these baselines none of them is #575's doing: each moves # on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y # Gasset, Ed' and 'abdul Salam, Ed' are two name words before the @@ -2420,7 +2420,9 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" +# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -2506,8 +2508,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index d47efb3b..f7b471f8 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1000,7 +1000,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -# 2026-10-01, #575: seven names join, the C1 examples of that change. +# 2026-10-01, #575: eight names join, the C1 examples of that change. # Against these baselines none of them is #575's doing: each moves # on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y # Gasset, Ed' and 'abdul Salam, Ed' are two name words before the @@ -1009,7 +1009,9 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" +# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -1095,8 +1097,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index f203f09f..37e52a7e 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -288,7 +288,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -# 2026-10-01, #575: seven names join, the C1 examples of that change. +# 2026-10-01, #575: eight names join, the C1 examples of that change. # Against these baselines none of them is #575's doing: each moves # on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y # Gasset, Ed' and 'abdul Salam, Ed' are two name words before the @@ -297,7 +297,9 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" +# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -383,8 +385,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a From e553d6e21de6c2ea163b1d510ba7814104028018 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 20:49:15 -0700 Subject: [PATCH 3/5] fix(C1): #575 second review -- title-particles excluded only when leading The first review round excluded title-particles and bound given-name particles from the count in every position. The review of that fix found it brought #575's own defect back inside a surname ('de St Pierre, Ed' and 'De St. Croix, Ed' lost the given name again), and that excluding 'Abu' gave up 2.0-2.3's reading of 'Abu Bakar, Ed' (given 'Ed', family 'Abu Bakar') rather than keeping it. Now a title-particle is no particle only where it opens the part, and a bound given-name particle is a particle. 'Freiherr von Berg, PhD' and 'St John, PhD' keep the title; 'de St Pierre, Ed' and 'Abu Bakar, Ed' read given 'Ed'. The agreement test compares both positions; a case row pins the position itself. Also: the five ledger comments quoted a C1 sentence the first round rewrote; the blast-radius recipe gains its credential filter and comparator; the decisions entry says why the review's boundary rows carry no shape tag, and lists Abu Bakar, PhD among the moves. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 +- docs/design/rules.md | 20 +++++----- docs/release_log.rst | 2 +- nameparser/_pipeline/_assign.py | 5 ++- nameparser/_pipeline/_vocab.py | 40 +++++++++++--------- tests/v2/cases.py | 26 +++++++++---- tests/v2/pipeline/test_classify.py | 12 ++++-- tools/differential/expected_since_1.4.0.toml | 4 +- tools/differential/expected_since_2.0.0.toml | 4 +- tools/differential/expected_since_2.1.0.toml | 4 +- tools/differential/expected_since_2.2.0.toml | 4 +- tools/differential/expected_since_2.3.0.toml | 4 +- 12 files changed, 76 insertions(+), 53 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index eca4865e..88b29133 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -884,10 +884,10 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). - A WORD THAT IS ALSO A TITLE OR A BOUND GIVEN NAME IS NO PARTICLE TO THE COUNT. Found by the docs review of the first draft, which counted every particle: `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`), and that draft read `Freiherr von Berg, PhD` as family 'Freiherr von Berg', `St John, PhD` as family 'St John' and `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', each against every release. In front of a name such a word reads as the title or the given-name join it also is, so all three keep master's reading. ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.2.0's and 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. + A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see position itself, which the `de St Pierre, Ed` case row pins. ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. `De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name. SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out). - BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the names already in the corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, followed by a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: every corpus name whose part before its first comma has `_vocab.surname_unit_count` 1 and more than one token. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. The move outside the corpora against 2.3.0 is `Van Johnson, Dr.`, above; `Freiherr von Berg, Ed` moves only against this cycle's master. + BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose part after it holds a suffix-vocabulary or dotted-credential word. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Moves outside the corpora against 2.3.0 are `Van Johnson, Dr.` and `Abu Bakar, PhD`, above; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index 9303fab2..bc7dcd8a 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1680,15 +1680,17 @@ C1. Rationale: a credential run after the comma means the name is in a given name, being the evidence that the part is a surname; a word after the one the particle attaches to is a second word, and a suffix word ends the particle's reach. A word that is also - title vocabulary or a bound given name is not a particle to this - count, reading in front of a name as the title (H1) or the - given-name join (P5) it also is. A connective join and a bound - given name are not one word here: whether P3 joins a connective - depends on the words of the whole name, which this count is part - of deciding, and a bound given name builds a given name rather - than a surname (P5). A title in front is a word to the count an - unambiguous credential takes, so 'Dr. van der Berg, PhD' reads as - its part does alone. + title vocabulary is not a particle to this count where it opens + the part, reading there as the title it also is (H1); inside a + surname it chains ('de St Pierre, Ed' reads given 'Ed'). A + connective join and a bound given-name pair are not one word + here: whether P3 joins a connective depends on the words of the + whole name, which this count is part of deciding, and a bound + given name builds a given name rather than a surname (P5) — a + word that is a particle as well counts as the particle ('Abu + Bakar, Ed' reads given 'Ed'). A title in front is a word to the + count an unambiguous credential takes, so 'Dr. van der Berg, PhD' + reads as its part does alone. For the ambiguous credential class — a bare acronym the vocabulary marks as also an ordinary name, and a word admitted to the class by shape, which S3 defines and bounds — the diff --git a/docs/release_log.rst b/docs/release_log.rst index 9616b7bc..16b12f5c 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -6,7 +6,7 @@ Release Log **Behavior Changes** - - **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where 1.4.0 through 2.3.0 gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name: ``Van Johnson, Dr.`` gives last ``Van Johnson``, where 2.2 and 2.3 gave first ``Van``. A given name in front still makes two words, so ``John van Buren, Ed`` reads suffix ``Ed`` as ``John Smith, Ed`` does. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575) + - **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where 1.4.0 through 2.3.0 gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does, and ``Abu Bakar, PhD`` gives last ``Abu Bakar`` where 1.4.0 through 2.3.0 gave first ``Abu``. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name: ``Van Johnson, Dr.`` gives last ``Van Johnson``, where 2.2 and 2.3 gave first ``Van``. A given name in front still makes two words, so ``John van Buren, Ed`` reads suffix ``Ed`` as ``John Smith, Ed`` does. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575) - **Fix a one-letter connective joining a name that gives no sign it is a connective.** ``HumanName("jose e maria santos")`` gives first ``jose``, middle ``e maria``, last ``santos``, where 1.4.0 through 2.3.0 gave first ``jose e maria``; and ``JUAN GARCIA Y LOPEZ`` gives last ``GARCIA Y LOPEZ``, where every release since 1.4.0 read the bare capital as an initial and gave middle ``GARCIA Y``. A single letter is an initial where the writing says so -- a bare Latin capital in a name that is not written wholly in one case -- and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there: ``e`` reads as an initial and ``y`` joins. Mixed-case input is untouched in both directions: ``Jose e Maria Santos`` still gives first ``Jose e Maria`` and ``Jose E Maria Santos`` still gives middle ``E Maria``. Short names move in the derived views rather than the fields, P3's three-word carve-out being unchanged: ``parse("john e smith").initials()`` is ``j. e. s.`` where 2.3.0 gave ``j. s.``, and ``HumanName("john e smith").capitalize()`` gives ``John E Smith`` where 2.3.0 gave ``John e Smith``; ``JUAN Y GARCIA`` moves its ``capitalize()`` the same way in reverse, giving ``Juan y Garcia``, while its initials do not move at all: ``parse(...).initials()`` is ``J. Y. G.``, what 2.3.0 gave and what 1.4.0's own view gave, the ``Y`` holding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461). ``HumanName.initials()`` agrees with the core on both -- see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (``Хосе И Мария Сантос`` still gives first ``Хосе И Мария``), and Arabic ``و`` never enters the rule, having no case to be written against. A ``Lexicon`` knob decides which letters are marked, so the reading is configurable rather than fixed. See the ``P3`` entry of ``docs/design/decisions.md`` (closes #383, closes #479) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index 2ae225c7..0b4cde76 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -1099,8 +1099,9 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: for i in piece} names = 0 start = 0 - for end in unit_ends([surname_unit_facts(tokens[i].tags) - for i in idx], chain=False): + for end in unit_ends([surname_unit_facts(tokens[i].tags, k == 0) + for k, i in enumerate(idx)], + chain=False): if not named.isdisjoint(idx[start:end]): names += 1 start = end diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 297320cc..531dca55 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -851,11 +851,14 @@ def end_of(i: int) -> int: #: The facts rules.md#C1's count before a comma reads (#575): a #: particle starts a surname unit, and a suffix word stops it. A word -#: that is ALSO title vocabulary ('Freiherr', 'St') or a bound -#: given-name head ('Abu') is not a particle here: in front of a name -#: it reads as the title or the given-name join it also is (rules.md -#: H1, P5), so 'Freiherr von Berg, PhD' keeps its title. Neither of -#: the other two joins `unit_ends` knows is read either. A bound +#: that is ALSO title vocabulary ('Freiherr', 'St') LEADING the part is +#: not a particle here: in front of a name it reads as the title it +#: also is (rules.md#H1), so 'Freiherr von Berg, PhD' keeps its title, +#: while inside a surname it chains as P1 and P2 chain it ('de St +#: Pierre' is one surname). A bound given-name head that is also a +#: particle ('Abu') stays a particle: 'Abu Bakar, Ed' is a surname +#: before the comma, as 2.0 through 2.3 read it. Neither of the other +#: two joins `unit_ends` knows is read. A bound #: given-name pair builds a GIVEN name, and P5 gives up a family word #: where the name has no other ('abdul Salam' alone is given 'abdul', #: family 'Salam'), so before a comma it is not one surname. Whether @@ -871,24 +874,23 @@ def end_of(i: int) -> int: _PARTICLE_ONLY = frozenset({"particle"}) _SUFFIX_ONLY = frozenset({"vocab:suffix"}) _NO_TAGS: frozenset[str] = frozenset() -_NOT_A_SURNAME_PARTICLE = frozenset({"vocab:title", "vocab:bound-given"}) - - -def surname_unit_facts(tags: Set[str]) -> frozenset[str]: +def surname_unit_facts(tags: Set[str], leading: bool) -> frozenset[str]: """`SURNAME_UNIT_TAGS` for one token, from classify's tags: the - reading assign's count before a comma takes.""" + reading assign's count before a comma takes. `leading` is whether + the token opens the part.""" particle = ("particle" in tags - and tags.isdisjoint(_NOT_A_SURNAME_PARTICLE)) + and not (leading and "vocab:title" in tags)) if "vocab:suffix" in tags: return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY return _PARTICLE_ONLY if particle else _NO_TAGS -def surname_unit_tags(text: str, lexicon: Lexicon) -> frozenset[str]: +def surname_unit_tags(text: str, lexicon: Lexicon, + leading: bool) -> frozenset[str]: """`surname_unit_facts` for one token from the vocabulary alone -- segment's view of classify's tags, built before classify runs, with - classify's own tests: particle, title and bound given-name - membership, `suffix_as_written`, and the period-joined derivation. + classify's own tests: particle and title membership, + `suffix_as_written`, and the period-joined derivation. Kept from drifting by test_classify.test_surname_unit_tags_agree_with_classify.""" n = _normalize(text) @@ -898,8 +900,8 @@ def surname_unit_tags(text: str, lexicon: Lexicon) -> frozenset[str]: suffix = suffix_as_written(n, text, lexicon) or ( "." in text and n not in lexicon.titles and period_joined_vocab(text, lexicon) == "suffix") - particle = (n in lexicon.particles and n not in lexicon.titles - and n not in lexicon.bound_given_names) + particle = n in lexicon.particles and not ( + leading and n in lexicon.titles) if suffix: return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY return _PARTICLE_ONLY if particle else _NO_TAGS @@ -919,7 +921,8 @@ def surname_unit_count(texts: Sequence[str], lexicon: Lexicon) -> int: # a title-particle passes it -- so it only ever skips work. if not any(_normalize(t) in lexicon.particles for t in texts): return len(texts) - return len(unit_ends([surname_unit_tags(t, lexicon) for t in texts], + return len(unit_ends([surname_unit_tags(t, lexicon, i == 0) + for i, t in enumerate(texts)], chain=False)) @@ -966,7 +969,8 @@ def name_word_count(texts: Sequence[str], lexicon: Lexicon, return sum(names) count = 0 start = 0 - for end in unit_ends([surname_unit_tags(t, lexicon) for t in texts], + for end in unit_ends([surname_unit_tags(t, lexicon, i == 0) + for i, t in enumerate(texts)], chain=False): if any(names[start:end]): count += 1 diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 4ddb9e0d..738c862d 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -1080,14 +1080,26 @@ def _check_cjk_shape_purity(self) -> None: "comma, so it stays a title in front of the surname. The " "first draft of #575 counted 'Freiherr von Berg' as one " "surname and folded the title into the family"), - Case("a_bound_given_particle_before_a_comma_is_still_a_given_name", + Case("a_bound_given_particle_surname_before_a_comma_is_one_name_word", "Abu Bakar, Ed", - {"given": "Abu", "family": "Bakar", "suffix": "Ed"}, - ambiguities=("suffix-or-name", "particle-or-given"), - notes="#575 boundary: 'abu' is a bound given-name head as well " - "as a particle, and in front of a name it reads as the " - "given-name join (P5), so it is not one surname with the " - "word after it"), + {"given": "Ed", "family": "Abu Bakar"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="'abu' is a bound given-name head and a particle; before " + "a comma it is the particle, so 'Abu Bakar' is one " + "surname, as 2.0 through 2.3 read it. 1.4.0 and this " + "cycle's count read given 'Abu', family 'Bakar', suffix " + "'Ed'. Untagged: its 1.4.0 diff is the rule's own"), + Case("a_title_particle_inside_a_surname_chains", + "de St Pierre, Ed", + {"given": "Ed", "family": "de St Pierre"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="#575: a title-particle stops being a particle only where " + "it LEADS the part; inside a surname it chains, so 'de St " + "Pierre' is one name word, as 2.3.0 read it. A draft that " + "excluded 'St' in every position read family 'de St " + "Pierre', suffix 'Ed', no given name"), Case("a_suffix_inside_the_part_stops_the_particle", "van Jr. Berg, Mr.", {"title": "Mr.", "given": "van", "middle": "Jr.", diff --git a/tests/v2/pipeline/test_classify.py b/tests/v2/pipeline/test_classify.py index 81c5ad89..457b547e 100644 --- a/tests/v2/pipeline/test_classify.py +++ b/tests/v2/pipeline/test_classify.py @@ -754,7 +754,11 @@ def _surname_unit_disagreements() -> list[str]: tags = _tags_by_text(f"Smith {word}, John", lexicon=lex).get(word) if tags is None: # tokenize split it; nothing to compare continue - if surname_unit_tags(word, lex) != surname_unit_facts(tags): + # both positions: the leading one is where a title-particle + # ('Freiherr', 'St') stops being a particle + if any(surname_unit_tags(word, lex, leading) + != surname_unit_facts(tags, leading) + for leading in (True, False)): disagree.append(word) return disagree @@ -773,9 +777,9 @@ def test_surname_unit_tags_agree_with_classify() -> None: # classify has tagged anything, while assign derives them from # classify's tags (`_vocab.surname_unit_facts`). Swept over every # single-word entry of the vocabularies a token can be tagged - # from, in three casings, plus the period shapes classify derives - # a suffix from; a title-particle ('Freiherr') and a bound - # given-name particle ('Abu') are in the sweep through those lists. + # from, in three casings and both positions, plus the period shapes + # classify derives a suffix from; the title-particles ('Freiherr', + # 'St') are in the sweep through those lists. assert _surname_unit_disagreements() == [] diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index ba52c291..84edcf5e 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -768,8 +768,8 @@ fields = ["title", "suffix"] [[change]] issue = "fix(#575) a particle surname before a comma is one name word" -# rules.md#C1: "a particle and the name word it attaches to are one -# word, as P2 joins them" -- so the part before the comma is one +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word" -- so the part before the comma is one # surname and the word after it is the given name, as 'Royce, Ed' # reads. 1.4.0 counted tokens: 'De La Cruz, Ed' and 'de la Cruz, Ma' # read suffix 'Ed'/'Ma' with no given name, and 'Van Buren, Ed' read diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 42383138..7c042eaf 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3896,8 +3896,8 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#575) a particle surname before a comma is one name word" -# rules.md#C1: "a particle and the name word it attaches to are one -# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' # does: family 'van der Berg', suffix 'PhD'. Every release counted # tokens and split the chain, given 'van', family 'der Berg', with a # particle-or-given report. diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 38363584..c0899981 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3807,8 +3807,8 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#575) a particle surname before a comma is one name word" -# rules.md#C1: "a particle and the name word it attaches to are one -# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' # does: family 'van der Berg', suffix 'PhD'. Every release counted # tokens and split the chain, given 'van', family 'der Berg', with a # particle-or-given report. diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index f7b471f8..4e4a2eb3 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2190,8 +2190,8 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#575) a particle surname before a comma is one name word" -# rules.md#C1: "a particle and the name word it attaches to are one -# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' # does: family 'van der Berg', suffix 'PhD'. Every release counted # tokens and split the chain, given 'van', family 'der Berg', with a # particle-or-given report. diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 37e52a7e..f005625e 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1451,8 +1451,8 @@ orders = ["DEFAULT"] [[change]] issue = "fix(#575) a particle surname before a comma is one name word" -# rules.md#C1: "a particle and the name word it attaches to are one -# word, as P2 joins them", so 'van der Berg, PhD' reads as 'Berg, PhD' +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' # does: family 'van der Berg', suffix 'PhD'. Every release counted # tokens and split the chain, given 'van', family 'der Berg', with a # particle-or-given report. From b4199f7de00450269d405ed41d6667400065f805 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 21:04:17 -0700 Subject: [PATCH 4/5] test(C1): pin assign's leading flag for #575; fix three doc claims The third review found assign's title-particle position flag unguarded: fixed to False or to True, every test passed. 'St John, Dr.' and 'von St Johann, PhD' now pin it one per direction, and both mutants fail. decisions.md#C1: the blast-radius recipe's second step now reproduces its "two" (ambiguous_class_candidate in the second segment; a filter on any suffix word keeps 17); the out-of-corpus moves are stated as the class they are, with examples measured against 2.3.0; the guard claim names which rows pin which count. rules.md#C1 interacts gains H1. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 ++-- docs/design/rules.md | 2 +- tests/v2/cases.py | 18 ++++++++++++++++++ 3 files changed, 21 insertions(+), 3 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 88b29133..ec31d2bf 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -884,10 +884,10 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). - A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see position itself, which the `de St Pierre, Ed` case row pins. ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. + A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. `De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name. SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out). - BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose part after it holds a suffix-vocabulary or dotted-credential word. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Moves outside the corpora against 2.3.0 are `Van Johnson, Dr.` and `Abu Bakar, PhD`, above; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. + BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index bc7dcd8a..68ba628b 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1920,7 +1920,7 @@ C1. Rationale: a credential run after the comma means the name is in V` reads the suffix and `Smith, John PhD I.` continues the run, while adding a suffix comma after either turns that same letter into the middle initial. - history: decisions.md#C1 · interacts: H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py + history: decisions.md#C1 · interacts: H1, H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py C2. Rationale: text beyond the recognized comma parts should be taken in without silent guessing. diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 738c862d..5be58a7b 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -1118,6 +1118,24 @@ def _check_cjk_shape_purity(self) -> None: "counts the title as a word, so the comma reads as a " "credential comma and the part reads as it does alone, " "P1's fork and all. Unchanged"), + # #575: assign's own count (the positional read after a comma + # followed by no name word) applies the title-particle exclusion + # by position too. These two pin that flag both ways -- a review + # mutant fixing it to False or True passed every other test. + Case("a_leading_title_particle_stays_a_title_before_a_title_only_comma", + "St John, Dr.", + {"title": "St Dr.", "family": "John"}, + notes="#575 boundary: 'St' opens the part, so it is the title " + "it also is and 'John' alone is the name: two words do " + "not stand before the comma, the positional read holds. " + "2.3.0's reading"), + Case("an_inner_title_particle_chains_before_a_credential_only_comma", + "von St Johann, PhD", + {"family": "von St Johann", "suffix": "PhD"}, + classification="fix(#575)", + notes="#575: 'St' inside the surname is a particle, so 'von St " + "Johann' is one name word and the family comma keeps it " + "whole. 2.3.0 read given 'von'"), Case("the_anchor_does_not_reach_across_a_maiden_clause", "Jane Doe Jr. nee Smith Ma", {"given": "Jane", "family": "Doe", "suffix": "Jr.", From 733cc17e7e468f9f651e9d90d1985e23379ead7f Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 21:14:22 -0700 Subject: [PATCH 5/5] docs(C1): Freiherr von Berg as the surname is the right reading (#575) Since 1919 a former German noble title is part of the legal surname, written between the given name and the particle, so 'Freiherr von Berg, Ed' reading family 'Freiherr von Berg' is correct, not a cost -- the cost is the 'Prof. Cruz, Ed' half of the same listing-form rule. rules.md#C1's Accepted block, decisions.md#C1 and the ledger comments say so, and particles.py records why 'freiherr' is a particle at all (added in 3e14ea20 without a stated reason). Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/design/rules.md | 9 +++++++-- nameparser/config/particles.py | 5 ++++- tools/differential/expected_since_1.4.0.toml | 10 +++++----- tools/differential/expected_since_2.0.0.toml | 2 +- tools/differential/expected_since_2.1.0.toml | 2 +- tools/differential/expected_since_2.2.0.toml | 2 +- tools/differential/expected_since_2.3.0.toml | 2 +- 8 files changed, 21 insertions(+), 13 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index ec31d2bf..037d5910 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -884,7 +884,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). - A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). ACCEPTED: `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed' — the surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed` and `Freiherr Berg, Ed` (rules.md#C1's Accepted line). A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. + A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed'. The surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed`. For `Prof.` that is a cost (rules.md#C1's Accepted line); for a German rank it is the right reading (Derek, 2026-10-01): since 1919 a former noble title is part of the legal surname, written between the given name and the particle (Karl-Theodor Freiherr von und zu Guttenberg), so `Freiherr von Berg` is the family name. That is also why `freiherr` is in PARTICLES at all — added with the German prefixes in 3e14ea20 (#18) without a stated reason, and the reason is this one. `Freiherr von Berg, PhD` keeps title 'Freiherr', the unambiguous credential's count taking the title as a word; both are readings of one name, and neither is changed here. A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. `De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name. SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out). BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. diff --git a/docs/design/rules.md b/docs/design/rules.md index 68ba628b..c7475e8a 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1877,8 +1877,13 @@ C1. Rationale: a credential run after the comma means the name is in counts once and the title is no name word, so the count reads the listing form, and the listing form keeps a title before the comma in the family, as it already did for any one-word surname ('Prof. - Cruz, Ed', 'Freiherr Berg, Ed'). The count had read 'von' as a - second name word and kept the title apart. + Cruz, Ed'). For a German rank that reading is the right one: + since 1919 a former noble title is part of the legal surname, + written between the given name and the particle, so 'Freiherr von + Berg' is the family name in 'Freiherr von Berg, Ed', as 2.0 + through 2.3 read it. Before an unambiguous credential the same + words keep the title apart ('Freiherr von Berg, PhD'), that count + taking the title as a word; both are readings of one name. "Freiherr von Berg, Ed" → family="Freiherr von Berg" Accepted: three or more initials run together with periods behind a surname of two words read as a credential. Initials are diff --git a/nameparser/config/particles.py b/nameparser/config/particles.py index 62696765..d8dff239 100644 --- a/nameparser/config/particles.py +++ b/nameparser/config/particles.py @@ -214,7 +214,10 @@ # {do, freiherr, st} as load-bearing for the emitter 'du', # Du is a Chinese surname (Du Fu), leading under a # family-first reading - 'freiherr', # German noble title; also in TITLES, load-bearing + 'freiherr', # German rank; since 1919 a former noble title is part + # of the legal surname, written before the particle + # ("Karl-Theodor Freiherr von und zu Guttenberg"), so + # it chains like one. Also in TITLES, load-bearing 'freiherrin', # as above 'heer', 'la', # La Shawn, La Toya: a given-name element. The Romance diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 84edcf5e..9e6fabe0 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -783,11 +783,11 @@ issue = "fix(#575) a particle surname before a comma is one name word" # # Literal; the probe 'John van Buren, Ed' (two name words: the # credential reading stays) is _MUST_NOT_MATCH. -# 'Freiherr von Berg, Ed' joins as rules.md#C1's accepted -# consequence: the surname counts once and 'Freiherr' is no name word, -# so the listing form reads it and keeps the title in the family, as -# it does for 'Prof. Cruz, Ed'. 1.4.0 read title 'Freiherr', first -# 'von Berg', suffix 'Ed'. +# 'Freiherr von Berg, Ed' joins, rules.md#C1's Accepted example: the +# surname counts once and 'Freiherr' is no name word, so the listing +# form reads it and keeps the rank in the family -- since 1919 part of +# the legal German surname. 1.4.0 read title 'Freiherr', first 'von +# Berg', suffix 'Ed'. name_regex = "^(?:De La Cruz, Ed|Freiherr von Berg, Ed|Van Buren, Ed|de la Cruz, Ma)$" fields = ["title", "given", "family", "suffix"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 7c042eaf..569fd613 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2533,7 +2533,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, # reads as these baselines read it and gains only the comma's report. name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index c0899981..8aaa93c7 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2420,7 +2420,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, # reads as these baselines read it and gains only the comma's report. name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 4e4a2eb3..0bb052e3 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1009,7 +1009,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, # reads as these baselines read it and gains only the comma's report. name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index f005625e..8f149d1d 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -297,7 +297,7 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # is one, and its capitals make 'MA' the credential, as 'Smith, MA'. # These baselines already read the particle surnames as one name; # what #575 fixed is this cycle's count, which had not. -# 'Freiherr von Berg, Ed', #575's accepted consequence in rules.md#C1, +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, # reads as these baselines read it and gains only the comma's report. name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"]