diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 8418c3b8..037d5910 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -880,6 +880,15 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had brought back a split of `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd'. That split is not only an intermediate state of #552: 2.2.0 and 2.3.0 shipped it (2.0.0 and 2.1.0 read family 'Smith vd Ma'), #530 removed it in this cycle, and #552 kept it removed. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them. The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. It is also 1.4.0's reading, verified against the released wheel the same day: 1.4.0 read suffix 'vd Ma' and 'VD MA', and 2.0.0 through 2.3.0 read the whole name as the family. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` before a comma and after a family comma, and the silent mixed-case `Doe, Jane PhD vd Ma`. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. +- 2026-10-01 (Derek), #575 — THE COUNT BEFORE THE COMMA COUNTS A PARTICLE SURNAME ONCE. Every count rules.md#C1 makes of the part before the comma — v1's "more than one word" for an unambiguous credential, the ambiguous class's name-word count (#289, #544), and assign's two-name-word test for the positional read after a comma followed by no name word (#296/#325) — counted tokens or pieces, so `De La Cruz` was three words and `van der Berg` two (the ambiguous `van` stands as its own piece, P1's fork). This cycle's count had therefore read `De La Cruz, Ed` as a credential comma with no given name and split `van der Berg, MA` into given 'van', family 'der Berg' — an unreleased regression against 2.3.0, which read both as the listing form — and 1.4.0 through 2.3.0 all split `van der Berg, PhD`. DECIDED: a particle run and the one name word it attaches to are one word in all three counts, so a particle surname reads as a one-word surname does. + THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words. + THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before. + ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`). + A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed'. The surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed`. For `Prof.` that is a cost (rules.md#C1's Accepted line); for a German rank it is the right reading (Derek, 2026-10-01): since 1919 a former noble title is part of the legal surname, written between the given name and the particle (Karl-Theodor Freiherr von und zu Guttenberg), so `Freiherr von Berg` is the family name. That is also why `freiherr` is in PARTICLES at all — added with the German prefixes in 3e14ea20 (#18) without a stated reason, and the reason is this one. `Freiherr von Berg, PhD` keeps title 'Freiherr', the unambiguous credential's count taking the title as a word; both are readings of one name, and neither is changed here. A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'. + `De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name. + SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out). + BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. + COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. ### T1 — separators, not joiners diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 1c8637fa..e818d42d 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -29,7 +29,7 @@ Problem shape. A rule needs to know how words were JOINED (chained titles, parti ## UNIT-PARTITION — count the units the joining rules built -Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all. How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_post_rules.py (`_unit_end`, `_units`); rules.md P1 is the counting rule, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. +Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma (`_vocab.name_word_count`, `surname_unit_count`, and assign's positional-read test), which run partly before classify and so build their facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember. ## MARK-DONT-STRIP — record the decision, keep the fact diff --git a/docs/design/rules.md b/docs/design/rules.md index d55988ff..c7475e8a 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -518,7 +518,8 @@ P3. Rationale: connective words ("y", "of the") bind name words into have gone the other way, and is reported. The joined part is ONE name word wherever another rule counts them, so a rule taking "one name word" takes the whole join and - never half of it. + never half of it — except C1's count of the words before a comma, + which counts particle surnames alone and states why. A connective counts as a name word wherever this rule counts them, whatever else the vocabulary says the word is, where it is placed to join. A word can be a connective and a generation at once — the @@ -1669,6 +1670,27 @@ C1. Rationale: a credential run after the comma means the name is in than one word precedes the comma; otherwise it reads as the listing form, the part before the comma being the family name. Only the part after the first comma decides. + Wherever this rule counts the words before the comma, a particle + run and the one name word it attaches to are one word, the reach + P1's fold counts rather than the whole of P2's chain: the listing + form puts a surname before the comma, and a particle surname is + one surname ('De La Cruz, Ed' reads as 'Royce, Ed' does, 'van der + Berg, PhD' as 'Berg, PhD'). Where that surname is the whole part, + the comma settles P1's fork for a leading particle that could be + a given name, being the evidence that the part is a surname; a + word after the one the particle attaches to is a second word, and + a suffix word ends the particle's reach. A word that is also + title vocabulary is not a particle to this count where it opens + the part, reading there as the title it also is (H1); inside a + surname it chains ('de St Pierre, Ed' reads given 'Ed'). A + connective join and a bound given-name pair are not one word + here: whether P3 joins a connective depends on the words of the + whole name, which this count is part of deciding, and a bound + given name builds a given name rather than a surname (P5) — a + word that is a particle as well counts as the particle ('Abu + Bakar, Ed' reads given 'Ed'). A title in front is a word to the + count an unambiguous credential takes, so 'Dr. van der Berg, PhD' + reads as its part does alone. For the ambiguous credential class — a bare acronym the vocabulary marks as also an ordinary name, and a word admitted to the class by shape, which S3 defines and bounds — the @@ -1694,9 +1716,10 @@ C1. Rationale: a credential run after the comma means the name is in company has it, and where nothing but words of both the suffix and the title vocabulary stands in front of them, those are the titles of the given part the initials open, as S2 reads them - there. Paired initials that are not the only shape word in the - part are a credential however little else speaks for them, since - no one writes a person's initials as two dotted groups. The same count reads a part of two or + there. Behind two or more name words, paired initials that are + not the only shape word in the part are a credential however + little else speaks for them, since no one writes a person's + initials as two dotted groups. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any @@ -1833,13 +1856,35 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, A.B." → given="A.B." "John Smith, A.B." → ambiguities=("suffix-or-name",) "John Smith, X.Y. P.Q." → suffix="X.Y. P.Q." - "De La Cruz, M.J. K.L." → ambiguities=("suffix-or-name",) + "De La Cruz, M.J. K.L." → given="M.J." + "De La Cruz, M.J. K.L." → ambiguities=("suffix-or-name", "suffix-or-name") + "De La Cruz, Ed" → given="Ed" + "van der Berg, MA" → family="van der Berg" + "van der Berg, MA" → suffix="MA" + "Van Buren, Ed" → family="Van Buren" + "van der Berg, PhD" → family="van der Berg" + "John van Buren, Ed" → suffix="Ed" · boundary + "Ortega y Gasset, Ed" → suffix="Ed" · boundary + "abdul Salam, Ed" → suffix="Ed" · boundary "John Smith, PhD X.Y." → suffix="PhD X.Y." "García Márquez, G.J" → given="G.J" "De La Cruz, M.J. PhD" → given="M.J." · boundary "García Márquez, Ms G.J." → given="G.J." · boundary "García Márquez, Ed G.J." → given="Ed" · boundary "John Smith, A.B. Ph.D." → given="A.B." · boundary + Accepted: a title in front of a particle surname, before the + credential class's word, is read into the family. The surname + counts once and the title is no name word, so the count reads the + listing form, and the listing form keeps a title before the comma + in the family, as it already did for any one-word surname ('Prof. + Cruz, Ed'). For a German rank that reading is the right one: + since 1919 a former noble title is part of the legal surname, + written between the given name and the particle, so 'Freiherr von + Berg' is the family name in 'Freiherr von Berg, Ed', as 2.0 + through 2.3 read it. Before an unambiguous credential the same + words keep the title apart ('Freiherr von Berg, PhD'), that count + taking the title as a word; both are readings of one name. + "Freiherr von Berg, Ed" → family="Freiherr von Berg" Accepted: three or more initials run together with periods behind a surname of two words read as a credential. Initials are conventionally written apart ('J. R. R.'), and a token of three or @@ -1880,7 +1925,7 @@ C1. Rationale: a credential run after the comma means the name is in V` reads the suffix and `Smith, John PhD I.` continues the run, while adding a suffix comma after either turns that same letter into the middle initial. - history: decisions.md#C1 · interacts: H2, P2, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py + history: decisions.md#C1 · interacts: H1, H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py C2. Rationale: text beyond the recognized comma parts should be taken in without silent guessing. diff --git a/docs/release_log.rst b/docs/release_log.rst index 9631c654..16b12f5c 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -6,6 +6,8 @@ Release Log **Behavior Changes** + - **Fix a particle surname before a comma being split when a credential follows it.** ``HumanName("van der Berg, PhD")`` gives last ``van der Berg``, suffix ``PhD``, where 1.4.0 through 2.3.0 gave first ``van``, last ``der Berg``. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as ``Berg, PhD`` does, and ``Abu Bakar, PhD`` gives last ``Abu Bakar`` where 1.4.0 through 2.3.0 gave first ``Abu``. The same count keeps a particle surname whole in front of a credential that is also a name: ``De La Cruz, Ed`` gives first ``Ed``, last ``De La Cruz``, as 2.3.0 read it, and ``van der Berg, MA`` gives last ``van der Berg``, suffix ``MA``, the reading ``Smith, MA`` gets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name: ``Van Johnson, Dr.`` gives last ``Van Johnson``, where 2.2 and 2.3 gave first ``Van``. A given name in front still makes two words, so ``John van Buren, Ed`` reads suffix ``Ed`` as ``John Smith, Ed`` does. A connective surname is not counted as one: ``Ortega y Gasset, PhD`` still gives first ``Ortega``, middle ``y``, as ``Ortega y Gasset`` reads on its own. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #575) + - **Fix a one-letter connective joining a name that gives no sign it is a connective.** ``HumanName("jose e maria santos")`` gives first ``jose``, middle ``e maria``, last ``santos``, where 1.4.0 through 2.3.0 gave first ``jose e maria``; and ``JUAN GARCIA Y LOPEZ`` gives last ``GARCIA Y LOPEZ``, where every release since 1.4.0 read the bare capital as an initial and gave middle ``GARCIA Y``. A single letter is an initial where the writing says so -- a bare Latin capital in a name that is not written wholly in one case -- and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there: ``e`` reads as an initial and ``y`` joins. Mixed-case input is untouched in both directions: ``Jose e Maria Santos`` still gives first ``Jose e Maria`` and ``Jose E Maria Santos`` still gives middle ``E Maria``. Short names move in the derived views rather than the fields, P3's three-word carve-out being unchanged: ``parse("john e smith").initials()`` is ``j. e. s.`` where 2.3.0 gave ``j. s.``, and ``HumanName("john e smith").capitalize()`` gives ``John E Smith`` where 2.3.0 gave ``John e Smith``; ``JUAN Y GARCIA`` moves its ``capitalize()`` the same way in reverse, giving ``Juan y Garcia``, while its initials do not move at all: ``parse(...).initials()`` is ``J. Y. G.``, what 2.3.0 gave and what 1.4.0's own view gave, the ``Y`` holding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461). ``HumanName.initials()`` agrees with the core on both -- see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (``Хосе И Мария Сантос`` still gives first ``Хосе И Мария``), and Arabic ``و`` never enters the rule, having no case to be written against. A ``Lexicon`` knob decides which letters are marked, so the reading is configurable rather than fixed. See the ``P3`` entry of ``docs/design/decisions.md`` (closes #383, closes #479) - **Fix HumanName.initials() reading a one-letter connective by vocabulary and written shape instead of by the parse.** ``HumanName("john e smith").initials()`` gives ``j. e. s.``, where every release from 1.4.0 through 2.3.0 gave ``j. s.``; ``JUAN GARCIA Y LOPEZ`` gives ``J. G. L.`` where 2.3.0 gave ``J. G. Y. L.``; ``JUAN Y GARCIA`` does not move at all this cycle, giving ``J. Y. G.`` on both surfaces as 2.3.0 and 1.4.0 did, since the connective-initials fix further down this list (closes #461) gives its ``Y`` an initial again. That first name read 1.4.0's way at 2.3.0 and only there: 2.0.0 through 2.2.0 already gave today's answer, by the unrelated bug the 2.3.0 note below records as fixed (the facade dropping a bare capital that is also a one-letter conjunction, #462), so against those three releases it does not move at all. The v1 facade decided whether a word was the connective by looking the word up and checking its shape, while ``parse(...).initials()`` read the tag the parse recorded -- so the change above, which reads a single letter in a one-case name from the vocabulary rather than from its case, moved one view and not the other. Both views of a parse now give the same answer. Mixed-case names are untouched on both, the writing having decided the letter: ``John E Smith`` is still ``J. E. S.`` and ``Scott E. Werner`` still ``S. E. W.``. So is a one-case name whose letter is outside the marked set -- ``maria y lopez`` is still ``m. l.``, ``y`` having joined before this release and after it. Two costs, and both match what ``capitalize()`` has always done: editing ``C.conjunctions`` after a name is parsed no longer changes its initials until ``full_name`` is assigned again, and a name restored from a pickle, copied with ``copy.copy``/``copy.deepcopy`` (the same state hooks), or built from keyword fields (``HumanName(first=..., middle=..., last=...)``) carries no tags, so its initials come from the vocabulary and can differ from a fresh parse of the same string. One private break, stated because a v1 subclass can hit it: an override of ``_process_initial`` written to v1's ``(name_part, firstname=False)`` signature now raises ``TypeError`` the first time ``initials()`` runs, since ``initials()`` passes the part's tokens. Such an override has to accept a ``tokens`` keyword *and pass it on* -- ``return super()._process_initial(name_part, firstname, tokens=tokens)`` -- to receive this fix. Widening the signature without forwarding still works, but on the pre-#528 STRING path: the token call hands the override the group's own text as ``name_part`` rather than an empty placeholder, so ``john e smith`` initials ``j. s.`` under such an override, not the ``j. e. s.`` above. A subclass overriding one of the public ``first_list``, ``middle_list`` or ``last_list`` properties keeps working too: that member takes the pre-2.4 vocabulary reading instead of the change above, while an un-overridden member still moves. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #528) @@ -22,7 +24,7 @@ Release Log - **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516) - - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``De La Cruz, M.J. K.L.`` gives last ``De La Cruz``, suffix ``M.J. K.L.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` + - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` - **Fix a credential ending the given part of a family-comma listing being read as a middle name in silence.** ``HumanName("Doe, John MA")`` gives first ``John``, last ``Doe``, suffix ``MA``, where 2.0 through 2.3 gave middle ``MA`` -- and 1.4.0 gave the suffix, so this restores v1's reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone: ``Doe, John Ma`` keeps middle ``Ma``, written the way a name is written, and ``Doe, John Ed`` keeps middle ``Ed``. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields -- a declined name re-rendered without its comma, ``John Ma Doe``, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent -- ``Doe, John MA Smith`` gives middle ``MA Smith`` and reports nothing -- while a credential run or a trailing title is transparent to it: ``Doe, John MA PhD`` gives suffix ``MA PhD`` and ``Doe, John MA Prof.`` gives title ``Prof.`` with suffix ``MA``. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further: ``Doe, John Prof. MA`` now gives title ``Prof.`` where it gave middle ``Prof. MA``, the trailing-title chain reaching a word the credential used to hide; and ``Doe, John van MA`` gives last ``van Doe`` with suffix ``MA`` where it gave middle ``van MA``, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read: ``DOE, MARY JO MA``, ``doe, john ma``, ``田中, 太郎 MA`` and ``김, 민준 MA`` all give a suffix. The unlisted dotted spelling moves with them without the parity claim -- ``Doe, John X.Y.Z.`` gives suffix ``X.Y.Z.`` where 1.4.0 and 2.3.0 both gave a middle name -- to match the comma-less ``John Doe X.Y.Z.``. One word is carved out: ``do`` is the only member of this class that is also a surname particle, so capitals decide it and, with nothing in front of it, the particle reading keeps every other spelling (a degree in front is the other exception, the next-but-one entry). ``Doe, John DO`` gives suffix ``DO``, while ``Doe, John do``, ``Doe, John Do``, ``DOE, JOHN DO`` and ``doe, john do`` are unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, so ``SMITH, JOHN DO`` keeps last ``DO SMITH`` as ``NASCIMENTO, EDSON ARANTES DO`` does -- right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See the ``S2`` and ``P6`` entries of ``docs/design/decisions.md`` (closes #531) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index 7785fe88..0b4cde76 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -72,6 +72,7 @@ from nameparser._lexicon import Lexicon from nameparser._pipeline._vocab import ( effective_script, is_suffix_lenient, resolve_script_set, + surname_unit_facts, unit_ends, ) from nameparser._pipeline._pieces import ( anchor_in_reach, credential_at_the_given_slot, given_slot_anchors, @@ -1079,9 +1080,33 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: f"read as " f"{'a credential' if suffix_here else 'a name'}", (i2,))) - if reading is not None and sum( - 1 for k, piece in enumerate(fam_pieces) - if not is_suffix_piece(piece, fam_tags[k], tokens)) > 1: + # The two name words are counted as UNITS (#575, + # mechanisms.md#UNIT-PARTITION): group leaves an ambiguous + # leading particle a piece of its own (P1's fork), so 'van der + # Berg, PhD' holds two name pieces and one surname, and read + # positionally lost 'van' to the given name. The comma is the + # evidence that settles that fork: the listing form puts a + # surname before it. The same facts as segment's count + # (`_vocab.surname_unit_facts`), over EVERY token, so a suffix + # word still stops a particle ('van Jr. Berg, Mr.' is two name + # words); a unit counts when a non-suffix piece holds a token + # of it. + positional = False + if reading is not None: + idx = [i for piece in fam_pieces for i in piece] + named = {i for k, piece in enumerate(fam_pieces) + if not is_suffix_piece(piece, fam_tags[k], tokens) + for i in piece} + names = 0 + start = 0 + for end in unit_ends([surname_unit_facts(tokens[i].tags, k == 0) + for k, i in enumerate(idx)], + chain=False): + if not named.isdisjoint(idx[start:end]): + names += 1 + start = end + positional = names > 1 + if positional: order = _assign_main(0, state, tokens, ambiguities) else: for k, piece in enumerate(fam_pieces): diff --git a/nameparser/_pipeline/_post_rules.py b/nameparser/_pipeline/_post_rules.py index ec030523..00f0cd99 100644 --- a/nameparser/_pipeline/_post_rules.py +++ b/nameparser/_pipeline/_post_rules.py @@ -28,7 +28,7 @@ AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure, WorkToken, _NEVER_FLIPPED, comma_bucket, copy_with, ) -from nameparser._pipeline._vocab import delimiter_cores +from nameparser._pipeline._vocab import delimiter_cores, unit_ends from nameparser._policy import PatronymicRule from nameparser._types import ( FOLDED_TAG, UNJOINED_CONJUNCTION_TAG, UNJOINED_TAG, AmbiguityKind, Role, @@ -256,40 +256,6 @@ def _retag(tokens: list[WorkToken], i: int, role: Role) -> None: # rule counts them" # rules.md#P5: "a recognized bound given-name word joins the word # after it into one given name" -def _unit_end(tokens: list[WorkToken], idx: list[int], i: int) -> int: - """One past the end of the unit starting at `idx[i]`. - - RECURSIVE, and that is the whole point: what a conjunction or a - bound given-name word joins is the next UNIT, not the next word. - Absorbing a single index instead strands a particle at the end of - the unit, severed from the words it chains -- "de la Vega y la - Vega" cut between `la` and `Vega`, reporting family - "de la Vega y la", which is the same defect as a bare particle - opening the given name, mirrored.""" - if "particle" in tokens[idx[i]].tags: - j = i - while j + 1 < len(idx) and "particle" in tokens[idx[j + 1]].tags: - j += 1 - # ... then the words it joins, stopping where the next - # particle starts a group of its own, at a suffix word (the - # stop _group's chain uses), or at a conjunction, which the - # shared loop below joins to the whole unit after it rather - # than to the one word after it. - while (j + 1 < len(idx) - and "particle" not in tokens[idx[j + 1]].tags - and "conjunction" not in tokens[idx[j + 1]].tags - and "vocab:suffix" not in tokens[idx[j + 1]].tags): - j += 1 - end = j + 1 - else: - end = i + 1 - if "vocab:bound-given" in tokens[idx[i]].tags and end < len(idx): - end = _unit_end(tokens, idx, end) - while end + 1 < len(idx) and "conjunction" in tokens[idx[end]].tags: - end = _unit_end(tokens, idx, end + 1) - return end - - def _units(tokens: list[WorkToken], idx: list[int]) -> list[list[int]]: """`idx` split into the units other rules COUNT: one name word each, except where another rule has already made several words one @@ -316,8 +282,7 @@ def _units(tokens: list[WorkToken], idx: list[int]) -> list[list[int]]: out separately.""" units: list[list[int]] = [] i = 0 - while i < len(idx): - end = _unit_end(tokens, idx, i) + for end in unit_ends([tokens[k].tags for k in idx]): units.append(list(idx[i:end])) i = end return units diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 2c2e3776..d926d225 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -56,6 +56,7 @@ ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, caps_shape_candidate, is_one_case, is_paired_initials, is_single_letter_numeral, is_wholly_suffix, name_word_count, + surname_unit_count, run_word_fold, ) from nameparser._types import AmbiguityKind @@ -391,9 +392,15 @@ def class_run(seg: tuple[int, ...]) -> bool: case_class() pre_comma_names = name_word_count(texts(groups[0]), state.lexicon, state.policy) + # rules.md#C1's "more than one word precedes the comma" counts a + # particle and the word it attaches to as one word (#575): v1's + # token count split 'van der Berg, PhD' into given 'van', family + # 'der Berg'. The token count stays first, so a one-token part + # never builds the units. structure = ( Structure.SUFFIX_COMMA - if ((suffixy(groups[1]) and len(groups[0]) > 1) + if ((suffixy(groups[1]) and len(groups[0]) > 1 + and surname_unit_count(texts(groups[0]), state.lexicon) > 1) or (pre_comma_names is not None and pre_comma_names >= 2)) else Structure.FAMILY_COMMA) ambiguities = list(state.ambiguities) diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 84fd6dcf..531dca55 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -43,7 +43,7 @@ import functools import re import unicodedata -from collections.abc import Callable, Iterable, Sequence +from collections.abc import Callable, Iterable, Sequence, Set from typing import Literal from nameparser._lexicon import ( @@ -775,6 +775,157 @@ def ambiguous_class_candidate(text: str, lexicon: Lexicon, return False +# mechanisms.md#UNIT-PARTITION: "A rule that counts name words counts +# those units, and takes each whole or not at all" -- C1's count before +# a comma being the stated exception (`SURNAME_UNIT_TAGS`). The walk is shared +# by two readers that hold the facts in different forms: post_rules +# reads classify's TAGS, and segment, which runs before classify has +# tagged anything, builds the same tag names from the vocabulary +# (`surname_unit_tags`). One walk over tag sets, so the two cannot +# disagree about where a unit ends (mechanisms.md#ONE-PREDICATE-PER- +# QUESTION) -- only, at worst, about a token's facts, which +# test_classify's agreement test pins. +def unit_ends(tags: Sequence[Set[str]], chain: bool = True) -> list[int]: + """The END (one past the last index) of each unit of `tags`, in + order, partitioning `range(len(tags))`: one name word each, except + where another rule has already made several words one name -- a + particle and the words it chains (P2), a conjunction-joined run + (P3), and a bound given-name word with the word it completes (P5). + Each element is the tag set of one token, read for "particle", + "conjunction", "vocab:bound-given" and "vocab:suffix". + + `chain` picks how far a particle reaches. True is P2's chain: the + particle run and every word after it up to the next stop ("van der + Berg Smith" is one unit). False is P1's fold reach: the particle + run and the ONE unit it attaches to, so "de Mesnil Jean" is two -- + the count rules.md#C1 needs before a comma, where under a + family-first order that part is family 'de Mesnil', given 'Jean' + (#575). + + RECURSIVE, and that is the whole point: what a conjunction or a + bound given-name word joins is the next UNIT, not the next word. + Absorbing a single index instead strands a particle at the end of + the unit, severed from the words it chains -- "de la Vega y la + Vega" cut between `la` and `Vega`, reporting family + "de la Vega y la", which is the same defect as a bare particle + opening the given name, mirrored.""" + n = len(tags) + + def end_of(i: int) -> int: + if "particle" in tags[i] and not chain: + j = i + 1 + while j < n and "particle" in tags[j]: + j += 1 + end = (end_of(j) if j < n and "conjunction" not in tags[j] + and "vocab:suffix" not in tags[j] else j) + elif "particle" in tags[i]: + j = i + while j + 1 < n and "particle" in tags[j + 1]: + j += 1 + # ... then the words it joins, stopping where the next + # particle starts a group of its own, at a suffix word (the + # stop _group's chain uses), or at a conjunction, which the + # shared loop below joins to the whole unit after it rather + # than to the one word after it. + while (j + 1 < n + and "particle" not in tags[j + 1] + and "conjunction" not in tags[j + 1] + and "vocab:suffix" not in tags[j + 1]): + j += 1 + end = j + 1 + else: + end = i + 1 + if "vocab:bound-given" in tags[i] and end < n: + end = end_of(end) + while end + 1 < n and "conjunction" in tags[end]: + end = end_of(end + 1) + return end + + ends: list[int] = [] + i = 0 + while i < n: + i = end_of(i) + ends.append(i) + return ends + + +#: The facts rules.md#C1's count before a comma reads (#575): a +#: particle starts a surname unit, and a suffix word stops it. A word +#: that is ALSO title vocabulary ('Freiherr', 'St') LEADING the part is +#: not a particle here: in front of a name it reads as the title it +#: also is (rules.md#H1), so 'Freiherr von Berg, PhD' keeps its title, +#: while inside a surname it chains as P1 and P2 chain it ('de St +#: Pierre' is one surname). A bound given-name head that is also a +#: particle ('Abu') stays a particle: 'Abu Bakar, Ed' is a surname +#: before the comma, as 2.0 through 2.3 read it. Neither of the other +#: two joins `unit_ends` knows is read. A bound +#: given-name pair builds a GIVEN name, and P5 gives up a family word +#: where the name has no other ('abdul Salam' alone is given 'abdul', +#: family 'Salam'), so before a comma it is not one surname. Whether +#: P3 joins a single-letter connective depends on the words of the +#: whole name -- 'Carod i Rovira, Josep' joins as four words where +#: 'Ortega y Gasset' alone, three words, does not -- and the count is +#: part of deciding what the whole name is, so segment could reach +#: P3's answer only by copying P3. Segment builds these facts from the +#: vocabulary (`surname_unit_tags`); assign derives them from +#: classify's tags (`surname_unit_facts`), so the two counts read one +#: set of facts. +SURNAME_UNIT_TAGS = frozenset({"particle", "vocab:suffix"}) +_PARTICLE_ONLY = frozenset({"particle"}) +_SUFFIX_ONLY = frozenset({"vocab:suffix"}) +_NO_TAGS: frozenset[str] = frozenset() +def surname_unit_facts(tags: Set[str], leading: bool) -> frozenset[str]: + """`SURNAME_UNIT_TAGS` for one token, from classify's tags: the + reading assign's count before a comma takes. `leading` is whether + the token opens the part.""" + particle = ("particle" in tags + and not (leading and "vocab:title" in tags)) + if "vocab:suffix" in tags: + return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY + return _PARTICLE_ONLY if particle else _NO_TAGS + + +def surname_unit_tags(text: str, lexicon: Lexicon, + leading: bool) -> frozenset[str]: + """`surname_unit_facts` for one token from the vocabulary alone -- + segment's view of classify's tags, built before classify runs, with + classify's own tests: particle and title membership, + `suffix_as_written`, and the period-joined derivation. + Kept from drifting by + test_classify.test_surname_unit_tags_agree_with_classify.""" + n = _normalize(text) + # classify's whole-token test, then its period-joined derivation + # for a token no whole-token title or suffix claimed ('JD.CPA'), + # asked only of a token with a period, which it needs + suffix = suffix_as_written(n, text, lexicon) or ( + "." in text and n not in lexicon.titles + and period_joined_vocab(text, lexicon) == "suffix") + particle = n in lexicon.particles and not ( + leading and n in lexicon.titles) + if suffix: + return SURNAME_UNIT_TAGS if particle else _SUFFIX_ONLY + return _PARTICLE_ONLY if particle else _NO_TAGS + + +def surname_unit_count(texts: Sequence[str], lexicon: Lexicon) -> int: + """How many units (`unit_ends`) the part before a comma holds, a + particle chain (P2) and a connective join (P3) each counting once: + rules.md#C1's count before the comma, for the credential reading + of a part that is wholly suffix words (#575), with a particle + reaching as P1's fold does (`unit_ends`'s `chain=False`). 'van der + Berg' is one surname, so 'van der Berg, PhD' reads as 'Berg, PhD' + does. A connective join is not one unit here: `SURNAME_UNIT_TAGS`.""" + # With no particle, every token is a unit of its own, so the count + # is the token count: the suffix tests and the walk are skipped for + # the commonest comma names ('John Smith, PhD'). A superset test -- + # a title-particle passes it -- so it only ever skips work. + if not any(_normalize(t) in lexicon.particles for t in texts): + return len(texts) + return len(unit_ends([surname_unit_tags(t, lexicon, i == 0) + for i, t in enumerate(texts)], + chain=False)) + + def name_word_count(texts: Sequence[str], lexicon: Lexicon, policy: Policy) -> int: """How many of these texts are NAME words -- not suffix @@ -792,16 +943,39 @@ def name_word_count(texts: Sequence[str], lexicon: Lexicon, leading peel reads it as a title, and that divergence is recorded once, at `is_title_shaped` itself (mechanisms.md#ONE-PREDICATE-PER-QUESTION). + + It counts UNITS holding a name word, not tokens (#575, + mechanisms.md#UNIT-PARTITION): a particle chain (P2) builds one + family name, so 'De La Cruz' is one name word and 'De La Cruz, Ed' + reads as 'Royce, Ed' does. Why the other two joins are not read + here: `SURNAME_UNIT_TAGS`. """ predicate = (is_suffix_lenient if policy.lenient_comma_suffixes else is_suffix_strict) - n = 0 + + # One pass: each token's name-word verdict, and whether any token + # is a particle. With none, every token is a unit of its own and + # the count is a sum, so the commonest comma names never build the + # units -- and pay no frame the token loop did not already pay. + names: list[bool] = [] + has_particle = False for text in texts: - if (predicate(text, lexicon) or _normalize(text) in lexicon.titles - or is_title_shaped(text)): - continue - n += 1 - return n + n = _normalize(text) + if n in lexicon.particles: + has_particle = True + names.append(not (predicate(text, lexicon) or n in lexicon.titles + or is_title_shaped(text))) + if not has_particle: + return sum(names) + count = 0 + start = 0 + for end in unit_ends([surname_unit_tags(t, lexicon, i == 0) + for i, t in enumerate(texts)], + chain=False): + if any(names[start:end]): + count += 1 + start = end + return count def is_wholly_suffix(texts: Sequence[str], lexicon: Lexicon, diff --git a/nameparser/config/particles.py b/nameparser/config/particles.py index 62696765..d8dff239 100644 --- a/nameparser/config/particles.py +++ b/nameparser/config/particles.py @@ -214,7 +214,10 @@ # {do, freiherr, st} as load-bearing for the emitter 'du', # Du is a Chinese surname (Du Fu), leading under a # family-first reading - 'freiherr', # German noble title; also in TITLES, load-bearing + 'freiherr', # German rank; since 1919 a former noble title is part + # of the legal surname, written before the particle + # ("Karl-Theodor Freiherr von und zu Guttenberg"), so + # it chains like one. Also in TITLES, load-bearing 'freiherrin', # as above 'heer', 'la', # La Shawn, La Toya: a given-name element. The Romance diff --git a/tests/v2/cases.py b/tests/v2/cases.py index d767d031..5be58a7b 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -992,6 +992,150 @@ def _check_cjk_shape_purity(self) -> None: "family around a suffix 'vd'. 2.3.0 read family 'Smith " "Ma', suffix 'vd', and 1.4.0 last 'Smith', suffix 'vd, " "Ma'"), + # #575: rules.md#C1's count before the comma counts a particle + # chain (P2) as ONE name word -- it builds one family name -- so a + # particle surname reads as a one-word surname does ('Royce, Ed', + # 'Berg, MA', 'Berg, PhD'). + Case("a_particle_surname_before_a_comma_is_one_name_word", + "De La Cruz, Ed", + {"given": "Ed", "family": "De La Cruz"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="the count read 'De La Cruz' as three name words and " + "flipped to a credential comma, leaving no given name. " + "2.3.0 read given 'Ed' too; 1.4.0 read suffix 'Ed'", + shape=2), + Case("a_lowercase_particle_surname_before_a_comma_is_one_name_word", + "de la Cruz, Ma", + {"given": "Ma", "family": "de la Cruz"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="as 'Cruz, Ma' reads. 2.3.0 read given 'Ma' too; 1.4.0 " + "read suffix 'Ma'", + shape=2), + Case("an_ambiguous_particle_chain_before_a_comma_stays_whole", + "van der Berg, MA", + {"family": "van der Berg", "suffix": "MA"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="one name word before the comma, so C1 reads the case " + "and capitals in a mixed-case name make 'MA' the " + "credential, as 'Berg, MA' reads. The count had split " + "the chain: given 'van', family 'der Berg'. 2.3.0 read " + "given 'MA', family 'van der Berg'", + shape=2), + Case("a_leading_particle_that_could_be_a_given_name_is_settled_by_the_comma", + "Van Buren, Ed", + {"given": "Ed", "family": "Van Buren"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="standing alone 'Van Buren' is P1's fork (given 'Van'); " + "the listing form puts a surname before the comma, so " + "the comma settles it. The count had read two name " + "words: given 'Van', family 'Buren', suffix 'Ed'. " + "2.3.0 read given 'Ed', family 'Van Buren'", + shape=2), + Case("a_connective_surname_is_not_one_name_word_before_a_comma", + "Ortega y Gasset, Ed", + {"given": "Ortega", "middle": "y", "family": "Gasset", + "suffix": "Ed"}, + ambiguities=("suffix-or-name",), + notes="#575 boundary: only a particle chain counts as one name " + "word. P3 leaves a single-letter connective in a " + "three-word name a name word, so 'Ortega y Gasset' is " + "three even standing alone (given 'Ortega', middle 'y'), " + "and the count follows it. 2.3.0 read given 'Ed'", + shape=3), + Case("a_particle_surname_before_an_unambiguous_credential_stays_whole", + "van der Berg, PhD", + {"family": "van der Berg", "suffix": "PhD"}, + classification="fix(#575)", + notes="as 'Berg, PhD' reads. Every release split the chain on " + "v1's token count: given 'van', family 'der Berg', " + "suffix 'PhD'", + shape=2), + Case("a_particle_surname_after_a_given_name_is_still_a_full_name", + "John van Buren, Ed", + {"given": "John", "family": "van Buren", "suffix": "Ed"}, + ambiguities=("suffix-or-name",), + notes="#575 boundary: a given name and a particle surname are " + "two name words, so the part after the comma is still " + "the credential run", + shape=3), + Case("a_bound_given_pair_before_a_comma_counts_as_two_name_words", + "abdul Salam, Ed", + {"given": "abdul", "family": "Salam", "suffix": "Ed"}, + ambiguities=("suffix-or-name",), + notes="#575 boundary: a bound pair (P5) builds a GIVEN name and " + "gives up a family word when the name has no other " + "('abdul Salam' alone is given 'abdul', family 'Salam'), " + "so it is not one surname before a comma", + shape=3), + Case("a_title_particle_before_a_comma_is_still_a_title", + "Freiherr von Berg, PhD", + {"title": "Freiherr", "family": "von Berg", "suffix": "PhD"}, + ambiguities=("particle-or-given",), + notes="#575 boundary: a word of both the title and the particle " + "vocabulary is not a particle to the count before the " + "comma, so it stays a title in front of the surname. The " + "first draft of #575 counted 'Freiherr von Berg' as one " + "surname and folded the title into the family"), + Case("a_bound_given_particle_surname_before_a_comma_is_one_name_word", + "Abu Bakar, Ed", + {"given": "Ed", "family": "Abu Bakar"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="'abu' is a bound given-name head and a particle; before " + "a comma it is the particle, so 'Abu Bakar' is one " + "surname, as 2.0 through 2.3 read it. 1.4.0 and this " + "cycle's count read given 'Abu', family 'Bakar', suffix " + "'Ed'. Untagged: its 1.4.0 diff is the rule's own"), + Case("a_title_particle_inside_a_surname_chains", + "de St Pierre, Ed", + {"given": "Ed", "family": "de St Pierre"}, + classification="fix(#575)", + ambiguities=("suffix-or-name",), + notes="#575: a title-particle stops being a particle only where " + "it LEADS the part; inside a surname it chains, so 'de St " + "Pierre' is one name word, as 2.3.0 read it. A draft that " + "excluded 'St' in every position read family 'de St " + "Pierre', suffix 'Ed', no given name"), + Case("a_suffix_inside_the_part_stops_the_particle", + "van Jr. Berg, Mr.", + {"title": "Mr.", "given": "van", "middle": "Jr.", + "family": "Berg"}, + ambiguities=("particle-or-given",), + notes="#575 boundary: the particle reaches no further than a " + "suffix word, so 'van' and 'Berg' are two name words and " + "the part keeps its positional read. A first draft " + "walked past the suffix and read family 'van Berg'"), + Case("a_title_in_front_keeps_the_particle_fork", + "Dr. van der Berg, PhD", + {"title": "Dr.", "given": "van", "family": "der Berg", + "suffix": "PhD"}, + ambiguities=("particle-or-given",), + notes="#575 boundary: the count for an unambiguous credential " + "counts the title as a word, so the comma reads as a " + "credential comma and the part reads as it does alone, " + "P1's fork and all. Unchanged"), + # #575: assign's own count (the positional read after a comma + # followed by no name word) applies the title-particle exclusion + # by position too. These two pin that flag both ways -- a review + # mutant fixing it to False or True passed every other test. + Case("a_leading_title_particle_stays_a_title_before_a_title_only_comma", + "St John, Dr.", + {"title": "St Dr.", "family": "John"}, + notes="#575 boundary: 'St' opens the part, so it is the title " + "it also is and 'John' alone is the name: two words do " + "not stand before the comma, the positional read holds. " + "2.3.0's reading"), + Case("an_inner_title_particle_chains_before_a_credential_only_comma", + "von St Johann, PhD", + {"family": "von St Johann", "suffix": "PhD"}, + classification="fix(#575)", + notes="#575: 'St' inside the surname is a particle, so 'von St " + "Johann' is one name word and the family comma keeps it " + "whole. 2.3.0 read given 'von'"), Case("the_anchor_does_not_reach_across_a_maiden_clause", "Jane Doe Jr. nee Smith Ma", {"given": "Jane", "family": "Doe", "suffix": "Jr.", @@ -2347,16 +2491,17 @@ def _check_cjk_shape_purity(self) -> None: "takes for a name. 1.4.0 read given 'X.Y.', middle " "'P.Q.'", shape=3), - Case("two_pairs_behind_a_two_word_surname_report_the_flip", + Case("two_pairs_behind_a_particle_surname_open_the_given_part", "De La Cruz, M.J. K.L.", - {"family": "De La Cruz", "suffix": "M.J. K.L."}, - classification="fix(#563)", - ambiguities=("suffix-or-name",), - notes="#563 second review round: the cost of the two-pair " - "line is a surname with no given name, so the flip it " - "makes is reported rather than silent (rules.md#C1). " - "1.4.0 read given 'M.J.', middle 'K.L.'", - shape=3), + {"given": "M.J.", "family": "De La Cruz", "suffix": "K.L."}, + classification="fix(#575)", + ambiguities=("suffix-or-name", "suffix-or-name"), + notes="'De La Cruz' is one name word (#575), so the count no " + "longer reaches the part and it reads as 'Cruz, M.J. " + "K.L.' does (Derek, 2026-10-01). #563 had read it as a " + "flipped credential run with no given name. 1.4.0 read " + "given 'M.J.', middle 'K.L.'", + shape=2), Case("a_class_member_in_front_of_paired_initials_does_not_speak", "García Márquez, Ed G.J.", {"given": "Ed", "family": "García Márquez", "suffix": "G.J."}, diff --git a/tests/v2/pipeline/test_assign.py b/tests/v2/pipeline/test_assign.py index 50af5ebd..0c0733d6 100644 --- a/tests/v2/pipeline/test_assign.py +++ b/tests/v2/pipeline/test_assign.py @@ -648,14 +648,28 @@ def test_non_title_post_comma_segment_is_untouched() -> None: def test_positional_segment_zero_reports_the_particle_fork() -> None: # The comma no longer fixed the family, so the leading ambiguous # particle IS a live fork again -- emitted at the site that decides - # it, per the ambiguity doctrine. - out = _assigned("Van Johnson, Dr.") + # it, per the ambiguity doctrine. Three words, since the read needs + # two name words counted as units (#575): the particle and the word + # it attaches to are one. + out = _assigned("Van Johnson Smith, Dr.") assert _by_role(out, Role.GIVEN) == "Van" - assert _by_role(out, Role.FAMILY) == "Johnson" + assert _by_role(out, Role.FAMILY) == "Smith" assert [a.kind for a in out.ambiguities] == \ [AmbiguityKind.PARTICLE_OR_GIVEN] +def test_a_lone_particle_surname_before_a_title_only_comma_is_the_family() -> None: + # #575: 'Van Johnson' is ONE name word before the comma -- a + # particle and the word it attaches to -- so there are not two to + # read positionally, and the listing form's surname reading + # settles P1's fork. This test pinned given 'Van', family + # 'Johnson' and a particle-or-given report until then. + out = _assigned("Van Johnson, Dr.") + assert _by_role(out, Role.FAMILY) == "Van Johnson" + assert _by_role(out, Role.GIVEN) == "" + assert out.ambiguities == () + + def test_a_credential_run_after_a_family_comma_reads_as_suffixes() -> None: # 'Smith, Jr.' -- the peel's whole-segment exception claimed this # even with 'jr' out of TITLES, because is_leading_title also diff --git a/tests/v2/pipeline/test_classify.py b/tests/v2/pipeline/test_classify.py index ed2666d8..457b547e 100644 --- a/tests/v2/pipeline/test_classify.py +++ b/tests/v2/pipeline/test_classify.py @@ -6,6 +6,7 @@ from nameparser._lexicon import Lexicon, _normalize from nameparser._pipeline import STAGES from nameparser._pipeline import _classify as _classify_module +from nameparser._pipeline import _vocab from nameparser._pipeline._classify import classify from nameparser._pipeline._extract import extract_delimited from nameparser._pipeline._segment import segment @@ -14,7 +15,8 @@ ) from nameparser._pipeline._tokenize import tokenize from nameparser._pipeline._vocab import ( - ambiguous_class_candidate, ambiguous_class_member, caps_shape_candidate, + ambiguous_class_candidate, ambiguous_class_member, + caps_shape_candidate, surname_unit_facts, surname_unit_tags, ) from nameparser._policy import Policy from nameparser._types import AmbiguityKind, Role @@ -731,3 +733,61 @@ def test_the_caps_shape_is_a_testable_predicate() -> None: assert (_tags_by_text("John Smith X.Y", policy=on)["X.Y"] == _tags_by_text("John Smith X.Y")["X.Y"]) + + +def _surname_unit_sweep() -> list[str]: + lex = Lexicon.default() + words: set[str] = {"Smith", "Ma", "M.D.", "Ph.D.", "JD.CPA", "Msc.Ed.", + "Lt.Gov.", "X.Y.Z.", "V.", "I", "v", "Jr.", "de."} + for field in ("particles", "suffix_acronyms", "suffix_words", + "conjunctions", "bound_given_names", "titles"): + for w in getattr(lex, field): + if " " not in w: + words |= {w, w.title(), w.upper()} + return sorted(words) + + +def _surname_unit_disagreements() -> list[str]: + lex = Lexicon.default() + disagree = [] + for word in _surname_unit_sweep(): + tags = _tags_by_text(f"Smith {word}, John", lexicon=lex).get(word) + if tags is None: # tokenize split it; nothing to compare + continue + # both positions: the leading one is where a title-particle + # ('Freiherr', 'St') stops being a particle + if any(surname_unit_tags(word, lex, leading) + != surname_unit_facts(tags, leading) + for leading in (True, False)): + disagree.append(word) + return disagree + + +#: The agreement test's recorded negative control: what it reports +#: with `surname_unit_tags`' period-joined mirror switched off -- the +#: two words its first run caught, which classify tags as suffixes +#: through that derivation alone. +_SURNAME_UNIT_CONTROL = ["JD.CPA", "Msc.Ed."] + + +def test_surname_unit_tags_agree_with_classify() -> None: + # #575: rules.md#C1's count before the comma reads two facts per + # token -- particle, suffix -- and `segment` must build them from + # the vocabulary (`_vocab.surname_unit_tags`), since it runs before + # classify has tagged anything, while assign derives them from + # classify's tags (`_vocab.surname_unit_facts`). Swept over every + # single-word entry of the vocabularies a token can be tagged + # from, in three casings and both positions, plus the period shapes + # classify derives a suffix from; the title-particles ('Freiherr', + # 'St') are in the sweep through those lists. + assert _surname_unit_disagreements() == [] + + +def test_the_surname_unit_agreement_test_can_fail( + monkeypatch: pytest.MonkeyPatch) -> None: + # The guard above, with the mirror it exists to hold switched off: + # `_vocab.surname_unit_tags` reads `period_joined_vocab` from its + # own module, so patching it there leaves classify untouched. + monkeypatch.setattr(_vocab, "period_joined_vocab", + lambda text, lexicon: None) + assert _surname_unit_disagreements() == _SURNAME_UNIT_CONTROL diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 94b647e9..80bf60fb 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1381,6 +1381,10 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: "fix(#289) a written case contrast decides a bare ambiguous acronym": ("JOHN SMITH MA", "ANH DO", "anh van do", "jack ma", "Jack Ma", "Jack Ma.", "毛泽东, MA", "Smith, A.B.C.", "Jack M.A."), + # #575: two name words before the comma, a given name and a + # particle surname, keep the credential reading. + "fix(#575) a particle surname before a comma is one name word": + ("John van Buren, Ed", "John van der Berg, PhD"), # #516's dotted rule must never reach the shapes the vocabulary # SETTLES (a whole-token match, or a surviving chunk claim), the # leading dotted runs rules.md#S2 excludes by position, the single @@ -2458,13 +2462,20 @@ class _LatinCopy(NamedTuple): "John de Ma", "John van der Berg Ma", r"Smith Jr\., MA", r"Smith Jr\., Ma", "Smith, MA", "abdul Smith Berg Ma", "abdul Smith Ma", "john smith, ma"}), - frozenset({"Davis Royce, Ed", r"Doe, Dr\. MA", "Doe, MA", "Doe, MA PhD", - r"Doe, Mr\. MA PhD", "Freiherr von Berg MA", "JOHN SMITH, MA", - "Jack MA", r"Jack MA\.", "Jack Wei Ma", r"John Prof\. MA", + # 2026-10-01, #575: the 2.x copy gains eight C1 examples, comma + # names whose post-comma word is a listed member -- no set copied. + frozenset({"Davis Royce, Ed", "De La Cruz, Ed", r"Doe, Dr\. MA", + "Doe, MA", "Doe, MA PhD", r"Doe, Mr\. MA PhD", + "Freiherr von Berg MA", "Freiherr von Berg, Ed", + "JOHN SMITH, MA", "Jack MA", + r"Jack MA\.", "Jack Wei Ma", r"John Prof\. MA", "John Smith Ma", "John Smith, Ed", "John Smith, MA", - "John Smith, Ma", "John de Ma", "John van der Berg Ma", + "John Smith, Ma", "John de Ma", "John van Buren, Ed", + "John van der Berg Ma", "Ortega y Gasset, Ed", r"Smith Jr\., MA", r"Smith Jr\., Ma", "Smith, MA", - "abdul Smith Berg Ma", "abdul Smith Ma", "john smith, ma"}), + "Van Buren, Ed", "abdul Salam, Ed", "abdul Smith Berg Ma", + "abdul Smith Ma", "de la Cruz, Ma", "john smith, ma", + "van der Berg, MA"}), frozenset({r"García Márquez, G\.J\.R\.", r"Jack X\.Y\.I\.", r"John Smith B\.Tech\.", r"John Smith C\.H\.A\.", r"John Smith E\.S\.Q\.", r"John Smith Q\.W\.E\.R\.T\.", @@ -3067,6 +3078,10 @@ class _LatinCopy(NamedTuple): frozenset({"Doe, Jane nee Smith PhD MEng", "Jane Doe nee Smith PhD MA", "Jane Doe nee Smith PhD MEng"}), frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), + # 2026-10-01, #575: rules.md#C1's particle-surname examples at + # 1.4.0, literal names that copy no set. + frozenset({"De La Cruz, Ed", "Freiherr von Berg, Ed", "Van Buren, Ed", + "de la Cruz, Ma"}), # 2026-10-01, #554: the run rule gains rules.md#C1's particle-and- # suffix examples, a vocabulary word in front of a member, which # copies no set either. @@ -3775,8 +3790,11 @@ def _claim(rule: dict) -> _Claim: # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, # Jr do', #554's rules.md#C1 examples. Reach, verified name by # name. + # 2026-10-01, #575: 414 -> 423, its nine rules.md#C1 and + # case-row names, every one a comma name. Reach, verified + # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(414, ('given', 'suffix', 'title'), "46663c4aa90a", None), + _Claim(423, ('given', 'suffix', 'title'), "f5edb96b96cb", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3875,8 +3893,15 @@ def _claim(rule: dict) -> _Claim: # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, # Jr do', #554's rules.md#C1 examples. Reach, verified name by # name. + # 2026-10-01, #575: 414 -> 423, its nine rules.md#C1 and + # case-row names, every one a comma name. Reach, verified + # name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(414, ('family', 'given'), "46663c4aa90a", None), + _Claim(423, ('family', 'given'), "f5edb96b96cb", None), + # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von + # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(4, ('family', 'given', 'suffix', 'title'), "30163564e03d", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4137,16 +4162,18 @@ def _claim(rule: dict) -> _Claim: # pinning that a trailing link does not count itself for an # unrelated 'y', which is in the corpus because its row # carries a shape tag. Reach again, verified name by name. + # 2026-10-01, #575: 104 -> 105, 'Ortega y Gasset, Ed'. Reach. "fix(initials-per-word) a connective run initials each word (facade, since 2.0.0)": - _Claim(104, ('_initials',), "bc7dfeba1da5", ('DEFAULT',)), + _Claim(105, ('_initials',), "231bbda6768a", ('DEFAULT',)), # 2026-09-19, #533: 41 -> 43. Two new corpus names opening # with a bound-given word, 'Berg, abdul MA' and 'Berg, abdul # nee Jones MA' -- the P5 pair this change added to record # that the clause form now agrees with the bare one. # 2026-09-26, #535: 43 -> 44, 'Berg, abdul nee Smith V' -- # another opening bound-given word. Reach again. + # 2026-10-01, #575: 44 -> 45, 'abdul Salam, Ed'. Reach. "fix(initials-per-word) a bound-given run initials each word (facade, since 2.0.0)": - _Claim(44, ('_initials',), "4d6f497bebf7", ('DEFAULT',)), + _Claim(45, ('_initials',), "690d9625f120", ('DEFAULT',)), # 2026-09-18: 109 -> 110. One new corpus name, # 'john van der berg ma' -- rules.md#P2's one-case contrast, # and a particle chain like every other member. @@ -4159,8 +4186,12 @@ def _claim(rule: dict) -> _Claim: # 2026-09-30, #563: 115 -> 117, 'De La Cruz, M.J. PhD' and # 'De La Cruz, M.J. K.L.' -- a particle chain, 'De La Cruz'. # Reach, verified name by name. + # 2026-10-01, #575: 117 -> 124, its seven particle-surname + # names ('De La Cruz, Ed', 'de la Cruz, Ma', 'van der Berg, + # MA', 'Van Buren, Ed', 'van der Berg, PhD', 'John van Buren, + # Ed', 'Freiherr von Berg, Ed'). Reach, verified name by name. "fix(initials-per-word) a particle chain inside a name part initials each word (facade, since 2.0.0)": - _Claim(117, ('_initials',), '0820bd80678d', ('DEFAULT',)), + _Claim(124, ('_initials',), 'bf190fecb582', ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": @@ -4960,7 +4991,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -5457,7 +5493,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6097,7 +6138,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6460,7 +6506,12 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Mr. MA PhD' -- 'Doe, Dr. MA' with a credential run # behind the member. Verified to be that name and no other; # no role joined the list. - _Claim(23, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "1f21a81740ea", ('DEFAULT',)), + # 2026-10-01, #575: 23 -> 31, the eight C1 examples the + # alternation gained. Reach, verified name by name. + _Claim(31, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "997497ae83e5", ('DEFAULT',)), + # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. + "fix(#575) a particle surname before a comma is one name word": + _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -8429,6 +8480,13 @@ def test_a_rule_reaching_no_corpus_name_says_why_it_is_kept() -> None: "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 2), ("fix(#296) a credential-only comma string reads a name and its postnominal", "fix(comma-precomma-family) pre-comma run reads as family, not given", 2), + # 2026-10-01, #575: both declared on the #575 rule, over its + # four names; it stands ahead of the two comma rules because + # each describes the opposite or the other half of the move. + ("fix(#575) a particle surname before a comma is one name word", + "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 4), + ("fix(#575) a particle surname before a comma is one name word", + "fix(comma-precomma-family) pre-comma run reads as family, not given", 4), ("fix(#296) a lone post-comma credential is a suffix", "fix(suffix-routing) a two-token name ending in the suffix word jr keeps it in `suffix`", 2), ("fix(#400/#274) bound-given join and maiden consumption in one name", diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index d816b00b..7151dfac 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -37,6 +37,7 @@ "Carod i Rovira, Josep" "Carod y de Rovira i" "Davis Royce, Ed" +"De La Cruz, Ed" "De La Cruz, M.J. K.L." "De La Cruz, M.J. PhD" "Del Toro" @@ -84,6 +85,7 @@ "Duke of Edinburgh" "Esq. Smith" "Freiherr von Berg MA" +"Freiherr von Berg, Ed" "Freiherr von Richthofen V" "Gal·la Serra" "Garcia" @@ -231,6 +233,7 @@ "John née Jones Smith MA" "John née Jones Smith Ma" "John née Jones Smith V" +"John van Buren, Ed" "John van Mc" "John van der Berg" "John van der Berg Ma" @@ -295,6 +298,7 @@ "Nguyen, Van Le" "Nguyễn Thị Minh Khai" "Nguyễn, Thị Vân" +"Ortega y Gasset, Ed" "Ph. D. Van Johnson" "Ph. D., John" "Prince of Wales Jr" @@ -363,6 +367,7 @@ "Smith. John" "Steven Hardman, MD, DO, DDS" "The Right Hon. the President of the Queen's Bench Division" +"Van Buren, Ed" "Van Johnson" "Vega, Juan de la" "Vincent van Gogh van Beethoven" @@ -379,6 +384,7 @@ "abdul Jr Smith Berg" "abdul Ma Smith" "abdul Ph. D. Smith Berg" +"abdul Salam, Ed" "abdul Sir Smith Berg" "abdul Smith Berg Ma" "abdul Smith Jr Ma" @@ -446,6 +452,8 @@ "tran lac" "van Berg Jan de" "van Gogh" +"van der Berg, MA" +"van der Berg, PhD" "van der Berg, abdul née Jones" "wang meng" "Иван Петрович Абрамович" diff --git a/tools/differential/corpus_shapes.jsonl b/tools/differential/corpus_shapes.jsonl index a7801994..7cbaf51b 100644 --- a/tools/differential/corpus_shapes.jsonl +++ b/tools/differential/corpus_shapes.jsonl @@ -139,6 +139,8 @@ {"name": "Carod i Rovira, Josep", "shape": 2} {"name": "DOE, JOHN MA", "shape": 2} {"name": "DOE, MARY JO MA", "shape": 2} +{"name": "De La Cruz, Ed", "shape": 2} +{"name": "De La Cruz, M.J. K.L.", "shape": 2} {"name": "Doe, Dr. John MA", "shape": 2} {"name": "Doe, Dr. MA", "shape": 2} {"name": "Doe, Dr. MA Smith", "shape": 2} @@ -222,11 +224,14 @@ {"name": "Smith, PhD MEng", "shape": 2} {"name": "Smith, PhD Ma", "shape": 2} {"name": "Smith, meng", "shape": 2} +{"name": "Van Buren, Ed", "shape": 2} +{"name": "de la Cruz, Ma", "shape": 2} {"name": "de la Vega, Juan", "shape": 2} {"name": "doe, jane v phd do", "shape": 2} {"name": "doe, john ma", "shape": 2} +{"name": "van der Berg, MA", "shape": 2} +{"name": "van der Berg, PhD", "shape": 2} {"name": "Davis Royce, Ed", "shape": 3} -{"name": "De La Cruz, M.J. K.L.", "shape": 3} {"name": "De La Cruz, M.J. PhD", "shape": 3} {"name": "Dr. John P. Doe-Ray, CLU, CFP, LUTC", "shape": 3} {"name": "García Márquez, Ed G.J.", "shape": 3} @@ -255,7 +260,10 @@ {"name": "John Smith, X.Y. P.Q.", "shape": 3} {"name": "John Smith, X.Y.Z.", "shape": 3} {"name": "John Smith, X.Y.Z. MA", "shape": 3} +{"name": "John van Buren, Ed", "shape": 3} +{"name": "Ortega y Gasset, Ed", "shape": 3} {"name": "Steven Hardman, MD, DO, DDS", "shape": 3} +{"name": "abdul Salam, Ed", "shape": 3} {"name": "john smith, ma", "shape": 3} {"name": "john smith, md ma", "shape": 3} {"name": "john smith, phd meng", "shape": 3} diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 34268a49..9e6fabe0 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -766,6 +766,52 @@ issue = "fix(#325) a credential run across a second comma reads as suffixes" name_regex = "(?i)^smith,\\s*jr\\.?,\\s*phd$" fields = ["title", "suffix"] +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word" -- so the part before the comma is one +# surname and the word after it is the given name, as 'Royce, Ed' +# reads. 1.4.0 counted tokens: 'De La Cruz, Ed' and 'de la Cruz, Ma' +# read suffix 'Ed'/'Ma' with no given name, and 'Van Buren, Ed' read +# given 'Van', family 'Buren', suffix 'Ed'. +# +# Ahead of the routing rule below on purpose: that rule's alternation +# reaches the two 'De La Cruz' spellings, and it describes the +# opposite move (a lone post-comma piece leaving `first`), so letting +# it absorb a piece ARRIVING in `given` would explain the diff with a +# rule that contradicts it. +# +# Literal; the probe 'John van Buren, Ed' (two name words: the +# credential reading stays) is _MUST_NOT_MATCH. +# 'Freiherr von Berg, Ed' joins, rules.md#C1's Accepted example: the +# surname counts once and 'Freiherr' is no name word, so the listing +# form reads it and keeps the rank in the family -- since 1919 part of +# the legal German surname. 1.4.0 read title 'Freiherr', first 'von +# Berg', suffix 'Ed'. +name_regex = "^(?:De La Cruz, Ed|Freiherr von Berg, Ed|Van Buren, Ed|de la Cruz, Ma)$" +fields = ["title", "given", "family", "suffix"] + +[[change.precedes_narrower]] +issue = "fix(comma-family) lone post-comma piece routes to suffix/title, not first" +why = """ +The later rule describes a lone post-comma piece LEAVING `first` for +`suffix` or `title`. These names move the other way: the word after +the comma ARRIVES in `given`, because the particle surname before it +is one name word (rules.md#C1, #575), and that is what this rule +names. +""" + +[[change.precedes_narrower]] +issue = "fix(comma-precomma-family) pre-comma run reads as family, not given" +why = """ +The later rule describes 1.4.0 handing a pre-comma word to `given` +that 2.x reads as family, which is half of these moves ('Van' leaving +`given` in 'Van Buren, Ed'). It says nothing about the word AFTER the +comma arriving in `given`, the other half, nor why the particle +surname is one name: that is rules.md#C1's count (#575), which this +rule names. +""" + [[change]] issue = "fix(comma-family) lone post-comma piece routes to suffix/title, not first" # 'Smith, Dr.' / 'Andrews, M.D.': v1 put the lone strict-suffix-or-title @@ -4020,8 +4066,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 209325e5..569fd613 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2524,7 +2524,18 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: eight names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -2610,8 +2621,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a @@ -3878,3 +3893,17 @@ issue = "fix(#296/#531) phd is not a prenominal, so behind a dual title it is th name_regex = "^Smith, MD PhD Ma$" fields = ["title", "given", "middle", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index d96f4faf..8aaa93c7 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2411,7 +2411,18 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: eight names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -2497,8 +2508,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a @@ -3789,3 +3804,17 @@ issue = "fix(#296/#531) phd is not a prenominal, so behind a dual title it is th name_regex = "^Smith, MD PhD Ma$" fields = ["title", "given", "middle", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index c99de14a..0bb052e3 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1000,7 +1000,18 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: eight names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -1086,8 +1097,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a @@ -2172,3 +2187,17 @@ issue = "fix(#544) a credential in front anchors a member after a one-word famil name_regex = "^Smith, PhD Ma$" fields = ["given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index c6103486..8f149d1d 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -288,7 +288,18 @@ issue = "fix(#289) a written case contrast decides a bare ambiguous acronym" # 'Ma' -- is gone; what the name still diffs on here is claimed by the # rule that explains it, and a literal left in this one would stand # ready to explain that cost coming back. -name_regex = "^(?:Davis Royce, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van der Berg Ma|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|abdul Smith Berg Ma|abdul Smith Ma|john smith, ma)$" +# 2026-10-01, #575: eight names join, the C1 examples of that change. +# Against these baselines none of them is #575's doing: each moves +# on this rule's own count and lean. 'John van Buren, Ed', 'Ortega y +# Gasset, Ed' and 'abdul Salam, Ed' are two name words before the +# comma and flip; 'De La Cruz, Ed', 'de la Cruz, Ma' and 'Van Buren, +# Ed' are one and gain only the comma's report; 'van der Berg, MA' +# is one, and its capitals make 'MA' the credential, as 'Smith, MA'. +# These baselines already read the particle surnames as one name; +# what #575 fixed is this cycle's count, which had not. +# 'Freiherr von Berg, Ed', rules.md#C1's Accepted example for #575, +# reads as these baselines read it and gains only the comma's report. +name_regex = "^(?:Davis Royce, Ed|De La Cruz, Ed|Doe, Dr\\. MA|Doe, MA|Doe, MA PhD|Doe, Mr\\. MA PhD|Freiherr von Berg MA|Freiherr von Berg, Ed|JOHN SMITH, MA|Jack MA|Jack MA\\.|Jack Wei Ma|John Prof\\. MA|John Smith Ma|John Smith, Ed|John Smith, MA|John Smith, Ma|John de Ma|John van Buren, Ed|John van der Berg Ma|Ortega y Gasset, Ed|Smith Jr\\., MA|Smith Jr\\., Ma|Smith, MA|Van Buren, Ed|abdul Salam, Ed|abdul Smith Berg Ma|abdul Smith Ma|de la Cruz, Ma|john smith, ma|van der Berg, MA)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -374,8 +385,12 @@ issue = "fix(#563) paired initials after a comma read as the given name unless a # G.J.', a title in front ("a suffix word in front of them that is # not also title vocabulary"), which moves nothing at any baseline. # The second period is optional ('García Márquez, G.J'). Two pairs -# speaking only for each other make the run and report it ('De La -# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# speaking only for each other make the run and report it behind two +# name words ('John Smith, X.Y. P.Q.'). 2026-10-01, #575: 'De La Cruz, +# M.J. K.L.' is no longer that case -- 'De La Cruz' is one name word, +# so the count no longer reaches it and it reads as 'Cruz, M.J. K.L.' +# does, given 'M.J.', suffix 'K.L.', reporting both; a listed member +# in front speaks for nothing # ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) # rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names # whose roles match the baseline's move only the report, which a @@ -1433,3 +1448,17 @@ issue = "fix(#544) a credential in front anchors a member after a one-word famil name_regex = "^Smith, PhD Ma$" fields = ["given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#575) a particle surname before a comma is one name word" +# rules.md#C1: "a particle run and the one name word it attaches to +# are one word", so 'van der Berg, PhD' reads as 'Berg, PhD' +# does: family 'van der Berg', suffix 'PhD'. Every release counted +# tokens and split the chain, given 'van', family 'der Berg', with a +# particle-or-given report. +# +# Literal; the probe 'John van der Berg, PhD' (two name words) is +# _MUST_NOT_MATCH. +name_regex = "^van der Berg, PhD$" +fields = ["given", "family", "_ambiguities"] +orders = ["DEFAULT"]