From 7fbb7934b110f4100125dfbac2ae9687b710cb2b Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 12:05:16 -0700 Subject: [PATCH 1/5] docs(C1): a capitals-settled run opened by a class word reports rules.md#C1 said a run whose every class word is written in capitals "reads whole in silence". That holds only behind another credential: a class word opening the part is the first word after the comma, whose decision C1 reports either way, so 'Doe, MA PhD' and 'John Smith, MA MA' report suffix-or-name on that word. The behavior is kept (Derek, 2026-10-01) and the statement corrected, with 'Doe, MA PhD', already a contract-tier corpus name, as its example. No parse moves. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 1 + docs/design/rules.md | 12 ++++++++---- tools/differential/corpus_rules.jsonl | 1 + 3 files changed, 10 insertions(+), 4 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index d0c84c37..f4fad24c 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -876,6 +876,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. +- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index f0b1de96..d1977b81 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1719,10 +1719,13 @@ C1. Rationale: a credential run after the comma means the name is in count leaves in the listing form reports as S2 reads the words in it: a word of this class read as the credential because a credential in front speaks for it reports, and one its own - capitals made the credential does not, so a run whose every such - word is written in capitals reads whole in silence - ('John Smith, PhD MA', 'Smith, PhD MA'), as does a part read as - titles before a lone given name ('Smith, Ms MD Ma'). It is one + capitals made the credential does not unless it opens the part, + where it is the first word after the comma and reports as that + word always does ('Doe, MA PhD', 'John Smith, MA MA'). So a run + whose every such word is written in capitals reads whole in + silence only behind another credential ('John Smith, PhD MA', + 'Smith, PhD MA'), as does a part read as titles before a lone + given name ('Smith, Ms MD Ma'). It is one of TWO places the comma's own decision is reported, the other being the word trailing the given part after it (S2), which is a second decision about a second word and never the same fork twice; an attachment decided after a family comma @@ -1775,6 +1778,7 @@ C1. Rationale: a credential run after the comma means the name is in "Smith, John V" → suffix="V" · boundary "Smith, Ph. D. Jr." → suffix="Ph. D. Jr." "Smith, MD PhD" → suffix="MD PhD" + "Doe, MA PhD" → ambiguities=("suffix-or-name",) "Smith, Dr." → title="Dr." "Smith, Dr. Jr." → suffix="Jr." "John Smith, Mr." → given="John" diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 895bb3e1..5701e1b6 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -62,6 +62,7 @@ "Doe, John X.Y.Z." "Doe, John van DO" "Doe, John van Ma" +"Doe, MA PhD" "Dr Jr" "Dr King Jr" "Dr." From 079ca28131454bb5b61a206f09fb04bca402e5a7 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 12:11:54 -0700 Subject: [PATCH 2/5] docs(C1): review round -- the release note's no-report claim, and the recipe The #544 release bullet still said a capitals-settled run reports nothing; it now says so only when another credential opens it, and an opening acronym reports on itself. The decisions bullet's grid is drawn with repetition, and 14 of its reporting runs are #562 particle pairs whose flip reports over the whole part. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/release_log.rst | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index f4fad24c..90d91be1 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -876,7 +876,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. -- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. +- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. ### T1 — separators, not joiners diff --git a/docs/release_log.rst b/docs/release_log.rst index f3df48df..f5b20a0d 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -14,7 +14,7 @@ Release Log - **Fix a bare trailing Meng or Lac being read as a credential and losing the family name: meng and lac are now acronyms that are also ordinary names.** ``HumanName("wang meng")`` gives first ``wang``, last ``meng``, and ``parse()`` reports a suffix-or-name ambiguity, where every release from 2.0.0 through 2.3.0 gave suffix ``meng`` and no last name; 1.4.0 read last ``meng``, so this is 1.4.0's answer plus the flag. ``li meng`` and ``tran lac`` move the same way, ``Wang, Meng`` gives first ``Meng``, last ``Wang`` again, and ``Parser(policy=Policy(name_order=FAMILY_FIRST)).parse("Wang Meng")`` gives given ``Meng`` where 2.0.0 through 2.3.0 gave family ``Wang``, suffix ``Meng`` and no given name. With a full name in front the credential reading stays: ``john smith meng`` and ``nguyen van lac`` keep suffix ``meng`` and ``lac``, now flagged. But a Title-case ``Nguyen Van Lac`` gives last ``Van Lac`` where every release gave last ``Van``, suffix ``Lac``. The cost is the marking's own, and it falls on the conventional spellings: a ``MEng`` or ``LAc`` written that way, in a name written in more than one case, is read the way ``John Smith Ma`` is (above), so ``John Smith MEng`` gives middle ``Smith``, last ``MEng``, where every release gave suffix ``MEng``. Where one of them LEADS a credential run nothing speaks for it: ``John Smith MEng PhD`` gives middle ``Smith``, last ``MEng``, suffix ``PhD``, where every release read suffix ``MEng PhD`` (``MEng, PhD`` through 2.2); behind a credential, or written with its periods, it is read with the run (the next entry). After a comma ``Smith, MEng`` and ``Smith, meng`` give first ``MEng`` and ``meng`` (1.4.0's reading, not the suffix 2.0 through 2.3 gave) and ``Smith, John MEng`` gives middle ``MEng`` (every release gave suffix ``MEng``), and a bracketed ``John Smith (MEng)`` falls through to nickname, where every release gave suffix ``MEng``. A lone credential after a comma behind a full name (``John Smith, MEng``) keeps the credential reading, and so does a run of them (the next entry). Meng is a common Chinese surname and given name, Lac a Vietnamese given name (``Nguyen Van Lac``) and a French surname; see the ``suffix-acronym-collisions`` entry of ``docs/design/decisions.md`` (closes #540) - - **Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.** ``HumanName("John Smith, Ed Ma")`` gives first ``John``, last ``Smith``, suffix ``Ed Ma``, where 2.0 through 2.3 gave first ``Ed``, middle ``Ma``, last ``John Smith`` -- 1.4.0's reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (``John Smith, MA``), and ``parse()`` reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report, the capitals having decided it (``John Smith, PhD MA``, ``John Smith, MS MA``). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so ``john smith, v phd ma`` gives first ``john``, last ``smith``, suffix ``v phd ma``, where 2.3 gave first ``v``, middle ``ma``, last ``john smith``. ``john smith, md ma`` and ``John Smith, Ms Ma`` move the same way, where 2.0 through 2.3 gave title ``md``/``Ms`` -- the second is the accepted cost, ``Ms`` read as the suffix word it also is, as ``John Smith, Ms`` alone already reads it -- while one name word before the comma keeps the listing form (``Smith, Ms Ma`` gives title ``Ms``, first ``Ma``). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential's company whatever its case: ``John Smith PhD MEng`` and ``Doe, Jane PhD MEng`` give suffix ``PhD MEng``, the fields every release gave (``PhD, MEng`` through 2.2), now reported, and ``doe, jane v phd do`` gives suffix ``v phd do`` where 2.3.0 gave last ``do doe`` -- a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: ``Wang Ma PhD`` keeps last ``Ma``. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: ``Smith, PhD Ma`` gives last ``Smith``, suffix ``PhD Ma``, where 2.3 gave first ``PhD``, middle ``Ma``. A title that is also a credential (``MD``, ``Ms``) opening that part stays a title and nothing in the part speaks, so ``Smith, MD PhD Ma`` keeps title ``MD``, first ``PhD``, middle ``Ma`` and ``Smith, Ms MD Ma`` title ``Ms MD``, first ``Ma``, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so ``Wang M.Eng.`` gives suffix ``M.Eng.``, as ``Wang M.A.`` does and as 2.0 through 2.3 did. See the #544 entry under ``S2`` in ``docs/design/decisions.md`` (closes #544) + - **Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.** ``HumanName("John Smith, Ed Ma")`` gives first ``John``, last ``Smith``, suffix ``Ed Ma``, where 2.0 through 2.3 gave first ``Ed``, middle ``Ma``, last ``John Smith`` -- 1.4.0's reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (``John Smith, MA``), and ``parse()`` reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report when another credential opens it, the capitals having decided it (``John Smith, PhD MA``, ``John Smith, MS MA``), while one opened by such an acronym reports on that first word, as ``John Smith, MA`` alone does (``John Smith, MA PhD``). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so ``john smith, v phd ma`` gives first ``john``, last ``smith``, suffix ``v phd ma``, where 2.3 gave first ``v``, middle ``ma``, last ``john smith``. ``john smith, md ma`` and ``John Smith, Ms Ma`` move the same way, where 2.0 through 2.3 gave title ``md``/``Ms`` -- the second is the accepted cost, ``Ms`` read as the suffix word it also is, as ``John Smith, Ms`` alone already reads it -- while one name word before the comma keeps the listing form (``Smith, Ms Ma`` gives title ``Ms``, first ``Ma``). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential's company whatever its case: ``John Smith PhD MEng`` and ``Doe, Jane PhD MEng`` give suffix ``PhD MEng``, the fields every release gave (``PhD, MEng`` through 2.2), now reported, and ``doe, jane v phd do`` gives suffix ``v phd do`` where 2.3.0 gave last ``do doe`` -- a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: ``Wang Ma PhD`` keeps last ``Ma``. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: ``Smith, PhD Ma`` gives last ``Smith``, suffix ``PhD Ma``, where 2.3 gave first ``PhD``, middle ``Ma``. A title that is also a credential (``MD``, ``Ms``) opening that part stays a title and nothing in the part speaks, so ``Smith, MD PhD Ma`` keeps title ``MD``, first ``PhD``, middle ``Ma`` and ``Smith, Ms MD Ma`` title ``Ms MD``, first ``Ma``, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so ``Wang M.Eng.`` gives suffix ``M.Eng.``, as ``Wang M.A.`` does and as 2.0 through 2.3 did. See the #544 entry under ``S2`` in ``docs/design/decisions.md`` (closes #544) - **Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run.** ``HumanName("John Smith, PhD DO DO")`` gives first ``John``, last ``Smith``, suffix ``PhD DO DO``, where 2.2 and 2.3 gave first ``PhD``, last ``DO DO John Smith`` and 2.0 and 2.1 gave title ``PhD``, first ``DO DO`` -- 1.4.0's reading, restored. ``DO``, ``MC`` and ``VD`` are credentials and surname particles at once, and two particles next to each other join into one particle run whatever their capitals, so the capitals no longer settle such a run silently: it is read by the count of name words before the comma, and ``parse()`` reports the call. ``John Smith, PhD vd DO`` gives suffix ``PhD vd DO`` the same way, and ``John Smith, MD DO DO`` gives suffix ``MD DO DO`` where 2.3 gave title ``MD``, first ``DO``, middle ``DO``. A single ``DO`` is still left to its capitals (``John Smith, PhD DO`` gives suffix ``PhD DO`` with no report), and one name word before the comma keeps the listing form (``Smith, PhD DO DO`` gives first ``PhD``, last ``DO DO Smith``, as 2.3 did). See the #562 entry under ``C1`` in ``docs/design/decisions.md`` (closes #562) From bb979e7b0d30a81f5b282af497eafa38971fb27f Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 12:28:56 -0700 Subject: [PATCH 3/5] docs(S2): point the first-slot precedence at C1's capitals exception S2 said the count decides the first slot after a family comma before case. That holds for one word ('John Smith, MA' flips by the count) but not for a run C1's shortcut settles on its capitals, which keeps the family comma and reports only an opening class word ('John Smith, MA MA' reports on the first 'MA' alone). Prose only, so the next reader does not rediscover it (Derek, 2026-10-01). Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/design/rules.md | 10 +++++++++- 2 files changed, 10 insertions(+), 2 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 90d91be1..37b3f3c0 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -876,7 +876,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. -- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. +- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index d1977b81..75a538ed 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1026,7 +1026,15 @@ S2. Rationale: generational suffixes and credentials are recognized to write in. After a family comma this evidence is SECOND at the FIRST slot after it: the count of name words before the comma decides there first (C1), and the case is read only where that - count leaves the word a name. At the trailing slot of the given + count leaves the word a name. C1 states the one exception: a part + of two or more words whose every word of this class is listed and + written in capitals in a mixed-case name, no two particles side + by side, is read on its capitals before the count is asked, so + the comma keeps its family reading, + and only a word of the class opening the part reports, as the + first word after the comma ('John Smith, MA MA' reports on the + first 'MA' alone, where a counted flip reports over the whole + part). At the trailing slot of the given part the comma has already settled the count, so the writing is the only evidence there is. Company is evidence that outranks both. A member of the ambiguous From e4d7b9ca7c5b55369783d2806b1f55c36690fcc7 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 12:39:06 -0700 Subject: [PATCH 4/5] docs(S2): a part beyond the second comma is not a reporting slot S2's list of slots where the ambiguous credential class reports ended with 'and the segments beyond it', and SUFFIX_OR_NAME's docstring said the same. No suffix-or-name report lands past the second comma: 'Smith, John, MA' and 'John Smith, Jr., MA' are silent, and such a part reports only C2's structural flag. Both now say so and point at C2; S2 gains C2 in interacts:. The silence is kept (Derek, 2026-10-01). Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/design/rules.md | 8 +++++--- nameparser/_types.py | 7 +++++-- 3 files changed, 11 insertions(+), 6 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 37b3f3c0..cdf4e75a 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -876,7 +876,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. -- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. +- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent, as at #530's merge cc78c960, where such a tail raised only C2's `comma-structure` flag), the one emitter that did reach a tail having been silenced by the 2026-09-28 #544 bullet above, so S2 now points at C2 instead. Derek, 2026-10-01: the silence is right. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index 75a538ed..041add8a 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -970,8 +970,10 @@ S2. Rationale: generational suffixes and credentials are recognized name — and at the slots that report, either reading carries the ambiguity flag. Those slots are the trailing slot of a name, the first slot after a family comma, the trailing slot of the GIVEN - part after that comma, the trailing slot of a maiden marker's - clause (M2), and the segments beyond it. A word this document + part after that comma, and the trailing slot of a maiden + marker's clause (M2). A part beyond the second comma is not one + of them: it is consumed wholly as suffixes and reports only its + own structural flag (C2). A word this document says is READ at one of those slots is not always a word one of them reports: where a reading moves a word out of the slot that asked about it, what reports is the slot it lands in, and that @@ -1187,7 +1189,7 @@ S2. Rationale: generational suffixes and credentials are recognized and unchanged (decisions.md#v1-xfail-triage: `king` stays a title, for the addressing forms). "Dr Jr" → suffix="Jr" - history: decisions.md#S2 · interacts: H1, H2, H3, H5, C1, S3, P2, P3, P5, P6, M2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_vocab.py + history: decisions.md#S2 · interacts: H1, H2, H3, H5, C1, C2, S3, P2, P3, P5, P6, M2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_vocab.py S3. Rationale: credentials are often written run together with periods; the chunks between the periods are what carry the diff --git a/nameparser/_types.py b/nameparser/_types.py index 66c92c5e..44368966 100644 --- a/nameparser/_types.py +++ b/nameparser/_types.py @@ -450,8 +450,11 @@ class AmbiguityKind(StrEnum): #: and this is the boundary rather than an omission to be read #: past. The emitters cover the trailing slot of a name, the FIRST #: PIECE after a family comma -- that piece and no further -- the - #: trailing slot of that listing's GIVEN part, the trailing slot - #: of a maiden marker's clause, and the extra segments beyond it. + #: trailing slot of that listing's GIVEN part, and the trailing + #: slot of a maiden marker's clause. A part beyond the second + #: comma is not among them: it reads wholly as suffixes and + #: reports only ``COMMA_STRUCTURE`` (rules.md#C2), so "Smith, + #: John, MA" is silent. #: Since 2.4 the given part's trailing slot reports whichever way #: it read the word: "Doe, John MA" reads suffix ``MA`` and says #: so, "Doe, John Ma" keeps middle ``Ma`` and says so too. Not in From b34a34e2254cbb3aa71ea768b770ecb13e1d900e Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 12:53:44 -0700 Subject: [PATCH 5/5] docs(S2): a tail reports no suffix-or-name, not only comma-structure Second review round: 'reports only COMMA_STRUCTURE' was false -- an unclosed delimiter in a tail adds unbalanced-delimiter ('John Smith, Jr., (Bob') and a maiden clause standing in one is read as one. S2 and the SUFFIX_OR_NAME docstring now say only that the tail is no reporting slot and leave the rest to C2. The decisions bullet's history is corrected too: at #530's merge a tail could still report suffix-or-name ('John Smith, Jr., PhD Do Ma'), silenced by the 2026-09-28 #544 bullet. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/design/rules.md | 4 ++-- nameparser/_types.py | 5 ++--- 3 files changed, 5 insertions(+), 6 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index cdf4e75a..25f2982d 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -876,7 +876,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-27 (Derek), #544 — THE NAME-WORD COUNT READS A RUN. The single-token rule this section records for the ambiguous class generalizes to a post-comma part of two or more words, every one suffix vocabulary or a class member (listed or by shape), at least one a member and none a single-letter roman numeral: behind two or more name words the part is the credential run, and the flip reports once over the whole part (`John Smith, Ed Ma`, `Jane Doe, MS LAc`). A title/suffix dual opening the part counts as suffix vocabulary there, and a run whose every member is listed and leans credential is left to the family-comma path, which already reads it whole. The forks and their measurements are the #544 entry under S2. - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. -- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent, as at #530's merge cc78c960, where such a tail raised only C2's `comma-structure` flag), the one emitter that did reach a tail having been silenced by the 2026-09-28 #544 bullet above, so S2 now points at C2 instead. Derek, 2026-10-01: the silence is right. +- 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent). It once did: at #530's merge cc78c960 `John Smith, Jr., PhD Do Ma` reported `suffix-or-name` on 'Ma' beside its `comma-structure` flag, from group's chain emitter, which the 2026-09-28 #544 bullet above silenced in a tail. S2 now points at C2 for what such a part does report, which is not only `comma-structure` — an unclosed delimiter there adds `unbalanced-delimiter` (`John Smith, Jr., (Bob`), and a maiden clause standing in it is read as one (`Smith, John, Jr nee Jones MA`, maiden 'Jones MA'). Derek, 2026-10-01: the silence is right. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index 041add8a..683fda0e 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -972,8 +972,8 @@ S2. Rationale: generational suffixes and credentials are recognized first slot after a family comma, the trailing slot of the GIVEN part after that comma, and the trailing slot of a maiden marker's clause (M2). A part beyond the second comma is not one - of them: it is consumed wholly as suffixes and reports only its - own structural flag (C2). A word this document + of them; what such a part does report is C2's to say. A word + this document says is READ at one of those slots is not always a word one of them reports: where a reading moves a word out of the slot that asked about it, what reports is the slot it lands in, and that diff --git a/nameparser/_types.py b/nameparser/_types.py index 44368966..93045aba 100644 --- a/nameparser/_types.py +++ b/nameparser/_types.py @@ -452,9 +452,8 @@ class AmbiguityKind(StrEnum): #: PIECE after a family comma -- that piece and no further -- the #: trailing slot of that listing's GIVEN part, and the trailing #: slot of a maiden marker's clause. A part beyond the second - #: comma is not among them: it reads wholly as suffixes and - #: reports only ``COMMA_STRUCTURE`` (rules.md#C2), so "Smith, - #: John, MA" is silent. + #: comma is not among them, so "Smith, John, MA" is silent; what + #: such a part does report is rules.md#C2's to say. #: Since 2.4 the given part's trailing slot reports whichever way #: it read the word: "Doe, John MA" reads suffix ``MA`` and says #: so, "Doe, John Ma" keeps middle ``Ma`` and says so too. Not in