From fc25880c7ec04a34cdd978d131080a446b6329e2 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 21:42:20 -0700 Subject: [PATCH 01/13] feat(S2): read an all-caps credential after a comma by default (#564) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Policy.unlisted_caps_suffixes becomes a CapsSuffixes StrEnum -- OFF / AFTER_COMMA / EVERYWHERE -- defaulting to AFTER_COMMA. The field was new in 2.4 and unreleased; a bool now raises a TypeError naming both replacements. The all-caps SURNAME convention ('Jean DUPONT', 'DUPONT, Jean') writes the capitals at the end of a name or before a comma, never after a comma behind a full name, so the default reads that one position: 'John Smith, XYZ' gives suffix 'XYZ', as do the corpus names 'Ahmad Jayadi, CHA', 'John Smith, RAI' and 'The Rt Hon Kenneth Clarke QC MP, HMG', all read as the given name at 2.3.0. 'John Smith, LEED AP' now reads suffix too, closing the #291 deviation marker C1 carried. A lone two-letter word is never admitted at the comma (the undotted form of #563's paired initials): 'García Márquez, MJ' keeps given 'MJ'. EVERYWHERE adds the trailing slots, the old True; OFF is 2.3's reading and the way to keep a capitalized given name behind a two-word surname ('García Márquez, GABRIEL', the accepted cost). The comma test is on by default, so it checks isupper() and the two-letter length in C before calling the shared predicate: ordinary comma names pay no frame. rules.md S2/C1/C2, decisions.md#S2, mechanisms.md, customize/usage/modules docs, the release log, and the ledgers at all five baselines follow; the gate exits 0 with radar counts unchanged. Co-Authored-By: Claude Opus 5.5 --- docs/customize.rst | 32 +++--- docs/design/decisions.md | 4 + docs/design/mechanisms.md | 2 +- docs/design/rules.md | 46 +++++--- docs/modules.rst | 3 + docs/release_log.rst | 2 +- docs/usage.rst | 8 +- nameparser/__init__.py | 2 + nameparser/_pipeline/_classify.py | 17 ++- nameparser/_pipeline/_segment.py | 16 ++- nameparser/_pipeline/_vocab.py | 15 +-- nameparser/_policy.py | 96 +++++++++++----- nameparser/_types.py | 4 +- tests/v2/cases.py | 112 ++++++++++++------- tests/v2/pipeline/test_classify.py | 49 ++++---- tests/v2/pipeline/test_vocab.py | 5 +- tests/v2/rules_doc.py | 10 +- tests/v2/test_facade_cases.py | 3 + tests/v2/test_ledger_guards.py | 69 ++++++++++-- tests/v2/test_policy.py | 35 ++++-- tests/v2/test_properties.py | 7 +- tests/v2/test_render.py | 3 +- tools/differential/compare.py | 19 +++- tools/differential/corpus_rules.jsonl | 3 + tools/differential/expected_since_1.4.0.toml | 26 ++++- tools/differential/expected_since_2.0.0.toml | 28 ++++- tools/differential/expected_since_2.1.0.toml | 28 ++++- tools/differential/expected_since_2.2.0.toml | 28 ++++- tools/differential/expected_since_2.3.0.toml | 20 ++++ 29 files changed, 508 insertions(+), 184 deletions(-) diff --git a/docs/customize.rst b/docs/customize.rst index 3a2d5aa3..78d6d9d5 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -529,20 +529,26 @@ listed below. ``"John Smith 1.4.2"`` keeps family ``1.4.2``; that retirement is not behind this switch). * - ``unlisted_caps_suffixes`` - - ``bool`` - - Reads an unlisted all-caps word of two or more letters, with no - period in it, in a name written in more than one case as a - credential where the position allows it: ``"John Smith XYZ"`` - gives suffix ``XYZ``, and since 2.4 so do the family-comma - form ``"Doe, John XYZ"`` and the word ending a maiden marker's + - ``CapsSuffixes`` + - Where an unlisted all-caps word of two or more letters, with no + period in it, in a name written in more than one case, reads as + a credential. ``CapsSuffixes.AFTER_COMMA``, the default, reads + it only in the part right after a comma with two or more name + words before it: ``"John Smith, XYZ"`` gives suffix ``XYZ``, + while ``"Smith, XYZ"`` keeps given ``XYZ`` and a lone two-letter + word, how initials are written, stays the given name + (``"García Márquez, MJ"``). The all-caps SURNAME convention + (``"Jean DUPONT"``, ``"DUPONT, Jean"``) never writes the + capitals there. ``CapsSuffixes.EVERYWHERE`` also reads the end + of a name, the given part's last word after a family comma + (``"Doe, John XYZ"``) and the word ending a maiden marker's clause (``"Jane Doe nee Smith XYZ"`` gives maiden ``Smith`` - with suffix ``XYZ``, where off it keeps maiden - ``Smith XYZ``). Defaults to ``False``, and - deliberately: - an all-caps surname is a real writing convention that shape - cannot separate from a credential, so ``"Jean Pierre DUPONT"`` - gives family ``Pierre``, suffix ``DUPONT`` with this on. Off, - nothing changes and nothing is reported. + with suffix ``XYZ``), where that convention does write them: + ``"Jean Pierre DUPONT"`` then gives family ``Pierre``, suffix + ``DUPONT``. ``CapsSuffixes.OFF`` reads none of them and reports + nothing, 2.3's reading -- and the way to keep a given name + written in capitals after a two-word surname, which the default + reads as a credential (``"García Márquez, GABRIEL"``). * - ``strip_emoji`` - ``bool`` - Excludes emoji from tokenization — they appear in no field or diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 037d5910..8828e8ae 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -686,6 +686,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SECOND REVIEW ROUND, same day, on the first round's fix commit. (a) TWO PAIRS (Derek): `De La Cruz, M.J. K.L.` read family 'De La Cruz', suffix 'M.J. K.L.' — no given name — in silence, the silent-flip rationale ("only paired initials are taken for a name") failing on exactly the case where each word speaking for a pair is itself a pair. The reading stands, two dotted groups not being how anyone writes a person's initials, but that flip now REPORTS: when the only words speaking for a pair are other pairs, every shape word in the part being one. `John Smith, X.Y. P.Q.` gains the report with it; `John Smith, X.Y.Z. G.J.` and `John Smith, PhD G.J. K.L.`, where something else speaks, stay silent. The same rule reports `García Márquez, Ms G.J. K.L.`, whose honorific the run takes, where the first round's fix took it in silence. (b) A CLASS MEMBER IN FRONT SPEAKS FOR NOTHING (Derek): `García Márquez, Ed G.J.`, `Ma G.J.` and `MA G.J.` flipped on the member in front, where S2's company lets only an unambiguous credential speak for a member. Now the comma keeps the family, given 'Ed' — and 'G.J.' then ends the given part, the slot (iii) left alone, so it reads suffix 'G.J.' as `Doe, John R.T.` does, both forks reported; 1.4.0 and 2.3.0 gave middle 'G.J.'. Resolving that slot is the no-comma question #563 leaves open. (c) WORDING, no behavior: C1's sentence counting a dual opening a part as a suffix word now defers to the pair's titles; the pair's report reaches it only where it opens the part, `Smith, Ms G.J.` having always been silent (the comma report's reach is the first post-comma piece, 2026-09-18); and the silence is stated as "no listed word takes part", which covers `García Márquez, PhD G.J.`, where "dotted-shape words alone" did not. SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. +- 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Implemented as "a part that is one word of fewer than three letters", so a run holding a longer word (`LEED AP`) is admitted as one with a longer member, mirroring #563, where one pair is a given name and two are a credential. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 256. The comma test is on by default now, so it asks a C-level `isupper()` of the part's first word and the lone-two-letter length test before calling `caps_shape_candidate`; a comma name with no all-caps word pays nothing. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index e818d42d..550baed6 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -55,7 +55,7 @@ Problem shape. "Which stage does X?" — asked before attributing behavior in pr ## ONE-PREDICATE-PER-QUESTION — one predicate answers it, and every other site calls that -Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from three sites that each needed the identical question answered — classify's own tag emission, this module's ambiguous_class_candidate, and _segment.py's multi-token run test — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold, because every one of the three callers' own FIRST conjunct is the caller-configured switch itself, `Policy.unlisted_caps_suffixes`, False by default, so the shared call is never reached at the default regardless of how many callers share it (decisions.md#S2)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. +Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from three sites that each needed the identical question answered — classify's own tag emission, this module's ambiguous_class_candidate, and _segment.py's multi-token run test — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: the trailing-position caller's first conjunct is the setting itself (`CapsSuffixes.EVERYWHERE`, not the default), and the comma run test — on by default since #564 — asks a C-level `isupper()` of the part's first word before calling, so a parse with no all-caps word after a comma never reaches it (decisions.md#S2, #C1)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. ## RENDER-HONORS-THE-PARSE — the parse decides it, the views honor it diff --git a/docs/design/rules.md b/docs/design/rules.md index c7475e8a..306b25e6 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1075,15 +1075,18 @@ S2. Rationale: generational suffixes and credentials are recognized one such shape, admitted by default (S3); an unlisted all-caps alphabetic word of two or more letters, standing in a suffix position of a mixed-case name and belonging to no wordlist, is - the other, admitted only under the caller switch the example - lines below name. That second shape is OFF by default because - French and Korean records write the SURNAME in capitals, so it - is a surname as often as it is a credential and only the caller - knows which corpus this is; the cost of turning it on is that a - three-word name gives up its family name to the acronym, while a - two-word name keeps it — there are no words to spare there, so - the class is considered and declined and only the fork is - reported. The dotted shape does not take every promise of the + the other. By default that second shape is admitted in one + position only, the part after a comma with two or more name + words before it (C1): French and Korean records write the + SURNAME in capitals, and that convention puts the capitals at + the end of a name or before a comma, never after a comma behind + a full name. Everywhere else it is admitted only under the + caller setting the example lines below name, being a surname + there as often as a credential, which only the caller can know; + the cost of that setting is that a three-word name gives up its + family name to the acronym, while a two-word name keeps it — + there are no words to spare there, so the class is considered + and declined and only the fork is reported. The dotted shape does not take every promise of the class with it: at the first slot after a comma, C1 reads paired initials as the given name unless something speaks for them, and makes a flip to the credential run in which no listed member takes @@ -1129,10 +1132,10 @@ S2. Rationale: generational suffixes and credentials are recognized "Doe, John DO Ed" → middle="DO Ed" · boundary "Doe, John DO Ed" → ambiguities=() · boundary "John Smith XYZ" → family="XYZ" - "John Smith XYZ" unlisted_caps_suffixes-on → suffix="XYZ" - "Jean DUPONT" unlisted_caps_suffixes-on → family="DUPONT" - "Jean Pierre DUPONT" unlisted_caps_suffixes-on → suffix="DUPONT" - "Jean Pierre DUPONT" unlisted_caps_suffixes-on → family="Pierre" + "John Smith XYZ" unlisted_caps_suffixes-everywhere → suffix="XYZ" + "Jean DUPONT" unlisted_caps_suffixes-everywhere → family="DUPONT" + "Jean Pierre DUPONT" unlisted_caps_suffixes-everywhere → suffix="DUPONT" + "Jean Pierre DUPONT" unlisted_caps_suffixes-everywhere → family="Pierre" "Jack Ma." → family="Ma." · boundary "Ph. D. Van Johnson" → family="Van Johnson" "Ph. D. Van Johnson" → title="Ph." @@ -1719,7 +1722,11 @@ C1. Rationale: a credential run after the comma means the name is in there. Behind two or more name words, paired initials that are not the only shape word in the part are a credential however little else speaks for them, since no one writes a person's - initials as two dotted groups. The same count reads a part of two or + initials as two dotted groups. A lone word of two capitals after + the comma is the same shape undotted ('García Márquez, MJ'): it + reads as the given name, and an unlisted all-caps word joins the + class there only at three letters or more, or in a run holding + such a word. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any @@ -1850,6 +1857,12 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, X.Y.Z." → suffix="X.Y.Z." "John Smith, X.Y.Z." → ambiguities=() "John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z." + "John Smith, XYZ" → suffix="XYZ" + "John Smith, XYZ" → ambiguities=("suffix-or-name",) + "John Smith, XYZ" unlisted_caps_suffixes-off → given="XYZ" + "John Smith, LEED AP" → suffix="LEED AP" + "Smith, XYZ" → given="XYZ" · boundary + "García Márquez, MJ" → given="MJ" · boundary "Smith, A.B." → given="A.B." · boundary "García Márquez, G.J." → given="G.J." "García Márquez, G.J." → family="García Márquez" @@ -1911,7 +1924,6 @@ C1. Rationale: a credential run after the comma means the name is in not structure — v1 applied the delimiter to the suffix-comma form alone, and that limitation is kept as parity: "Smith, RN - CRNA" reads given "RN" under the policy as without it. - "John Smith, LEED AP" → family="Smith" deviates: #291 (today: family="John Smith") Accepted: the further-comma qualifier carries no example line of its own. It discriminates PAIRS and spans both branches, so exemplifying it means a with-comma partner for each — every one @@ -1945,7 +1957,7 @@ C2. Rationale: text beyond the recognized comma parts should be whether the name is written in one case so that nothing leans at all, or the member is written the way a name is written. And S2's other by-shape half, the unlisted all-caps word, does not - reach here under its switch either: the shape a tail segment is + reach here in any setting: the shape a tail segment is recognized by is the dotted one alone. A part the parse consumes wholly as suffixes raises no report about reading a word of it as a name, whether it is the part @@ -1963,7 +1975,7 @@ C2. Rationale: text beyond the recognized comma parts should be "John Smith, MD, Ma" → ambiguities=("comma-structure",) · boundary "Steven Hardman, MD, DO, DDS" → ambiguities=() "STEVEN HARDMAN, MD, DO, DDS" → ambiguities=("comma-structure",) · boundary - "John Smith, MD, XYZ" unlisted_caps_suffixes-on → ambiguities=("comma-structure",) + "John Smith, MD, XYZ" unlisted_caps_suffixes-everywhere → ambiguities=("comma-structure",) Accepted: the no-name-reading clause carries no example line of its own. What it moves is a report with no field beside it, and an example line would enter the rules corpus for that report diff --git a/docs/modules.rst b/docs/modules.rst index 55febcab..c9953b13 100644 --- a/docs/modules.rst +++ b/docs/modules.rst @@ -100,6 +100,9 @@ Configuration .. autoclass:: nameparser.PatronymicRule :members: +.. autoclass:: nameparser.CapsSuffixes + :members: + .. autoclass:: nameparser.Script :members: diff --git a/docs/release_log.rst b/docs/release_log.rst index 16b12f5c..9f1995a3 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word of three or more letters, or a run holding one, in the part right after a comma behind two or more name words: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` gives suffix ``LEED AP`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), and so does a lone two-letter word (``García Márquez, MJ``), the way initials are written. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/docs/usage.rst b/docs/usage.rst index 42bd99bd..194b14d2 100644 --- a/docs/usage.rst +++ b/docs/usage.rst @@ -853,9 +853,11 @@ Two ``Policy`` switches extend the same class to words the vocabulary does not hold. ``unlisted_dotted_suffixes`` (on by default) reads a token of two or more period-separated chunks the same way — ``parse("John Smith X.Y.Z.").suffix`` is ``'X.Y.Z.'`` — and -``unlisted_caps_suffixes`` (off by default) does the same for an -unlisted all-caps word, which is opt-in because an all-caps surname is -written that way too. See :doc:`customize` for both. +``unlisted_caps_suffixes`` does the same for an unlisted all-caps word, +by default only after a comma behind a full name, where an all-caps +surname is never written — ``parse("John Smith, XYZ").suffix`` is +``'XYZ'`` — and at the end of a name only on request, since an +all-caps surname is written there. See :doc:`customize` for both. A reading the vocabulary settles on its own is not a guess and reports nothing — periods make ``M.A.`` unambiguously a credential: diff --git a/nameparser/__init__.py b/nameparser/__init__.py index 19afbd93..c21f26fc 100644 --- a/nameparser/__init__.py +++ b/nameparser/__init__.py @@ -16,6 +16,7 @@ FAMILY_FIRST_GIVEN_LAST, GIVEN_FIRST, UNSET, + CapsSuffixes, PatronymicRule, Policy, PolicyPatch, @@ -40,6 +41,7 @@ "Span", "Role", "Token", "Ambiguity", "AmbiguityKind", "ParsedName", "STABLE_TAGS", "Segmentation", "Segmenter", "Lexicon", "Policy", "PolicyPatch", "PatronymicRule", "Script", "UNSET", + "CapsSuffixes", "GIVEN_FIRST", "FAMILY_FIRST", "FAMILY_FIRST_GIVEN_LAST", "DEFAULT_NICKNAME_DELIMITERS", "DEFAULT_SCRIPT_ORDERS", "Locale", "Parser", "parse", "parser_for", diff --git a/nameparser/_pipeline/_classify.py b/nameparser/_pipeline/_classify.py index 76ae08a6..51eff84d 100644 --- a/nameparser/_pipeline/_classify.py +++ b/nameparser/_pipeline/_classify.py @@ -51,6 +51,7 @@ from nameparser._lexicon import _normalize +from nameparser._policy import CapsSuffixes from nameparser._pipeline._state import ( AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, ParseState, PendingAmbiguity, WorkToken, copy_with, @@ -216,14 +217,18 @@ def _tags_for(token: WorkToken, n: str, state: ParseState, tags.add(SHAPE_ACRONYM_TAG) if state.policy.unlisted_dotted_suffixes: tags.add(AMBIGUOUS_ACRONYM_TAG) - elif (state.policy.unlisted_caps_suffixes and token.role is None + elif (state.policy.unlisted_caps_suffixes is CapsSuffixes.EVERYWHERE + and token.role is None and caps_shape_candidate(token.text, lex, state.policy, one_case)): - # #516's all-caps half, OPT-IN: an unlisted word written - # in capitals inside a mixed-case name. The policy conjunct - # comes FIRST and stays a plain attribute read -- False by - # default, so `caps_shape_candidate` is never CALLED at the - # default and sharing its body costs the default nothing + # #516's all-caps half, OPT-IN at this slot: an unlisted + # word written in capitals inside a mixed-case name. Only + # EVERYWHERE tags it -- the comma position the default + # reads is decided in `segment` from the text alone (#564), + # and the tag is what the comma-less trailing slot reads. + # The policy conjunct comes FIRST and stays a plain + # attribute read, so `caps_shape_candidate` is never CALLED + # at the default and sharing its body costs it nothing # (that is why this half is a call where the dotted branch # above stays inline: the dotted caller has no such cheap # first conjunct to hide behind). The predicate's own diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index d926d225..5eb42562 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -49,6 +49,7 @@ from nameparser._lexicon import _normalize from nameparser._pipeline._pieces import own_words +from nameparser._policy import CapsSuffixes from nameparser._pipeline._state import ( ParseState, PendingAmbiguity, Structure, comma_bucket, copy_with, ) @@ -249,7 +250,20 @@ def class_run(seg: tuple[int, ...]) -> bool: # left, and the verdict is `case_class() is False` directly rather # than a second walk that could only reach the same answer (a # quality-review finding: the walk was provably redundant). - if (not candidate and state.policy.unlisted_caps_suffixes and groups[1] + # + # #564: on by default (`CapsSuffixes.AFTER_COMMA`), so two cheap + # C-level conjuncts go before the call: the first word must be + # written in capitals at all, which every word of the run must + # be, and a LONE two-letter word is declined -- it is how a + # person's initials are written, and two words before the comma + # may be one surname ('García Márquez, MJ'), the case #563 decides + # for the dotted 'M.J.' (rules.md#C1). A run holding a longer word + # ('LEED AP') is not that shape. + first = state.tokens[groups[1][0]].text if groups[1] else "" + if (not candidate + and state.policy.unlisted_caps_suffixes is not CapsSuffixes.OFF + and first.isupper() + and not (len(groups[1]) == 1 and len(first) < 3) and all(caps_shape_candidate(state.tokens[i].text, state.lexicon, state.policy, one_case=False) diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 531dca55..b200dfa3 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -49,7 +49,7 @@ from nameparser._lexicon import ( FULL_STOPS, Lexicon, _VOCAB_FIELDS, _normalize, ) -from nameparser._policy import (Policy, Script, _JA_SCRIPTS, _NO_INITIALS, +from nameparser._policy import (CapsSuffixes, Policy, Script, _JA_SCRIPTS, _NO_INITIALS, _SCRIPT_RANGES, _script_matcher) from nameparser._pipeline._state import WorkToken, comma_bucket @@ -653,11 +653,11 @@ def run_word_fold( # `_segment.py`'s multi-token run test, and this module's own unit # tests), where it had been spelled three times over (quality-review # finding). The usual objection to sharing -- a call costing every -# default-policy parse a frame it cannot use -- does not apply: every -# caller's own first conjunct is `policy.unlisted_caps_suffixes`, -# False by default, so neither this call nor the loop inside it is -# ever reached at the default (confirmed against the 412/449 frame -# band and the default comma harness). +# default-policy parse a frame it cannot use -- does not apply: the +# trailing position is behind `CapsSuffixes.EVERYWHERE` (classify's +# first conjunct), and since #564 the comma run test, which IS on by +# default, asks a C-level `isupper()` of the part's first word before +# calling, so a comma name with no all-caps word never reaches it. def caps_shape_candidate(text: str, lexicon: Lexicon, policy: Policy, one_case: bool | None) -> bool: """Whether TEXT is an UNLISTED all-caps credential candidate: two @@ -695,7 +695,8 @@ def caps_shape_candidate(text: str, lexicon: Lexicon, policy: Policy, `_segment.py`'s run test relies on exactly that rather than calling both. """ - if not (policy.unlisted_caps_suffixes and one_case is False + if not (policy.unlisted_caps_suffixes is not CapsSuffixes.OFF + and one_case is False and len(text) >= 2 and text.isalpha() and text.isupper()): return False n = _normalize(text) diff --git a/nameparser/_policy.py b/nameparser/_policy.py index 8aabc580..ac72dd04 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -12,7 +12,9 @@ from enum import Enum, StrEnum, auto from typing import Any -from nameparser._types import Role, _guarded_getstate, _guarded_setstate +from nameparser._types import ( + Role, _coerce_enum, _guarded_getstate, _guarded_setstate, +) class PatronymicRule(StrEnum): @@ -30,6 +32,25 @@ class PatronymicRule(StrEnum): TURKIC = "turkic" +class CapsSuffixes(StrEnum): + """Where ``Policy.unlisted_caps_suffixes`` reads an unlisted + all-caps word as a credential (#516, #564).""" + + #: Nowhere: such a word is name material in every position, as + #: 2.3 read it. + OFF = "off" + #: Only after a comma, behind two or more name words ("John Smith, + #: XYZ" gives suffix ``XYZ``). The default: the all-caps SURNAME + #: convention ("Jean DUPONT", "DUPONT, Jean") never writes the + #: capitals there. + AFTER_COMMA = "after-comma" + #: After a comma and ending a name with no comma ("John Smith XYZ" + #: gives suffix ``XYZ``), where the all-caps surname convention + #: does write them -- "Jean Pierre DUPONT" then gives family + #: ``Pierre``, suffix ``DUPONT``. + EVERYWHERE = "everywhere" + + class Script(StrEnum): """Writing systems the parser can key SCRIPT-CONDITIONAL behavior on: per-script name order (``Policy.script_orders``) and @@ -686,35 +707,36 @@ class Policy: #: claimed; the roman-chunk retirement (rules.md#S3) is not #: behind this switch, and still reports the fork. unlisted_dotted_suffixes: bool = True - #: Reads an UNLISTED all-caps word of two or more letters, with no - #: period in it, in a name written in more than one case as a - #: credential where the position allows it: with this on, - #: "John Smith XYZ" gives suffix ``XYZ`` and "Smith, XYZ" still - #: gives given ``XYZ``, the same words-to-spare rule the rest of - #: the class takes. A listed member keeps its own case lean - #: regardless of this switch ("Jack MA" still gives suffix ``MA`` - #: on or off), and the roman-numeral fork still claims a bare - #: numeral first either way ("Jack VI" is unaffected by this - #: switch, on or off). OFF BY DEFAULT, and the asymmetry with - #: ``unlisted_dotted_suffixes`` is deliberate: an all-caps surname - #: is a real writing convention that shape cannot separate from a - #: credential -- "Jean Pierre DUPONT" gives family ``Pierre``, - #: suffix ``DUPONT`` with this on, and a swallowed family name is - #: the worse failure. The two-word "Jean DUPONT" and "Minjun KIM" - #: read as family names at the default and KEEP that family with - #: this on too (one word before the credential is never enough, - #: the same words-to-spare rule above) -- but a genuine candidate - #: this switch does not move still gains the fork's report: it - #: was a real fork the parser considered and declined, and that - #: is reported even where the reading did not change. Off, - #: nothing changes and nothing is reported. A digit anywhere + #: Where an UNLISTED all-caps word of two or more letters, with no + #: period in it, in a name written in more than one case, reads as + #: a credential (:class:`CapsSuffixes`). A listed member keeps its + #: own case lean in every setting ("Jack MA" gives suffix ``MA``), + #: and the roman-numeral fork claims a bare numeral first ("Jack + #: VI" is unaffected). + #: + #: ``AFTER_COMMA``, the default, reads it only in the part after a + #: comma with two or more name words before it: "John Smith, XYZ" + #: gives suffix ``XYZ``, while "Smith, XYZ" keeps given ``XYZ`` (one + #: name word) and "García Márquez, MJ" keeps given ``MJ`` -- a lone + #: two-letter word is how a person's initials are written, and two + #: words before the comma may be one surname, the case #563 decides + #: for "M.J." too. The all-caps SURNAME convention ("Jean DUPONT", + #: "DUPONT, Jean") writes the capitals at the end of a name or + #: before a comma, never after one behind a full name, which is why + #: this position is on by default and the others are not. + #: + #: ``EVERYWHERE`` adds the end of a name with no comma, the + #: convention's own position: "John Smith XYZ" gives suffix ``XYZ``, + #: and "Jean Pierre DUPONT" gives family ``Pierre``, suffix + #: ``DUPONT`` -- the cost, a swallowed family name being the worse + #: failure. The two-word "Jean DUPONT" and "Minjun KIM" keep their + #: family even then (one word before the credential is never + #: enough), but gain the fork's report. ``OFF`` reads no such word + #: as a credential anywhere and reports nothing, as 2.3 did. A digit #: disqualifies the token and a single capital stays an initial. - #: ``isupper()`` is script-agnostic, so this is the same - #: convention and the same reason for being off in ANY script - #: that has a case contrast at all, not just Latin -- an all-caps - #: Cyrillic surname ("Иван ИВАНОВ") or an accented Latin one - #: ("Jean ÉCOLE") joins this class exactly as an ASCII one does. - unlisted_caps_suffixes: bool = False + #: ``isupper()`` is script-agnostic, so the same holds in any script + #: with a case contrast ("Иван ИВАНОВ", "Jean ÉCOLE"). + unlisted_caps_suffixes: CapsSuffixes = CapsSuffixes.AFTER_COMMA # in the class body so @dataclass(slots=True) keeps them __getstate__ = _guarded_getstate @@ -821,6 +843,18 @@ def __post_init__(self) -> None: raise TypeError( f"{flag} must be a bool, got {value!r}" ) + # A bool is the likeliest wrong value: the field was a bool + # flag until #564, and `True` is an `int` the enum would read + # as no member. Name both spellings that replace it. + caps = self.unlisted_caps_suffixes + if isinstance(caps, bool): + raise TypeError( + f"unlisted_caps_suffixes must be a CapsSuffixes, got " + f"{caps!r}; use CapsSuffixes.EVERYWHERE for the old " + f"True and CapsSuffixes.OFF for the old False") + object.__setattr__(self, "unlisted_caps_suffixes", _coerce_enum( + caps, CapsSuffixes, "unlisted_caps_suffixes value", + "values")) def __repr__(self) -> str: # Bounded: only fields that deviate from the default are shown @@ -855,7 +889,7 @@ def patched(self, patch: PolicyPatch) -> Policy: #: Every bool-valued Policy field, read off the dataclass rather than #: listed: `__post_init__`'s bool check sweeps this, so a flag added to #: the class above is validated the day it lands and cannot ship -#: unchecked the way `unlisted_dotted_suffixes` and +#: unchecked the way `unlisted_dotted_suffixes` and the then-bool #: `unlisted_caps_suffixes` did. `from __future__ import annotations` #: makes every annotation a string, so the comparison is against the #: SPELLING "bool" -- which is also what a reader of the class body @@ -912,7 +946,7 @@ class PolicyPatch: # sequences equal, so a field inserted on one side would have to be # inserted on the other and both would re-bind together. unlisted_dotted_suffixes: bool | _Unset = UNSET - unlisted_caps_suffixes: bool | _Unset = UNSET + unlisted_caps_suffixes: CapsSuffixes | _Unset = UNSET # in the class body so @dataclass(slots=True) keeps them __getstate__ = _guarded_getstate diff --git a/nameparser/_types.py b/nameparser/_types.py index 93045aba..59e4fac2 100644 --- a/nameparser/_types.py +++ b/nameparser/_types.py @@ -155,8 +155,8 @@ def __add__(self, other: object) -> NoReturn: # type: ignore[override] #: tag classify writes on a token it admits to the credential reading #: by its SHAPE rather than by listed vocabulary (`Policy. #: unlisted_dotted_suffixes` is the first emitter, `Policy. -#: unlisted_caps_suffixes` the second, classify writing the tag from -#: both branches). Defined here, at the bottom of the graph, because +#: unlisted_caps_suffixes` at `CapsSuffixes.EVERYWHERE` the second, +#: classify writing the tag from both branches). Defined here, at the bottom of the graph, because #: a render view reads it as well as the pipeline: case repair writes #: such a suffix in capitals as it does a listed acronym (#459), and #: _render may not import _pipeline. One constant, not a string diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 5be58a7b..516537c3 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -35,6 +35,7 @@ import unicodedata from dataclasses import dataclass +from nameparser._policy import CapsSuffixes from nameparser import (FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST, GIVEN_FIRST, Policy) # Not in nameparser.__all__: _order_repr renders a name_order for an @@ -522,16 +523,18 @@ def _check_cjk_shape_purity(self) -> None: "'Smith' becomes a middle name. The shape-plus-position " "heuristic that would recover it is a parking-lot " "bullet of decisions.md#suffix-acronym-collisions"), - Case("removed_credential_after_a_comma_reads_as_the_given_name", + Case("removed_credential_after_a_comma_reads_by_its_capitals", "Ahmad Jayadi, CHA", - {"given": "CHA", "family": "Ahmad Jayadi"}, - classification="fix(#342)", - notes="the comma form moves the OTHER way and is why the " - "ledger rule declares four fields rather than two. " - "With 'cha' gone the comma is an ordinary family " - "comma (C1): the pre-comma run is the family and the " - "post-comma word is the given name. 'John Smith, RAI' " - "is the same shape and moves with it"), + {"given": "Ahmad", "family": "Jayadi", "suffix": "CHA"}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="with 'cha' gone from the vocabulary (#342) the word is " + "unlisted, and since #564 an unlisted all-caps word " + "after a comma behind two name words is a credential by " + "default (rules.md#C1, S2), so the comma is a suffix " + "comma again and reports the call. 2.3.0 read given " + "'CHA', family 'Ahmad Jayadi'; 'John Smith, RAI' moves " + "with it"), Case("removed_credential_loses_the_dotted_spelling_too", "John Smith C.H.A.", {"given": "John", "family": "Smith", "suffix": "C.H.A."}, @@ -2799,7 +2802,7 @@ def _check_cjk_shape_purity(self) -> None: Case("the_caps_shape_never_reaches_a_tail_segment", "John Smith, MD, XYZ", {"given": "John", "family": "Smith", "suffix": "MD, XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("comma-structure",), classification="fix(#516)", notes="the SECOND way C2's quiet is narrow, and the half its " @@ -2832,7 +2835,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_surname_is_swallowed_with_the_switch_on", "Jean Pierre DUPONT", {"given": "Jean", "family": "Pierre", "suffix": "DUPONT"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("suffix-or-name",), notes="the cost of the switch, pinned so nobody turns it on " @@ -2849,7 +2852,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_surname_reports_but_does_not_move_at_two_words", "Jean DUPONT", {"given": "Jean", "family": "DUPONT"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name",), notes="one name word is never enough to spend the credential " "reading, so the family stays 'DUPONT' -- but the fork " @@ -2874,7 +2877,7 @@ def _check_cjk_shape_purity(self) -> None: Case("unlisted_caps_reads_by_position_with_the_switch_on", "John Smith XYZ", {"given": "John", "family": "Smith", "suffix": "XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("suffix-or-name",), notes="what the switch buys: the shape reads, and the " @@ -2887,17 +2890,42 @@ def _check_cjk_shape_purity(self) -> None: "where a fork could be said to exist, because the " "reading was never on offer (#516)", shape=1), + # #564: the comma position is read by default + # (CapsSuffixes.AFTER_COMMA); the trailing one stays opt-in. + Case("an_all_caps_word_after_a_full_name_and_a_comma_is_a_credential", + "John Smith, XYZ", + {"given": "John", "family": "Smith", "suffix": "XYZ"}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="the all-caps SURNAME convention never writes the " + "capitals after a comma behind a full name, so the " + "default reads the comma position (rules.md#C1, S2). " + "2.3.0 read given 'XYZ', family 'John Smith'"), + Case("the_comma_caps_reading_can_be_switched_off", + "John Smith, XYZ", + {"given": "XYZ", "family": "John Smith"}, + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.OFF), + notes="#564: CapsSuffixes.OFF is 2.3.0's reading, and the one " + "way to keep a given name written in capitals after a " + "two-word surname ('García Márquez, GABRIEL')"), + Case("a_lone_two_letter_caps_word_after_a_comma_is_the_given_name", + "García Márquez, MJ", + {"given": "MJ", "family": "García Márquez"}, + notes="#564 boundary: two capitals alone are how initials are " + "written undotted, and two words before the comma may " + "be one surname -- the case #563 decides for 'M.J.'. No " + "report: the comma position declines it outright"), Case("the_caps_comma_count_needs_two_name_words", "John Smith, XYZ", {"given": "John", "family": "Smith", "suffix": "XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("suffix-or-name",), notes="the comma structure moves with this half too, on the " "same NAME-word count"), Case("the_caps_comma_count_declines_at_one_word", "Smith, XYZ", {"given": "XYZ", "family": "Smith"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name",), notes="its control, switch and all: one name word before the " "comma is never enough, so the capitalised word stays " @@ -2905,7 +2933,7 @@ def _check_cjk_shape_purity(self) -> None: Case("one_case_input_never_reaches_the_caps_switch", "JOHN SMITH XYZ", {"given": "JOHN", "middle": "SMITH", "family": "XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the one-case control, and it needs the switch ON to " "mean anything: capitals against capitals are no " "contrast, so the shape never fires and the row is " @@ -2913,7 +2941,7 @@ def _check_cjk_shape_purity(self) -> None: Case("suffix_vocabulary_never_reaches_the_caps_switch", "John Smith MC", {"given": "John", "family": "Smith", "suffix": "MC"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="UNLISTED is the load-bearing word: 'mc' is suffix " "vocabulary, so the whole-token lookup claims it before " "any shape reading and this row reads the same with " @@ -2921,7 +2949,7 @@ def _check_cjk_shape_purity(self) -> None: Case("the_caps_comma_count_reaches_a_multi_word_run", "John Smith, LEED AP", {"given": "John", "family": "Smith", "suffix": "LEED AP"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("suffix-or-name",), notes="the first time this class reaches the comma form as " @@ -2938,7 +2966,7 @@ def _check_cjk_shape_purity(self) -> None: Case("the_caps_comma_multi_word_run_declines_at_one_word", "Smith, LEED AP", {"given": "LEED", "family": "Smith", "suffix": "AP"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#531)", ambiguities=("suffix-or-name", "suffix-or-name"), notes="the one-pre-comma-word twin of the row above, and the " @@ -2969,13 +2997,13 @@ def _check_cjk_shape_purity(self) -> None: # (`Case.shape`'s own docstring). Case("caps_switch_does_not_silence_the_listed_lean", "Jack MA", {"given": "Jack", "suffix": "MA"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name", "given-or-family"), notes="the switch must not touch a LISTED member's own " "#289 lean: identical to the default reading"), Case("caps_switch_does_not_silence_the_comma_lean", "Smith, MA", {"family": "Smith", "suffix": "MA"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name",), notes="the comma-form twin of the row above, same guarantee"), # #516 review round: UNLISTED means in no wordlist at all, not @@ -2985,14 +3013,14 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_switch_does_not_claim_a_capitalized_particle", "John Smith DE", {"given": "John", "middle": "Smith", "family": "DE"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="'de' is a particle, not merely absent from suffix " "vocabulary -- identical to the default reading with " "the switch on"), Case("caps_switch_does_not_claim_a_capitalized_particle_phrase", "John Smith, DE LA", {"family": "John Smith DE LA"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the multi-token twin: 'DE' and 'LA' are both " "particles, so the RUN test (#516's own 'LEED AP' " "shape) must decline them too -- identical to the " @@ -3006,7 +3034,7 @@ def _check_cjk_shape_purity(self) -> None: # and the reason these controls read identically on or off. Case("caps_switch_does_not_move_a_roman_numeral", "Jack VI", {"given": "Jack", "suffix": "VI"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name", "given-or-family"), notes="the roman-numeral fork claims 'VI' before this " "switch's lean/count is consulted -- identical to the " @@ -3014,19 +3042,19 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_switch_does_not_move_a_roman_numeral_with_words_to_spare", "John Smith VI", {"given": "John", "family": "Smith", "suffix": "VI"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name",), notes="the words-to-spare twin of the row above, same " "mechanism, same guarantee"), Case("caps_switch_does_not_move_a_title_floor_control", "Mr XXX", {"title": "Mr", "family": "XXX"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="one piece behind a title never reaches the peel's " "`k >= 2` floor -- identical to the default reading"), Case("caps_switch_does_not_reach_delimited_content", "Andrew Perkins (XYZ)", {"given": "Andrew", "family": "Perkins", "nickname": "XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="delimited content is decided by the clause escape, " "never at the trailing slot -- identical to the " "default reading"), @@ -3039,7 +3067,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_switch_run_test_declines_a_pure_listed_run", "John Smith, Ed Ma", {"given": "John", "family": "Smith", "suffix": "Ed Ma"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), ambiguities=("suffix-or-name",), notes="'Ed' and 'Ma' are both LISTED ambiguous members, " "Title-case (leans NAME, #289) -- the caps run test " @@ -3062,14 +3090,14 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_switch_does_not_claim_a_one_case_maiden_marker", "JOHN SMITH NEE", {"given": "JOHN", "middle": "SMITH", "family": "NEE"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="one case, and a maiden marker with no clause to open " "(nothing follows it) -- identical to the default " "reading either way"), Case("caps_switch_does_not_claim_a_mixed_case_maiden_marker", "John Smith NEE", {"given": "John", "middle": "Smith", "family": "NEE"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the mixed-case twin: 'NEE' is unlisted by the caps " "shape test's OWN membership check too, so this one " "was already declining before this fix -- pinned " @@ -3080,7 +3108,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_switch_reads_the_name_level_case_past_a_clause", "née JONES XYZ", {"given": "née", "middle": "JONES", "family": "XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the positive half of the two NEE rows above, and the " "row that fails if classify's caps branch is reverted " "to `one_case_own`: a marker OPENING the name leaves " @@ -3093,7 +3121,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_one_case_comma_declines_a_single_token", "JOHN SMITH, XYZ", {"given": "XYZ", "family": "JOHN SMITH"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the comma control for the one-case gate: two name " "words before the comma would flip the structure for a " "LISTED member, but the caps class needs a case " @@ -3103,7 +3131,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_one_case_comma_declines_a_run", "JOHN SMITH, LEED AP", {"given": "LEED", "middle": "AP", "family": "JOHN SMITH"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the multi-token twin: segment's run test asks " "`caps_shape_candidate` of every token and then reads " "the case fact ONCE, so a one-case name declines the " @@ -3111,7 +3139,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_run_needs_every_token_not_any", "John Smith, LEED BA", {"given": "LEED", "family": "John Smith", "suffix": "BA"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#531)", ambiguities=("suffix-or-name", "suffix-or-name"), notes="pins `all()` rather than `any()`: 'BA' is a LISTED " @@ -3174,14 +3202,14 @@ def _check_cjk_shape_purity(self) -> None: "word and reporting all three forks"), Case("caps_run_declines_a_bound_given_head", "John Smith, ABDUL AP", {"given": "ABDUL AP", "family": "John Smith"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the exclusion end to end rather than at the " "predicate: 'ABDUL' is bound-given vocabulary, so the " "run is no caps run and the part after the comma is " "the given name it would be at the default"), Case("caps_run_declines_a_conjunction", "John Smith, AND AP", {"given": "AND AP", "family": "John Smith"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="the same end to end for `conjunctions`, the row " "test_classify.py's predicate table names as the one " "that genuinely exercises that arm ('Y' declines at " @@ -3189,7 +3217,7 @@ def _check_cjk_shape_purity(self) -> None: Case("caps_switch_leaves_a_capitalized_title_a_title", "John Smith, MR", {"title": "MR", "given": "John", "family": "Smith"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), notes="`titles` end to end: the post-comma part holds no " "name word, so C1's no-name-word clause keeps the " "pre-comma positional read and 'MR' is the title it " @@ -3273,7 +3301,7 @@ def _check_cjk_shape_purity(self) -> None: Case("two_caps_credentials_peel_as_a_run", "John MA XYZ", {"given": "John", "suffix": "MA XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("given-or-family", "suffix-or-name", "suffix-or-name"), @@ -3288,7 +3316,7 @@ def _check_cjk_shape_purity(self) -> None: Case("the_caps_shape_is_script_agnostic_cyrillic", "Иван Петр ИВАНОВ", {"given": "Иван", "family": "Петр", "suffix": "ИВАНОВ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("suffix-or-name",), notes="the switch's docstring claims `isupper()` is " @@ -3331,7 +3359,7 @@ def _check_cjk_shape_purity(self) -> None: Case("the_caps_shape_is_script_agnostic_accented", "Jean Pierre ÉCOLE", {"given": "Jean", "family": "Pierre", "suffix": "ÉCOLE"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#516)", ambiguities=("suffix-or-name",), notes="the other half of the same claim: a non-ASCII LATIN " @@ -4368,7 +4396,7 @@ def _check_cjk_shape_purity(self) -> None: "Jane Doe nee Smith XYZ", {"given": "Jane", "family": "Doe", "suffix": "XYZ", "maiden": "Smith"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#533)", ambiguities=("suffix-or-name",), notes="the opt-in class reaches this slot like any other, " @@ -5487,7 +5515,7 @@ def _check_cjk_shape_purity(self) -> None: Case("the_caps_switch_reaches_the_trailing_slot", "Doe, John XYZ", {"given": "John", "family": "Doe", "suffix": "XYZ"}, - policy=Policy(unlisted_caps_suffixes=True), + policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), classification="fix(#531)", ambiguities=("suffix-or-name",), notes="the by-shape class reaches this slot through the " diff --git a/tests/v2/pipeline/test_classify.py b/tests/v2/pipeline/test_classify.py index 457b547e..6af86d2d 100644 --- a/tests/v2/pipeline/test_classify.py +++ b/tests/v2/pipeline/test_classify.py @@ -2,6 +2,7 @@ import pytest +from nameparser._policy import CapsSuffixes from nameparser import Parser from nameparser._lexicon import Lexicon, _normalize from nameparser._pipeline import STAGES @@ -193,8 +194,12 @@ def test_mixed_case_keeps_todays_rule_verbatim() -> None: def test_a_trailing_uppercase_suffix_makes_the_name_mixed_case() -> None: # the gate reads the whole name's text, so 'john e jones, III' is # mixed case and keeps today's reading -- the boundary that keeps - # a v1 corpus name still. - out = _classified("john e jones, III") + # a v1 corpus name still. 'iii' is listed here as it is in the + # shipped vocabulary: unlisted, 'III' is an all-caps word after a + # comma behind two name words, which the default reads as a + # credential run since #564 and reports -- a different fork. + lex = dataclasses.replace(_LEX, suffix_words=_LEX.suffix_words | {"iii"}) + out = _classified_with("john e jones, III", lex) assert "conjunction" in _tags(out, "e") assert out.ambiguities == () @@ -530,7 +535,7 @@ def test_no_listed_member_ever_carries_the_shape_tag() -> None: "John Smith BA", "John Smith X.Y.Z.")), (dotted, ("Jack A.B.", "John Smith A.B.", "Smith, A.B."))): - for policy in (Policy(), Policy(unlisted_caps_suffixes=True), + for policy in (Policy(), Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), Policy(unlisted_dotted_suffixes=False)): for text in texts: for word, tags in _tags_by_text(text, lexicon=lex, @@ -573,26 +578,26 @@ def test_delimited_content_never_joins_the_shape_class() -> None: # -- the predicate classify calls, and since the review round the # only route into the caps half at all. ("XYZ", Policy(), None), - ("XYZ", Policy(unlisted_caps_suffixes=True), False), - ("MC", Policy(unlisted_caps_suffixes=True), False), - ("X", Policy(unlisted_caps_suffixes=True), False), - ("XY2", Policy(unlisted_caps_suffixes=True), False), - ("DUPONT", Policy(unlisted_caps_suffixes=True), False), + ("XYZ", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("MC", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("X", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("XY2", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("DUPONT", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), # #516 review round, F1/F1b: LISTED members must keep the LEAN # (SHAPE_ACRONYM_TAG must NOT ride beside their membership tag), # and UNLISTED means in no wordlist at all -- a particle, an # ambiguous particle and a conjunction, capitalized, must not # join the class either. - ("MA", Policy(unlisted_caps_suffixes=True), False), - ("BA", Policy(unlisted_caps_suffixes=True), False), - ("DE", Policy(unlisted_caps_suffixes=True), False), - ("Y", Policy(unlisted_caps_suffixes=True), False), - ("VAN", Policy(unlisted_caps_suffixes=True), False), + ("MA", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("BA", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("DE", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("Y", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("VAN", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), # 'Y' declines at the `len(text) >= 2` shape gate and never # actually reaches the conjunction-exclusion check at all; # 'AND' is the row that genuinely exercises it (quality-review # finding). - ("AND", Policy(unlisted_caps_suffixes=True), False), + ("AND", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), # #516 review round (second finding): the four wordlists the # first fix round's exclusion left with no row of their own # (titles, given_name_titles, bound_given_names, suffix_words -- @@ -601,11 +606,11 @@ def test_delimited_content_never_joins_the_shape_class() -> None: # `vocab:title` guard and this elif's own check, and the row # still pins that it declines either way), and the maiden-marker # gap itself. - ("SIR", Policy(unlisted_caps_suffixes=True), False), - ("AUNT", Policy(unlisted_caps_suffixes=True), False), - ("ABDUL", Policy(unlisted_caps_suffixes=True), False), - ("JR", Policy(unlisted_caps_suffixes=True), False), - ("NEE", Policy(unlisted_caps_suffixes=True), False), + ("SIR", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("AUNT", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("ABDUL", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("JR", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), + ("NEE", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), False), ]) def test_ambiguous_class_candidate_agrees_with_the_tag( word: str, policy: Policy, one_case: bool | None) -> None: @@ -665,7 +670,7 @@ def test_the_caps_branch_reads_the_name_level_case_not_the_own_span( Without the wrapper the reading below holds; with it, 'XYZ' joins the class and the family name is lost. """ - on = Policy(unlisted_caps_suffixes=True) + on = Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE) parser = Parser(policy=on) name = parser.parse("née JONES XYZ") assert (name.given, name.middle, name.family) == ("née", "JONES", "XYZ") @@ -692,7 +697,7 @@ def test_the_caps_shape_is_silent_until_its_switch_is_on() -> None: # the dotted one: there is no fork to report while a caller has # not asked for the reading (#516). assert _tags_by_text("John Smith XYZ")["XYZ"] == frozenset() - on = Policy(unlisted_caps_suffixes=True) + on = Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE) tags = _tags_by_text("John Smith XYZ", policy=on)["XYZ"] assert "shape:acronym" in tags and "vocab:suffix-ambiguous" in tags @@ -703,7 +708,7 @@ def test_the_caps_shape_is_a_testable_predicate() -> None: # vocabulary only in the SHIPPED lexicon (this module's own _LEX # is deliberately minimal and does not carry it), so that one # assertion uses Lexicon.default() rather than the module fixture. - on = Policy(unlisted_caps_suffixes=True) + on = Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE) # a digit anywhere disqualifies it assert "shape:acronym" not in _tags_by_text("John Smith XY2", policy=on)["XY2"] # a single capital stays what it is today, an initial diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index 248a1378..d48cef13 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -2,6 +2,7 @@ import pytest +from nameparser._policy import CapsSuffixes from nameparser import Parser from nameparser._lexicon import ( Lexicon, _VOCAB_FIELDS, _normalize, _title_key, @@ -470,7 +471,7 @@ def test_a_callers_own_conjunction_marker_keeps_its_word_a_name() -> None: # 'ZZQ' with the switch on and the default vocabulary; one # wordlist entry is the whole difference (2026-09-18 review # round). - on = Policy(unlisted_caps_suffixes=True) + on = Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE) plain = Parser(policy=on).parse("John Smith ZZQ") assert (plain.family, plain.suffix) == ("Smith", "ZZQ") listed = Parser( @@ -494,7 +495,7 @@ def test_the_caps_exclusion_covers_every_vocabulary_field() -> None: # loop adds the same unlisted word to ONE field and checks the # predicate declines it. A field whose exclusion is dropped fails # here by name. - on = Policy(unlisted_caps_suffixes=True) + on = Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE) base = Lexicon.default() assert caps_shape_candidate("ZZQX", base, on, False) for field in _VOCAB_FIELDS: diff --git a/tests/v2/rules_doc.py b/tests/v2/rules_doc.py index f0525b24..44477afa 100644 --- a/tests/v2/rules_doc.py +++ b/tests/v2/rules_doc.py @@ -109,7 +109,7 @@ def has_boundary_or_waiver(self) -> bool: from nameparser._policy import ( # noqa: E402 - FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST, Policy) + FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST, CapsSuffixes, Policy) #: Named policies example annotations may reference. Grown as #: extraction demands; each addition is a diff to this dict only. @@ -128,9 +128,13 @@ def has_boundary_or_waiver(self) -> bool: #: implementation-free, so the annotation slot is the one place #: the doc can put the caller-facing name of the switch its prose #: describes. The suffix says which way the field is set, since - #: one is on by default and the other off. + #: one is on by default and the caps field (#564) has three values, + #: so its suffix names the value set. "unlisted_dotted_suffixes-off": Policy(unlisted_dotted_suffixes=False), - "unlisted_caps_suffixes-on": Policy(unlisted_caps_suffixes=True), + "unlisted_caps_suffixes-everywhere": Policy( + unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), + "unlisted_caps_suffixes-off": Policy( + unlisted_caps_suffixes=CapsSuffixes.OFF), #: The delimiter switch (#206), named after its Policy FIELD for #: the reason above and carrying the DELIMITER in the suffix, #: because this field's value is a set rather than a flag: a rule diff --git a/tests/v2/test_facade_cases.py b/tests/v2/test_facade_cases.py index 0620e0cc..ce2a474f 100644 --- a/tests/v2/test_facade_cases.py +++ b/tests/v2/test_facade_cases.py @@ -120,6 +120,9 @@ "suffix_vocabulary_never_reaches_the_caps_switch", "the_caps_comma_count_reaches_a_multi_word_run", "the_caps_comma_multi_word_run_declines_at_one_word", + # #564: CapsSuffixes.OFF has no v1 spelling either; the facade + # reads the field's default, AFTER_COMMA, like the core. + "the_comma_caps_reading_can_be_switched_off", # #516 review round: the F1/F1b/F2/F5 regression-guard rows, all # under the same non-default policy. "caps_switch_does_not_silence_the_listed_lean", diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 80bf60fb..0bf3420a 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1385,6 +1385,9 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: # particle surname, keep the credential reading. "fix(#575) a particle surname before a comma is one name word": ("John van Buren, Ed", "John van der Berg, PhD"), + # #564: one name word, a lone two-letter word, and no comma. + "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": + ("Smith, XYZ", "García Márquez, MJ", "John Smith XYZ"), # #516's dotted rule must never reach the shapes the vocabulary # SETTLES (a whole-token match, or a surviving chunk claim), the # leading dotted runs rules.md#S2 excludes by position, the single @@ -3078,6 +3081,13 @@ class _LatinCopy(NamedTuple): frozenset({"Doe, Jane nee Smith PhD MEng", "Jane Doe nee Smith PhD MA", "Jane Doe nee Smith PhD MEng"}), frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), + # 2026-10-01, #564: the comma caps rule, literal names that copy no + # set, at 1.4.0 and in the 2.x copy. + frozenset({"John Smith, LEED AP", "John Smith, XYZ", + "The Rt Hon Kenneth Clarke QC MP, HMG"}), + frozenset({"Ahmad Jayadi, CHA", "John Smith, LEED AP", + "John Smith, RAI", "John Smith, XYZ", + "The Rt Hon Kenneth Clarke QC MP, HMG"}), # 2026-10-01, #575: rules.md#C1's particle-surname examples at # 1.4.0, literal names that copy no set. frozenset({"De La Cruz, Ed", "Freiherr von Berg, Ed", "Van Buren, Ed", @@ -3575,7 +3585,9 @@ def _claim(rule: dict) -> _Claim: # at this baseline 'Aishwarya Rai' reaches the rule and is # explained by parity instead, so the gate heading reads 4. "fix(#342) rai and cha left the credential acronyms, so a trailing Rai or CHA is a name word": - _Claim(5, ('family', 'given', 'middle', 'suffix'), "c3d76812da97", None), + # 2026-10-01, #564: `given` leaves the roles -- the comma pair + # reads suffix again by its capitals, so no name here moves it. + _Claim(5, ('family', 'middle', 'suffix'), "c3d76812da97", None), "fix(A2) content-free input names nobody, so every role empties": _Claim(5, ('given',), "1af8d718688b", None), "fix(#335) a marker-led clause leaves the one name word its bare reading": @@ -3794,7 +3806,10 @@ def _claim(rule: dict) -> _Claim: # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(423, ('given', 'suffix', 'title'), "f5edb96b96cb", None), + # 2026-10-01, #564: 423 -> 426, 'John Smith, XYZ', 'Smith, + # XYZ' and 'García Márquez, MJ', #564's rules.md#C1 examples. + # Reach, verified name by name. + _Claim(426, ('given', 'suffix', 'title'), "684dd90b64f3", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3830,7 +3845,8 @@ def _claim(rule: dict) -> _Claim: # explained by neither: 1.4.0 already reads the given name # for both, so there is no diff here to explain. "fix(#296) a lone post-comma credential is a suffix": - _Claim(23, ('family', 'given', 'suffix', 'title'), "54c1ae9911e1", None), + # 2026-10-01, #564: 23 -> 24, 'Smith, XYZ'. Reach. + _Claim(24, ('family', 'given', 'suffix', 'title'), "eae9a2bb02b0", None), # 2026-09-27, #544: 6 -> 7; gains 'Smith, PhD MEng'. # 2026-09-28, #544: 7 -> 8; gains 'Smith, PhD Ma'. "fix(#325) a split credential followed by another suffix after a one-word family comma reads as suffixes": @@ -3897,11 +3913,18 @@ def _claim(rule: dict) -> _Claim: # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(423, ('family', 'given'), "f5edb96b96cb", None), + # 2026-10-01, #564: 423 -> 426, 'John Smith, XYZ', 'Smith, + # XYZ' and 'García Márquez, MJ', #564's rules.md#C1 examples. + # Reach, verified name by name. + _Claim(426, ('family', 'given'), "684dd90b64f3", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": _Claim(4, ('family', 'given', 'suffix', 'title'), "30163564e03d", None), + # 2026-10-01, #564: new, 3; 'John Smith, LEED AP', 'John Smith, + # XYZ', 'The Rt Hon Kenneth Clarke QC MP, HMG'. + "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": + _Claim(3, ('family', 'given', 'middle', 'suffix', 'title'), "1926a01dc71c", ('DEFAULT',)), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4657,7 +4680,9 @@ def _claim(rule: dict) -> _Claim: # so a widening taking one shape alone would change the roles # here before it reached the gate. "fix(#342) rai and cha left the credential acronyms, so a trailing Rai or CHA is a name word": - _Claim(5, ('family', 'given', 'middle', 'suffix'), "c3d76812da97", None), + # 2026-10-01, #564: `given` leaves the roles -- the comma pair + # reads suffix again by its capitals, so no name here moves it. + _Claim(5, ('family', 'middle', 'suffix'), "c3d76812da97", None), # The compound rule, at the two baselines where 'abdul Smith # Jr V' already diffs {family, given} under fix(#401) and the # widened diff leaves that rule's `fields`. Three roles here @@ -4820,7 +4845,8 @@ def _claim(rule: dict) -> _Claim: # explained by neither: at this baseline the diff moves # `given`, a field outside this rule's own ('suffix', 'title'). "fix(#296) a lone post-comma credential is a suffix": - _Claim(23, ('suffix', 'title'), "54c1ae9911e1", None), + # 2026-10-01, #564: 23 -> 24, 'Smith, XYZ'. Reach. + _Claim(24, ('suffix', 'title'), "eae9a2bb02b0", None), # 2026-09-27, #544: 6 -> 7; gains 'Smith, PhD MEng'. # 2026-09-28, #544: 7 -> 8; gains 'Smith, PhD Ma'. "fix(#325) a split credential followed by another suffix after a one-word family comma reads as suffixes": @@ -4997,6 +5023,11 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), + # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, + # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon + # Kenneth Clarke QC MP, HMG'. + "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -5387,7 +5418,9 @@ def _claim(rule: dict) -> _Claim: # so a widening taking one shape alone would change the roles # here before it reached the gate. "fix(#342) rai and cha left the credential acronyms, so a trailing Rai or CHA is a name word": - _Claim(5, ('family', 'given', 'middle', 'suffix'), "c3d76812da97", None), + # 2026-10-01, #564: `given` leaves the roles -- the comma pair + # reads suffix again by its capitals, so no name here moves it. + _Claim(5, ('family', 'middle', 'suffix'), "c3d76812da97", None), # The four one-name CJK rules, literal-anchored, at the two # baselines where the render is the whole of what moved. A # reach of 1 is one _CORPUS_CLAIMS cannot police on its own -- @@ -5499,6 +5532,11 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), + # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, + # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon + # Kenneth Clarke QC MP, HMG'. + "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -5838,7 +5876,9 @@ def _claim(rule: dict) -> _Claim: # so a widening taking one shape alone would change the roles # here before it reached the gate. "fix(#342) rai and cha left the credential acronyms, so a trailing Rai or CHA is a name word": - _Claim(5, ('family', 'given', 'middle', 'suffix'), "c3d76812da97", None), + # 2026-10-01, #564: `given` leaves the roles -- the comma pair + # reads suffix again by its capitals, so no name here moves it. + _Claim(5, ('family', 'middle', 'suffix'), "c3d76812da97", None), # The four one-name CJK rules, literal-anchored, at the two # baselines where the render is the whole of what moved. A # reach of 1 is one _CORPUS_CLAIMS cannot police on its own -- @@ -5982,7 +6022,8 @@ def _claim(rule: dict) -> _Claim: # explained by neither: at this baseline the diff moves # `given`, a field outside this rule's own ('suffix', 'title'). "fix(#296) a lone post-comma credential is a suffix": - _Claim(23, ('suffix', 'title'), "54c1ae9911e1", None), + # 2026-10-01, #564: 23 -> 24, 'Smith, XYZ'. Reach. + _Claim(24, ('suffix', 'title'), "eae9a2bb02b0", None), # 2026-09-27, #544: 6 -> 7; gains 'Smith, PhD MEng'. # 2026-09-28, #544: 7 -> 8; gains 'Smith, PhD Ma'. "fix(#325) a split credential followed by another suffix after a one-word family comma reads as suffixes": @@ -6144,6 +6185,11 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), + # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, + # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon + # Kenneth Clarke QC MP, HMG'. + "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6512,6 +6558,11 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), + # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, + # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon + # Kenneth Clarke QC MP, HMG'. + "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index b1ca4c6c..e2c626e1 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -3,6 +3,7 @@ import pytest +from nameparser._policy import CapsSuffixes from nameparser._policy import ( DEFAULT_SCRIPT_ORDERS, FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST, GIVEN_FIRST, PatronymicRule, Policy, PolicyPatch, Script, UNSET, @@ -817,17 +818,27 @@ def test_unlisted_dotted_suffixes_off_reads_name_material() -> None: assert p.parse("Doe, John Msc.Ed.").suffix == "Msc.Ed." -def test_unlisted_caps_suffixes_is_a_validated_bool_defaulting_off() -> None: - # #516's all-caps half is OPT-IN, and the asymmetry with the - # dotted half is the whole decision: an all-caps surname is a real - # writing convention that shape cannot separate from a credential. - assert Policy().unlisted_caps_suffixes is False - assert Policy(unlisted_caps_suffixes=True).unlisted_caps_suffixes is True - with pytest.raises(TypeError, match="unlisted_caps_suffixes"): - Policy(unlisted_caps_suffixes="yes") # type: ignore[arg-type] +def test_unlisted_caps_suffixes_is_a_validated_enum_defaulting_after_comma( +) -> None: + # #564: the all-caps half reads the comma position by default and + # the trailing one only on request -- an all-caps SURNAME is a real + # convention there, and never after a comma behind a full name. + assert Policy().unlisted_caps_suffixes is CapsSuffixes.AFTER_COMMA + # a plain string is coerced at runtime; the annotation names what + # the field stores, so the type checker wants the member + assert (Policy(unlisted_caps_suffixes="everywhere") # type: ignore[arg-type] + .unlisted_caps_suffixes is CapsSuffixes.EVERYWHERE) assert Policy().patched( - PolicyPatch(unlisted_caps_suffixes=True) - ).unlisted_caps_suffixes is True + PolicyPatch(unlisted_caps_suffixes=CapsSuffixes.OFF) + ).unlisted_caps_suffixes is CapsSuffixes.OFF + with pytest.raises(ValueError, match="after-comma, everywhere"): + Policy(unlisted_caps_suffixes="yes") # type: ignore[arg-type] + # the old bool spelling: a TypeError naming both replacements, in + # a form that type-checks when pasted (AGENTS.md) + for old, new in ((True, "CapsSuffixes.EVERYWHERE"), + (False, "CapsSuffixes.OFF")): + with pytest.raises(TypeError, match=new.replace(".", r"\.")): + Policy(unlisted_caps_suffixes=old) # type: ignore[arg-type] def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: @@ -839,7 +850,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # comma candidate pays for the fact it forces: re-measured # 2026-09-18 on this tree, same-interpreter harness (the resolved # `sys.executable` used on both sides, `Parser().parse` and - # `Parser(policy=Policy(unlisted_caps_suffixes=True)).parse`, mean + # `Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)).parse`, mean # of 50 after one warm-up parse) -- `"Smith, John"` is +7 # (206 -> 213), `"Smith, XYZ"` (a real candidate) is +40 # (205 -> 245). An earlier round's comment here read +6/+39 @@ -856,7 +867,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # as a silent drift. from nameparser import Parser - on = Parser(policy=Policy(unlisted_caps_suffixes=True)) + on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) off = Parser() # what it buys assert on.parse("John Smith XYZ").suffix == "XYZ" diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index 4a4a5272..9f75c8d9 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -16,6 +16,7 @@ from hypothesis import given, settings from hypothesis import strategies as st +from nameparser._policy import CapsSuffixes from nameparser import ( DEFAULT_SCRIPT_ORDERS, FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST, GIVEN_FIRST, HumanName, Lexicon, Parser, PatronymicRule, Policy, @@ -555,7 +556,7 @@ def test_a_maiden_clause_does_not_change_how_a_trailing_word_reads( # the allowlist's structural argument rests on. markers = ("nee", "née", "geb.") policies = (("default", Policy()), - ("caps", Policy(unlisted_caps_suffixes=True)), + ("caps", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)), ("nodot", Policy(unlisted_dotted_suffixes=False))) def side(name: ParsedName, word: str) -> str: @@ -698,7 +699,7 @@ def _maiden_clause_grid() -> list[tuple[str, Parser, str]]: "Smith V MA", "MA", "MA PhD", "MA ba", "Smith Jones MA", "Smith PhD", "Smith Jr", "Jones Smith Ma", "Smith MA PhD") policies = (("default", Policy()), - ("caps", Policy(unlisted_caps_suffixes=True)), + ("caps", Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)), ("nodot", Policy(unlisted_dotted_suffixes=False))) parsers = [(label, Parser(policy=p)) for label, p in policies] texts: list[str] = [] @@ -1063,7 +1064,7 @@ def test_no_two_ambiguities_name_the_same_token_span() -> None: "Jones Smith") parsers = [Parser(), Parser(policy=Policy(unlisted_dotted_suffixes=False)), - Parser(policy=Policy(unlisted_caps_suffixes=True))] + Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE))] failures = [] multi = 0 for head in heads: diff --git a/tests/v2/test_render.py b/tests/v2/test_render.py index fd5924db..023488ac 100644 --- a/tests/v2/test_render.py +++ b/tests/v2/test_render.py @@ -4,6 +4,7 @@ import pytest +from nameparser._policy import CapsSuffixes from nameparser import FAMILY_FIRST, HumanName, Parser, Policy, parse from nameparser._lexicon import Lexicon from nameparser.config import Constants @@ -774,7 +775,7 @@ def test_a_credential_read_by_its_shape_repairs_to_capitals() -> None: assert str(off.capitalized(off.parse("john smith x.y.z."))) \ == "John Smith X.y.z." # the opt-in all-caps half writes the same tag - caps = Parser(policy=Policy(unlisted_caps_suffixes=True)) + caps = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) assert str(caps.capitalized(caps.parse("John Smith XYZ"), force=True)) == "John Smith XYZ" assert str(parse("John Smith XYZ").capitalized(force=True)) \ diff --git a/tools/differential/compare.py b/tools/differential/compare.py index 814fa18b..8e40fa80 100644 --- a/tools/differential/compare.py +++ b/tools/differential/compare.py @@ -2152,7 +2152,10 @@ class _ShapeMismatch(NamedTuple): "Jane van der Berg 旧姓 Jones": ("family", "maiden"), "Janey née Jones": ("family", "given", "maiden", "middle"), "John Smith Rev.": ("family", "middle", "title"), - "John Smith, RAI": ("family", "given", "suffix"), + # 'John Smith, RAI' left 2026-10-01 (#564): an unlisted all-caps + # word after a comma behind two name words is a credential by + # default, so the name reads suffix 'RAI' as 1.4.0 did -- parity, + # and a watched row with no diff to watch. "John V": ("family", "suffix"), "John of the Doe": ("_initials",), "Jong van der": ("_initials",), @@ -2193,7 +2196,10 @@ class _ShapeMismatch(NamedTuple): "Janey née Jones": ("family", "given"), "Joe E. Smith": ("_initials",), "John Smith Rev.": ("family", "middle", "title"), - "John Smith, RAI": ("family", "given", "suffix"), + # 2026-10-01 (#564): the default now reads 'RAI' as the + # credential again by its capitals, as this baseline read it + # by vocabulary, so only the comma's report differs. + "John Smith, RAI": ("_ambiguities",), "John, Smith, Dr.": ("_ambiguities",), "Jong van der": ("_initials",), "Jong, van der": ("_initials",), @@ -2239,7 +2245,10 @@ class _ShapeMismatch(NamedTuple): "Janey née Jones": ("family", "given"), "Joe E. Smith": ("_initials",), "John Smith Rev.": ("family", "middle", "title"), - "John Smith, RAI": ("family", "given", "suffix"), + # 2026-10-01 (#564): the default now reads 'RAI' as the + # credential again by its capitals, as this baseline read it + # by vocabulary, so only the comma's report differs. + "John Smith, RAI": ("_ambiguities",), "John, Smith, Dr.": ("_ambiguities",), "Jong van der": ("_initials",), "Jong, van der": ("_initials",), @@ -2269,7 +2278,9 @@ class _ShapeMismatch(NamedTuple): "E Anne D,Leonardo": ("_initials",), "Joe E. Smith": ("_initials",), "John Smith Rev.": ("family", "middle", "title"), - "John Smith, RAI": ("family", "given", "suffix"), + # 2026-10-01 (#564): reads suffix 'RAI' again by its capitals, + # as this baseline read it by vocabulary; the report differs. + "John Smith, RAI": ("_ambiguities",), "Jose E. Maria Santos": ("_initials",), "Lala Lajpat Rai": ("family", "middle", "suffix"), "Smith, John E, III, Jr": ("_initials",), diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 7151dfac..d63657d9 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -94,6 +94,7 @@ "García Márquez, G.J" "García Márquez, G.J." "García Márquez, G.J.R." +"García Márquez, MJ" "García Márquez, Ms G.J." "Hans „Erster“ und “Zweiter” Müller" "Hassan Mohamad Ali" @@ -228,6 +229,7 @@ "John Smith, V." "John Smith, X.Y. P.Q." "John Smith, X.Y.Z." +"John Smith, XYZ" "John Smith, vd Ma" "John and Jane Smith" "John née Jones Smith MA" @@ -363,6 +365,7 @@ "Smith, PhD MEng" "Smith, PhD Ma" "Smith, Sr." +"Smith, XYZ" "Smith, de Mesnil Jean" "Smith. John" "Steven Hardman, MD, DO, DDS" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 9e6fabe0..3331f0da 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -1302,6 +1302,12 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # 'John Smith RAI' {middle, family, suffix} same # 'John Smith, RAI' {given, family, suffix} an ordinary family comma (C1) # 'Ahmad Jayadi, CHA' {given, family, suffix} same +# 2026-10-01, #564: the comma pair reads suffix again by its capitals +# (an unlisted all-caps word after a comma behind two name words), so +# it no longer moves {given, family, suffix} here; at 1.4.0 it reaches +# parity, at 2.x only the comma's report differs (the fix(#564) rule). +# `given` therefore leaves `fields`, which the OVER-DECLARED check +# requires: no name here moves it any more. # The comma pair is why `fields` declares four roles rather than the # two the pre-2.3 fix(#342) rule declared: with no credential after # the comma, C1 reads the pre-comma run as the family and the @@ -1325,7 +1331,7 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # reach at 5 with its digest, and _MUST_NOT_MATCH carries the # boundaries the paragraph above argues. name_regex = "^(?:Ahmad Jayadi, CHA|Aishwarya Rai|John Smith RAI|John Smith, RAI|Lala Lajpat Rai)$" -fields = ["family", "given", "middle", "suffix"] +fields = ["family", "middle", "suffix"] [[change]] issue = "fix(#397) accepted: a two-word Catalan link has no name word to its right and stays the generation" @@ -4849,3 +4855,21 @@ issue = "fix(#296) phd is not a prenominal, so behind a dual title it is the giv # _MUST_NOT_MATCH. name_regex = "^Smith, MD PhD Ma$" fields = ["title", "given", "middle"] + +[[change]] +issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" +# rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to +# CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more +# name words before it reads an unlisted all-caps word of three or +# more letters, or a run holding one, as the credential run and +# reports the call. The all-caps SURNAME convention never writes the +# capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the +# C1 examples (the latter closing the #291 deviation the doc carried); +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. +# +# Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, +# MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the +# trailing position stays opt-in) are _MUST_NOT_MATCH. +name_regex = "^(?:John Smith, LEED AP|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +fields = ["title", "given", "middle", "family", "suffix"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 569fd613..ad138dda 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -307,6 +307,12 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # 'John Smith RAI' {middle, family, suffix} same # 'John Smith, RAI' {given, family, suffix} an ordinary family comma (C1) # 'Ahmad Jayadi, CHA' {given, family, suffix} same +# 2026-10-01, #564: the comma pair reads suffix again by its capitals +# (an unlisted all-caps word after a comma behind two name words), so +# it no longer moves {given, family, suffix} here; at 1.4.0 it reaches +# parity, at 2.x only the comma's report differs (the fix(#564) rule). +# `given` therefore leaves `fields`, which the OVER-DECLARED check +# requires: no name here moves it any more. # The comma pair is why `fields` declares four roles rather than the # two the pre-2.3 fix(#342) rule declared: with no credential after # the comma, C1 reads the pre-comma run as the family and the @@ -330,7 +336,7 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # reach at 5 with its digest, and _MUST_NOT_MATCH carries the # boundaries the paragraph above argues. name_regex = "^(?:Ahmad Jayadi, CHA|Aishwarya Rai|John Smith RAI|John Smith, RAI|Lala Lajpat Rai)$" -fields = ["family", "given", "middle", "suffix"] +fields = ["family", "middle", "suffix"] [[change]] issue = "fix(#271/#272/#298) native-script CJK: family-first order, hangul segmentation, the kana license and the dots" @@ -3907,3 +3913,23 @@ issue = "fix(#575) a particle surname before a comma is one name word" name_regex = "^van der Berg, PhD$" fields = ["given", "family", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" +# rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to +# CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more +# name words before it reads an unlisted all-caps word of three or +# more letters, or a run holding one, as the credential run and +# reports the call. The all-caps SURNAME convention never writes the +# capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the +# C1 examples (the latter closing the #291 deviation the doc carried); +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read +# suffix again, as this baseline read them by vocabulary before #342 +# removed both words, and differ only by the comma's report. +# +# Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, +# MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the +# trailing position stays opt-in) are _MUST_NOT_MATCH. +name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 8aaa93c7..c2af9221 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -363,6 +363,12 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # 'John Smith RAI' {middle, family, suffix} same # 'John Smith, RAI' {given, family, suffix} an ordinary family comma (C1) # 'Ahmad Jayadi, CHA' {given, family, suffix} same +# 2026-10-01, #564: the comma pair reads suffix again by its capitals +# (an unlisted all-caps word after a comma behind two name words), so +# it no longer moves {given, family, suffix} here; at 1.4.0 it reaches +# parity, at 2.x only the comma's report differs (the fix(#564) rule). +# `given` therefore leaves `fields`, which the OVER-DECLARED check +# requires: no name here moves it any more. # The comma pair is why `fields` declares four roles rather than the # two the pre-2.3 fix(#342) rule declared: with no credential after # the comma, C1 reads the pre-comma run as the family and the @@ -386,7 +392,7 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # reach at 5 with its digest, and _MUST_NOT_MATCH carries the # boundaries the paragraph above argues. name_regex = "^(?:Ahmad Jayadi, CHA|Aishwarya Rai|John Smith RAI|John Smith, RAI|Lala Lajpat Rai)$" -fields = ["family", "given", "middle", "suffix"] +fields = ["family", "middle", "suffix"] [[change]] issue = "fix(#369) a given-name title licenses the bound given-name join with one word to spare" @@ -3818,3 +3824,23 @@ issue = "fix(#575) a particle surname before a comma is one name word" name_regex = "^van der Berg, PhD$" fields = ["given", "family", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" +# rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to +# CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more +# name words before it reads an unlisted all-caps word of three or +# more letters, or a run holding one, as the credential run and +# reports the call. The all-caps SURNAME convention never writes the +# capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the +# C1 examples (the latter closing the #291 deviation the doc carried); +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read +# suffix again, as this baseline read them by vocabulary before #342 +# removed both words, and differ only by the comma's report. +# +# Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, +# MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the +# trailing position stays opt-in) are _MUST_NOT_MATCH. +name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 0bb052e3..43dfe0c6 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -352,6 +352,12 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # 'John Smith RAI' {middle, family, suffix} same # 'John Smith, RAI' {given, family, suffix} an ordinary family comma (C1) # 'Ahmad Jayadi, CHA' {given, family, suffix} same +# 2026-10-01, #564: the comma pair reads suffix again by its capitals +# (an unlisted all-caps word after a comma behind two name words), so +# it no longer moves {given, family, suffix} here; at 1.4.0 it reaches +# parity, at 2.x only the comma's report differs (the fix(#564) rule). +# `given` therefore leaves `fields`, which the OVER-DECLARED check +# requires: no name here moves it any more. # The comma pair is why `fields` declares four roles rather than the # two the pre-2.3 fix(#342) rule declared: with no credential after # the comma, C1 reads the pre-comma run as the family and the @@ -375,7 +381,7 @@ issue = "fix(#342) rai and cha left the credential acronyms, so a trailing Rai o # reach at 5 with its digest, and _MUST_NOT_MATCH carries the # boundaries the paragraph above argues. name_regex = "^(?:Ahmad Jayadi, CHA|Aishwarya Rai|John Smith RAI|John Smith, RAI|Lala Lajpat Rai)$" -fields = ["family", "given", "middle", "suffix"] +fields = ["family", "middle", "suffix"] [[change]] issue = "fix(#462) the facade keeps an initial-shaped conjunction letter" @@ -2201,3 +2207,23 @@ issue = "fix(#575) a particle surname before a comma is one name word" name_regex = "^van der Berg, PhD$" fields = ["given", "family", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" +# rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to +# CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more +# name words before it reads an unlisted all-caps word of three or +# more letters, or a run holding one, as the credential run and +# reports the call. The all-caps SURNAME convention never writes the +# capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the +# C1 examples (the latter closing the #291 deviation the doc carried); +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read +# suffix again, as this baseline read them by vocabulary before #342 +# removed both words, and differ only by the comma's report. +# +# Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, +# MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the +# trailing position stays opt-in) are _MUST_NOT_MATCH. +name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 8f149d1d..d3645efc 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1462,3 +1462,23 @@ issue = "fix(#575) a particle surname before a comma is one name word" name_regex = "^van der Berg, PhD$" fields = ["given", "family", "_ambiguities"] orders = ["DEFAULT"] + +[[change]] +issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" +# rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to +# CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more +# name words before it reads an unlisted all-caps word of three or +# more letters, or a run holding one, as the credential run and +# reports the call. The all-caps SURNAME convention never writes the +# capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the +# C1 examples (the latter closing the #291 deviation the doc carried); +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' move +# from this baseline's family comma (given 'RAI'/'CHA', the words +# removed from the vocabulary by #342) to the credential run. +# +# Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, +# MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the +# trailing position stays opt-in) are _MUST_NOT_MATCH. +name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] +orders = ["DEFAULT"] From 8d1abb5c972e93291162dd696caaca39f7fcf7ff Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 22:04:47 -0700 Subject: [PATCH 02/13] fix(S2): #564 review round -- mixed runs, case repair, frame guard, docs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Mixed runs (Derek): an unlisted all-caps word is a by-shape member of C1's run beside listed credentials, so 'John Smith, PhD XYZ' and 'John Smith, XYZ Jr.' read as credential runs where adding a listed credential had flipped the part back to a name. A two-letter caps word in a run is read as #563 reads a pair ('García Márquez, MJ PhD' keeps given 'MJ'; 'MJ JK' flips). - Case repair: classify tags a caps word in the part a suffix comma opened under the default too, so capitalized(force=True) keeps 'XYZ' in 'John Smith, XYZ' as EVERYWHERE does. - Frames: the comma test needs two words before the comma first, so 'Smith, JOHN' pays nothing (it paid +25); classify skips a listed member before the caps predicate. - rules.md#C1 and the release log state the two-letter rule as the code does; new examples pin it and the mixed run. Stale "off by default" prose fixed in tests, the case table, mechanisms.md (which also named a caller that does not exist) and a guard comment; decisions.md#C1 points at the S2 entry; the old recipe is translated. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 6 +- docs/design/mechanisms.md | 2 +- docs/design/rules.md | 13 ++-- docs/release_log.rst | 2 +- nameparser/_pipeline/_classify.py | 35 ++++++--- nameparser/_pipeline/_segment.py | 53 +++++++++---- tests/v2/cases.py | 68 +++++++++-------- tests/v2/pipeline/test_classify.py | 17 +++-- tests/v2/test_ledger_guards.py | 78 +++++++++++--------- tests/v2/test_policy.py | 19 +++-- tools/differential/corpus_rules.jsonl | 3 + tools/differential/expected_since_1.4.0.toml | 5 +- tools/differential/expected_since_2.0.0.toml | 7 +- tools/differential/expected_since_2.1.0.toml | 7 +- tools/differential/expected_since_2.2.0.toml | 7 +- tools/differential/expected_since_2.3.0.toml | 7 +- 16 files changed, 211 insertions(+), 118 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 8828e8ae..d93ce542 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,9 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Implemented as "a part that is one word of fewer than three letters", so a run holding a longer word (`LEED AP`) is admitted as one with a longer member, mirroring #563, where one pair is a given name and two are a credential. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read as #563 reads a pair: with only a credential behind it, it is still the given name (`García Márquez, MJ PhD`, as `De La Cruz, M.J. PhD`), while a credential in front or a second such word makes the run (`García Márquez, PhD MJ`, `García Márquez, MJ JK`). The first draft's rules.md and release-log wording said "three letters or more, or a run holding such a word", which the code never did; both reviews caught it. (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. Mixed case is still required, as for the all-caps run. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 256. The comma test is on by default now, so it asks a C-level `isupper()` of the part's first word and the lone-two-letter length test before calling `caps_shape_candidate`; a comma name with no all-caps word pays nothing. + CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render alike. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 258. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Classify skips a listed member before calling the caps predicate. The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) @@ -893,6 +894,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out). BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. +- 2026-10-01 (Derek), #564 — AN UNLISTED ALL-CAPS WORD IS A CREDENTIAL IN THE PART AFTER A COMMA BEHIND TWO NAME WORDS, BY DEFAULT. The decision, its two-letter rule and the mixed-run membership are recorded under decisions.md#S2 (#564), where the caps shape's switch has always been decided; this bullet is C1's pointer to it. ### T1 — separators, not joiners diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 550baed6..66f82e37 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -55,7 +55,7 @@ Problem shape. "Which stage does X?" — asked before attributing behavior in pr ## ONE-PREDICATE-PER-QUESTION — one predicate answers it, and every other site calls that -Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from three sites that each needed the identical question answered — classify's own tag emission, this module's ambiguous_class_candidate, and _segment.py's multi-token run test — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: the trailing-position caller's first conjunct is the setting itself (`CapsSuffixes.EVERYWHERE`, not the default), and the comma run test — on by default since #564 — asks a C-level `isupper()` of the part's first word before calling, so a parse with no all-caps word after a comma never reaches it (decisions.md#S2, #C1)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. +Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from the sites that each needed the identical question answered — classify's own tag emission and _segment.py's two comma tests (the all-caps run and, since #564, the #544 run test's by-shape member), with its unit tests — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: the trailing-position caller's first conjunct is the setting itself (`CapsSuffixes.EVERYWHERE`, not the default), and the comma run test — on by default since #564 — asks a C-level `isupper()` of the part's first word before calling, so a parse with no all-caps word after a comma never reaches it (decisions.md#S2, #C1)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. ## RENDER-HONORS-THE-PARSE — the parse decides it, the views honor it diff --git a/docs/design/rules.md b/docs/design/rules.md index 306b25e6..73d258ea 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1722,11 +1722,11 @@ C1. Rationale: a credential run after the comma means the name is in there. Behind two or more name words, paired initials that are not the only shape word in the part are a credential however little else speaks for them, since no one writes a person's - initials as two dotted groups. A lone word of two capitals after - the comma is the same shape undotted ('García Márquez, MJ'): it - reads as the given name, and an unlisted all-caps word joins the - class there only at three letters or more, or in a run holding - such a word. The same count reads a part of two or + initials as two dotted groups. An unlisted word of two capitals + after the comma is that shape undotted and is read as paired + initials are: alone it is the given name ('García Márquez, MJ'), + and so it is with only a credential behind it, while a credential + in front of it or a second such word makes the run. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any @@ -1863,6 +1863,9 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, LEED AP" → suffix="LEED AP" "Smith, XYZ" → given="XYZ" · boundary "García Márquez, MJ" → given="MJ" · boundary + "García Márquez, MJ PhD" → given="MJ" · boundary + "García Márquez, MJ JK" → suffix="MJ JK" + "John Smith, PhD XYZ" → suffix="PhD XYZ" "Smith, A.B." → given="A.B." · boundary "García Márquez, G.J." → given="G.J." "García Márquez, G.J." → family="García Márquez" diff --git a/docs/release_log.rst b/docs/release_log.rst index 9f1995a3..1719540b 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word of three or more letters, or a run holding one, in the part right after a comma behind two or more name words: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` gives suffix ``LEED AP`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), and so does a lone two-letter word (``García Márquez, MJ``), the way initials are written. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), and a two-letter word reads as initials do: alone, or with only a credential behind it, it is the given name (``García Márquez, MJ``, ``García Márquez, MJ PhD``). ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/nameparser/_pipeline/_classify.py b/nameparser/_pipeline/_classify.py index 51eff84d..3c644782 100644 --- a/nameparser/_pipeline/_classify.py +++ b/nameparser/_pipeline/_classify.py @@ -54,7 +54,7 @@ from nameparser._policy import CapsSuffixes from nameparser._pipeline._state import ( AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, ParseState, PendingAmbiguity, - WorkToken, copy_with, + Structure, WorkToken, copy_with, ) from nameparser._types import AmbiguityKind, Role from nameparser._pipeline._vocab import ( @@ -76,7 +76,7 @@ # spare" def _tags_for(token: WorkToken, n: str, state: ParseState, marker_tag: str | None, one_case_own: bool, - one_case: bool) -> frozenset[str]: + one_case: bool, comma_run: bool = False) -> frozenset[str]: """`n` is _normalize(token.text), folded once by the caller and shared with the marker pass; `marker_tag` is what that pass decided for this token, or None. The marker DECISION is entirely @@ -217,15 +217,23 @@ def _tags_for(token: WorkToken, n: str, state: ParseState, tags.add(SHAPE_ACRONYM_TAG) if state.policy.unlisted_dotted_suffixes: tags.add(AMBIGUOUS_ACRONYM_TAG) - elif (state.policy.unlisted_caps_suffixes is CapsSuffixes.EVERYWHERE + elif ((state.policy.unlisted_caps_suffixes is CapsSuffixes.EVERYWHERE + or comma_run) and token.role is None + # a listed member is no candidate (the predicate excludes + # every wordlist); asked first, in C, so 'John Smith, MA' + # pays no frame for the comma part's own credential + and n not in lex.suffix_acronyms_ambiguous and caps_shape_candidate(token.text, lex, state.policy, one_case)): - # #516's all-caps half, OPT-IN at this slot: an unlisted - # word written in capitals inside a mixed-case name. Only - # EVERYWHERE tags it -- the comma position the default - # reads is decided in `segment` from the text alone (#564), - # and the tag is what the comma-less trailing slot reads. + # #516's all-caps half: an unlisted word written in capitals + # inside a mixed-case name. EVERYWHERE tags it in every + # slot; the default tags it only in the part a suffix comma + # opened (`comma_run`), the position `segment` reads from + # the text alone (#564), so the reading there carries the + # same marks under either setting -- case repair keeps + # 'XYZ' in 'John Smith, XYZ' (rules.md#R4) rather than + # title-casing a credential the default admitted. # The policy conjunct comes FIRST and stays a plain # attribute read, so `caps_shape_candidate` is never CALLED # at the default and sharing its body costs it nothing @@ -268,11 +276,20 @@ def classify(state: ParseState) -> ParseState: # per token, and the fork and its emitter then agree with the case # class they consult. No extra frame -- it is one more boolean in a # comprehension that already walks every token. + # #564: the part a suffix comma opened, where the default admits + # the caps shape; empty for any other structure or setting. + comma_run: frozenset[int] = ( + frozenset(state.segments[1]) + if (state.structure is Structure.SUFFIX_COMMA + and len(state.segments) > 1 + and state.policy.unlisted_caps_suffixes is CapsSuffixes.AFTER_COMMA) + else frozenset()) tokens = tuple( copy_with( t, tags=_tags_for(t, folded[i], state, marker_tags.get(i), one_case_own=one_case and i < clause_at - and t.role is None, one_case=one_case)) + and t.role is None, one_case=one_case, + comma_run=i in comma_run)) for i, t in enumerate(state.tokens)) # Delimited content whose vocabulary cannot settle it: extract's # escape sends an UNambiguous suffix straight through ("(MBA)" -> diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 5eb42562..9c81c264 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -228,9 +228,9 @@ def class_run(seg: tuple[int, ...]) -> bool: # 'LEED AP' is two separate all-caps words, not one glued acronym # -- and the one membership test in this class that NEEDS the case # fact to answer membership at all: 'XYZ' is only credential-shaped - # where the name contrasts it. Gated on the switch (default off, so - # a non-candidate comma name never enters this branch) and tried - # only where the single-token test above already declined. + # where the name contrasts it. Gated on the setting (on by default + # since #564, behind the cheap conjuncts below) and tried only + # where the single-token test above already declined. # # `caps_shape_candidate` directly, not the un-narrowed # `ambiguous_class_candidate`: the run is a property of the CAPS @@ -251,10 +251,13 @@ def class_run(seg: tuple[int, ...]) -> bool: # than a second walk that could only reach the same answer (a # quality-review finding: the walk was provably redundant). # - # #564: on by default (`CapsSuffixes.AFTER_COMMA`), so two cheap - # C-level conjuncts go before the call: the first word must be - # written in capitals at all, which every word of the run must - # be, and a LONE two-letter word is declined -- it is how a + # #564: on by default (`CapsSuffixes.AFTER_COMMA`), so three cheap + # C-level conjuncts go before the call: two or more words before + # the comma, a necessary condition for the two NAME words the flip + # needs (as at the run test below -- 'Smith, JOHN', the commonest + # record format, never reaches the call); the first word written + # in capitals at all, which every word of the run must be; and a + # LONE two-letter word declined -- it is how a # person's initials are written, and two words before the comma # may be one surname ('García Márquez, MJ'), the case #563 decides # for the dotted 'M.J.' (rules.md#C1). A run holding a longer word @@ -262,6 +265,7 @@ def class_run(seg: tuple[int, ...]) -> bool: first = state.tokens[groups[1][0]].text if groups[1] else "" if (not candidate and state.policy.unlisted_caps_suffixes is not CapsSuffixes.OFF + and len(groups[0]) >= 2 and first.isupper() and not (len(groups[1]) == 1 and len(first) < 3) and all(caps_shape_candidate(state.tokens[i].text, @@ -318,6 +322,8 @@ def class_run(seg: tuple[int, ...]) -> bool: # 'MD MD ... G.J. G.J. ...' quadratic. unspoken_pair = False any_listed = False + caps_member = False + caps_on = state.policy.unlisted_caps_suffixes is not CapsSuffixes.OFF shaped = 0 pairs = 0 lexicon = state.lexicon @@ -328,15 +334,30 @@ def class_run(seg: tuple[int, ...]) -> bool: is_member = (fold == "member" or (fold == "ask" and ambiguous_class_candidate( text, lexicon, state.policy))) + # #564: an unlisted all-caps word is a member BY SHAPE too, + # so a run mixing it with listed credentials ('PhD XYZ', + # 'XYZ Jr.') reads as the all-caps run alone already does, + # rather than more evidence for a credential producing a + # name reading. `isupper()` first, in C; the predicate + # declines every listed word, so it never re-admits one. + caps = (not is_member and caps_on and text.isupper() + and caps_shape_candidate(text, lexicon, state.policy, + one_case=False)) + if caps: + is_member = caps_member = True if is_member: - # LISTED is exactly "no period" for a member, as at the - # single-token test above - listed = "." not in text + # LISTED is exactly "no period" for a listed or dotted + # member, as at the single-token test above; a caps + # member is by shape + listed = "." not in text and not caps if listed: any_listed = True else: shaped += 1 - if is_paired_initials(text): + # two capitals are paired initials undotted, read + # as #563 reads 'M.J.': 'García Márquez, MJ PhD' + # keeps given 'MJ' as 'De La Cruz, M.J. PhD' does + if is_paired_initials(text) or (caps and len(text) == 2): pairs += 1 if pairs == 1 and all( _normalize(w) in lexicon.titles @@ -392,8 +413,14 @@ def class_run(seg: tuple[int, ...]) -> bool: and not (settled and case_class() is False and all(ambiguous_lean(t, False) == "credential" - for t in members))) - flip_reports = candidate and (any_listed or pair_only) + for t in members)) + # the caps shape needs the contrast, as the + # all-caps run above does: one-case input + # leans nothing + and not (caps_member and case_class() is not False)) + # a caps member reports its flip as the all-caps run does + flip_reports = candidate and (any_listed or pair_only + or caps_member) # Computed only where `candidate` is true, alongside `case_class()` # -- the same lazy gate: a non-candidate comma name never counts # its pre-comma words either. Hoisted to a local because the diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 516537c3..364b4b04 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2813,12 +2813,14 @@ def _check_cjk_shape_purity(self) -> None: "the same one the default policy raises. Folding " "`caps_shape_candidate` into the class run would make " "this row fail"), - # #516's all-caps half, and it is OPT-IN. The rows come in pairs: - # the same name under the DEFAULT policy, where nothing moves and - # nothing is reported, and under Policy(unlisted_caps_suffixes= - # True), where the shape reads. The French and Korean names are - # why the default is off (mechanisms.md#VOCABULARY-EXERCISES-FORKS - # -- each pair pins the switch, not the words). A row that sets + # #516's all-caps half. Its trailing slots are OPT-IN + # (CapsSuffixes.EVERYWHERE); since #564 the default reads the part + # after a comma behind two name words, which the all-caps surname + # convention never writes in. The rows come in pairs: the same name + # under the DEFAULT policy and under EVERYWHERE. The French and + # Korean names are why the trailing slots stay off + # (mechanisms.md#VOCABULARY-EXERCISES-FORKS -- each pair pins the + # setting, not the words). A row that sets # the non-default policy carries no `shape=` tag: the contract # corpus (build_shapes_corpus.py) keys only on (shape, text), with # no policy of its own, so admitting one of these texts under a @@ -2955,14 +2957,10 @@ def _check_cjk_shape_purity(self) -> None: notes="the first time this class reaches the comma form as " "MORE than one token: 'LEED' and 'AP' are two separate " "all-caps words, and every token in the run must be a " - "candidate for the run itself to be one -- rules.md#C1's " - "`deviates: #291` line comes true ONLY under this " - "switch. At the DEFAULT this exact text still reads " - "given 'LEED', middle 'AP', family 'John Smith' -- the " - "deviation stands there unchanged -- so nothing in " - "this arc may remove rules.md#C1's `deviates: #291` " - "marker on this switch's account; the switch only " - "narrows what makes the deviation true"), + "candidate for the run itself to be one. Since #564 the " + "DEFAULT reads it the same way, the comma position being " + "on by default, and rules.md#C1's `deviates: #291` " + "marker came off with that change"), Case("the_caps_comma_multi_word_run_declines_at_one_word", "Smith, LEED AP", {"given": "LEED", "family": "Smith", "suffix": "AP"}, @@ -3137,25 +3135,31 @@ def _check_cjk_shape_purity(self) -> None: "the case fact ONCE, so a one-case name declines the " "whole run rather than per token"), Case("caps_run_needs_every_token_not_any", - "John Smith, LEED BA", - {"given": "LEED", "family": "John Smith", "suffix": "BA"}, + "John Smith, LEED Jones", + {"given": "LEED", "middle": "Jones", "family": "John Smith"}, policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE), - classification="fix(#531)", - ambiguities=("suffix-or-name", "suffix-or-name"), - notes="pins `all()` rather than `any()`: 'BA' is a LISTED " - "ambiguous acronym, so the caps shape test excludes it " - "and the run is no caps run -- the structure stays the " - "listing form. The first report is assign's post-comma " - "one, fired on 'LEED' alone, which carries the shape " - "tag from classify whatever segment made of the run. " - "#531 moves 'BA' itself: this is a family-comma name " - "whose given part now ends in a class member, and the " - "member is written in CAPITALS in a name written in " - "more than one case, so it leans credential (#289) and " - "reads suffix where it read middle -- its own second " - "report. What #516 pins here is untouched: the caps " - "run test still declines, and the structure is still " - "the listing form"), + classification="fix(#516)", + ambiguities=("suffix-or-name",), + notes="pins `all()` rather than `any()` in the all-caps run " + "test: a name word in the part leaves it no caps run, so " + "the structure stays the listing form, and no other " + "route reaches it (the #544 run test breaks on the name " + "word). The report is assign's post-comma one, fired on " + "'LEED', which carries the shape tag from classify. Until " + "#564 this row used 'John Smith, LEED BA', whose listed " + "'BA' the run test now reads beside a caps member as one " + "credential run (the row below)"), + Case("a_caps_word_and_a_listed_credential_make_one_run", + "John Smith, LEED BA", + {"given": "John", "family": "Smith", "suffix": "LEED BA"}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="#564: an unlisted all-caps word is a member of C1's run " + "by shape, so beside a listed credential it reads as the " + "run it is -- as 'John Smith, PhD XYZ' and 'John Smith, " + "XYZ Jr.' do. Until #564 the listing form read given " + "'LEED' with the switch on, and more evidence for a " + "credential produced a name reading"), Case("the_comma_count_counts_names_not_words_behind_a_title", "Mr Smith, Ma", {"given": "Ma", "family": "Mr Smith"}, diff --git a/tests/v2/pipeline/test_classify.py b/tests/v2/pipeline/test_classify.py index 6af86d2d..e201c317 100644 --- a/tests/v2/pipeline/test_classify.py +++ b/tests/v2/pipeline/test_classify.py @@ -572,7 +572,8 @@ def test_delimited_content_never_joins_the_shape_class() -> None: ("Xyz.", Policy(), None), ("MA", Policy(), None), ("田.中.", Policy(), None), - # #516's caps half: OFF is silent (matches the tag either way, + # #516's caps half: the default is silent here, outside a suffix + # comma's part (#564), and so is OFF (matches the tag either way, # since the token never carries the fact); ON needs the REAL # `one_case` fact, which these rows hand to `caps_shape_candidate` # -- the predicate classify calls, and since the review round the @@ -680,9 +681,10 @@ def test_the_caps_branch_reads_the_name_level_case_not_the_own_span( def reverted(token: WorkToken, n: str, state: ParseState, marker_tag: str | None, one_case_own: bool, - one_case: bool) -> frozenset[str]: + one_case: bool, comma_run: bool = False) -> frozenset[str]: return real(token, n, state, marker_tag, - one_case_own=one_case_own, one_case=one_case_own) + one_case_own=one_case_own, one_case=one_case_own, + comma_run=comma_run) monkeypatch.setattr(_classify_module, "_tags_for", reverted) broken = Parser(policy=on).parse("née JONES XYZ") @@ -692,10 +694,11 @@ def reverted(token: WorkToken, n: str, state: ParseState, def test_the_caps_shape_is_silent_until_its_switch_is_on() -> None: - # OFF (the default) emits NOTHING -- not the membership tag and - # not the shape tag either, which is where this half differs from - # the dotted one: there is no fork to report while a caller has - # not asked for the reading (#516). + # The default emits NOTHING at the trailing slot -- not the + # membership tag and not the shape tag either, which is where this + # half differs from the dotted one: there is no fork to report + # while a caller has not asked for the reading there (#516; the + # default reads only a suffix comma's part since #564). assert _tags_by_text("John Smith XYZ")["XYZ"] == frozenset() on = Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE) tags = _tags_by_text("John Smith XYZ", policy=on)["XYZ"] diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 0bf3420a..b8e2f805 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1427,10 +1427,13 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: # name word before the comma keeps the listing form, chain or not. "fix(#562) a comma credential run holding a particle chain reads by the name-word count": ("John Smith, PhD DO", "Smith, PhD DO DO", "John Smith, PhD MA"), - # Policy.unlisted_caps_suffixes is OFF by default, so its whole - # population is a probe here: the corpora run at the default, and - # a rule of this arc reaching one of these names would mean the - # DEFAULT changed. 'Mr XXX' is the boundary from the other side -- + # The caps shape's trailing slots are behind CapsSuffixes.EVERYWHERE, + # not the default, so their population is a probe here: the corpora + # run at the default, and a rule of this arc reaching one of these + # names would mean the default reached a trailing slot. 'John Smith, + # LEED AP' moves at the default since #564, but at the comma and + # under fix(#564), which is why THIS rule must still not reach it. + # 'Mr XXX' is the boundary from the other side -- # one piece behind a title never reaches the peel's two-piece # floor, switch or no switch. "fix(#289/#516) the ambiguous credential class reports at slots " @@ -3083,9 +3086,11 @@ class _LatinCopy(NamedTuple): frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), # 2026-10-01, #564: the comma caps rule, literal names that copy no # set, at 1.4.0 and in the 2.x copy. - frozenset({"John Smith, LEED AP", "John Smith, XYZ", + frozenset({"García Márquez, MJ JK", "John Smith, LEED AP", + "John Smith, PhD XYZ", "John Smith, XYZ", "The Rt Hon Kenneth Clarke QC MP, HMG"}), - frozenset({"Ahmad Jayadi, CHA", "John Smith, LEED AP", + frozenset({"Ahmad Jayadi, CHA", "García Márquez, MJ JK", + "John Smith, LEED AP", "John Smith, PhD XYZ", "John Smith, RAI", "John Smith, XYZ", "The Rt Hon Kenneth Clarke QC MP, HMG"}), # 2026-10-01, #575: rules.md#C1's particle-surname examples at @@ -3806,10 +3811,11 @@ def _claim(rule: dict) -> _Claim: # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - # 2026-10-01, #564: 423 -> 426, 'John Smith, XYZ', 'Smith, - # XYZ' and 'García Márquez, MJ', #564's rules.md#C1 examples. - # Reach, verified name by name. - _Claim(426, ('given', 'suffix', 'title'), "684dd90b64f3", None), + # 2026-10-01, #564: 423 -> 429, 'John Smith, XYZ', 'Smith, + # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', + # 'García Márquez, MJ JK' and 'John Smith, PhD XYZ', #564's + # rules.md#C1 examples. Reach, verified name by name. + _Claim(429, ('given', 'suffix', 'title'), "2c336f3d3ae0", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3913,18 +3919,20 @@ def _claim(rule: dict) -> _Claim: # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - # 2026-10-01, #564: 423 -> 426, 'John Smith, XYZ', 'Smith, - # XYZ' and 'García Márquez, MJ', #564's rules.md#C1 examples. - # Reach, verified name by name. - _Claim(426, ('family', 'given'), "684dd90b64f3", None), + # 2026-10-01, #564: 423 -> 429, 'John Smith, XYZ', 'Smith, + # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', + # 'García Márquez, MJ JK' and 'John Smith, PhD XYZ', #564's + # rules.md#C1 examples. Reach, verified name by name. + _Claim(429, ('family', 'given'), "2c336f3d3ae0", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": _Claim(4, ('family', 'given', 'suffix', 'title'), "30163564e03d", None), - # 2026-10-01, #564: new, 3; 'John Smith, LEED AP', 'John Smith, - # XYZ', 'The Rt Hon Kenneth Clarke QC MP, HMG'. + # 2026-10-01, #564: new, 5; 'García Márquez, MJ JK', 'John + # Smith, LEED AP', 'John Smith, PhD XYZ', 'John Smith, XYZ', + # 'The Rt Hon Kenneth Clarke QC MP, HMG'. "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": - _Claim(3, ('family', 'given', 'middle', 'suffix', 'title'), "1926a01dc71c", ('DEFAULT',)), + _Claim(5, ('family', 'given', 'middle', 'suffix', 'title'), "0de364bee3e3", ('DEFAULT',)), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -5023,11 +5031,12 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), - # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, - # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon - # Kenneth Clarke QC MP, HMG'. + # 2026-10-01, #564: new, 7; 'Ahmad Jayadi, CHA', 'García + # Márquez, MJ JK', 'John Smith, LEED AP', 'John Smith, PhD XYZ', + # 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon Kenneth + # Clarke QC MP, HMG'. "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "66be5f479c54", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -5532,11 +5541,12 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), - # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, - # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon - # Kenneth Clarke QC MP, HMG'. + # 2026-10-01, #564: new, 7; 'Ahmad Jayadi, CHA', 'García + # Márquez, MJ JK', 'John Smith, LEED AP', 'John Smith, PhD XYZ', + # 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon Kenneth + # Clarke QC MP, HMG'. "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "66be5f479c54", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6185,11 +6195,12 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), - # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, - # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon - # Kenneth Clarke QC MP, HMG'. + # 2026-10-01, #564: new, 7; 'Ahmad Jayadi, CHA', 'García + # Márquez, MJ JK', 'John Smith, LEED AP', 'John Smith, PhD XYZ', + # 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon Kenneth + # Clarke QC MP, HMG'. "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "66be5f479c54", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the @@ -6558,11 +6569,12 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #575: new, 1; 'van der Berg, PhD'. "fix(#575) a particle surname before a comma is one name word": _Claim(1, ('_ambiguities', 'family', 'given'), "d449a9b43779", ('DEFAULT',)), - # 2026-10-01, #564: new, 5; 'Ahmad Jayadi, CHA', 'John Smith, - # LEED AP', 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon - # Kenneth Clarke QC MP, HMG'. + # 2026-10-01, #564: new, 7; 'Ahmad Jayadi, CHA', 'García + # Márquez, MJ JK', 'John Smith, LEED AP', 'John Smith, PhD XYZ', + # 'John Smith, RAI', 'John Smith, XYZ', 'The Rt Hon Kenneth + # Clarke QC MP, HMG'. "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "b17ff67704b1", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "66be5f479c54", ('DEFAULT',)), # #516's alternation. Literal-anchored to the by-shape movers, # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index e2c626e1..f684c70c 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -865,18 +865,25 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # they are reported here, dated, so a reader who turns the switch # on knows what it costs and a later re-measurement does not read # as a silent drift. + # + # 2026-10-01, #564: the default is now AFTER_COMMA, which reads the + # comma position, so `default` below is no longer "off" and these + # assertions are about the TRAILING slot EVERYWHERE adds. Default + # costs, same harness against master: 'Smith, John' 183 -> 183, + # 'Smith, XYZ' 182 -> 182 (the comma test needs two words before + # the comma), 'John Smith, XYZ' 251 -> 258 (decisions.md#S2). from nameparser import Parser on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) - off = Parser() + default = Parser() # what it buys assert on.parse("John Smith XYZ").suffix == "XYZ" assert on.parse("John Smith, XYZ").suffix == "XYZ" - # what it costs, and why the default is off + # what it costs, and why the trailing slot is not the default assert on.parse("Jean Pierre DUPONT").suffix == "DUPONT" - assert off.parse("Jean Pierre DUPONT").family == "DUPONT" - # the default emits nothing at all - assert off.parse("John Smith XYZ").ambiguities == () + assert default.parse("Jean Pierre DUPONT").family == "DUPONT" + # the default emits nothing at the trailing slot + assert default.parse("John Smith XYZ").ambiguities == () # the boundaries: one case, one letter, a digit, and vocabulary. # 'John Smith X' is NOT the single-letter control -- 'X' is a bare # roman numeral (rules.md#S2's numeral fork) and reads as suffix @@ -885,6 +892,6 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # is the actual boundary (a single capital never satisfies the # `len(text) >= 2` half of the shape test, on or off). assert on.parse("JOHN SMITH XYZ").family == "XYZ" - assert on.parse("John Smith Z").family == off.parse("John Smith Z").family + assert on.parse("John Smith Z").family == default.parse("John Smith Z").family assert on.parse("John Smith XY2").family == "XY2" assert on.parse("John Smith MC").suffix == "MC" diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index d63657d9..391fb8e9 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -95,6 +95,8 @@ "García Márquez, G.J." "García Márquez, G.J.R." "García Márquez, MJ" +"García Márquez, MJ JK" +"García Márquez, MJ PhD" "García Márquez, Ms G.J." "Hans „Erster“ und “Zweiter” Müller" "Hassan Mohamad Ali" @@ -225,6 +227,7 @@ "John Smith, PhD DO DO" "John Smith, PhD MEng" "John Smith, PhD X.Y." +"John Smith, PhD XYZ" "John Smith, PhD vd DO" "John Smith, V." "John Smith, X.Y. P.Q." diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 3331f0da..ce6e0269 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4866,10 +4866,13 @@ issue = "fix(#564) an unlisted all-caps word after a comma behind two name words # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); # 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. +# 'John Smith, PhD XYZ' is a caps word in C1's run beside a listed +# credential, and 'García Márquez, MJ JK' two words of two capitals, +# read as #563 reads two pairs (both C1 examples). # # Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, # MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the # trailing position stays opt-in) are _MUST_NOT_MATCH. -name_regex = "^(?:John Smith, LEED AP|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +name_regex = "^(?:García Márquez, MJ JK|John Smith, LEED AP|John Smith, PhD XYZ|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" fields = ["title", "given", "middle", "family", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index ad138dda..3d6e5a57 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3923,13 +3923,16 @@ issue = "fix(#564) an unlisted all-caps word after a comma behind two name words # reports the call. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); -# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. +# 'John Smith, PhD XYZ' is a caps word in C1's run beside a listed +# credential, and 'García Márquez, MJ JK' two words of two capitals, +# read as #563 reads two pairs (both C1 examples). 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read # suffix again, as this baseline read them by vocabulary before #342 # removed both words, and differ only by the comma's report. # # Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, # MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the # trailing position stays opt-in) are _MUST_NOT_MATCH. -name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +name_regex = "^(?:Ahmad Jayadi, CHA|García Márquez, MJ JK|John Smith, LEED AP|John Smith, PhD XYZ|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index c2af9221..55819f32 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3834,13 +3834,16 @@ issue = "fix(#564) an unlisted all-caps word after a comma behind two name words # reports the call. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); -# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. +# 'John Smith, PhD XYZ' is a caps word in C1's run beside a listed +# credential, and 'García Márquez, MJ JK' two words of two capitals, +# read as #563 reads two pairs (both C1 examples). 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read # suffix again, as this baseline read them by vocabulary before #342 # removed both words, and differ only by the comma's report. # # Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, # MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the # trailing position stays opt-in) are _MUST_NOT_MATCH. -name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +name_regex = "^(?:Ahmad Jayadi, CHA|García Márquez, MJ JK|John Smith, LEED AP|John Smith, PhD XYZ|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 43dfe0c6..3ed62be6 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2217,13 +2217,16 @@ issue = "fix(#564) an unlisted all-caps word after a comma behind two name words # reports the call. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); -# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. +# 'John Smith, PhD XYZ' is a caps word in C1's run beside a listed +# credential, and 'García Márquez, MJ JK' two words of two capitals, +# read as #563 reads two pairs (both C1 examples). 'John Smith, RAI' and 'Ahmad Jayadi, CHA' read # suffix again, as this baseline read them by vocabulary before #342 # removed both words, and differ only by the comma's report. # # Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, # MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the # trailing position stays opt-in) are _MUST_NOT_MATCH. -name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +name_regex = "^(?:Ahmad Jayadi, CHA|García Márquez, MJ JK|John Smith, LEED AP|John Smith, PhD XYZ|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index d3645efc..8f197df2 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1472,13 +1472,16 @@ issue = "fix(#564) an unlisted all-caps word after a comma behind two name words # reports the call. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); -# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. 'John Smith, RAI' and 'Ahmad Jayadi, CHA' move +# 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. +# 'John Smith, PhD XYZ' is a caps word in C1's run beside a listed +# credential, and 'García Márquez, MJ JK' two words of two capitals, +# read as #563 reads two pairs (both C1 examples). 'John Smith, RAI' and 'Ahmad Jayadi, CHA' move # from this baseline's family comma (given 'RAI'/'CHA', the words # removed from the vocabulary by #342) to the credential run. # # Literal; the probes 'Smith, XYZ' (one name word), 'García Márquez, # MJ' (a lone two-letter word) and 'John Smith XYZ' (no comma: the # trailing position stays opt-in) are _MUST_NOT_MATCH. -name_regex = "^(?:Ahmad Jayadi, CHA|John Smith, LEED AP|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" +name_regex = "^(?:Ahmad Jayadi, CHA|García Márquez, MJ JK|John Smith, LEED AP|John Smith, PhD XYZ|John Smith, RAI|John Smith, XYZ|The Rt Hon Kenneth Clarke QC MP, HMG)$" fields = ["title", "given", "middle", "family", "suffix", "_ambiguities"] orders = ["DEFAULT"] From 38d41894ff298a0b633c8a63d51b5e8f5765f1d3 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 22:26:34 -0700 Subject: [PATCH 03/13] fix(S2): #564 second review -- the name's contrast, repair guard, costs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - An all-caps word joins C1's run only where the NAME before the comma has a lowercase letter: the first fix let a mixed-case credential supply the contrast, so 'LLOYD WEBBER, ANDREW PhD' lost its given name. 'García Márquez, JUAN Jr.' (a mixed-case name) stays in the accepted cost, now named with a case row. - rules.md#C1 points two capitals at #563's own sentences instead of restating them; the restatement contradicted them. Measured by the review: MJ and M.J. flip identically across 4036 inputs. - rules.md#R4 names the caps shape; a forced-repair test guards the classify tagging (an R4 example could not: R5 leaves mixed case unrepaired), with its negative control run. - Every caller of the caps predicate asks, in C, what it would decline anyway, so a listed credential pays no frame ('John Smith, CPA' and 'John Smith, Ph. D.' are back to master's counts). - decisions.md#S2, the release log and mechanisms.md corrected to match, including the third-comma-part render and the cost figures. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 6 +++--- docs/design/mechanisms.md | 2 +- docs/design/rules.md | 16 ++++++++++------ docs/release_log.rst | 2 +- nameparser/_pipeline/_classify.py | 16 +++++++++------- nameparser/_pipeline/_segment.py | 27 +++++++++++++++++++++------ tests/v2/cases.py | 20 ++++++++++++++++++++ tests/v2/test_render.py | 18 ++++++++++++++++++ 8 files changed, 83 insertions(+), 24 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index d93ce542..7742c056 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read as #563 reads a pair: with only a credential behind it, it is still the given name (`García Márquez, MJ PhD`, as `De La Cruz, M.J. PhD`), while a credential in front or a second such word makes the run (`García Márquez, PhD MJ`, `García Márquez, MJ JK`). The first draft's rules.md and release-log wording said "three letters or more, or a run holding such a word", which the code never did; both reviews caught it. (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. Mixed case is still required, as for the all-caps run. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured by swapping `MJ` ↔ `M.J.` across 4036 inputs, the flip decision never differed (the fix-commit review) — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's: a lowercase letter in the part before the comma. The first fix asked the whole name's case, and a mixed-case credential supplied the contrast itself, so an all-caps record lost its given name — `LLOYD WEBBER, ANDREW PhD` read suffix 'ANDREW PhD' and `GARCÍA MÁRQUEZ, GABRIEL Jr.` suffix 'GABRIEL Jr.' (the fix-commit review); both keep the given name now. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. - CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render alike. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 258. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Classify skips a listed member before calling the caps predicate. The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 258. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 66f82e37..3a061b0e 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -55,7 +55,7 @@ Problem shape. "Which stage does X?" — asked before attributing behavior in pr ## ONE-PREDICATE-PER-QUESTION — one predicate answers it, and every other site calls that -Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from the sites that each needed the identical question answered — classify's own tag emission and _segment.py's two comma tests (the all-caps run and, since #564, the #544 run test's by-shape member), with its unit tests — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: the trailing-position caller's first conjunct is the setting itself (`CapsSuffixes.EVERYWHERE`, not the default), and the comma run test — on by default since #564 — asks a C-level `isupper()` of the part's first word before calling, so a parse with no all-caps word after a comma never reaches it (decisions.md#S2, #C1)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. +Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from the sites that each needed the identical question answered — classify's own tag emission and _segment.py's two comma tests (the all-caps run and, since #564, the #544 run test's by-shape member), with its unit tests — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: every caller asks first, in C, what the predicate would decline anyway — the setting and position, then alphabetic capitals and no listed suffix word — so only an unlisted all-caps word reaches the call, at the default as under any setting (decisions.md#S2)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. ## RENDER-HONORS-THE-PARSE — the parse decides it, the views honor it diff --git a/docs/design/rules.md b/docs/design/rules.md index 73d258ea..0b31acab 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1723,10 +1723,13 @@ C1. Rationale: a credential run after the comma means the name is in not the only shape word in the part are a credential however little else speaks for them, since no one writes a person's initials as two dotted groups. An unlisted word of two capitals - after the comma is that shape undotted and is read as paired - initials are: alone it is the given name ('García Márquez, MJ'), - and so it is with only a credential behind it, while a credential - in front of it or a second such word makes the run. The same count reads a part of two or + after the comma is that shape undotted and is read exactly as + paired initials are, by the sentences above ('García Márquez, MJ' + and 'García Márquez, MJ PhD' keep given 'MJ'). An unlisted + all-caps word joins the class in such a part only where the name + before the comma is written with a lowercase letter: the contrast + is the name's, so a record written wholly in capitals keeps its + given name beside a credential written in mixed case. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any @@ -2548,8 +2551,9 @@ R4. Rationale: case repair is a display concern, applied only on mark one, R5 defers to it. A credential acronym the exceptions map does not carry is an initialism, so a single-case word the parse put in the suffix role from the acronym vocabulary, or read as a - credential by its dotted shape alone (S3), repairs to its - all-caps spelling rather than a title-cased one, and that repair + credential by its dotted shape alone (S3) or by its capitals + (S2), repairs to its all-caps spelling rather than a title-cased + one, and that repair outranks the Mac/Mc convention where a word fits both (MCSE, not McSe). A roman numeral the parse put in the suffix role is written in capitals the way a generation is written, whether or diff --git a/docs/release_log.rst b/docs/release_log.rst index 1719540b..5e661771 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), and a two-letter word reads as initials do: alone, or with only a credential behind it, it is the given name (``García Márquez, MJ``, ``García Márquez, MJ PhD``). ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and a record written wholly in capitals keeps its given name beside a mixed-case credential (``LLOYD WEBBER, ANDREW PhD``). ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/nameparser/_pipeline/_classify.py b/nameparser/_pipeline/_classify.py index 3c644782..a08caa79 100644 --- a/nameparser/_pipeline/_classify.py +++ b/nameparser/_pipeline/_classify.py @@ -220,10 +220,11 @@ def _tags_for(token: WorkToken, n: str, state: ParseState, elif ((state.policy.unlisted_caps_suffixes is CapsSuffixes.EVERYWHERE or comma_run) and token.role is None - # a listed member is no candidate (the predicate excludes - # every wordlist); asked first, in C, so 'John Smith, MA' - # pays no frame for the comma part's own credential + # what the predicate would decline anyway, asked first in + # C: a listed member ('John Smith, MA'), and a word that + # is not alphabetic capitals ('John Smith, Ph. D.') and n not in lex.suffix_acronyms_ambiguous + and token.text.isalpha() and token.text.isupper() and caps_shape_candidate(token.text, lex, state.policy, one_case)): # #516's all-caps half: an unlisted word written in capitals @@ -234,10 +235,11 @@ def _tags_for(token: WorkToken, n: str, state: ParseState, # same marks under either setting -- case repair keeps # 'XYZ' in 'John Smith, XYZ' (rules.md#R4) rather than # title-casing a credential the default admitted. - # The policy conjunct comes FIRST and stays a plain - # attribute read, so `caps_shape_candidate` is never CALLED - # at the default and sharing its body costs it nothing - # (that is why this half is a call where the dotted branch + # The setting and the position come FIRST, then C-level + # checks of what the predicate would decline, so outside + # a suffix comma's part the default never calls it and + # inside one it calls it only for an unlisted all-caps + # word (that is why this half is a call where the dotted branch # above stays inline: the dotted caller has no such cheap # first conjunct to hide behind). The predicate's own # docstring carries the whole-vocabulary roster and what diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 9c81c264..86ef4f72 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -266,7 +266,9 @@ def class_run(seg: tuple[int, ...]) -> bool: if (not candidate and state.policy.unlisted_caps_suffixes is not CapsSuffixes.OFF and len(groups[0]) >= 2 - and first.isupper() + and first.isalpha() and first.isupper() + and first.lower() not in state.lexicon.suffix_acronyms + and first.lower() not in state.lexicon.suffix_words and not (len(groups[1]) == 1 and len(first) < 3) and all(caps_shape_candidate(state.tokens[i].text, state.lexicon, state.policy, @@ -340,7 +342,15 @@ def class_run(seg: tuple[int, ...]) -> bool: # rather than more evidence for a credential producing a # name reading. `isupper()` first, in C; the predicate # declines every listed word, so it never re-admits one. - caps = (not is_member and caps_on and text.isupper() + # The C-level prechecks are exactly what the predicate + # would decline (it needs `isalpha()`/`isupper()` and + # excludes every wordlist; `lower()` is `_normalize` for an + # alphabetic word), so a listed credential ('MD', 'CPA') + # never pays for the call. + caps = (not is_member and caps_on and text.isalpha() + and text.isupper() + and text.lower() not in lexicon.suffix_acronyms + and text.lower() not in lexicon.suffix_words and caps_shape_candidate(text, lexicon, state.policy, one_case=False)) if caps: @@ -414,10 +424,15 @@ def class_run(seg: tuple[int, ...]) -> bool: and all(ambiguous_lean(t, False) == "credential" for t in members)) - # the caps shape needs the contrast, as the - # all-caps run above does: one-case input - # leans nothing - and not (caps_member and case_class() is not False)) + # the caps shape needs the contrast, and it has + # to come from the NAME: beside a mixed-case + # credential the credential's own lowercase + # would supply it, and an all-caps record + # ('LLOYD WEBBER, ANDREW PhD') would lose its + # given name to it + and not (caps_member and not any( + ch.islower() for i in groups[0] + for ch in state.tokens[i].text))) # a caps member reports its flip as the all-caps run does flip_reports = candidate and (any_listed or pair_only or caps_member) diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 364b4b04..d5692054 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2917,6 +2917,26 @@ def _check_cjk_shape_purity(self) -> None: "written undotted, and two words before the comma may " "be one surname -- the case #563 decides for 'M.J.'. No " "report: the comma position declines it outright"), + Case("an_all_caps_record_keeps_its_given_name_beside_a_mixed_credential", + "LLOYD WEBBER, ANDREW PhD", + {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "PhD"}, + notes="#564 boundary: an all-caps word joins C1's run only where " + "the NAME before the comma has a lowercase letter -- here " + "the only lowercase is the credential's own, and a record " + "written wholly in capitals keeps its given name. A " + "review-round draft let 'PhD' supply the contrast and " + "read suffix 'ANDREW PhD'"), + Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", + "García Márquez, JUAN Jr.", + {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="#564's accepted cost (Derek): a given name written in " + "capitals behind a two-word surname without particles, " + "in a mixed-case name, reads as the credential -- with a " + "generation behind it as alone ('García Márquez, " + "GABRIEL'). Reported, and CapsSuffixes.OFF keeps given " + "'JUAN'"), Case("the_caps_comma_count_needs_two_name_words", "John Smith, XYZ", {"given": "John", "family": "Smith", "suffix": "XYZ"}, diff --git a/tests/v2/test_render.py b/tests/v2/test_render.py index 023488ac..b92bed28 100644 --- a/tests/v2/test_render.py +++ b/tests/v2/test_render.py @@ -1520,3 +1520,21 @@ def test_initials_separator_is_honored_on_the_live_token_path() -> None: == "J. PD." assert HumanName("Ph. D., John", initials_separator="") \ .initials_list() == ["J", "PD"] + + +def test_a_caps_credential_after_a_comma_keeps_its_capitals_in_every_setting( +) -> None: + # A credential read by its capitals (S2) repairs to its all-caps + # spelling (R4 in rules.md). #564 made the comma reading a default + # decided in `segment` from the text, and classify's shape tag -- + # what repair reads -- has to follow it there too, or the default + # title-cases a credential EVERYWHERE keeps: the first draft gave + # 'John Smith Xyz' at the default and 'John Smith XYZ' under + # EVERYWHERE. Forced, since R5 leaves a mixed-case name alone. + for setting in (CapsSuffixes.AFTER_COMMA, CapsSuffixes.EVERYWHERE): + parser = Parser(policy=Policy(unlisted_caps_suffixes=setting)) + for text, rendered in (("John Smith, XYZ", "John Smith XYZ"), + ("John Smith, LEED AP", "John Smith LEED AP"), + ("John Smith, PhD XYZ", "John Smith PhD XYZ")): + name = parser.parse(text) + assert str(name.capitalized(force=True)) == rendered, (setting, text) From b26beff9f11085b95274cd68c3a293a04233d16e Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 22:43:58 -0700 Subject: [PATCH 04/13] fix(S2): #564 third review -- one name_contrast over the name's own words MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The second round's contrast test read every token before the comma, so a maiden clause ('LLOYD WEBBER née Smith, ANDREW PhD') or a title ('Mr LLOYD WEBBER, ANDREW') supplied the lowercase and an all-caps record lost its given name again; it also cost a frame per character on an all-caps record. name_contrast now reads the name's own words before the comma (own_words, shared with case_class's walk rather than a second spelling of it), titles and particles aside, in one C-level comparison, and both caps branches ask it. The all-caps run's later words get the same C-level prechecks as its first, so a listed credential never reaches the caps predicate. rules.md#C1 limits "exactly as paired initials" to mixed-case names and states the contrast as the name's own words; decisions.md#S2 records both wrong drafts, the 'De La Cruz García, MARÍA' accepted case, and the cost (all-caps record +24 constant, no longer per character). Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 +- docs/design/rules.md | 18 ++++---- nameparser/_pipeline/_segment.py | 71 +++++++++++++++++++++++--------- tests/v2/cases.py | 18 ++++++++ tests/v2/test_policy.py | 2 +- 5 files changed, 82 insertions(+), 31 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 7742c056..92f33831 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured by swapping `MJ` ↔ `M.J.` across 4036 inputs, the flip decision never differed (the fix-commit review) — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's: a lowercase letter in the part before the comma. The first fix asked the whole name's case, and a mixed-case credential supplied the contrast itself, so an all-caps record lost its given name — `LLOYD WEBBER, ANDREW PhD` read suffix 'ANDREW PhD' and `GARCÍA MÁRQUEZ, GABRIEL Jr.` suffix 'GABRIEL Jr.' (the fix-commit review); both keep the given name now. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured by swapping `MJ` ↔ `M.J.` across 4036 inputs whose name before the comma is written in mixed case, the flip decision never differed (the fix-commit review); in a one-case name `MJ` is no class member at all, the caps shape needing the contrast, while `M.J.` still is, so `GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run (the review of the next round) — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's: a lowercase letter among the name's own words before the comma (`own_words`, so no maiden clause or delimited content), titles and particles aside, which a record capitalizing its surname still writes in lowercase. Two drafts got this wrong. The first asked the whole name's case, so a mixed-case credential supplied the contrast itself and `LLOYD WEBBER, ANDREW PhD` read suffix 'ANDREW PhD'; the second asked every token before the comma, so a maiden clause supplied it and `LLOYD WEBBER née Smith, ANDREW PhD` lost its given name the same way, as did `Mr LLOYD WEBBER, ANDREW` through its title (the two fix-commit reviews). One `name_contrast` now answers it for both caps branches, reading the own-words walk `case_class` already makes rather than a second spelling of it. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 258. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 264, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 362, a constant +24 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 677; the second draft's per-character contrast test had made that +350). The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index 0b31acab..74708f83 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1722,14 +1722,16 @@ C1. Rationale: a credential run after the comma means the name is in there. Behind two or more name words, paired initials that are not the only shape word in the part are a credential however little else speaks for them, since no one writes a person's - initials as two dotted groups. An unlisted word of two capitals - after the comma is that shape undotted and is read exactly as - paired initials are, by the sentences above ('García Márquez, MJ' - and 'García Márquez, MJ PhD' keep given 'MJ'). An unlisted - all-caps word joins the class in such a part only where the name - before the comma is written with a lowercase letter: the contrast - is the name's, so a record written wholly in capitals keeps its - given name beside a credential written in mixed case. The same count reads a part of two or + initials as two dotted groups. In a name written in more than one + case, an unlisted word of two capitals after the comma is that + shape undotted and is read exactly as paired initials are, by the + sentences above ('García Márquez, MJ' and 'García Márquez, MJ PhD' + keep given 'MJ'). An unlisted all-caps word joins the class in + such a part only where the name's own words before the comma, + titles and particles aside, carry a lowercase letter: the contrast + is the name's, so a record that writes its surname in capitals + keeps its given name beside a mixed-case credential, a lowercase + title or particle, or a maiden clause. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 86ef4f72..438ab8d4 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -104,21 +104,52 @@ def segment(state: ParseState) -> ParseState: # field two sites would not otherwise share). one_case = state.one_case + # The own-words walk, once per parse: `case_class` reads the words + # and `name_contrast` (#564) the clause boundary. + # Inline in both rather than a third helper: a wrapper is a frame + # every case-forcing comma name ('John Smith, MA') would pay. + own_span: tuple[list[str], int] | None = None + def case_class() -> bool: - nonlocal one_case + nonlocal one_case, own_span if one_case is None: - # `own, _` rather than `[0]`: the second element is the - # maiden clause's start index, which this stage has no use - # for, and saying so by name is what stops a reader having - # to go and look up what a bare subscript dropped. - own, _ = own_words(state.tokens, state.comma_offsets, - state.lexicon.maiden_markers) - one_case = is_one_case(own) + if own_span is None: + own_span = own_words(state.tokens, state.comma_offsets, + state.lexicon.maiden_markers) + one_case = is_one_case(own_span[0]) return one_case def texts(seg: tuple[int, ...]) -> list[str]: return [state.tokens[i].text for i in seg] + # #564: the contrast the caps shape needs is the NAME's -- its own + # words before the comma (no maiden clause, no delimited content, + # as `own_words` defines them for `case_class` above), less titles + # and particles, which a record writing its surname in capitals + # still writes in lowercase ('Mr LLOYD WEBBER', 'de GAULLE'). + # Neither the credential's own lowercase nor a clause's may supply + # it, or an all-caps record loses its given name ('LLOYD WEBBER, + # ANDREW PhD', 'LLOYD WEBBER née Smith, ANDREW PhD'). Asked only + # once a caps word is in hand, and in one C-level comparison. + def name_contrast() -> bool: + # no lowercase before the comma at all: an all-caps record, + # settled in C before the walk + before = "".join([state.tokens[i].text for i in groups[0]]) + if before == before.upper(): + return False + nonlocal own_span + if own_span is None: + own_span = own_words(state.tokens, state.comma_offsets, + state.lexicon.maiden_markers) + clause_at = own_span[1] + lex = state.lexicon + words = "".join( + tok.text for i in groups[0] + if i < clause_at and (tok := state.tokens[i]).role is None + and (n := _normalize(tok.text)) not in lex.titles + and n not in lex.particles) + return words != words.upper() + # Inlined rather than built on `texts` (measured, #289/#516's # eager-gate fix round): every comma parse calls `suffixy` at # least once, and a `texts(seg)` indirection costs a SECOND frame @@ -270,11 +301,13 @@ def class_run(seg: tuple[int, ...]) -> bool: and first.lower() not in state.lexicon.suffix_acronyms and first.lower() not in state.lexicon.suffix_words and not (len(groups[1]) == 1 and len(first) < 3) - and all(caps_shape_candidate(state.tokens[i].text, - state.lexicon, state.policy, - one_case=False) + and all((t := state.tokens[i].text).isalpha() and t.isupper() + and t.lower() not in state.lexicon.suffix_acronyms + and t.lower() not in state.lexicon.suffix_words + and caps_shape_candidate(t, state.lexicon, state.policy, + one_case=False) for i in groups[1])): - candidate = case_class() is False + candidate = name_contrast() flip_reports = candidate # rules.md#C1: "The same count reads a part of two or more words as # the credential run when every word of it is a suffix word or a @@ -342,11 +375,11 @@ def class_run(seg: tuple[int, ...]) -> bool: # rather than more evidence for a credential producing a # name reading. `isupper()` first, in C; the predicate # declines every listed word, so it never re-admits one. - # The C-level prechecks are exactly what the predicate - # would decline (it needs `isalpha()`/`isupper()` and - # excludes every wordlist; `lower()` is `_normalize` for an - # alphabetic word), so a listed credential ('MD', 'CPA') - # never pays for the call. + # The C-level prechecks decline only what the predicate + # would decline too (it needs `isalpha()`/`isupper()` and + # excludes every wordlist, and a word whose `lower()` is a + # listed entry is one of those), so a listed credential + # ('MD', 'CPA') never pays for the call. caps = (not is_member and caps_on and text.isalpha() and text.isupper() and text.lower() not in lexicon.suffix_acronyms @@ -430,9 +463,7 @@ def class_run(seg: tuple[int, ...]) -> bool: # would supply it, and an all-caps record # ('LLOYD WEBBER, ANDREW PhD') would lose its # given name to it - and not (caps_member and not any( - ch.islower() for i in groups[0] - for ch in state.tokens[i].text))) + and not (caps_member and not name_contrast())) # a caps member reports its flip as the all-caps run does flip_reports = candidate and (any_listed or pair_only or caps_member) diff --git a/tests/v2/cases.py b/tests/v2/cases.py index d5692054..1b3ee055 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2926,6 +2926,24 @@ def _check_cjk_shape_purity(self) -> None: "written wholly in capitals keeps its given name. A " "review-round draft let 'PhD' supply the contrast and " "read suffix 'ANDREW PhD'"), + Case("a_maiden_clause_supplies_no_contrast_for_the_caps_shape", + "LLOYD WEBBER née Smith, ANDREW PhD", + {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "PhD", + "maiden": "Smith"}, + notes="#564 boundary: the contrast is the name's OWN words " + "(own_words): a maiden clause's lowercase is not the " + "name's, so the all-caps record keeps its given name. A " + "review-round draft counted every pre-comma token and " + "read suffix 'ANDREW PhD'"), + Case("a_title_supplies_no_contrast_for_the_caps_shape", + "Mr LLOYD WEBBER, ANDREW PhD", + {"given": "ANDREW", "family": "Mr LLOYD WEBBER", "suffix": "PhD"}, + notes="#564 boundary: titles and particles are written in " + "lowercase in a record that capitalizes its surname, so " + "neither counts as the name's contrast and the given name " + "stays. The title in the family is the listing form's " + "long-standing reading of a title before the comma " + "('Prof. Cruz, Ed'), master's too"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index f684c70c..808192d1 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -871,7 +871,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # assertions are about the TRAILING slot EVERYWHERE adds. Default # costs, same harness against master: 'Smith, John' 183 -> 183, # 'Smith, XYZ' 182 -> 182 (the comma test needs two words before - # the comma), 'John Smith, XYZ' 251 -> 258 (decisions.md#S2). + # the comma), 'John Smith, XYZ' 251 -> 264 (decisions.md#S2). from nameparser import Parser on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) From b783f668d117520ec483f07629b258df212e5ea8 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 22:54:58 -0700 Subject: [PATCH 05/13] fix(S2): #564 fourth review -- the contrast comes from unlisted name words MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The third round excluded titles and particles by list and missed a generation ('LLOYD WEBBER Jr., ANDREW', the common SURNAME Jr., GIVEN record), a connective ('GARCÍA y LÓPEZ, ANDREW') and a title recognized by shape ('Insp. LLOYD WEBBER, ANDREW'): each lost its given name. The test is inverted now: the contrast comes only from the name's own words before the comma that no wordlist claims and no period marks, through _vocab.in_any_wordlist -- the caps predicate's own "unlisted", lifted out so both share it. Case rows pin the generation, connective and later-comma-part shapes. rules.md#C1 states MJ = M.J. as what it is, the comma's decision where the name carries the contrast; decisions.md#S2 replaces the relabelled 4036 figure with a measurement on this tree (0 of 1295 with the contrast, 325 of 925 without, by design) and its recompute. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 ++-- docs/design/rules.md | 21 +++++++++++---------- nameparser/_pipeline/_segment.py | 24 ++++++++++++++---------- nameparser/_pipeline/_vocab.py | 16 +++++++++++++--- tests/v2/cases.py | 21 +++++++++++++++++++++ tests/v2/test_policy.py | 2 +- 6 files changed, 62 insertions(+), 26 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 92f33831..f1063b6e 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured by swapping `MJ` ↔ `M.J.` across 4036 inputs whose name before the comma is written in mixed case, the flip decision never differed (the fix-commit review); in a one-case name `MJ` is no class member at all, the caps shape needing the contrast, while `M.J.` still is, so `GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run (the review of the next round) — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's: a lowercase letter among the name's own words before the comma (`own_words`, so no maiden clause or delimited content), titles and particles aside, which a record capitalizing its surname still writes in lowercase. Two drafts got this wrong. The first asked the whole name's case, so a mixed-case credential supplied the contrast itself and `LLOYD WEBBER, ANDREW PhD` read suffix 'ANDREW PhD'; the second asked every token before the comma, so a maiden clause supplied it and `LLOYD WEBBER née Smith, ANDREW PhD` lost its given name the same way, as did `Mr LLOYD WEBBER, ANDREW` through its title (the two fix-commit reviews). One `name_contrast` now answers it for both caps branches, reading the own-words walk `case_class` already makes rather than a second spelling of it. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's: a lowercase letter among the name's own words before the comma (`own_words`, so no maiden clause or delimited content), titles and particles aside, which a record capitalizing its surname still writes in lowercase. Three drafts got this wrong, each by an exclusion too narrow. The first asked the whole name's case, so a mixed-case credential supplied the contrast and `LLOYD WEBBER, ANDREW PhD` read suffix 'ANDREW PhD'; the second asked every token before the comma, so a maiden clause or a title supplied it (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`); the third excluded titles and particles by list and missed a generation (`LLOYD WEBBER Jr., ANDREW`, the common `SURNAME Jr., GIVEN` format), a connective (`GARCÍA y LÓPEZ, ANDREW`) and a title recognized by shape (`Insp. LLOYD WEBBER, ANDREW`) (the three fix-commit reviews). The shipped test inverts it: the contrast comes only from the name's own words that no wordlist claims (`_vocab.in_any_wordlist`, the caps predicate's own "unlisted" lifted out so the two share it) and no period marks. One `name_contrast` answers it for both caps branches, reading the own-words walk `case_class` makes. It reads only the part before the comma, so a later comma part no longer supplies the contrast either: `JOHN SMITH, ANDREW, Jr.` keeps given 'ANDREW', as master did, where the whole-name test had flipped it. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 264, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 362, a constant +24 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 677; the second draft's per-character contrast test had made that +350). The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 268, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs a fold and a wordlist test per word before the comma, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index 74708f83..b239a997 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1722,16 +1722,17 @@ C1. Rationale: a credential run after the comma means the name is in there. Behind two or more name words, paired initials that are not the only shape word in the part are a credential however little else speaks for them, since no one writes a person's - initials as two dotted groups. In a name written in more than one - case, an unlisted word of two capitals after the comma is that - shape undotted and is read exactly as paired initials are, by the - sentences above ('García Márquez, MJ' and 'García Márquez, MJ PhD' - keep given 'MJ'). An unlisted all-caps word joins the class in - such a part only where the name's own words before the comma, - titles and particles aside, carry a lowercase letter: the contrast - is the name's, so a record that writes its surname in capitals - keeps its given name beside a mixed-case credential, a lowercase - title or particle, or a maiden clause. The same count reads a part of two or + initials as two dotted groups. An unlisted all-caps word joins + the class in such a part only where the name carries the contrast: + a lowercase letter in one of its own words before the comma that + no wordlist claims and no period marks, so a record that writes + its surname in capitals keeps its given name beside a mixed-case + credential, or a title, particle, connective, generation or maiden + clause written in lowercase. Where the name carries it, an + unlisted word of two capitals after the comma is the paired + initials' shape undotted and is decided at this comma exactly as + they are, by the sentences above ('García Márquez, MJ' and 'García + Márquez, MJ PhD' keep given 'MJ'). The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 438ab8d4..eeea0e57 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -56,7 +56,8 @@ from nameparser._pipeline._vocab import ( ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, caps_shape_candidate, is_one_case, is_paired_initials, - is_single_letter_numeral, is_wholly_suffix, name_word_count, + in_any_wordlist, is_single_letter_numeral, is_wholly_suffix, + name_word_count, surname_unit_count, run_word_fold, ) @@ -124,13 +125,16 @@ def texts(seg: tuple[int, ...]) -> list[str]: # #564: the contrast the caps shape needs is the NAME's -- its own # words before the comma (no maiden clause, no delimited content, - # as `own_words` defines them for `case_class` above), less titles - # and particles, which a record writing its surname in capitals - # still writes in lowercase ('Mr LLOYD WEBBER', 'de GAULLE'). - # Neither the credential's own lowercase nor a clause's may supply - # it, or an all-caps record loses its given name ('LLOYD WEBBER, - # ANDREW PhD', 'LLOYD WEBBER née Smith, ANDREW PhD'). Asked only - # once a caps word is in hand, and in one C-level comparison. + # as `own_words` defines them for `case_class` above), and of those + # only the words no wordlist claims and no period marks: a record + # writing its surname in capitals still writes its titles, + # particles, connectives and generations as it likes ('Mr LLOYD + # WEBBER', 'de GAULLE', 'GARCÍA y LÓPEZ', 'LLOYD WEBBER Jr.', + # 'Insp. LLOYD WEBBER'). Neither those nor the credential's own + # lowercase nor a clause's may supply it, or an all-caps record + # loses its given name. Asked only once a caps word is in hand: + # an all-caps part is settled in one C-level comparison, and past + # that each word before the comma costs a fold. def name_contrast() -> bool: # no lowercase before the comma at all: an all-caps record, # settled in C before the walk @@ -146,8 +150,8 @@ def name_contrast() -> bool: words = "".join( tok.text for i in groups[0] if i < clause_at and (tok := state.tokens[i]).role is None - and (n := _normalize(tok.text)) not in lex.titles - and n not in lex.particles) + and "." not in tok.text + and not in_any_wordlist(_normalize(tok.text), lex)) return words != words.upper() # Inlined rather than built on `texts` (measured, #289/#516's diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index b200dfa3..c0bce6c8 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -699,11 +699,21 @@ def caps_shape_candidate(text: str, lexicon: Lexicon, policy: Policy, and one_case is False and len(text) >= 2 and text.isalpha() and text.isupper()): return False - n = _normalize(text) + return not in_any_wordlist(_normalize(text), lexicon) + + +def in_any_wordlist(n: str, lexicon: Lexicon) -> bool: + """Whether the folded word `n` is in ANY of the lexicon's wordlists + (`_lexicon._VOCAB_FIELDS`, the whole roster). The caps shape's + "unlisted" (`caps_shape_candidate`), and the test of which words + before a comma are the NAME's own (#564's `name_contrast` in + `_segment.py`): a word some wordlist claims -- title, particle, + connective, suffix, bound given name -- is not one, so it neither + joins the caps class nor supplies the name's case contrast.""" for field in _VOCAB_FIELDS: if n in getattr(lexicon, field): - return False - return True + return True + return False # The comma form's own candidate test (rules.md#C1, decisions.md#S2). diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 1b3ee055..38c5d1f6 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2944,6 +2944,27 @@ def _check_cjk_shape_purity(self) -> None: "stays. The title in the family is the listing form's " "long-standing reading of a title before the comma " "('Prof. Cruz, Ed'), master's too"), + Case("a_generation_supplies_no_contrast_for_the_caps_shape", + "LLOYD WEBBER Jr., ANDREW", + {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "Jr."}, + notes="#564 boundary: only the name's own words that no " + "wordlist claims and no period marks supply the contrast; " + "'Jr.' is neither, so the 'SURNAME Jr., GIVEN' record " + "keeps its given name. A review-round draft excluded " + "titles and particles only and read suffix 'Jr., ANDREW'"), + Case("a_connective_supplies_no_contrast_for_the_caps_shape", + "GARCÍA y LÓPEZ, ANDREW", + {"given": "ANDREW", "family": "GARCÍA y LÓPEZ"}, + notes="#564 boundary: a lowercase connective in a record that " + "capitalizes its surnames is the record's convention, " + "not the name's contrast"), + Case("a_later_comma_part_supplies_no_contrast_for_the_caps_shape", + "JOHN SMITH, ANDREW, Jr.", + {"given": "ANDREW", "family": "JOHN SMITH", "suffix": "Jr."}, + notes="#564 boundary: the contrast is read before the comma; " + "a generation in a third part is not the name's. The " + "whole-name case test of an earlier draft read suffix " + "'ANDREW, Jr.'; master keeps given 'ANDREW' too"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index 808192d1..e69abcd0 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -871,7 +871,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # assertions are about the TRAILING slot EVERYWHERE adds. Default # costs, same harness against master: 'Smith, John' 183 -> 183, # 'Smith, XYZ' 182 -> 182 (the comma test needs two words before - # the comma), 'John Smith, XYZ' 251 -> 264 (decisions.md#S2). + # the comma), 'John Smith, XYZ' 251 -> 268 (decisions.md#S2). from nameparser import Parser on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) From 4ba12baf94d806de2f7d50f46199670d5f528313 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 23:16:12 -0700 Subject: [PATCH 06/13] fix(S2): #564 -- the name's contrast is a Title-case name word (Derek) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four review rounds showed that "a lowercase letter somewhere, less some exclusions" cannot separate an all-caps record from a mixed-case name: each draft leaked a class the next found, last glued particles and unlisted lowercase words ("GISCARD d'ESTAING, VALÉRY", 'HAFEZ al-ASSAD, BASHAR', 'LLOYD McDONALD, RONALD', 'LLOYD ap RHYS, DAFYDD'), all losing the given name. Derek's criterion: the name carries the contrast only through one of its own words before the comma written in Title case -- S2's "written the way a name is written" -- with no period and not claimed by the vocabulary as non-name text. A lowercase-only word never counts, so no list has to know 'ap' or 'thi'. That question gets its own predicate, _vocab.claimed_as_non_name, which leaves the surname and bound-given lists out: a draft sharing in_any_wordlist made a caller's surname list switch the reading off. Accepted with it: an all-lowercase name ('john smith, XYZ') reads the listing form, as do Mc/Mac-only names; an unlisted Title-case title ('Doña') still carries the contrast. rules.md#C1, the Policy docstring, customize.rst, the release log and decisions.md#S2 say so; case rows pin the glued particle, the unlisted lowercase word and the lowercase name, and a unit test the surname list. Co-Authored-By: Claude Opus 5.5 --- docs/customize.rst | 6 +++-- docs/design/decisions.md | 4 +-- docs/design/rules.md | 15 ++++++----- docs/release_log.rst | 2 +- nameparser/_pipeline/_segment.py | 42 +++++++++++++++---------------- nameparser/_pipeline/_vocab.py | 29 ++++++++++++++++----- nameparser/_policy.py | 6 +++-- tests/v2/cases.py | 21 ++++++++++++++++ tests/v2/pipeline/test_segment.py | 12 +++++++++ tests/v2/test_policy.py | 2 +- 10 files changed, 98 insertions(+), 41 deletions(-) diff --git a/docs/customize.rst b/docs/customize.rst index 78d6d9d5..e9c5a520 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -531,8 +531,10 @@ listed below. * - ``unlisted_caps_suffixes`` - ``CapsSuffixes`` - Where an unlisted all-caps word of two or more letters, with no - period in it, in a name written in more than one case, reads as - a credential. ``CapsSuffixes.AFTER_COMMA``, the default, reads + period in it, reads as a credential. The name must contrast it + with a word written in Title case (``Smith``): a record written + wholly in capitals, or wholly in lowercase, keeps every word a + name word. ``CapsSuffixes.AFTER_COMMA``, the default, reads it only in the part right after a comma with two or more name words before it: ``"John Smith, XYZ"`` gives suffix ``XYZ``, while ``"Smith, XYZ"`` keeps given ``XYZ`` and a lone two-letter diff --git a/docs/design/decisions.md b/docs/design/decisions.md index f1063b6e..292cbcdf 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's: a lowercase letter among the name's own words before the comma (`own_words`, so no maiden clause or delimited content), titles and particles aside, which a record capitalizing its surname still writes in lowercase. Three drafts got this wrong, each by an exclusion too narrow. The first asked the whole name's case, so a mixed-case credential supplied the contrast and `LLOYD WEBBER, ANDREW PhD` read suffix 'ANDREW PhD'; the second asked every token before the comma, so a maiden clause or a title supplied it (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`); the third excluded titles and particles by list and missed a generation (`LLOYD WEBBER Jr., ANDREW`, the common `SURNAME Jr., GIVEN` format), a connective (`GARCÍA y LÓPEZ, ANDREW`) and a title recognized by shape (`Insp. LLOYD WEBBER, ANDREW`) (the three fix-commit reviews). The shipped test inverts it: the contrast comes only from the name's own words that no wordlist claims (`_vocab.in_any_wordlist`, the caps predicate's own "unlisted" lifted out so the two share it) and no period marks. One `name_contrast` answers it for both caps branches, reading the own-words walk `case_class` makes. It reads only the part before the comma, so a later comma part no longer supplies the contrast either: `JOHN SMITH, ANDREW, Jr.` keeps given 'ANDREW', as master did, where the whole-name test had flipped it. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) written in TITLE CASE, with no period, and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek, after four review rounds). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and finally glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD McDONALD, RONALD`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. Title case is S2's own "written the way a name is written", and a word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ'; a name whose only name words are Mc/Mac-cased (`McDonald MacLeod, XYZ`, not Title case to `istitle()`) does too; and an unlisted title written in Title case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 268, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs a fold and a wordlist test per word before the comma, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 265, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs, per word before the comma, a Title-case test in C and a fold and wordlist test where it passes, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index b239a997..5d9be83b 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1723,12 +1723,15 @@ C1. Rationale: a credential run after the comma means the name is in not the only shape word in the part are a credential however little else speaks for them, since no one writes a person's initials as two dotted groups. An unlisted all-caps word joins - the class in such a part only where the name carries the contrast: - a lowercase letter in one of its own words before the comma that - no wordlist claims and no period marks, so a record that writes - its surname in capitals keeps its given name beside a mixed-case - credential, or a title, particle, connective, generation or maiden - clause written in lowercase. Where the name carries it, an + the class in such a part only where the name carries the + contrast: one of its own words before the comma written in Title + case, the way a name is written (S2), with no period and not + claimed by the vocabulary as a title, particle, connective, + credential or generation. A word written wholly in lowercase never + carries it, so a record that writes its surname in capitals keeps + its given name whatever else it writes in lowercase or beside it + ("GISCARD d'ESTAING, VALÉRY"), and a name written wholly in + lowercase reads the listing form. Where the name carries it, an unlisted word of two capitals after the comma is the paired initials' shape undotted and is decided at this comma exactly as they are, by the sentences above ('García Márquez, MJ' and 'García diff --git a/docs/release_log.rst b/docs/release_log.rst index 5e661771..b8af0558 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and a record written wholly in capitals keeps its given name beside a mixed-case credential (``LLOYD WEBBER, ANDREW PhD``). ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own written in Title case, so a record that writes its surname in capitals keeps its given name whatever it writes in lowercase (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index eeea0e57..741c8bf5 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -56,7 +56,7 @@ from nameparser._pipeline._vocab import ( ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, caps_shape_candidate, is_one_case, is_paired_initials, - in_any_wordlist, is_single_letter_numeral, is_wholly_suffix, + claimed_as_non_name, is_single_letter_numeral, is_wholly_suffix, name_word_count, surname_unit_count, run_word_fold, @@ -123,21 +123,22 @@ def case_class() -> bool: def texts(seg: tuple[int, ...]) -> list[str]: return [state.tokens[i].text for i in seg] - # #564: the contrast the caps shape needs is the NAME's -- its own - # words before the comma (no maiden clause, no delimited content, - # as `own_words` defines them for `case_class` above), and of those - # only the words no wordlist claims and no period marks: a record - # writing its surname in capitals still writes its titles, - # particles, connectives and generations as it likes ('Mr LLOYD - # WEBBER', 'de GAULLE', 'GARCÍA y LÓPEZ', 'LLOYD WEBBER Jr.', - # 'Insp. LLOYD WEBBER'). Neither those nor the credential's own - # lowercase nor a clause's may supply it, or an all-caps record - # loses its given name. Asked only once a caps word is in hand: - # an all-caps part is settled in one C-level comparison, and past - # that each word before the comma costs a fold. + # #564 (Derek): the contrast the caps shape needs is the NAME's, + # and the name carries it only if one of its own words before the + # comma (no maiden clause, no delimited content, as `own_words` + # defines them for `case_class` above) is written in Title case -- + # S2's "written the way a name is written" -- and is not claimed as + # a title, particle, connective, credential or generation + # (`claimed_as_non_name`); a word with a period is an abbreviation + # or an initial, not one. A lowercase-only word never counts, so a + # record that writes its surname in capitals keeps its given name + # whatever else it writes in lowercase ('LLOYD ap RHYS', 'HAFEZ + # al-ASSAD', "GISCARD d'ESTAING", 'LLOYD WEBBER née Smith'), and + # so does a mixed-case credential ('LLOYD WEBBER, ANDREW PhD'). + # Asked only once a caps word is in hand: a part with no lowercase + # at all is settled in one C-level comparison, and past that each + # word before the comma costs a fold and a wordlist test. def name_contrast() -> bool: - # no lowercase before the comma at all: an all-caps record, - # settled in C before the walk before = "".join([state.tokens[i].text for i in groups[0]]) if before == before.upper(): return False @@ -147,12 +148,11 @@ def name_contrast() -> bool: state.lexicon.maiden_markers) clause_at = own_span[1] lex = state.lexicon - words = "".join( - tok.text for i in groups[0] - if i < clause_at and (tok := state.tokens[i]).role is None - and "." not in tok.text - and not in_any_wordlist(_normalize(tok.text), lex)) - return words != words.upper() + return any( + i < clause_at and (tok := state.tokens[i]).role is None + and tok.text.istitle() and "." not in tok.text + and not claimed_as_non_name(_normalize(tok.text), lex) + for i in groups[0]) # Inlined rather than built on `texts` (measured, #289/#516's # eager-gate fix round): every comma parse calls `suffixy` at diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index c0bce6c8..9e8f4fc0 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -704,18 +704,35 @@ def caps_shape_candidate(text: str, lexicon: Lexicon, policy: Policy, def in_any_wordlist(n: str, lexicon: Lexicon) -> bool: """Whether the folded word `n` is in ANY of the lexicon's wordlists - (`_lexicon._VOCAB_FIELDS`, the whole roster). The caps shape's - "unlisted" (`caps_shape_candidate`), and the test of which words - before a comma are the NAME's own (#564's `name_contrast` in - `_segment.py`): a word some wordlist claims -- title, particle, - connective, suffix, bound given name -- is not one, so it neither - joins the caps class nor supplies the name's case contrast.""" + (`_lexicon._VOCAB_FIELDS`, the whole roster): the caps shape's + "unlisted" (`caps_shape_candidate`), where a word a caller listed + even as a SURNAME must never become a credential.""" for field in _VOCAB_FIELDS: if n in getattr(lexicon, field): return True return False +#: The wordlists that claim a word as something OTHER than name text. +#: `surnames` and `bound_given_names` are left out: they say a word IS +#: a name word ('Kim', 'Abdul'), the opposite claim. +_NON_NAME_FIELDS = tuple(f for f in _VOCAB_FIELDS + if f not in ("surnames", "bound_given_names")) + + +def claimed_as_non_name(n: str, lexicon: Lexicon) -> bool: + """Whether a wordlist claims the folded word `n` as a title, + particle, connective, credential, generation, maiden marker or + honorific -- not name text. #564's `name_contrast` (`_segment.py`) + asks it: such a word does not supply the name's case contrast. A + different question from `in_any_wordlist`'s, which a surname list + must also answer yes to.""" + for field in _NON_NAME_FIELDS: + if n in getattr(lexicon, field): + return True + return False + + # The comma form's own candidate test (rules.md#C1, decisions.md#S2). def ambiguous_class_candidate(text: str, lexicon: Lexicon, policy: Policy) -> bool: diff --git a/nameparser/_policy.py b/nameparser/_policy.py index ac72dd04..d088a9cc 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -708,8 +708,10 @@ class Policy: #: behind this switch, and still reports the fork. unlisted_dotted_suffixes: bool = True #: Where an UNLISTED all-caps word of two or more letters, with no - #: period in it, in a name written in more than one case, reads as - #: a credential (:class:`CapsSuffixes`). A listed member keeps its + #: period in it, reads as a credential (:class:`CapsSuffixes`). The + #: name must contrast it: a word of the name written in Title case + #: ("Smith") -- a record written wholly in capitals or wholly in + #: lowercase keeps every word a name word. A listed member keeps its #: own case lean in every setting ("Jack MA" gives suffix ``MA``), #: and the roman-numeral fork claims a bare numeral first ("Jack #: VI" is unaffected). diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 38c5d1f6..78a7aa9f 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2965,6 +2965,27 @@ def _check_cjk_shape_purity(self) -> None: "a generation in a third part is not the name's. The " "whole-name case test of an earlier draft read suffix " "'ANDREW, Jr.'; master keeps given 'ANDREW' too"), + Case("a_glued_particle_supplies_no_contrast_for_the_caps_shape", + "GISCARD d'ESTAING, VALÉRY", + {"given": "VALÉRY", "family": "GISCARD d'ESTAING"}, + notes="#564 (Derek): the name carries the contrast only through " + "a word written in Title case that no wordlist claims as " + "non-name text. A lowercase particle glued to a capitalized " + "surname is not one -- the French convention the trailing " + "slot stays off for. A draft counting any lowercase letter " + "read suffix 'VALÉRY'"), + Case("an_unlisted_lowercase_word_supplies_no_contrast_for_the_caps_shape", + "LLOYD ap RHYS, DAFYDD", + {"given": "DAFYDD", "family": "LLOYD ap RHYS"}, + notes="#564 boundary: 'ap' is in no wordlist, but a word written " + "wholly in lowercase never carries the contrast, so no " + "list has to know it"), + Case("an_all_lowercase_name_carries_no_contrast_for_the_caps_shape", + "john smith, XYZ", + {"given": "XYZ", "family": "john smith"}, + notes="#564 (Derek): a lowercase-only name has no Title-case " + "word, so the default keeps the listing form; the cost of " + "the Title-case criterion, accepted with it"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, diff --git a/tests/v2/pipeline/test_segment.py b/tests/v2/pipeline/test_segment.py index 836bbb3b..030559d1 100644 --- a/tests/v2/pipeline/test_segment.py +++ b/tests/v2/pipeline/test_segment.py @@ -349,3 +349,15 @@ def test_the_run_test_reads_lenient_comma_suffixes() -> None: policy=dataclasses.replace(Policy(), lenient_comma_suffixes=False)) out = segment(tokenize(extract_delimited(state))) assert out.structure is Structure.FAMILY_COMMA + + +def test_a_listed_surname_still_carries_the_name_contrast() -> None: + # #564: a word a caller lists as a SURNAME is name text, so it + # carries the contrast the comma caps reading needs + # (`_vocab.claimed_as_non_name` leaves the surname and bound-given + # lists out); `in_any_wordlist`, the caps shape's own "unlisted", + # is a different question. A draft shared one predicate for both + # and this read given 'XYZ'. + from nameparser import Lexicon, Parser + parser = Parser(lexicon=Lexicon.default().add(surnames={"smith", "jones"})) + assert parser.parse("Smith Jones, XYZ").suffix == "XYZ" diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index e69abcd0..1771121b 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -871,7 +871,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # assertions are about the TRAILING slot EVERYWHERE adds. Default # costs, same harness against master: 'Smith, John' 183 -> 183, # 'Smith, XYZ' 182 -> 182 (the comma test needs two words before - # the comma), 'John Smith, XYZ' 251 -> 268 (decisions.md#S2). + # the comma), 'John Smith, XYZ' 251 -> 265 (decisions.md#S2). from nameparser import Parser on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) From 3e03c08ec0380f45028c83e36b9f2d8aaac1ebcf Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 23:24:41 -0700 Subject: [PATCH 07/13] fix(S2): #564 -- a lone capital is not Title case; record the full cost MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `istitle()` accepts a single capital, so an undotted initial carried the name's contrast and a capitalized-surname record with lowercase particles lost its given name ('de GAULLE C, CHARLES', 'de la O GARCÍA, MARÍA'). Title case now needs a lowercase letter behind the capital. decisions.md#S2 records the whole class istitle() rejects when every name word is of it (Mc/Mac and other interior capitals, the Dutch IJ, lowercase-led elisions, a vocabulary claim beside one) rather than Mc/Mac alone; two case notes and the cost comment state the Title-case criterion; the Policy docstring and customize.rst add the vocabulary exclusion. Co-Authored-By: Claude Opus 5.5 --- docs/customize.rst | 7 ++++--- docs/design/decisions.md | 2 +- nameparser/_pipeline/_segment.py | 12 ++++++++---- nameparser/_policy.py | 3 ++- tests/v2/cases.py | 25 +++++++++++++++++-------- 5 files changed, 32 insertions(+), 17 deletions(-) diff --git a/docs/customize.rst b/docs/customize.rst index e9c5a520..7064e6f7 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -532,9 +532,10 @@ listed below. - ``CapsSuffixes`` - Where an unlisted all-caps word of two or more letters, with no period in it, reads as a credential. The name must contrast it - with a word written in Title case (``Smith``): a record written - wholly in capitals, or wholly in lowercase, keeps every word a - name word. ``CapsSuffixes.AFTER_COMMA``, the default, reads + with a word written in Title case (``Smith``) that the vocabulary + does not claim as a title, particle or credential: a record + written wholly in capitals, or wholly in lowercase, keeps every + word a name word. ``CapsSuffixes.AFTER_COMMA``, the default, reads it only in the part right after a comma with two or more name words before it: ``"John Smith, XYZ"`` gives suffix ``XYZ``, while ``"Smith, XYZ"`` keeps given ``XYZ`` and a lone two-letter diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 292cbcdf..ea8a9cb6 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,7 +687,7 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) written in TITLE CASE, with no period, and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek, after four review rounds). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and finally glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD McDONALD, RONALD`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. Title case is S2's own "written the way a name is written", and a word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ'; a name whose only name words are Mc/Mac-cased (`McDonald MacLeod, XYZ`, not Title case to `istitle()`) does too; and an unlisted title written in Title case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) written in TITLE CASE, with no period, and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek, after four review rounds). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and finally glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD McDONALD, RONALD`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. Title case is S2's own "written the way a name is written", and a word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ'; a name whose every name word fails `istitle()` does too — Mc/Mac and other interior capitals (`McDonald MacLeod, XYZ`, `DiCaprio LaBeouf, XYZ`), the Dutch IJ (`IJzerman IJsselmeer, XYZ`), lowercase-led elisions and hyphens (`al-Rashid al-Hassan, XYZ`, `d'Estaing d'Orléans, XYZ`), or a word the vocabulary claims beside one of those (`Abd al-Rahman al-Sudais, XYZ`, `abd` being a credential acronym) — while one Title-case word among them is enough (`Ahmed al-Rashid, XYZ` flips); a lone capital is not Title case (`istitle()` accepts 'A', and a draft taking it lost the given name of `de GAULLE C, CHARLES` and `de la O GARCÍA, MARÍA`); and an unlisted title written in Title case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 265, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs, per word before the comma, a Title-case test in C and a fold and wordlist test where it passes, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 741c8bf5..acf48669 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -130,14 +130,17 @@ def texts(seg: tuple[int, ...]) -> list[str]: # S2's "written the way a name is written" -- and is not claimed as # a title, particle, connective, credential or generation # (`claimed_as_non_name`); a word with a period is an abbreviation - # or an initial, not one. A lowercase-only word never counts, so a + # or an initial, not one, and nor is a lone capital ('de GAULLE C'), + # which `istitle()` alone accepts -- Title case needs a lowercase + # letter behind the capital. A lowercase-only word never counts, so a # record that writes its surname in capitals keeps its given name # whatever else it writes in lowercase ('LLOYD ap RHYS', 'HAFEZ # al-ASSAD', "GISCARD d'ESTAING", 'LLOYD WEBBER née Smith'), and # so does a mixed-case credential ('LLOYD WEBBER, ANDREW PhD'). # Asked only once a caps word is in hand: a part with no lowercase - # at all is settled in one C-level comparison, and past that each - # word before the comma costs a fold and a wordlist test. + # at all is settled in one C-level comparison; past that the walk + # costs about three frames a word before the comma, a fold and a + # wordlist test only where the Title-case test in C passes. def name_contrast() -> bool: before = "".join([state.tokens[i].text for i in groups[0]]) if before == before.upper(): @@ -150,7 +153,8 @@ def name_contrast() -> bool: lex = state.lexicon return any( i < clause_at and (tok := state.tokens[i]).role is None - and tok.text.istitle() and "." not in tok.text + and tok.text.istitle() and not tok.text.isupper() + and "." not in tok.text and not claimed_as_non_name(_normalize(tok.text), lex) for i in groups[0]) diff --git a/nameparser/_policy.py b/nameparser/_policy.py index d088a9cc..c4afac31 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -710,7 +710,8 @@ class Policy: #: Where an UNLISTED all-caps word of two or more letters, with no #: period in it, reads as a credential (:class:`CapsSuffixes`). The #: name must contrast it: a word of the name written in Title case - #: ("Smith") -- a record written wholly in capitals or wholly in + #: ("Smith") that the vocabulary does not claim as a title, particle + #: or credential -- a record written wholly in capitals or wholly in #: lowercase keeps every word a name word. A listed member keeps its #: own case lean in every setting ("Jack MA" gives suffix ``MA``), #: and the roman-numeral fork claims a bare numeral first ("Jack diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 78a7aa9f..798ee790 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2921,11 +2921,12 @@ def _check_cjk_shape_purity(self) -> None: "LLOYD WEBBER, ANDREW PhD", {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "PhD"}, notes="#564 boundary: an all-caps word joins C1's run only where " - "the NAME before the comma has a lowercase letter -- here " - "the only lowercase is the credential's own, and a record " - "written wholly in capitals keeps its given name. A " - "review-round draft let 'PhD' supply the contrast and " - "read suffix 'ANDREW PhD'"), + "the NAME carries the contrast, a word of its own before " + "the comma written in Title case -- here the only " + "lowercase is the credential's own, and a record written " + "wholly in capitals keeps its given name. A review-round " + "draft let 'PhD' supply the contrast and read suffix " + "'ANDREW PhD'"), Case("a_maiden_clause_supplies_no_contrast_for_the_caps_shape", "LLOYD WEBBER née Smith, ANDREW PhD", {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "PhD", @@ -2947,9 +2948,10 @@ def _check_cjk_shape_purity(self) -> None: Case("a_generation_supplies_no_contrast_for_the_caps_shape", "LLOYD WEBBER Jr., ANDREW", {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "Jr."}, - notes="#564 boundary: only the name's own words that no " - "wordlist claims and no period marks supply the contrast; " - "'Jr.' is neither, so the 'SURNAME Jr., GIVEN' record " + notes="#564 boundary: only the name's own words written in " + "Title case, unclaimed by the vocabulary and unmarked by " + "a period, supply the contrast; 'Jr.' is none of those, " + "so the 'SURNAME Jr., GIVEN' record " "keeps its given name. A review-round draft excluded " "titles and particles only and read suffix 'Jr., ANDREW'"), Case("a_connective_supplies_no_contrast_for_the_caps_shape", @@ -2986,6 +2988,13 @@ def _check_cjk_shape_purity(self) -> None: notes="#564 (Derek): a lowercase-only name has no Title-case " "word, so the default keeps the listing form; the cost of " "the Title-case criterion, accepted with it"), + Case("a_lone_capital_supplies_no_contrast_for_the_caps_shape", + "de GAULLE C, CHARLES", + {"given": "CHARLES", "family": "de GAULLE C"}, + notes="#564 boundary: an undotted initial is a single capital, " + "which `istitle()` accepts but Title case does not -- it " + "needs a lowercase letter behind the capital. A draft " + "missing that read suffix 'CHARLES' with no given name"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, From a16927c66f8b39d9d5307f45d4c8491e58ba011b Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Fri, 2 Oct 2026 00:08:00 -0700 Subject: [PATCH 08/13] fix(S2): #564 -- the name's contrast is a capital followed by lowercase (Derek) Title case (`istitle()`) read too narrowly: interior capitals and elisions ('DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing", 'McDonald') are mixed-case names and kept a given 'XYZ'. Derek's criterion: a word of the name carries the contrast if it holds a capital directly followed by a lowercase letter, a leading Mc/Mac skipped where a capital follows it (_vocab.written_as_a_name). A lowercase prefix glued to a surname written in capitals ("d'ESTAING", 'al-ASSAD', 'McDONALD') has no such pair, so those records keep their given names; a lone capital and a lowercase-only word have none either. Weighed and declined, with measurements in decisions.md#S2: "contains a capital" (any lowercase word in an all-caps record then makes every word a contrast) and "not all capitals but contains a capital" (takes d'ESTAING and al-ASSAD). The accepted costs shrink to all-lowercase names and unlisted mixed-case titles. A parametrized test pins the predicate; case rows pin McDONALD and DiCaprio. Co-Authored-By: Claude Opus 5.5 --- docs/customize.rst | 3 +- docs/design/decisions.md | 4 +-- docs/design/rules.md | 6 ++-- docs/release_log.rst | 2 +- nameparser/_pipeline/_segment.py | 29 +++++++++--------- nameparser/_pipeline/_vocab.py | 16 ++++++++++ nameparser/_policy.py | 5 ++-- tests/v2/cases.py | 50 ++++++++++++++++++++++---------- tests/v2/pipeline/test_vocab.py | 18 +++++++++++- tests/v2/test_policy.py | 2 +- 10 files changed, 96 insertions(+), 39 deletions(-) diff --git a/docs/customize.rst b/docs/customize.rst index 7064e6f7..4ea1d92b 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -532,7 +532,8 @@ listed below. - ``CapsSuffixes`` - Where an unlisted all-caps word of two or more letters, with no period in it, reads as a credential. The name must contrast it - with a word written in Title case (``Smith``) that the vocabulary + with a word holding a capital followed by a lowercase letter + (``Smith``, ``DiCaprio``) that the vocabulary does not claim as a title, particle or credential: a record written wholly in capitals, or wholly in lowercase, keeps every word a name word. ``CapsSuffixes.AFTER_COMMA``, the default, reads diff --git a/docs/design/decisions.md b/docs/design/decisions.md index ea8a9cb6..ffc9d00a 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) written in TITLE CASE, with no period, and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek, after four review rounds). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and finally glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD McDONALD, RONALD`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. Title case is S2's own "written the way a name is written", and a word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ'; a name whose every name word fails `istitle()` does too — Mc/Mac and other interior capitals (`McDonald MacLeod, XYZ`, `DiCaprio LaBeouf, XYZ`), the Dutch IJ (`IJzerman IJsselmeer, XYZ`), lowercase-led elisions and hyphens (`al-Rashid al-Hassan, XYZ`, `d'Estaing d'Orléans, XYZ`), or a word the vocabulary claims beside one of those (`Abd al-Rahman al-Sudais, XYZ`, `abd` being a credential acronym) — while one Title-case word among them is enough (`Ahmed al-Rashid, XYZ` flips); a lone capital is not Title case (`istitle()` accepts 'A', and a draft taking it lost the given name of `de GAULLE C, CHARLES` and `de la O GARCÍA, MARÍA`); and an unlisted title written in Title case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital directly followed by a lowercase letter, a leading `Mc`/`Mac` skipped where a capital follows it (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and chose the capital-then-lowercase pair with the Mc/Mac skip: in `d'ESTAING`, `al-ASSAD` and `McDONALD` no capital is followed by a lowercase letter once the prefix is set aside, while `McDonald` keeps `Do` and `Mack` its own `Ma`. A word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', and an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 265, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs, per word before the comma, a Title-case test in C and a fold and wordlist test where it passes, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 268, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs, per word before the comma, a case-pair test and a fold and wordlist test where it passes, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index 5d9be83b..da7de369 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1724,8 +1724,10 @@ C1. Rationale: a credential run after the comma means the name is in little else speaks for them, since no one writes a person's initials as two dotted groups. An unlisted all-caps word joins the class in such a part only where the name carries the - contrast: one of its own words before the comma written in Title - case, the way a name is written (S2), with no period and not + contrast: one of its own words before the comma written the way a + name is written in mixed case, a capital directly followed by a + lowercase letter (a leading Mc or Mac before a capital aside), + with no period and not claimed by the vocabulary as a title, particle, connective, credential or generation. A word written wholly in lowercase never carries it, so a record that writes its surname in capitals keeps diff --git a/docs/release_log.rst b/docs/release_log.rst index b8af0558..8f891db1 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own written in Title case, so a record that writes its surname in capitals keeps its given name whatever it writes in lowercase (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital followed by a lowercase letter (``Smith``, ``DiCaprio``), so a record that writes its surname in capitals keeps its given name whatever it writes beside it (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD McDONALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index acf48669..1a0d2aa3 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -57,6 +57,7 @@ ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, caps_shape_candidate, is_one_case, is_paired_initials, claimed_as_non_name, is_single_letter_numeral, is_wholly_suffix, + written_as_a_name, name_word_count, surname_unit_count, run_word_fold, @@ -126,21 +127,22 @@ def texts(seg: tuple[int, ...]) -> list[str]: # #564 (Derek): the contrast the caps shape needs is the NAME's, # and the name carries it only if one of its own words before the # comma (no maiden clause, no delimited content, as `own_words` - # defines them for `case_class` above) is written in Title case -- - # S2's "written the way a name is written" -- and is not claimed as - # a title, particle, connective, credential or generation + # defines them for `case_class` above) is written the way a name is + # written in mixed case -- a capital directly followed by a + # lowercase letter (`written_as_a_name`) -- and is not claimed as a + # title, particle, connective, credential or generation # (`claimed_as_non_name`); a word with a period is an abbreviation - # or an initial, not one, and nor is a lone capital ('de GAULLE C'), - # which `istitle()` alone accepts -- Title case needs a lowercase - # letter behind the capital. A lowercase-only word never counts, so a - # record that writes its surname in capitals keeps its given name - # whatever else it writes in lowercase ('LLOYD ap RHYS', 'HAFEZ - # al-ASSAD', "GISCARD d'ESTAING", 'LLOYD WEBBER née Smith'), and - # so does a mixed-case credential ('LLOYD WEBBER, ANDREW PhD'). + # or an initial, not one. A lone capital ('de GAULLE C') or a + # lowercase-only word has no such pair, and nor has a lowercase + # prefix glued to a capitalized surname, so a record that writes + # its surname in capitals keeps its given name whatever else it + # writes ('LLOYD ap RHYS', 'HAFEZ al-ASSAD', "GISCARD d'ESTAING", + # 'LLOYD McDONALD', 'LLOYD WEBBER née Smith'), and so does a + # mixed-case credential ('LLOYD WEBBER, ANDREW PhD'). # Asked only once a caps word is in hand: a part with no lowercase # at all is settled in one C-level comparison; past that the walk - # costs about three frames a word before the comma, a fold and a - # wordlist test only where the Title-case test in C passes. + # costs a few frames a word before the comma, a fold and a + # wordlist test only where the case test passes. def name_contrast() -> bool: before = "".join([state.tokens[i].text for i in groups[0]]) if before == before.upper(): @@ -153,8 +155,7 @@ def name_contrast() -> bool: lex = state.lexicon return any( i < clause_at and (tok := state.tokens[i]).role is None - and tok.text.istitle() and not tok.text.isupper() - and "." not in tok.text + and "." not in tok.text and written_as_a_name(tok.text) and not claimed_as_non_name(_normalize(tok.text), lex) for i in groups[0]) diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 9e8f4fc0..95b60ef5 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -720,6 +720,22 @@ def in_any_wordlist(n: str, lexicon: Lexicon) -> bool: if f not in ("surnames", "bound_given_names")) +def written_as_a_name(text: str) -> bool: + """Whether TEXT is written the way a name is written in mixed case: + a capital directly followed by a lowercase letter ('Smith', + 'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing") -- #564's test + for the name's case contrast (Derek). A word written wholly in one + case has no such pair, and nor does a lowercase prefix glued to a + surname written in capitals ("d'ESTAING", 'al-ASSAD'). A leading + 'Mc'/'Mac' is skipped where a capital follows it, so 'McDONALD' + has no pair while 'McDonald' keeps 'Do' and 'Mack' its own 'Ma'.""" + if text.startswith("Mc") and text[2:3].isupper(): + text = text[2:] + elif text.startswith("Mac") and text[3:4].isupper(): + text = text[3:] + return any(a.isupper() and b.islower() for a, b in zip(text, text[1:])) + + def claimed_as_non_name(n: str, lexicon: Lexicon) -> bool: """Whether a wordlist claims the folded word `n` as a title, particle, connective, credential, generation, maiden marker or diff --git a/nameparser/_policy.py b/nameparser/_policy.py index c4afac31..8ccd77e0 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -709,8 +709,9 @@ class Policy: unlisted_dotted_suffixes: bool = True #: Where an UNLISTED all-caps word of two or more letters, with no #: period in it, reads as a credential (:class:`CapsSuffixes`). The - #: name must contrast it: a word of the name written in Title case - #: ("Smith") that the vocabulary does not claim as a title, particle + #: name must contrast it: a word of the name with a capital followed + #: by a lowercase letter ("Smith", "DiCaprio") that the vocabulary + #: does not claim as a title, particle #: or credential -- a record written wholly in capitals or wholly in #: lowercase keeps every word a name word. A listed member keeps its #: own case lean in every setting ("Jack MA" gives suffix ``MA``), diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 798ee790..61d326b7 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2922,7 +2922,8 @@ def _check_cjk_shape_purity(self) -> None: {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "PhD"}, notes="#564 boundary: an all-caps word joins C1's run only where " "the NAME carries the contrast, a word of its own before " - "the comma written in Title case -- here the only " + "the comma with a capital followed by a lowercase letter " + "-- here the only " "lowercase is the credential's own, and a record written " "wholly in capitals keeps its given name. A review-round " "draft let 'PhD' supply the contrast and read suffix " @@ -2948,9 +2949,10 @@ def _check_cjk_shape_purity(self) -> None: Case("a_generation_supplies_no_contrast_for_the_caps_shape", "LLOYD WEBBER Jr., ANDREW", {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "Jr."}, - notes="#564 boundary: only the name's own words written in " - "Title case, unclaimed by the vocabulary and unmarked by " - "a period, supply the contrast; 'Jr.' is none of those, " + notes="#564 boundary: only the name's own words written as a " + "name is (a capital followed by a lowercase letter), " + "unclaimed by the vocabulary and unmarked by a period, " + "supply the contrast; 'Jr.' is none of those, " "so the 'SURNAME Jr., GIVEN' record " "keeps its given name. A review-round draft excluded " "titles and particles only and read suffix 'Jr., ANDREW'"), @@ -2971,11 +2973,12 @@ def _check_cjk_shape_purity(self) -> None: "GISCARD d'ESTAING, VALÉRY", {"given": "VALÉRY", "family": "GISCARD d'ESTAING"}, notes="#564 (Derek): the name carries the contrast only through " - "a word written in Title case that no wordlist claims as " - "non-name text. A lowercase particle glued to a capitalized " - "surname is not one -- the French convention the trailing " - "slot stays off for. A draft counting any lowercase letter " - "read suffix 'VALÉRY'"), + "a word with a capital followed by a lowercase letter that " + "no wordlist claims as non-name text. A lowercase particle " + "glued to a capitalized surname has no such pair -- the " + "French convention the trailing slot stays off for. A " + "draft counting any lowercase letter read suffix " + "'VALÉRY'"), Case("an_unlisted_lowercase_word_supplies_no_contrast_for_the_caps_shape", "LLOYD ap RHYS, DAFYDD", {"given": "DAFYDD", "family": "LLOYD ap RHYS"}, @@ -2985,16 +2988,33 @@ def _check_cjk_shape_purity(self) -> None: Case("an_all_lowercase_name_carries_no_contrast_for_the_caps_shape", "john smith, XYZ", {"given": "XYZ", "family": "john smith"}, - notes="#564 (Derek): a lowercase-only name has no Title-case " - "word, so the default keeps the listing form; the cost of " - "the Title-case criterion, accepted with it"), + notes="#564 (Derek): a lowercase-only name has no capital " + "followed by a lowercase letter, so the default keeps the " + "listing form; the cost of the criterion, accepted with " + "it"), Case("a_lone_capital_supplies_no_contrast_for_the_caps_shape", "de GAULLE C, CHARLES", {"given": "CHARLES", "family": "de GAULLE C"}, notes="#564 boundary: an undotted initial is a single capital, " - "which `istitle()` accepts but Title case does not -- it " - "needs a lowercase letter behind the capital. A draft " - "missing that read suffix 'CHARLES' with no given name"), + "with no lowercase letter behind it. A draft testing " + "`istitle()`, which accepts 'C', read suffix 'CHARLES' " + "with no given name"), + Case("a_mc_prefix_on_a_capitalized_surname_supplies_no_contrast", + "LLOYD McDONALD, RONALD", + {"given": "RONALD", "family": "LLOYD McDONALD"}, + notes="#564 (Derek): a leading 'Mc'/'Mac' is skipped where a " + "capital follows it, so a record writing 'McDONALD' keeps " + "its given name, while 'McDonald' still carries the " + "contrast through 'Do'"), + Case("an_interior_capital_carries_the_contrast", + "DiCaprio LaBeouf, XYZ", + {"given": "DiCaprio", "family": "LaBeouf", "suffix": "XYZ"}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="#564 (Derek): a capital followed by a lowercase letter " + "anywhere in the word is the contrast, so an interior " + "capital counts ('DiCaprio', 'IJzerman', 'al-Rashid'). " + "The `istitle()` draft read given 'XYZ' here"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index d48cef13..fcd4df46 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -13,7 +13,7 @@ effective_script, is_initial, is_initial_shaped, is_one_case, is_single_letter_numeral, is_suffix_lenient, is_suffix_strict, is_title_shaped, is_wholly_suffix, - maiden_marker_run, name_word_count, period_joined_vocab, + maiden_marker_run, name_word_count, period_joined_vocab, written_as_a_name, resolve_script_set, run_word_fold, single_script, ) from nameparser._policy import (Policy, Script, _NO_INITIALS, @@ -890,3 +890,19 @@ def test_is_single_letter_numeral() -> None: assert is_single_letter_numeral(text), text for text in ("II", "IV", "Jr", "B", "", ".", "Ma"): assert not is_single_letter_numeral(text), text + + +@pytest.mark.parametrize("text, expected", [ + ("Smith", True), ("DiCaprio", True), ("IJzerman", True), + ("al-Rashid", True), ("d'Estaing", True), ("McDonald", True), + ("MacLeod", True), ("Mack", True), ("Macy", True), + ("SMITH", False), ("smith", False), ("C", False), ("ap", False), + ("d'ESTAING", False), ("al-ASSAD", False), ("McDONALD", False), + ("MacDONALD", False), +]) +def test_written_as_a_name(text: str, expected: bool) -> None: + # #564 (Derek): the name's case contrast is a capital directly + # followed by a lowercase letter, a leading Mc/Mac skipped where a + # capital follows it -- so a lowercase prefix glued to a surname + # written in capitals carries none, and an interior capital does. + assert written_as_a_name(text) is expected diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index 1771121b..e69abcd0 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -871,7 +871,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # assertions are about the TRAILING slot EVERYWHERE adds. Default # costs, same harness against master: 'Smith, John' 183 -> 183, # 'Smith, XYZ' 182 -> 182 (the comma test needs two words before - # the comma), 'John Smith, XYZ' 251 -> 265 (decisions.md#S2). + # the comma), 'John Smith, XYZ' 251 -> 268 (decisions.md#S2). from nameparser import Parser on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) From bbe6b0f62e1507f680e92bcc54955f1c03a36e3d Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Fri, 2 Oct 2026 00:22:33 -0700 Subject: [PATCH 09/13] fix(S2): #564 -- the name's contrast: a capital and a lowercase last letter (Derek) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The capital-then-lowercase test leaked a whole class: any capitalized prefix glued to a surname written in capitals supplies the pair ('LLOYD FitzGERALD, RONALD', 'PAOLO DeVITO, MARCO', 'JEAN LaFLEUR, PIERRE', 'DICK VanDYKE, JOHN', 'PAUL DuBOIS, JEAN', 'JEAN St-PIERRE, MARC', 'LLOYD SMITH-McDONALD, RONALD'), as does ß ('MÜLLER WEIß, HANS'); each lost its given name, the Mc/Mac skip being one member of the class. It also missed titlecase digraphs ('Džokić Ljubić, XYZ') and cost a frame per letter. Derek's criterion: a word carries the contrast if it holds a capital and its last letter is lowercase, a trailing ß aside. A surname written in capitals ends in one whatever is glued in front, so no prefix list is needed; 'McDonald', 'Džokić' and decomposed accents pass. Two C-level checks, so the per-letter cost is gone (a record with lowercase particles is a constant +22 over master at any length). Measured against the previous commit: exactly those eight prefixes and Džokić moved, no corpus or case-table name, OFF still equal to master. rules.md#C1, the Policy docstring, customize.rst and the release log drop the false universal; decisions.md#S2 records the chosen option's own leaks and the accepted hyphenated case; case rows and a broader parametrized test pin it. Co-Authored-By: Claude Opus 5.5 --- docs/customize.rst | 2 +- docs/design/decisions.md | 4 ++-- docs/design/rules.md | 18 ++++++++-------- docs/release_log.rst | 2 +- nameparser/_pipeline/_segment.py | 23 +++++++++++---------- nameparser/_pipeline/_vocab.py | 22 +++++++++----------- nameparser/_policy.py | 6 +++--- tests/v2/cases.py | 35 +++++++++++++++++++------------- tests/v2/pipeline/test_vocab.py | 15 ++++++++------ tests/v2/test_policy.py | 2 +- 10 files changed, 69 insertions(+), 60 deletions(-) diff --git a/docs/customize.rst b/docs/customize.rst index 4ea1d92b..8b4a08ef 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -532,7 +532,7 @@ listed below. - ``CapsSuffixes`` - Where an unlisted all-caps word of two or more letters, with no period in it, reads as a credential. The name must contrast it - with a word holding a capital followed by a lowercase letter + with a word holding a capital and ending in a lowercase letter (``Smith``, ``DiCaprio``) that the vocabulary does not claim as a title, particle or credential: a record written wholly in capitals, or wholly in lowercase, keeps every diff --git a/docs/design/decisions.md b/docs/design/decisions.md index ffc9d00a..80315b29 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital directly followed by a lowercase letter, a leading `Mc`/`Mac` skipped where a capital follows it (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and chose the capital-then-lowercase pair with the Mc/Mac skip: in `d'ESTAING`, `al-ASSAD` and `McDONALD` no capital is followed by a lowercase letter once the prefix is set aside, while `McDonald` keeps `Do` and `Mack` its own `Ma`. A word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', and an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast. ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and ending in a lowercase letter, a trailing `ß` aside (`_vocab.written_as_a_name`, two C-level checks), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list; `McDonald`, `Džokić` and a name typed with decomposed accents (`Élodie`) pass; a trailing `ß` is set aside (`WEIß` is written in capitals, `Weiß` is not). Measured 2026-10-02 over the corpora, every case-table text and a grid of 28 prefixes × 6 post-comma parts against the capital-then-lowercase commit: exactly those eight capitalized-surname prefixes and `Džokić Ljubić` moved, no corpus or case-table name, and OFF still equals master (0 of 4212). A word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so does a hyphenated surname whose last part is written in mixed case inside a record otherwise in capitals (`JOHN SMITH-Jones, XYZ` reads suffix `XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 268, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs, per word before the comma, a case-pair test and a fold and wordlist test where it passes, paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 266, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs a frame per word before the comma (the generator) and a fold and wordlist test only where the C-level case test passes, paid only once a caps word is in hand: a record with lowercase particles stays a constant above master whatever its length (`'de ' + 'GAULLE'*k + ', CHARLES'` 253 → 275, 343 → 365 and 631 → 653 at k = 1, 16, 64), where the capital-then-lowercase draft grew a frame per letter. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index da7de369..ce2eb939 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1725,15 +1725,15 @@ C1. Rationale: a credential run after the comma means the name is in initials as two dotted groups. An unlisted all-caps word joins the class in such a part only where the name carries the contrast: one of its own words before the comma written the way a - name is written in mixed case, a capital directly followed by a - lowercase letter (a leading Mc or Mac before a capital aside), - with no period and not - claimed by the vocabulary as a title, particle, connective, - credential or generation. A word written wholly in lowercase never - carries it, so a record that writes its surname in capitals keeps - its given name whatever else it writes in lowercase or beside it - ("GISCARD d'ESTAING, VALÉRY"), and a name written wholly in - lowercase reads the listing form. Where the name carries it, an + name is written in mixed case, holding a capital and ending in a + lowercase letter, with no period and not claimed by the vocabulary + as a title, particle, connective, credential or generation. A + surname written in capitals ends in a capital whatever is glued in + front of it, and a word written wholly in lowercase holds none, so + such a record keeps its given name beside its lowercase particles, + titles and clauses ("GISCARD d'ESTAING, VALÉRY", 'LLOYD + FitzGERALD, RONALD'), and a name written wholly in lowercase reads + the listing form. Where the name carries it, an unlisted word of two capitals after the comma is the paired initials' shape undotted and is decided at this comma exactly as they are, by the sentences above ('García Márquez, MJ' and 'García diff --git a/docs/release_log.rst b/docs/release_log.rst index 8f891db1..54a494d5 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital followed by a lowercase letter (``Smith``, ``DiCaprio``), so a record that writes its surname in capitals keeps its given name whatever it writes beside it (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD McDONALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital and ending in a lowercase letter (``Smith``, ``DiCaprio``). A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index 1a0d2aa3..e0b212a7 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -128,21 +128,22 @@ def texts(seg: tuple[int, ...]) -> list[str]: # and the name carries it only if one of its own words before the # comma (no maiden clause, no delimited content, as `own_words` # defines them for `case_class` above) is written the way a name is - # written in mixed case -- a capital directly followed by a - # lowercase letter (`written_as_a_name`) -- and is not claimed as a + # written in mixed case -- a capital in it and its last letter + # lowercase (`written_as_a_name`) -- and is not claimed as a # title, particle, connective, credential or generation # (`claimed_as_non_name`); a word with a period is an abbreviation - # or an initial, not one. A lone capital ('de GAULLE C') or a - # lowercase-only word has no such pair, and nor has a lowercase - # prefix glued to a capitalized surname, so a record that writes - # its surname in capitals keeps its given name whatever else it - # writes ('LLOYD ap RHYS', 'HAFEZ al-ASSAD', "GISCARD d'ESTAING", - # 'LLOYD McDONALD', 'LLOYD WEBBER née Smith'), and so does a - # mixed-case credential ('LLOYD WEBBER, ANDREW PhD'). + # or an initial, not one. A lone capital ('de GAULLE C'), a + # lowercase-only word and a surname written in capitals with + # anything glued in front ("d'ESTAING", 'FitzGERALD', 'McDONALD') + # all fail it, so such a record keeps its given name beside its + # lowercase particles, clauses and titles ('LLOYD ap RHYS', 'HAFEZ + # al-ASSAD', 'LLOYD WEBBER née Smith') and beside a mixed-case + # credential ('LLOYD WEBBER, ANDREW PhD'). # Asked only once a caps word is in hand: a part with no lowercase # at all is settled in one C-level comparison; past that the walk - # costs a few frames a word before the comma, a fold and a - # wordlist test only where the case test passes. + # costs a frame a word before the comma (the generator), plus a + # call, a fold and a wordlist test only where the C-level case + # test passes. def name_contrast() -> bool: before = "".join([state.tokens[i].text for i in groups[0]]) if before == before.upper(): diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 95b60ef5..e80f1027 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -722,18 +722,16 @@ def in_any_wordlist(n: str, lexicon: Lexicon) -> bool: def written_as_a_name(text: str) -> bool: """Whether TEXT is written the way a name is written in mixed case: - a capital directly followed by a lowercase letter ('Smith', - 'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing") -- #564's test - for the name's case contrast (Derek). A word written wholly in one - case has no such pair, and nor does a lowercase prefix glued to a - surname written in capitals ("d'ESTAING", 'al-ASSAD'). A leading - 'Mc'/'Mac' is skipped where a capital follows it, so 'McDONALD' - has no pair while 'McDonald' keeps 'Do' and 'Mack' its own 'Ma'.""" - if text.startswith("Mc") and text[2:3].isupper(): - text = text[2:] - elif text.startswith("Mac") and text[3:4].isupper(): - text = text[3:] - return any(a.isupper() and b.islower() for a, b in zip(text, text[1:])) + it holds a capital (or titlecase letter) and its last letter is + lowercase -- #564's test for the name's case contrast (Derek). + 'Smith', 'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing", + 'McDonald', 'Džokić' pass. A surname written in capitals fails + whatever is glued in front of it ("d'ESTAING", 'al-ASSAD', + 'McDONALD', 'FitzGERALD', 'DeVITO', 'St-PIERRE'), and so do a lone + capital and a lowercase-only word. A trailing 'ß' is set aside, + having no single capital form: 'WEIß' is written in capitals and + 'Weiß' is not. Two C-level checks, so a word costs no frame.""" + return text != text.lower() and text.rstrip("ß")[-1:].islower() def claimed_as_non_name(n: str, lexicon: Lexicon) -> bool: diff --git a/nameparser/_policy.py b/nameparser/_policy.py index 8ccd77e0..1622d0eb 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -709,9 +709,9 @@ class Policy: unlisted_dotted_suffixes: bool = True #: Where an UNLISTED all-caps word of two or more letters, with no #: period in it, reads as a credential (:class:`CapsSuffixes`). The - #: name must contrast it: a word of the name with a capital followed - #: by a lowercase letter ("Smith", "DiCaprio") that the vocabulary - #: does not claim as a title, particle + #: name must contrast it: a word of the name holding a capital and + #: ending in a lowercase letter ("Smith", "DiCaprio") that the + #: vocabulary does not claim as a title, particle #: or credential -- a record written wholly in capitals or wholly in #: lowercase keeps every word a name word. A listed member keeps its #: own case lean in every setting ("Jack MA" gives suffix ``MA``), diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 61d326b7..7a17a67c 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2922,7 +2922,7 @@ def _check_cjk_shape_purity(self) -> None: {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "PhD"}, notes="#564 boundary: an all-caps word joins C1's run only where " "the NAME carries the contrast, a word of its own before " - "the comma with a capital followed by a lowercase letter " + "the comma holding a capital and ending in lowercase " "-- here the only " "lowercase is the credential's own, and a record written " "wholly in capitals keeps its given name. A review-round " @@ -2950,7 +2950,7 @@ def _check_cjk_shape_purity(self) -> None: "LLOYD WEBBER Jr., ANDREW", {"given": "ANDREW", "family": "LLOYD WEBBER", "suffix": "Jr."}, notes="#564 boundary: only the name's own words written as a " - "name is (a capital followed by a lowercase letter), " + "name is (a capital, and a lowercase last letter), " "unclaimed by the vocabulary and unmarked by a period, " "supply the contrast; 'Jr.' is none of those, " "so the 'SURNAME Jr., GIVEN' record " @@ -2973,9 +2973,9 @@ def _check_cjk_shape_purity(self) -> None: "GISCARD d'ESTAING, VALÉRY", {"given": "VALÉRY", "family": "GISCARD d'ESTAING"}, notes="#564 (Derek): the name carries the contrast only through " - "a word with a capital followed by a lowercase letter that " - "no wordlist claims as non-name text. A lowercase particle " - "glued to a capitalized surname has no such pair -- the " + "a word holding a capital and ending in lowercase that no " + "wordlist claims as non-name text. A surname written in " + "capitals ends in a capital whatever is glued to it -- the " "French convention the trailing slot stays off for. A " "draft counting any lowercase letter read suffix " "'VALÉRY'"), @@ -2988,8 +2988,8 @@ def _check_cjk_shape_purity(self) -> None: Case("an_all_lowercase_name_carries_no_contrast_for_the_caps_shape", "john smith, XYZ", {"given": "XYZ", "family": "john smith"}, - notes="#564 (Derek): a lowercase-only name has no capital " - "followed by a lowercase letter, so the default keeps the " + notes="#564 (Derek): a lowercase-only name holds no capital, " + "so the default keeps the " "listing form; the cost of the criterion, accepted with " "it"), Case("a_lone_capital_supplies_no_contrast_for_the_caps_shape", @@ -2999,21 +2999,28 @@ def _check_cjk_shape_purity(self) -> None: "with no lowercase letter behind it. A draft testing " "`istitle()`, which accepts 'C', read suffix 'CHARLES' " "with no given name"), + Case("a_glued_capitalized_prefix_supplies_no_contrast", + "LLOYD FitzGERALD, RONALD", + {"given": "RONALD", "family": "LLOYD FitzGERALD"}, + notes="#564 (Derek): a surname written in capitals ends in a " + "capital whatever is glued in front of it ('FitzGERALD', " + "'McDONALD', 'DeVITO', 'St-PIERRE'), so the record keeps " + "its given name. The capital-then-lowercase draft took " + "'Fi' as the contrast and read suffix 'RONALD'"), Case("a_mc_prefix_on_a_capitalized_surname_supplies_no_contrast", "LLOYD McDONALD, RONALD", {"given": "RONALD", "family": "LLOYD McDONALD"}, - notes="#564 (Derek): a leading 'Mc'/'Mac' is skipped where a " - "capital follows it, so a record writing 'McDONALD' keeps " - "its given name, while 'McDonald' still carries the " - "contrast through 'Do'"), + notes="#564: 'McDONALD' ends in a capital, so no Mc/Mac case is " + "needed; 'McDonald' ends lowercase and carries the " + "contrast"), Case("an_interior_capital_carries_the_contrast", "DiCaprio LaBeouf, XYZ", {"given": "DiCaprio", "family": "LaBeouf", "suffix": "XYZ"}, classification="fix(#564)", ambiguities=("suffix-or-name",), - notes="#564 (Derek): a capital followed by a lowercase letter " - "anywhere in the word is the contrast, so an interior " - "capital counts ('DiCaprio', 'IJzerman', 'al-Rashid'). " + notes="#564 (Derek): a capital anywhere in a word ending " + "lowercase is the contrast, so an interior capital counts " + "('DiCaprio', 'IJzerman', 'al-Rashid'). " "The `istitle()` draft read given 'XYZ' here"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index fcd4df46..281f197a 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -895,14 +895,17 @@ def test_is_single_letter_numeral() -> None: @pytest.mark.parametrize("text, expected", [ ("Smith", True), ("DiCaprio", True), ("IJzerman", True), ("al-Rashid", True), ("d'Estaing", True), ("McDonald", True), - ("MacLeod", True), ("Mack", True), ("Macy", True), + ("MacLeod", True), ("Mack", True), ("O'Neil", True), + ("Džokić", True), ("E\u0301lodie", True), ("Weiß", True), ("SMITH", False), ("smith", False), ("C", False), ("ap", False), ("d'ESTAING", False), ("al-ASSAD", False), ("McDONALD", False), - ("MacDONALD", False), + ("MacDONALD", False), ("FitzGERALD", False), ("DeVITO", False), + ("LaFLEUR", False), ("St-PIERRE", False), ("SMITH-McDONALD", False), + ("WEIß", False), ("O'NEIL", False), ]) def test_written_as_a_name(text: str, expected: bool) -> None: - # #564 (Derek): the name's case contrast is a capital directly - # followed by a lowercase letter, a leading Mc/Mac skipped where a - # capital follows it -- so a lowercase prefix glued to a surname - # written in capitals carries none, and an interior capital does. + # #564 (Derek): the name's case contrast is a word holding a + # capital and ending in a lowercase letter -- a surname written in + # capitals ends in one whatever is glued in front of it, and a + # trailing ß, which has no single capital, is set aside. assert written_as_a_name(text) is expected diff --git a/tests/v2/test_policy.py b/tests/v2/test_policy.py index e69abcd0..a612cd56 100644 --- a/tests/v2/test_policy.py +++ b/tests/v2/test_policy.py @@ -871,7 +871,7 @@ def test_unlisted_caps_suffixes_on_reads_an_all_caps_word() -> None: # assertions are about the TRAILING slot EVERYWHERE adds. Default # costs, same harness against master: 'Smith, John' 183 -> 183, # 'Smith, XYZ' 182 -> 182 (the comma test needs two words before - # the comma), 'John Smith, XYZ' 251 -> 268 (decisions.md#S2). + # the comma), 'John Smith, XYZ' 251 -> 266 (decisions.md#S2). from nameparser import Parser on = Parser(policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)) From 07c5662727954cf4ae5db1d9e0e1987e2d841df0 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Fri, 2 Oct 2026 00:41:24 -0700 Subject: [PATCH 10/13] fix(S2): #564 -- the contrast reads a word's last LETTER, not character MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The previous commit tested the last character, so a name typed with decomposed accents ending in an accented letter lost the contrast ('José André, XYZ', 'Lê Thị Hà, XYZ', 'René Noé, XYZ' read given where their composed spellings read suffix), as did "Jones'", 'Smith2' and a glued 'Smith)'. written_as_a_name now composes the word (NFC), passes over trailing non-letters, and passes over a lowercase letter with no single capital form (ß, ĸ), which generalizes the ß rule. One call, no generator. Measured against the previous commit: the four decomposed-accent names and 'JOHN Smith)' regain the contrast and the ĸ record keeps its given name; nothing else moves, and OFF equals master on 4292 inputs. Also from the review: rules.md#C1 states the last-letter reading and pins 'MÜLLER WEIß, HANS'; decisions.md#S2 records the true blast radius of the previous commit (seven prefixes plus ß and Džokić, and 'JOHN O'NEILL's' / 'JEAN-pierre DUPONT' / 'MARY-kate OLSEN', now accepted costs) with a recompute recipe; the cost prose says what it is, about four frames per word before the comma and constant in a word's letters; the fix(#564) ledger comments drop the wording decisions.md calls wrong. Co-Authored-By: Claude Opus 5.5 --- docs/customize.rst | 2 +- docs/design/decisions.md | 4 +-- docs/design/rules.md | 5 +++- docs/release_log.rst | 2 +- nameparser/_pipeline/_segment.py | 9 +++--- nameparser/_pipeline/_vocab.py | 30 ++++++++++++++++---- nameparser/_policy.py | 4 +-- tests/v2/cases.py | 10 +++++++ tests/v2/pipeline/test_vocab.py | 11 +++++-- tests/v2/test_ledger_guards.py | 18 ++++++------ tools/differential/corpus_rules.jsonl | 1 + tools/differential/expected_since_1.4.0.toml | 6 ++-- tools/differential/expected_since_2.0.0.toml | 6 ++-- tools/differential/expected_since_2.1.0.toml | 6 ++-- tools/differential/expected_since_2.2.0.toml | 6 ++-- tools/differential/expected_since_2.3.0.toml | 6 ++-- 16 files changed, 84 insertions(+), 42 deletions(-) diff --git a/docs/customize.rst b/docs/customize.rst index 8b4a08ef..d4713f16 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -532,7 +532,7 @@ listed below. - ``CapsSuffixes`` - Where an unlisted all-caps word of two or more letters, with no period in it, reads as a credential. The name must contrast it - with a word holding a capital and ending in a lowercase letter + with a word holding a capital whose last letter is lowercase (``Smith``, ``DiCaprio``) that the vocabulary does not claim as a title, particle or credential: a record written wholly in capitals, or wholly in lowercase, keeps every diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 80315b29..36ed5bd3 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,10 +687,10 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and ending in a lowercase letter, a trailing `ß` aside (`_vocab.written_as_a_name`, two C-level checks), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list; `McDonald`, `Džokić` and a name typed with decomposed accents (`Élodie`) pass; a trailing `ß` is set aside (`WEIß` is written in capitals, `Weiß` is not). Measured 2026-10-02 over the corpora, every case-table text and a grid of 28 prefixes × 6 post-comma parts against the capital-then-lowercase commit: exactly those eight capitalized-surname prefixes and `Džokić Ljubić` moved, no corpus or case-table name, and OFF still equals master (0 of 4212). A word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so does a hyphenated surname whose last part is written in mixed case inside a record otherwise in capitals (`JOHN SMITH-Jones, XYZ` reads suffix `XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and whose last LETTER is lowercase (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list, and `McDonald` and `Džokić` pass. "Last letter" is the last LETTER, not character: the word is composed (NFC) first, trailing non-letters are passed over, and so is a lowercase letter with no single capital form (`ß`, `ĸ`), which is no evidence of case — `WEIß` is written in capitals, `Weiß` is not. The first cut tested the last character, and its review found every name typed with decomposed accents and ending in an accented letter losing the contrast (`José André, XYZ`, `Lê Thị Hà, XYZ`, `René Noé, XYZ`, `Chloé Zoé, MBA XYZ`, read given where their composed spellings read suffix), and `Jones'`, `Smith2`, `Smith)` with them. Measured 2026-10-02 against the capital-then-lowercase commit, over the corpora, every case-table text and a grid of prefixes × six post-comma parts: the seven capitalized-surname prefixes (`FitzGERALD`, `DeVITO`, `LaFLEUR`, `VanDYKE`, `DuBOIS`, `St-PIERRE`, `SMITH-McDONALD`), `MÜLLER WEIß` and `Džokić Ljubić` moved as intended, and so did `JOHN O'NEILL's`, `JEAN-pierre DUPONT` and `MARY-kate OLSEN` (accepted below); no corpus or case-table name moved, and OFF equals master on all 4292 inputs. Recompute: parse each input under `CapsSuffixes.OFF` on this tree and under `unlisted_caps_suffixes=False` on `git archive origin/master`, and under each setting on this tree and on the parent commit, comparing `as_dict()`, the ambiguity kinds and `capitalized(force=True)`. A word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so do two writings in an otherwise all-caps record that end in a lowercase letter: a glued possessive or plural (`JOHN O'NEILL's, XYZ`, `JOHN SMITHs, XYZ` read suffix `XYZ`) and a hyphenated part written in mixed case or lowercase (`JOHN SMITH-Jones, XYZ`, `JEAN-pierre DUPONT, XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. - COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 266, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast costs a frame per word before the comma (the generator) and a fold and wordlist test only where the C-level case test passes, paid only once a caps word is in hand: a record with lowercase particles stays a constant above master whatever its length (`'de ' + 'GAULLE'*k + ', CHARLES'` 253 → 275, 343 → 365 and 631 → 653 at k = 1, 16, 64), where the capital-then-lowercase draft grew a frame per letter. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. + COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 266, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast is linear in the words before the comma, about four frames a word (the generator, the case test's call, and the own-words walk's fold and marker test; `'GAULLE ' + 'ap '*k + 'GAULLE, CHARLES'` is +7, +26, +38, +54 over master at k = 0, 1, 4, 8), and constant in a word's letters (`'de ' + 'GAULLE'*k + ', CHARLES'` 253 → 275, 343 → 365 and 631 → 653 at k = 1, 16, 64), where the capital-then-lowercase draft grew a frame per letter; all paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index ce2eb939..3077b181 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1726,7 +1726,9 @@ C1. Rationale: a credential run after the comma means the name is in the class in such a part only where the name carries the contrast: one of its own words before the comma written the way a name is written in mixed case, holding a capital and ending in a - lowercase letter, with no period and not claimed by the vocabulary + lowercase letter (its last letter, past any trailing mark, and + setting aside a lowercase letter with no capital of its own, as ß + has none), with no period and not claimed by the vocabulary as a title, particle, connective, credential or generation. A surname written in capitals ends in a capital whatever is glued in front of it, and a word written wholly in lowercase holds none, so @@ -1874,6 +1876,7 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, LEED AP" → suffix="LEED AP" "Smith, XYZ" → given="XYZ" · boundary "García Márquez, MJ" → given="MJ" · boundary + "MÜLLER WEIß, HANS" → given="HANS" · boundary "García Márquez, MJ PhD" → given="MJ" · boundary "García Márquez, MJ JK" → suffix="MJ JK" "John Smith, PhD XYZ" → suffix="PhD XYZ" diff --git a/docs/release_log.rst b/docs/release_log.rst index 54a494d5..766196d7 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -22,7 +22,7 @@ Release Log - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital and ending in a lowercase letter (``Smith``, ``DiCaprio``). A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) + - **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital whose last letter is lowercase (``Smith``, ``DiCaprio``); a name typed with decomposed accents reads as its composed spelling. A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564) - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index e0b212a7..fb6cc763 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -140,10 +140,11 @@ def texts(seg: tuple[int, ...]) -> list[str]: # al-ASSAD', 'LLOYD WEBBER née Smith') and beside a mixed-case # credential ('LLOYD WEBBER, ANDREW PhD'). # Asked only once a caps word is in hand: a part with no lowercase - # at all is settled in one C-level comparison; past that the walk - # costs a frame a word before the comma (the generator), plus a - # call, a fold and a wordlist test only where the C-level case - # test passes. + # at all is settled in one C-level comparison; past that it is + # linear in the words before the comma, about four frames a word + # (the generator, the case test's call, and the own-words walk's + # fold and marker test), plus a fold and a wordlist test where the + # case test passes -- and constant in a word's letters. def name_contrast() -> bool: before = "".join([state.tokens[i].text for i in groups[0]]) if before == before.upper(): diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index e80f1027..875fe57a 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -722,16 +722,36 @@ def in_any_wordlist(n: str, lexicon: Lexicon) -> bool: def written_as_a_name(text: str) -> bool: """Whether TEXT is written the way a name is written in mixed case: - it holds a capital (or titlecase letter) and its last letter is + it holds a capital (or titlecase letter) and its last LETTER is lowercase -- #564's test for the name's case contrast (Derek). 'Smith', 'DiCaprio', 'IJzerman', 'al-Rashid', "d'Estaing", 'McDonald', 'Džokić' pass. A surname written in capitals fails whatever is glued in front of it ("d'ESTAING", 'al-ASSAD', 'McDONALD', 'FitzGERALD', 'DeVITO', 'St-PIERRE'), and so do a lone - capital and a lowercase-only word. A trailing 'ß' is set aside, - having no single capital form: 'WEIß' is written in capitals and - 'Weiß' is not. Two C-level checks, so a word costs no frame.""" - return text != text.lower() and text.rstrip("ß")[-1:].islower() + capital and a lowercase-only word. + + The last letter is found, not the last character: the word is + composed first (NFC), so a name typed with decomposed accents reads + as its composed spelling ('André' ends in 'é', not in the combining + accent), and trailing non-letters are passed over ("Jones'", + 'Smith2', 'Smith)'). A lowercase letter with no single capital form + is passed over too, being no evidence of case: 'WEIß' and 'KAĸ' are + written in capitals, 'Weiß' is not. One call and no generator, so a + word costs one frame.""" + if text == text.lower(): + return False + word = unicodedata.normalize("NFC", text) + i = len(word) + while i: + i -= 1 + ch = word[i] + if not ch.isalpha(): + continue + upper = ch.upper() + if ch.islower() and (upper == ch or len(upper) != 1): + continue + return ch.islower() + return False def claimed_as_non_name(n: str, lexicon: Lexicon) -> bool: diff --git a/nameparser/_policy.py b/nameparser/_policy.py index 1622d0eb..057183ef 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -709,8 +709,8 @@ class Policy: unlisted_dotted_suffixes: bool = True #: Where an UNLISTED all-caps word of two or more letters, with no #: period in it, reads as a credential (:class:`CapsSuffixes`). The - #: name must contrast it: a word of the name holding a capital and - #: ending in a lowercase letter ("Smith", "DiCaprio") that the + #: name must contrast it: a word of the name holding a capital whose + #: last letter is lowercase ("Smith", "DiCaprio") that the #: vocabulary does not claim as a title, particle #: or credential -- a record written wholly in capitals or wholly in #: lowercase keeps every word a name word. A listed member keeps its diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 7a17a67c..fdc07022 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -3022,6 +3022,16 @@ def _check_cjk_shape_purity(self) -> None: "lowercase is the contrast, so an interior capital counts " "('DiCaprio', 'IJzerman', 'al-Rashid'). " "The `istitle()` draft read given 'XYZ' here"), + Case("a_name_typed_with_decomposed_accents_carries_the_contrast", + "Jose\u0301 Andre\u0301, XYZ", + {"given": "Jose\u0301", "family": "Andre\u0301", "suffix": "XYZ"}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="#564: the contrast test reads a word's last LETTER after " + "composing it, so 'André' typed with a combining accent " + "ends in 'é' and reads as its composed spelling does. A " + "draft reading the last character took the accent and " + "kept given 'XYZ'"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index 281f197a..8e3e281f 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -896,7 +896,10 @@ def test_is_single_letter_numeral() -> None: ("Smith", True), ("DiCaprio", True), ("IJzerman", True), ("al-Rashid", True), ("d'Estaing", True), ("McDonald", True), ("MacLeod", True), ("Mack", True), ("O'Neil", True), - ("Džokić", True), ("E\u0301lodie", True), ("Weiß", True), + ("Džokić", True), ("E\u0301lodie", True), ("Andre\u0301", True), + ("Ha\u0300", True), ("Jones'", True), ("Smith2", True), + ("Smith)", True), ("Weiß", True), ("KAĸ", False), + ("ANDRE\u0301", False), ("SMITH", False), ("smith", False), ("C", False), ("ap", False), ("d'ESTAING", False), ("al-ASSAD", False), ("McDONALD", False), ("MacDONALD", False), ("FitzGERALD", False), ("DeVITO", False), @@ -906,6 +909,8 @@ def test_is_single_letter_numeral() -> None: def test_written_as_a_name(text: str, expected: bool) -> None: # #564 (Derek): the name's case contrast is a word holding a # capital and ending in a lowercase letter -- a surname written in - # capitals ends in one whatever is glued in front of it, and a - # trailing ß, which has no single capital, is set aside. + # capitals ends in one whatever is glued in front of it. The last + # LETTER: composed first, so a decomposed accent at the end counts + # as its letter; trailing non-letters and a lowercase letter with + # no single capital (ß, ĸ) are passed over. assert written_as_a_name(text) is expected diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index b8e2f805..6f693fed 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -3811,11 +3811,12 @@ def _claim(rule: dict) -> _Claim: # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - # 2026-10-01, #564: 423 -> 429, 'John Smith, XYZ', 'Smith, + # 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith, # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', - # 'García Márquez, MJ JK' and 'John Smith, PhD XYZ', #564's - # rules.md#C1 examples. Reach, verified name by name. - _Claim(429, ('given', 'suffix', 'title'), "2c336f3d3ae0", None), + # 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER + # WEIß, HANS', #564's rules.md#C1 examples. Reach, verified + # name by name. + _Claim(430, ('given', 'suffix', 'title'), "dcfb3a9638a9", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3919,11 +3920,12 @@ def _claim(rule: dict) -> _Claim: # case-row names, every one a comma name. Reach, verified # name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - # 2026-10-01, #564: 423 -> 429, 'John Smith, XYZ', 'Smith, + # 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith, # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', - # 'García Márquez, MJ JK' and 'John Smith, PhD XYZ', #564's - # rules.md#C1 examples. Reach, verified name by name. - _Claim(429, ('family', 'given'), "2c336f3d3ae0", None), + # 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER + # WEIß, HANS', #564's rules.md#C1 examples. Reach, verified + # name by name. + _Claim(430, ('family', 'given'), "dcfb3a9638a9", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 391fb8e9..b9ae7ff5 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -294,6 +294,7 @@ "Mr. Jack and Jill" "Mr. Johnson" "Mrs. Garcia" +"MÜLLER WEIß, HANS" "Ménil Christophe de" "Ménil de" "Nguyen Thi Van" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index ce6e0269..01cee6ec 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4860,9 +4860,9 @@ fields = ["title", "given", "middle"] issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" # rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to # CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more -# name words before it reads an unlisted all-caps word of three or -# more letters, or a run holding one, as the credential run and -# reports the call. The all-caps SURNAME convention never writes the +# name words before it reads an unlisted all-caps word as the +# credential, alone or in C1's run, where the name contrasts it, and +# reports the call; a lone two-letter word is declined as initials. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); # 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 3d6e5a57..8e71d300 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3918,9 +3918,9 @@ orders = ["DEFAULT"] issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" # rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to # CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more -# name words before it reads an unlisted all-caps word of three or -# more letters, or a run holding one, as the credential run and -# reports the call. The all-caps SURNAME convention never writes the +# name words before it reads an unlisted all-caps word as the +# credential, alone or in C1's run, where the name contrasts it, and +# reports the call; a lone two-letter word is declined as initials. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); # 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 55819f32..d7346982 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3829,9 +3829,9 @@ orders = ["DEFAULT"] issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" # rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to # CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more -# name words before it reads an unlisted all-caps word of three or -# more letters, or a run holding one, as the credential run and -# reports the call. The all-caps SURNAME convention never writes the +# name words before it reads an unlisted all-caps word as the +# credential, alone or in C1's run, where the name contrasts it, and +# reports the call; a lone two-letter word is declined as initials. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); # 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 3ed62be6..14296fae 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2212,9 +2212,9 @@ orders = ["DEFAULT"] issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" # rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to # CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more -# name words before it reads an unlisted all-caps word of three or -# more letters, or a run holding one, as the credential run and -# reports the call. The all-caps SURNAME convention never writes the +# name words before it reads an unlisted all-caps word as the +# credential, alone or in C1's run, where the name contrasts it, and +# reports the call; a lone two-letter word is declined as initials. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); # 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 8f197df2..7c08354f 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1467,9 +1467,9 @@ orders = ["DEFAULT"] issue = "fix(#564) an unlisted all-caps word after a comma behind two name words is a credential by default" # rules.md#C1 and S2 (#564): Policy.unlisted_caps_suffixes defaults to # CapsSuffixes.AFTER_COMMA, so the part after a comma with two or more -# name words before it reads an unlisted all-caps word of three or -# more letters, or a run holding one, as the credential run and -# reports the call. The all-caps SURNAME convention never writes the +# name words before it reads an unlisted all-caps word as the +# credential, alone or in C1's run, where the name contrasts it, and +# reports the call; a lone two-letter word is declined as initials. The all-caps SURNAME convention never writes the # capitals there. 'John Smith, XYZ' and 'John Smith, LEED AP' are the # C1 examples (the latter closing the #291 deviation the doc carried); # 'The Rt Hon Kenneth Clarke QC MP, HMG' is a radar corpus name. From d7f4dac878ce675224d1bdede487e567c38f06d6 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Fri, 2 Oct 2026 01:05:21 -0700 Subject: [PATCH 11/13] fix(S2): #564 -- a caseless letter is no evidence of case; a runnable recipe MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The last-letter scan stopped at a caseless letter, so a mixed-case name whose words end in one lost the contrast: 'Asmāʾ Wafāʾ, XYZ' and 'JOHN Jonesʼ, XYZ' (U+02BC) kept given 'XYZ' where 'JOHN Jones', XYZ' read suffix. written_as_a_name now passes over every letter without case, as it passes over non-letters, so the rule is one sentence: the last cased letter, read after composing the word, setting aside a lowercase letter whose capital is not a single letter. rules.md#C1 states exactly that. decisions.md#S2's recipe compared against "the parent commit" while its numbers were measured against a16927c6, and its grid was unnamed. It now names both comparators by SHA and lists every input, and was run as written: against a16927c6 exactly the intended classes move, and OFF equals master on all 4286 inputs. Correction to the previous commit's message: "nothing else moves" also missed 'JOHN SMIfi' (a ligature), which moved there and is unchanged here. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/design/rules.md | 9 +++++---- nameparser/_pipeline/_vocab.py | 22 ++++++++++++---------- tests/v2/cases.py | 9 +++++++++ tests/v2/pipeline/test_vocab.py | 2 ++ 5 files changed, 29 insertions(+), 15 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 36ed5bd3..c441204b 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,7 +687,7 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and whose last LETTER is lowercase (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list, and `McDonald` and `Džokić` pass. "Last letter" is the last LETTER, not character: the word is composed (NFC) first, trailing non-letters are passed over, and so is a lowercase letter with no single capital form (`ß`, `ĸ`), which is no evidence of case — `WEIß` is written in capitals, `Weiß` is not. The first cut tested the last character, and its review found every name typed with decomposed accents and ending in an accented letter losing the contrast (`José André, XYZ`, `Lê Thị Hà, XYZ`, `René Noé, XYZ`, `Chloé Zoé, MBA XYZ`, read given where their composed spellings read suffix), and `Jones'`, `Smith2`, `Smith)` with them. Measured 2026-10-02 against the capital-then-lowercase commit, over the corpora, every case-table text and a grid of prefixes × six post-comma parts: the seven capitalized-surname prefixes (`FitzGERALD`, `DeVITO`, `LaFLEUR`, `VanDYKE`, `DuBOIS`, `St-PIERRE`, `SMITH-McDONALD`), `MÜLLER WEIß` and `Džokić Ljubić` moved as intended, and so did `JOHN O'NEILL's`, `JEAN-pierre DUPONT` and `MARY-kate OLSEN` (accepted below); no corpus or case-table name moved, and OFF equals master on all 4292 inputs. Recompute: parse each input under `CapsSuffixes.OFF` on this tree and under `unlisted_caps_suffixes=False` on `git archive origin/master`, and under each setting on this tree and on the parent commit, comparing `as_dict()`, the ambiguity kinds and `capitalized(force=True)`. A word written wholly in lowercase never carries it, so no list has to know `ap` or `thi`. The predicate asks a different question from the caps shape's "unlisted" (`in_any_wordlist`): the surname and bound-given lists claim a word AS name text, so a caller's surname list carries the contrast (a draft sharing one predicate read `Smith Jones, XYZ` as given 'XYZ' under `add(surnames={"smith", "jones"})`). ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so do two writings in an otherwise all-caps record that end in a lowercase letter: a glued possessive or plural (`JOHN O'NEILL's, XYZ`, `JOHN SMITHs, XYZ` read suffix `XYZ`) and a hyphenated part written in mixed case or lowercase (`JOHN SMITH-Jones, XYZ`, `JEAN-pierre DUPONT, XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and whose last LETTER is lowercase (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list, and `McDonald` and `Džokić` pass. "Last letter" is the last letter that carries CASE, not the last character: the word is composed (NFC) first; every non-letter and every caseless letter is passed over (`Jones'`, `Smith2`, `Smith)`, `Jonesʼ` with U+02BC, the transliteration letters in `Wafāʾ`, a trailing `李`); and so is a lowercase letter whose capital is not one character (`ß` → `SS`, `ĸ` with none, the `fi` ligature), which is no evidence of case — `WEIß` is written in capitals, `Weiß` is not. Two cuts got this wrong: the first tested the last character, and its review found every name typed with decomposed accents and ending in an accented letter losing the contrast (`José André, XYZ`, `Lê Thị Hà, XYZ`, `René Noé, XYZ`, `Chloé Zoé, MBA XYZ` read given where their composed spellings read suffix), with `Jones'`, `Smith2` and `Smith)`; the second skipped non-letters but stopped at a caseless letter, so `Asmāʾ Wafāʾ, XYZ` and `JOHN Jonesʼ, XYZ` kept given 'XYZ' (its review). MEASURED 2026-10-02 against `a16927c6`, the capital-then-lowercase commit: the seven capitalized-surname prefixes (`FitzGERALD`, `DeVITO`, `LaFLEUR`, `VanDYKE`, `DuBOIS`, `St-PIERRE`, `SMITH-McDONALD`), `MÜLLER WEIß` and `KUJAĸ de PETERSEN` keep their given names, `Džokić Ljubić` gains the contrast, `JOHN O'NEILL's` and `JEAN-pierre DUPONT` gain one (accepted below), and nothing else moves; the decomposed-accent names, `Asmāʾ Wafāʾ` and `JOHN Jonesʼ` read as `a16927c6` reads them. OFF on this tree equals `unlisted_caps_suffixes=False` on `7394078a` (origin/master at the time) on all 4286 inputs. RECIPE, the comparator being those two commits (`git archive` each): the inputs are every name in `tools/differential/corpus*.jsonl`, every double-quoted string literal in `tests/v2/cases.py` holding a comma (fragments of case notes included, harmlessly), and each of these prefixes joined by ", " to each of the parts `XYZ`, `ANDREW`, `ANDREW PhD`, `PhD XYZ`, `MJ`, `XYZ Jr.` — `John Smith`, `DiCaprio LaBeouf`, `IJzerman IJsselmeer`, `al-Rashid al-Hassan`, `d'Estaing d'Orléans`, `McDonald MacLeod`, `Džokić Ljubić`, `Asmāʾ Wafāʾ`, `John Smith-Jones`, `john smith`, `García y López`, `LLOYD WEBBER`, `GISCARD d'ESTAING`, `HAFEZ al-ASSAD`, `LLOYD McDONALD`, `LLOYD FitzGERALD`, `PAOLO DeVITO`, `JEAN LaFLEUR`, `DICK VanDYKE`, `PAUL DuBOIS`, `JEAN St-PIERRE`, `LLOYD SMITH-McDONALD`, `MÜLLER WEIß`, `KUJAĸ de PETERSEN`, `LLOYD ap RHYS`, `de GAULLE C`, `Mr LLOYD WEBBER`, `LLOYD WEBBER née Smith`, `JOHN O'NEILL's`, `JOHN SMITHs`, `JEAN-pierre DUPONT`, `JOHN Jones'`, `JOHN Jonesʼ`, `JOHN Smith)`, `JOHN Smith2`, and the NFD forms of `José André`, `Lê Thị Hà`, `René Noé`, `Chloé Zoé`. Parse each under OFF, AFTER_COMMA and EVERYWHERE and compare `as_dict()`, the ambiguity kinds and `capitalized(force=True)`; the population is a fact of the tree it runs on (4286 here), since the corpora and the case table grow. ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so do two writings in an otherwise all-caps record that end in a lowercase letter: a glued possessive or plural (`JOHN O'NEILL's, XYZ`, `JOHN SMITHs, XYZ` read suffix `XYZ`) and a hyphenated part written in mixed case or lowercase (`JOHN SMITH-Jones, XYZ`, `JEAN-pierre DUPONT, XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 266, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast is linear in the words before the comma, about four frames a word (the generator, the case test's call, and the own-words walk's fold and marker test; `'GAULLE ' + 'ap '*k + 'GAULLE, CHARLES'` is +7, +26, +38, +54 over master at k = 0, 1, 4, 8), and constant in a word's letters (`'de ' + 'GAULLE'*k + ', CHARLES'` 253 → 275, 343 → 365 and 631 → 653 at k = 1, 16, 64), where the capital-then-lowercase draft grew a frame per letter; all paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA. diff --git a/docs/design/rules.md b/docs/design/rules.md index 3077b181..d90afc95 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1725,10 +1725,11 @@ C1. Rationale: a credential run after the comma means the name is in initials as two dotted groups. An unlisted all-caps word joins the class in such a part only where the name carries the contrast: one of its own words before the comma written the way a - name is written in mixed case, holding a capital and ending in a - lowercase letter (its last letter, past any trailing mark, and - setting aside a lowercase letter with no capital of its own, as ß - has none), with no period and not claimed by the vocabulary + name is written in mixed case, holding a capital with its last + cased letter lowercase — read after composing the word, passing + over every non-letter and caseless letter and any lowercase letter + whose capital is not a single letter, as ß's is not — with no + period and not claimed by the vocabulary as a title, particle, connective, credential or generation. A surname written in capitals ends in a capital whatever is glued in front of it, and a word written wholly in lowercase holds none, so diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 875fe57a..b5e6ef5e 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -730,14 +730,16 @@ def written_as_a_name(text: str) -> bool: 'McDONALD', 'FitzGERALD', 'DeVITO', 'St-PIERRE'), and so do a lone capital and a lowercase-only word. - The last letter is found, not the last character: the word is - composed first (NFC), so a name typed with decomposed accents reads - as its composed spelling ('André' ends in 'é', not in the combining - accent), and trailing non-letters are passed over ("Jones'", - 'Smith2', 'Smith)'). A lowercase letter with no single capital form - is passed over too, being no evidence of case: 'WEIß' and 'KAĸ' are - written in capitals, 'Weiß' is not. One call and no generator, so a - word costs one frame.""" + The letter read is the last one that carries case evidence, not the + last character. The word is composed first (NFC), so a name typed + with decomposed accents reads as its composed spelling ('André' + ends in 'é', not in the combining accent). Passed over, as no + evidence of case: every non-letter ("Jones'", 'Smith2', 'Smith)'), + every caseless letter ("Jonesʼ" with U+02BC, 'Wafāʾ', 'Smith李'), + and a lowercase letter whose capital is not a single character + ('ß' -> 'SS', 'ĸ' with none): 'WEIß' and 'KAĸ' are written in + capitals, 'Weiß' is not. One call and no generator, so a word costs + one frame.""" if text == text.lower(): return False word = unicodedata.normalize("NFC", text) @@ -745,8 +747,8 @@ def written_as_a_name(text: str) -> bool: while i: i -= 1 ch = word[i] - if not ch.isalpha(): - continue + if not (ch.isupper() or ch.islower() or ch.istitle()): + continue # a non-letter or a caseless letter upper = ch.upper() if ch.islower() and (upper == ch or len(upper) != 1): continue diff --git a/tests/v2/cases.py b/tests/v2/cases.py index fdc07022..58ce7bcc 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -3032,6 +3032,15 @@ def _check_cjk_shape_purity(self) -> None: "ends in 'é' and reads as its composed spelling does. A " "draft reading the last character took the accent and " "kept given 'XYZ'"), + Case("a_trailing_caseless_letter_is_no_evidence_against_the_contrast", + "Asmāʾ Wafāʾ, XYZ", + {"given": "Asmāʾ", "family": "Wafāʾ", "suffix": "XYZ"}, + classification="fix(#564)", + ambiguities=("suffix-or-name",), + notes="#564: the contrast test reads the last letter that has " + "case, passing over the caseless transliteration letter " + "'ʾ' as it passes over an apostrophe. A draft stopping at " + "it kept given 'XYZ'"), Case("a_capitalized_given_name_behind_a_two_word_surname_is_the_accepted_cost", "García Márquez, JUAN Jr.", {"given": "García", "family": "Márquez", "suffix": "JUAN Jr."}, diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index 8e3e281f..3442447f 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -899,6 +899,8 @@ def test_is_single_letter_numeral() -> None: ("Džokić", True), ("E\u0301lodie", True), ("Andre\u0301", True), ("Ha\u0300", True), ("Jones'", True), ("Smith2", True), ("Smith)", True), ("Weiß", True), ("KAĸ", False), + ("Jones\u02bc", True), ("Wafāʾ", True), ("Smith李", True), + ("SMITHʾ", False), ("ANDRE\u0301", False), ("SMITH", False), ("smith", False), ("C", False), ("ap", False), ("d'ESTAING", False), ("al-ASSAD", False), ("McDONALD", False), From 2479eabb6c594615936328e8ddd8398df651e23b Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Fri, 2 Oct 2026 01:14:46 -0700 Subject: [PATCH 12/13] fix(S2): #564 -- no unreachable fall-through in written_as_a_name MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Codecov flagged the trailing `return False` as the PR's one uncovered line. It was dead: any word failing `text == text.lower()` holds an uppercase or titlecase letter (checked over every code point, and NFC keeps one), and the scan returns on reaching it. The capital test now sits on the scan's return instead, so a word with no letter carrying case evidence ('李', '2', 'ß', '') falls through and is pinned by new rows. Same answers as before on 1,421,107 inputs (the corpora and case-table words, every code point, 300,000 random mixes); no frame added, the gate unchanged at all five baselines. Co-Authored-By: Claude Opus 5.5 --- nameparser/_pipeline/_vocab.py | 6 +++--- tests/v2/pipeline/test_vocab.py | 2 ++ 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index b5e6ef5e..70bf9520 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -740,8 +740,6 @@ def written_as_a_name(text: str) -> bool: ('ß' -> 'SS', 'ĸ' with none): 'WEIß' and 'KAĸ' are written in capitals, 'Weiß' is not. One call and no generator, so a word costs one frame.""" - if text == text.lower(): - return False word = unicodedata.normalize("NFC", text) i = len(word) while i: @@ -752,7 +750,9 @@ def written_as_a_name(text: str) -> bool: upper = ch.upper() if ch.islower() and (upper == ch or len(upper) != 1): continue - return ch.islower() + # the capital is asked here, of the word that has a letter to + # read; one with none falls through below ('李', '2', 'ß') + return ch.islower() and text != text.lower() return False diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index 3442447f..5d814deb 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -907,6 +907,8 @@ def test_is_single_letter_numeral() -> None: ("MacDONALD", False), ("FitzGERALD", False), ("DeVITO", False), ("LaFLEUR", False), ("St-PIERRE", False), ("SMITH-McDONALD", False), ("WEIß", False), ("O'NEIL", False), + # no letter carrying case evidence at all + ("李", False), ("2", False), ("ß", False), ("", False), ]) def test_written_as_a_name(text: str, expected: bool) -> None: # #564 (Derek): the name's case contrast is a word holding a From 71fea582cd6d7bcef1df372ec50e19d86c0768fb Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Fri, 2 Oct 2026 01:24:06 -0700 Subject: [PATCH 13/13] docs(S2): #564 -- the recipe states its extraction regex Review of d7f4dac8/2479eabb: the recipe's 4286 came from a regex ("([^"\n]{2,60})") the prose did not state; reading its words literally gave 4273 (tokenize) or 4284 (an escape-aware regex), same conclusions. The prose now names the regex and records the exact-literal count beside it. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index c441204b..9f4d4d8d 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -687,7 +687,7 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. - 2026-10-01 (Derek), #564 — THE ALL-CAPS HALF READS THE COMMA POSITION BY DEFAULT, AND THE SWITCH HAS THREE SETTINGS. Supersedes the default of the 2026-09-14 entry above (its reasoning stands for the positions it was argued over). That entry turned the whole caps half off because French and Korean records write the SURNAME in capitals; but the convention writes them at the end of a name (`Jean DUPONT`) or before a comma (`DUPONT, Jean`), never after a comma behind a full name, so the reason for the off default never reached the comma position and that position was switched off with it. The corpus held three names of exactly that shape — `Ahmad Jayadi, CHA`, `John Smith, RAI`, `The Rt Hon Kenneth Clarke QC MP, HMG` — all credentials, all read as the given name at 2.3.0. - DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and whose last LETTER is lowercase (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list, and `McDonald` and `Džokić` pass. "Last letter" is the last letter that carries CASE, not the last character: the word is composed (NFC) first; every non-letter and every caseless letter is passed over (`Jones'`, `Smith2`, `Smith)`, `Jonesʼ` with U+02BC, the transliteration letters in `Wafāʾ`, a trailing `李`); and so is a lowercase letter whose capital is not one character (`ß` → `SS`, `ĸ` with none, the `fi` ligature), which is no evidence of case — `WEIß` is written in capitals, `Weiß` is not. Two cuts got this wrong: the first tested the last character, and its review found every name typed with decomposed accents and ending in an accented letter losing the contrast (`José André, XYZ`, `Lê Thị Hà, XYZ`, `René Noé, XYZ`, `Chloé Zoé, MBA XYZ` read given where their composed spellings read suffix), with `Jones'`, `Smith2` and `Smith)`; the second skipped non-letters but stopped at a caseless letter, so `Asmāʾ Wafāʾ, XYZ` and `JOHN Jonesʼ, XYZ` kept given 'XYZ' (its review). MEASURED 2026-10-02 against `a16927c6`, the capital-then-lowercase commit: the seven capitalized-surname prefixes (`FitzGERALD`, `DeVITO`, `LaFLEUR`, `VanDYKE`, `DuBOIS`, `St-PIERRE`, `SMITH-McDONALD`), `MÜLLER WEIß` and `KUJAĸ de PETERSEN` keep their given names, `Džokić Ljubić` gains the contrast, `JOHN O'NEILL's` and `JEAN-pierre DUPONT` gain one (accepted below), and nothing else moves; the decomposed-accent names, `Asmāʾ Wafāʾ` and `JOHN Jonesʼ` read as `a16927c6` reads them. OFF on this tree equals `unlisted_caps_suffixes=False` on `7394078a` (origin/master at the time) on all 4286 inputs. RECIPE, the comparator being those two commits (`git archive` each): the inputs are every name in `tools/differential/corpus*.jsonl`, every double-quoted string literal in `tests/v2/cases.py` holding a comma (fragments of case notes included, harmlessly), and each of these prefixes joined by ", " to each of the parts `XYZ`, `ANDREW`, `ANDREW PhD`, `PhD XYZ`, `MJ`, `XYZ Jr.` — `John Smith`, `DiCaprio LaBeouf`, `IJzerman IJsselmeer`, `al-Rashid al-Hassan`, `d'Estaing d'Orléans`, `McDonald MacLeod`, `Džokić Ljubić`, `Asmāʾ Wafāʾ`, `John Smith-Jones`, `john smith`, `García y López`, `LLOYD WEBBER`, `GISCARD d'ESTAING`, `HAFEZ al-ASSAD`, `LLOYD McDONALD`, `LLOYD FitzGERALD`, `PAOLO DeVITO`, `JEAN LaFLEUR`, `DICK VanDYKE`, `PAUL DuBOIS`, `JEAN St-PIERRE`, `LLOYD SMITH-McDONALD`, `MÜLLER WEIß`, `KUJAĸ de PETERSEN`, `LLOYD ap RHYS`, `de GAULLE C`, `Mr LLOYD WEBBER`, `LLOYD WEBBER née Smith`, `JOHN O'NEILL's`, `JOHN SMITHs`, `JEAN-pierre DUPONT`, `JOHN Jones'`, `JOHN Jonesʼ`, `JOHN Smith)`, `JOHN Smith2`, and the NFD forms of `José André`, `Lê Thị Hà`, `René Noé`, `Chloé Zoé`. Parse each under OFF, AFTER_COMMA and EVERYWHERE and compare `as_dict()`, the ambiguity kinds and `capitalized(force=True)`; the population is a fact of the tree it runs on (4286 here), since the corpora and the case table grow. ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so do two writings in an otherwise all-caps record that end in a lowercase letter: a glued possessive or plural (`JOHN O'NEILL's, XYZ`, `JOHN SMITHs, XYZ` read suffix `XYZ`) and a hyphenated part written in mixed case or lowercase (`JOHN SMITH-Jones, XYZ`, `JEAN-pierre DUPONT, XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. + DECIDED, with Derek's calls marked. (1) `Policy.unlisted_caps_suffixes` becomes a `CapsSuffixes` StrEnum, OFF / AFTER_COMMA / EVERYWHERE, default AFTER_COMMA (Derek: an enum over two booleans or a mixed `bool | "comma"`, so every state is named and no meaningless combination exists). The field was new in 2.4 and unreleased, so no deprecation is owed; a bool raises a TypeError naming both replacements. The third value exists because redefining "off" would have left no setting that keeps a given name written in capitals (Derek's question, `García Márquez, GABRIEL`). (2) AFTER_COMMA reads only the part right after the first comma, behind two or more name words — the comma's own structure decision, which segment makes from the text before classify runs. The given part's last word after a family comma (`Doe, John XYZ`) and a maiden clause's last word are trailing slots and stay behind EVERYWHERE, as the comma-less end of a name does. (3) A lone two-letter word is never admitted at the comma, in any setting (Derek): it is the undotted form of #563's paired initials, and two words before the comma may be one surname, so `García Márquez, MJ` keeps given 'MJ'. Inside a longer part a two-letter word is read exactly as #563 reads the dotted pair — measured 2026-10-01 on the shipped tree by swapping `MJ` ↔ `M.J.` in runs of one to three words from {MJ, PhD, XYZ, CPA, Jr., MA, LEED, III} after seven name prefixes whose own words carry the contrast (`García Márquez`, `Smith Jones`, `Mr García Márquez`, `de García Márquez`, `García Márquez née Smith`, `John st Smith`, `García y López`): the comma's structure decision differed on 0 of 1295. After five all-caps prefixes it differed on 325 of 925, by design, `MJ` being no class member without the contrast while the dotted shape is case-free (`GARCÍA MÁRQUEZ, CPA MJ` keeps given 'CPA' where `CPA M.J.` makes the run). Recompute: run segment and read `ParseState.structure` for each pair. The equivalence is the comma's decision only: at the given part's last word after a family comma a dotted `M.J.` is a credential by shape (S3) where `MJ` is one only under EVERYWHERE, so `García Márquez, MA M.J.` reads suffix 'M.J.' and `MA MJ` middle 'MJ' — so `García Márquez, MJ PhD` keeps given 'MJ' as `De La Cruz, M.J. PhD` does, and C1 says so by pointing at #563's own sentences rather than restating them. Two restatements were wrong: the first draft's "three letters or more, or a run holding such a word" (both reviews), and its replacement, "with only a credential behind it ... while a credential in front makes the run", which #563's vocabulary rules contradict (`John Smith, MA MJ` reads given 'MA'; `García Márquez, MJ XYZ` makes the run). (3b) An unlisted all-caps word is a by-shape member of C1's run (#544), beside listed credentials, not only in a run of all-caps words (Derek, after review): before, `John Smith, XYZ` read suffix while `John Smith, PhD XYZ` read given 'PhD' and `John Smith, XYZ Jr.` given 'XYZ' — more evidence for a credential producing a name reading, already true under the old `True` and made a default by this change. The contrast has to be the NAME's, and it is carried only by one of the name's own words before the comma (`own_words`, so no maiden clause or delimited content) holding a capital and whose last LETTER is lowercase (`_vocab.written_as_a_name`), with no period and not claimed by a wordlist as a title, particle, connective, credential, generation, maiden marker or honorific (`_vocab.claimed_as_non_name`) (Derek). Every earlier draft asked for "a lowercase letter somewhere, less some exclusions", and each leaked a class the next review found: a mixed-case credential (`LLOYD WEBBER, ANDREW PhD`), a maiden clause or title (`LLOYD WEBBER née Smith, ANDREW PhD`, `Mr LLOYD WEBBER, ANDREW`), a generation, connective or shape title (`LLOYD WEBBER Jr., ANDREW`, `GARCÍA y LÓPEZ, ANDREW`, `Insp. LLOYD WEBBER, ANDREW`), and glued particles and unlisted lowercase words (`GISCARD d'ESTAING, VALÉRY`, `HAFEZ al-ASSAD, BASHAR`, `LLOYD ap RHYS, DAFYDD`) — no closed list separates an all-caps record from a mixed-case name. A Title-case draft (`str.istitle()`) closed those but read too narrowly, missing interior capitals and elisions (`DiCaprio`, `IJzerman`, `al-Rashid`, `d'Estaing`, `McDonald` — every one a mixed-case name that kept a given 'XYZ') and taking a lone capital as Title case (`de GAULLE C, CHARLES` lost its given name). Derek weighed "contains a capital" (an all-caps record with any lowercase word then always carries the contrast) and "not all capitals but contains a capital" (which takes `d'ESTAING` and `al-ASSAD`, the lowercase standing before the capitals) and first chose a capital directly followed by a lowercase letter, with a leading `Mc`/`Mac` skipped. Its review found the skip was one member of a class: any capitalized prefix glued to a surname written in capitals supplies such a pair (`LLOYD FitzGERALD, RONALD`, `PAOLO DeVITO, MARCO`, `JEAN LaFLEUR, PIERRE`, `DICK VanDYKE, JOHN`, `PAUL DuBOIS, JEAN`, `JEAN St-PIERRE, MARC`, and `LLOYD SMITH-McDONALD, RONALD`, the skip applying only at a word's start), as does `ß`, which has no capital (`MÜLLER WEIß, HANS`) — each lost its given name — while titlecase digraphs failed it (`Džokić Ljubić, XYZ`) and it cost a frame per letter. The shipped test, Derek's: a capital in the word and its last letter lowercase. A surname written in capitals ends in a capital whatever is glued in front of it, so every one of those records keeps its given name with no prefix list, and `McDonald` and `Džokić` pass. "Last letter" is the last letter that carries CASE, not the last character: the word is composed (NFC) first; every non-letter and every caseless letter is passed over (`Jones'`, `Smith2`, `Smith)`, `Jonesʼ` with U+02BC, the transliteration letters in `Wafāʾ`, a trailing `李`); and so is a lowercase letter whose capital is not one character (`ß` → `SS`, `ĸ` with none, the `fi` ligature), which is no evidence of case — `WEIß` is written in capitals, `Weiß` is not. Two cuts got this wrong: the first tested the last character, and its review found every name typed with decomposed accents and ending in an accented letter losing the contrast (`José André, XYZ`, `Lê Thị Hà, XYZ`, `René Noé, XYZ`, `Chloé Zoé, MBA XYZ` read given where their composed spellings read suffix), with `Jones'`, `Smith2` and `Smith)`; the second skipped non-letters but stopped at a caseless letter, so `Asmāʾ Wafāʾ, XYZ` and `JOHN Jonesʼ, XYZ` kept given 'XYZ' (its review). MEASURED 2026-10-02 against `a16927c6`, the capital-then-lowercase commit: the seven capitalized-surname prefixes (`FitzGERALD`, `DeVITO`, `LaFLEUR`, `VanDYKE`, `DuBOIS`, `St-PIERRE`, `SMITH-McDONALD`), `MÜLLER WEIß` and `KUJAĸ de PETERSEN` keep their given names, `Džokić Ljubić` gains the contrast, `JOHN O'NEILL's` and `JEAN-pierre DUPONT` gain one (accepted below), and nothing else moves; the decomposed-accent names, `Asmāʾ Wafāʾ` and `JOHN Jonesʼ` read as `a16927c6` reads them. OFF on this tree equals `unlisted_caps_suffixes=False` on `7394078a` (origin/master at the time) on all 4286 inputs. RECIPE, the comparator being those two commits (`git archive` each): the inputs are every name in `tools/differential/corpus*.jsonl`, every match of the regex `"([^"\n]{2,60})"` in `tests/v2/cases.py` holding a comma (a text match, not a literal parser, so fragments of case notes are included, harmlessly; extracting the literals exactly with `tokenize` gave 4273 inputs and the same conclusions), and each of these prefixes joined by ", " to each of the parts `XYZ`, `ANDREW`, `ANDREW PhD`, `PhD XYZ`, `MJ`, `XYZ Jr.` — `John Smith`, `DiCaprio LaBeouf`, `IJzerman IJsselmeer`, `al-Rashid al-Hassan`, `d'Estaing d'Orléans`, `McDonald MacLeod`, `Džokić Ljubić`, `Asmāʾ Wafāʾ`, `John Smith-Jones`, `john smith`, `García y López`, `LLOYD WEBBER`, `GISCARD d'ESTAING`, `HAFEZ al-ASSAD`, `LLOYD McDONALD`, `LLOYD FitzGERALD`, `PAOLO DeVITO`, `JEAN LaFLEUR`, `DICK VanDYKE`, `PAUL DuBOIS`, `JEAN St-PIERRE`, `LLOYD SMITH-McDONALD`, `MÜLLER WEIß`, `KUJAĸ de PETERSEN`, `LLOYD ap RHYS`, `de GAULLE C`, `Mr LLOYD WEBBER`, `LLOYD WEBBER née Smith`, `JOHN O'NEILL's`, `JOHN SMITHs`, `JEAN-pierre DUPONT`, `JOHN Jones'`, `JOHN Jonesʼ`, `JOHN Smith)`, `JOHN Smith2`, and the NFD forms of `José André`, `Lê Thị Hà`, `René Noé`, `Chloé Zoé`. Parse each under OFF, AFTER_COMMA and EVERYWHERE and compare `as_dict()`, the ambiguity kinds and `capitalized(force=True)`; the population is a fact of the tree it runs on (4286 here), since the corpora and the case table grow. ACCEPTED with it (Derek): a name written wholly in lowercase reads the listing form, `john smith, XYZ` keeping given 'XYZ', an unlisted title written in mixed case (`Doña GARCÍA LÓPEZ, MARÍA`) still carries the contrast, and so do two writings in an otherwise all-caps record that end in a lowercase letter: a glued possessive or plural (`JOHN O'NEILL's, XYZ`, `JOHN SMITHs, XYZ` read suffix `XYZ`) and a hyphenated part written in mixed case or lowercase (`JOHN SMITH-Jones, XYZ`, `JEAN-pierre DUPONT, XYZ`). ACCEPTED with (4): in a mixed-case name a capitalized given name followed by a generation goes the same way as one alone, `García Márquez, JUAN Jr.` reading suffix 'JUAN Jr.', reported; and behind a particle-led surname of two name words the flip leaves no given name at all, `De La Cruz García, MARÍA` reading family 'De La Cruz García', suffix 'MARÍA', because P1 reads a never-given particle's part as all surname. (4) ACCEPTED (Derek): a given name written in capitals behind a two-word surname without particles reads as the credential, `García Márquez, GABRIEL` giving given 'García', family 'Márquez', suffix 'GABRIEL', reported; capitalizing the given and not the family is not a convention, and OFF is the escape hatch. #575 was the prerequisite: before it, the count read `Van Buren, MARTIN` and `De La Cruz, MARIA` as two or three name words and this default would have taken their given names too. WHAT MOVES, measured 2026-10-01 with the gate at all five baselines. The three corpus names above, plus the C1 examples `John Smith, XYZ` and `John Smith, LEED AP` — the latter closing a `deviates: #291` marker rules.md had carried under a closed issue — classified by `fix(#564)` in every ledger; the radar-unclassified count is what it was before the change at every baseline. `John Smith, RAI` and `Ahmad Jayadi, CHA` read suffix again by their capitals, as they did by vocabulary before #342 removed both words: parity at 1.4.0, only the comma's report at 2.0 through 2.2, so the #342 rule's `fields` lose `given` (the OVER-DECLARED check) and the watched shape for `John Smith, RAI` is re-recorded at those four baselines. `Smith, XYZ` keeps given 'XYZ' and, at the default, reports nothing; EVERYWHERE still reports the declined fork there, as it did. CASE REPAIR FOLLOWS THE READING. The comma decision is segment's, made from the text, while case repair keeps a word in capitals only where classify wrote the shape tag (rules.md#R4); the first draft tagged only under EVERYWHERE, so `parse("John Smith, XYZ").capitalized(force=True)` rendered 'John Smith Xyz' at the default and 'John Smith XYZ' under EVERYWHERE (found by the docs review, axis 5). Classify now tags a caps-shaped word in the part a suffix comma opened under AFTER_COMMA too, so both settings render that part alike, and rules.md#R4 names the caps shape beside the dotted one; `test_render`'s forced-repair test fails with the tagging removed (an R4 example could not witness it, R5 leaving a mixed-case name unrepaired unless forced). A caps word in a third or later comma part is still tagged only under EVERYWHERE, so `John Smith, MD, XYZ` renders 'Xyz' at the default, as master does. COST, `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `Smith, JOHN` 183 → 183, `Smith, XYZ` 182 → 182, `John Smith, PhD` 217 → 217, `John Smith, MA` 252 → 252, `John Smith, Ph. D.` 254 → 254, `John Smith, CPA` 217 → 217, `John Smith, MD PhD` 260 → 260, `John Smith, MBA CPA` 261 → 261, `Dr. Juan de la Vega III` 366 → 366, `John Smith, XYZ` 251 → 266, and an all-caps record with a capitalized given name and a credential, `LLOYD WEBBER, ANDREW PhD`, 338 → 364, a constant +26 that does not grow with the name's length (`'LLOYD'*64 + ' WEBBER, ANDREW PhD'` 653 → 679; the second draft's per-character contrast test had made that +350). Past the all-caps check the contrast is linear in the words before the comma, about four frames a word (the generator, the case test's call, and the own-words walk's fold and marker test; `'GAULLE ' + 'ap '*k + 'GAULLE, CHARLES'` is +7, +26, +38, +54 over master at k = 0, 1, 4, 8), and constant in a word's letters (`'de ' + 'GAULLE'*k + ', CHARLES'` 253 → 275, 343 → 365 and 631 → 653 at k = 1, 16, 64), where the capital-then-lowercase draft grew a frame per letter; all paid only once a caps word is in hand. The comma test is on by default, so three C-level checks go before any call: two or more words before the comma (the first draft lacked it, and `Smith, JOHN`, a common record format, paid +25 for a flip it can never make — the code review), the first word in capitals, and the lone-two-letter length. Every caller of the caps predicate asks first, in C, what it would decline anyway — alphabetic capitals, not a listed suffix word — so a listed credential never pays for the call (the review of the first fix found `John Smith, CPA` +4 and `John Smith, Ph. D.` +5 before these). The 2026-09-14 entry's recompute recipe predates the enum: its `on` is `Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE)` now, and its `off=Parser()` is `Policy(unlisted_caps_suffixes=CapsSuffixes.OFF)`, `Parser()` being AFTER_COMMA.