Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -502,7 +502,15 @@ listed below.
ending a maiden marker's clause, also since 2.4:
``"Jane Doe nee Smith X.Y.Z."`` gives maiden ``Smith`` with
suffix ``X.Y.Z.``, where ``False`` keeps maiden
``Smith X.Y.Z.``.
``Smith X.Y.Z.``. Two single letters right after a comma are
the exception: they are how a person's initials are written,
and the words before the comma may be one surname of two
words, so ``"García Márquez, G.J."`` keeps given ``G.J.``
(reported) unless an unambiguous post-nominal in front of them
that is not also a title, or another unlisted dotted word beside
them, says otherwise (``"John Smith, PhD X.Y."`` gives suffix
``PhD X.Y.``, while ``"García Márquez, Ms G.J."`` gives title
``Ms``, given ``G.J.``).
Case is irrelevant — the periods are the signal.
Whole-token vocabulary still wins (``M.A.``, ``Ph.D.``), and a
single trailing period is not this shape
Expand Down
5 changes: 5 additions & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -681,6 +681,11 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f
ADDENDUM 2026-09-28 (Derek). (e) A PARTICLE ANCHORS NOTHING, a fifth boundary beside the four WHAT MAY ANCHOR lists: a word of both the particle and the unambiguous suffix vocabulary (`vd` and `mc` in the shipped lexicon; `do` sits in the ambiguous half, so it is a member and never an anchor) heads the family name behind it (P2), and anchoring there split the family around a suffix — `Smith vd Ma, John` read family 'Smith Ma', suffix 'vd'; `Jan vd Ma` suffix 'vd Ma' and no family; `Smith Mc Ma, John` and `D. Mc Ba Ed, Smith` the same way. Each reads as e10e83b4 reads it again (family 'Smith vd Ma', 'vd Ma', 'Smith Mc Ma', 'D. Mc Ba Ed'). The exclusion is of the word in front only: a particle MEMBER behind a credential is still spoken for (`doe, jane v phd do` keeps suffix 'v phd do'), and a C1 run holding the particle is C1's reading rather than the anchor's (`Jan Berg, vd Ma` reads suffix 'vd Ma', as `Jan Berg, vd` reads suffix 'vd'). Measured over the grid of MEASURED above with `vd` and `Mc` added to its words (127,578 texts, 382,734 parses), against the tree before this addendum: 9,426 parses (3,142 texts) move, every one holding `vd` or `Mc` directly in front of a member; 8,940 return to e10e83b4's reading exactly, and the other 486 all hold `M.Eng.`, whose own change is (3) — 406 of them take e10e83b4's fields and differ only in the pick `M.Eng.` no longer reports, and 80 take the fields e10e83b4 gives the same text with `PhD` in `M.Eng.`'s place (`Jane Doe nee Smith M.Eng. vd Ma` reads middle 'Doe M.Eng.', family 'vd Ma', as `Jane Doe nee Smith PhD vd Ma` reads middle 'Doe PhD'). The differential corpora move 0 of 1441 names, and the grid of MEASURED without the two words moves 0 parses, so every figure there stands. THE PASS IS ASKED ONLY WHERE A SUFFIX PIECE IS IN REACH: `_pieces.anchor_in_reach` walks back through the lone members in front of a declined member on tags alone and answers False wherever the pass must — the first piece past the members is no suffix piece, or there is none — and it is asked only before the pass exists, one lookup answering after that, so a run of members stays linear. An ordinary name ending in a declined member stops paying for the pass (frames, py3.11, before → after, e10e83b4 in brackets): `John Smith Ma` 250 → 246 (244), `Smith, John Ma` 285 → 279 (275), `Doe, Jane MA do` 374 → 369 (363), `Doe, John Q. Ma` 328 → 320 (316), `Jan vd Ma` 308 → 243 (238); a name the pass does read pays the test on top, `John Smith PhD MEng` 326 → 329 and `Doe, Jane PhD MEng` 359 → 362. On the no-comma peel the reach starts past the walk's leading piece, which never anchors — so a by-shape member, which reaches the pass only at the peel's second position (a lean of None with a word to spare is consumed first), finds nothing in reach, and the by-shape guard that stood there was deleted as unreachable. `segment_suffix_reading` calls `credential_anchors` (`first_kept=False`, a comma part's first piece being a run member like any other) in place of the inline copy NO MECHANISMS describes, and the keep-in-step note goes with it. The pass is built only where the reach test finds a suffix piece in front of a declined member, which no member opening the part has, and a part whose name word stands ahead of any member returns before one is asked: a plain comma name pays nothing (`Smith, John` 182, `Doe, John MA` 283, `Smith, J. Q.` 246, `Smith, Ed` and `Smith, Ma` 213, `Smith, Ed John` 270, all as before the call replaced the copy; built on every declined member instead, the last three paid 215, 215 and 273), a title in front pays the reach test alone (`Smith, Dr. Ma` 253 → 254), and a part the company decides pays three frames over the copy (`Smith, PhD MEng` 273 → 276, `Smith, PhD Ma` 271 → 274; `Smith, PhD Ma John` 330 → 335, its pass built for 'Ma' before 'John' ends the part). THE TWO-INPUT CHECK's second control reads 42, not 30, with no change to the property: the walk's leading piece is now held out of the pass twice on the no-comma peel, by `credential_anchors` and by the reach test, either alone suffices, and with both removed the by-shape heads the deleted guard held back (`PhD X.Y.Z.`) fail beside the listed ones; the first control is unmoved at 48. And the 2026-09-28 C2 bullet's "a TAIL segment, which assign reads wholly as suffixes" holds outside a maiden clause standing in it — `Jane Doe, PhD, Jr nee van Ma` keeps maiden 'van Ma' — which rules.md#C2's statement now says. The reach test's one-way exactness is held by tests/v2/pipeline/test_pieces.py's `test_anchor_in_reach_never_hides_an_anchor`, over the case table's texts and runs of one to three words behind two heads; its recorded negative control, the reach test reading a `vocab:suffix` token as no suffix piece, fails at 878 member positions.
- 2026-09-27 (#544) — THE 2026-09-15 ACCEPTED ITEM (ii) IS REVERSED, and the bullet stands as it landed. `abdul Smith Jr Ma` reads given 'abdul', family 'Smith', suffix 'Jr Ma' — 2.3.0's reading — because the unambiguous 'Jr' in front of the Title-case 'Ma' anchors it (the entry above): the peel takes both, and P5's reserve, now seeing the family the join would take, declines the join. Item (i) stands. The case row is now `a_credential_in_front_anchors_a_declined_pick`, and rules.md#S2's Accepted block names the shapes that still keep the company out of reach.
- 2026-09-27 (Derek), #544 — THE #531 PAIRING GAINS A SECOND EXCEPTION, and CAPITALS DECIDE FOR `do` above stands as it landed. That bullet let only a positive credential lean override P6's attachment at the given slot; an unambiguous credential IN FRONT of the member now overrides it as well, the degree being a second and stronger signal: `doe, jane v phd do` reads suffix 'v phd do' and reports `suffix-or-name` where it read family 'do doe' and reported P6's fork. The one-case record the pairing protects has nothing in front of its particle, so `NASCIMENTO, EDSON ARANTES DO` still reads family 'DO NASCIMENTO' and reports `particle-or-given`.
- 2026-09-30 (Derek), #563 — PAIRED INITIALS ARE THE EXCEPTION TO C1'S NAME-WORD COUNT, AND A FLIP ON THE DOTTED SHAPE ALONE IS SILENT. The 2026-09-14 (Derek) entry above, #516's DOTTED HALF IS A SWITCH, read `John Smith, A.B.` as suffix 'A.B.' on the count of two name words before the comma. The count cannot tell a given name and a family name from ONE surname written in two words, and the dotted shape's own most common member is a person's initials, so the unreleased tree read `García Márquez, G.J.` as given 'García', family 'Márquez', suffix 'G.J.', `De La Cruz, M.J.` with no given name at all, and `van der Berg, A.J.` as given 'van' — where 1.4.0 through 2.3.0 read each as initials. A particle check would have caught only the last two; `García Márquez` and `Lloyd Webber` look exactly like `John Smith`. What the parser CAN see is the token, so the line is drawn there: two single letters joined by a period, with or without one after the second (`_vocab.is_paired_initials`; `García Márquez, G.J` reads as `G.J.` does), read as the given name after the comma and report the fork, and only an UNAMBIGUOUS suffix word IN FRONT of them that is not also title vocabulary, or a second word the class admits by dotted shape in the same part, makes them the credential run (`John Smith, PhD X.Y.`, `John Smith, X.Y. P.Q.`). Three letters or more (`X.Y.Z.`) and any longer chunk (`B.Tech.`) read by the count as before. Three choices made with it (Derek, 2026-09-30): (i) FRONT ONLY — a word behind the pair does not speak, so `De La Cruz, M.J. PhD` keeps given 'M.J.' and `John Smith, A.B. PhD` reads given 'A.B.', the direction rules.md#S2's company clause already takes; a LISTED dotted word behind does not speak either (`John Smith, A.B. Ph.D.` → given 'A.B.'), the second-word exception being for words the class admits by shape; (ii) SILENT FLIP — a flip resting only on words the class admits by dotted shape reports nothing, `B.Tech.` included, because the only such word a reader takes for a name is a pair of initials, and the comma flips a pair only where a word beside it has already said otherwise; a LISTED member in the same part still reports the flip (`John Smith, PhD MEng`, `John Smith, X.Y.Z. MA`); (iii) THE COMMA SLOT ONLY — the trailing slot (`John Smith R.T.`), the given part's trailing slot after a family comma (`Doe, John R.T.`) and the maiden clause keep #516's reading and its report; #563 leaves the no-comma question open. A generational word in front speaks like any unambiguous suffix word (`García Márquez, Jr. G.J.` → suffix 'Jr. G.J.'), as S2's company clause has it — unless it is also a title, so `García Márquez, Sr G.J.` reads title 'Sr', given 'G.J.'. ACCEPTED: `García Márquez, G.J.R.` reads suffix 'G.J.R.' — initials are conventionally written apart, as separate words this rule never reads (`García Márquez, G. J. R.` gives given 'G.', middle 'J. R.'), and three letters run together are far likelier a credential. REVERSED BY THIS ENTRY rather than edited, for the dotted half at the first slot after a comma only: the #516 SWITCH entry's `John Smith, A.B.` → suffix and its "reports either way"; the 2026-09-14 NO NEW AmbiguityKind entry's "Every decision at these slots emits" SUFFIX_OR_NAME; and rules.md#C1's "A decision either way at this comma is reported". ALSO FIXED: the trailing slot's report described a by-shape pick in the listed member's words ("written without periods is both a post-nominal"), false of `X.Y.Z.` on both counts; it now says the word is shaped like a post-nominal but listed in no vocabulary.
REVIEW ROUND, same day. The docs review found the first version let a word of both the title and the suffix vocabulary speak for the pair, since C1 counts such a word opening the part as a suffix word: `García Márquez, Ms G.J.` read given 'García', family 'Márquez', suffix 'Ms G.J.' with NO report — #563's own defect, made silent — where 2.3.0 read title 'Ms', given 'G.J.'. The twelve title/suffix duals (`ms`, `md`, `sr`, `lt`, `sa`, `ra`, `vc`, `cpl`, `cpo`, `cpt`, `csm`, `sgm`) all reached it. Fixed by letting only a suffix word that is not title vocabulary speak: in front of a pair, a dual is the title of the given part the pair opens, as S2 already reads a dual in that part's title run, and `Ms G.J.`, `MD G.J.`, `Lt G.J.` now read title plus given as every release did. The title lookup runs only once a pair is met. The code review's frame count found the single-token path asking `ambiguous_class_member` a second time, +2 frames on every reporting comma name (`John Smith, MA`, 252 → 254, profiler call events); for a word past the candidate test LISTED is exactly "no period" (`ambiguous_class_member` declines any period), which is now asked inline. Measured after both fixes against master, same counter: `Smith, John` 183 and `John Smith, MA` 252 on both trees, `John Smith, X.Y.Z.` 295 → 286 (the report it no longer builds), and `García Márquez, G.J.` 338 → 350, the family-comma reading and its report costing more than the flip did.
SECOND REVIEW ROUND, same day, on the first round's fix commit. (a) TWO PAIRS (Derek): `De La Cruz, M.J. K.L.` read family 'De La Cruz', suffix 'M.J. K.L.' — no given name — in silence, the silent-flip rationale ("only paired initials are taken for a name") failing on exactly the case where each word speaking for a pair is itself a pair. The reading stands, two dotted groups not being how anyone writes a person's initials, but that flip now REPORTS: when the only words speaking for a pair are other pairs, every shape word in the part being one. `John Smith, X.Y. P.Q.` gains the report with it; `John Smith, X.Y.Z. G.J.` and `John Smith, PhD G.J. K.L.`, where something else speaks, stay silent. The same rule reports `García Márquez, Ms G.J. K.L.`, whose honorific the run takes, where the first round's fix took it in silence. (b) A CLASS MEMBER IN FRONT SPEAKS FOR NOTHING (Derek): `García Márquez, Ed G.J.`, `Ma G.J.` and `MA G.J.` flipped on the member in front, where S2's company lets only an unambiguous credential speak for a member. Now the comma keeps the family, given 'Ed' — and 'G.J.' then ends the given part, the slot (iii) left alone, so it reads suffix 'G.J.' as `Doe, John R.T.` does, both forks reported; 1.4.0 and 2.3.0 gave middle 'G.J.'. Resolving that slot is the no-comma question #563 leaves open. (c) WORDING, no behavior: C1's sentence counting a dual opening a part as a suffix word now defers to the pair's titles; the pair's report reaches it only where it opens the part, `Smith, Ms G.J.` having always been silent (the comma report's reach is the first post-comma piece, 2026-09-18); and the silence is stated as "no listed word takes part", which covers `García Márquez, PhD G.J.`, where "dotted-shape words alone" did not.
SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal.
MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole.

### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343)

Expand Down
Loading
Loading