Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -880,6 +880,15 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py):
- 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had brought back a split of `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd'. That split is not only an intermediate state of #552: 2.2.0 and 2.3.0 shipped it (2.0.0 and 2.1.0 read family 'Smith vd Ma'), #530 removed it in this cycle, and #552 kept it removed. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them.
The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. It is also 1.4.0's reading, verified against the released wheel the same day: 1.4.0 read suffix 'vd Ma' and 'VD MA', and 2.0.0 through 2.3.0 read the whole name as the family. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks.
#563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` before a comma and after a family comma, and the silent mixed-case `Doe, Jane PhD vd Ma`. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`.
- 2026-10-01 (Derek), #575 — THE COUNT BEFORE THE COMMA COUNTS A PARTICLE SURNAME ONCE. Every count rules.md#C1 makes of the part before the comma — v1's "more than one word" for an unambiguous credential, the ambiguous class's name-word count (#289, #544), and assign's two-name-word test for the positional read after a comma followed by no name word (#296/#325) — counted tokens or pieces, so `De La Cruz` was three words and `van der Berg` two (the ambiguous `van` stands as its own piece, P1's fork). This cycle's count had therefore read `De La Cruz, Ed` as a credential comma with no given name and split `van der Berg, MA` into given 'van', family 'der Berg' — an unreleased regression against 2.3.0, which read both as the listing form — and 1.4.0 through 2.3.0 all split `van der Berg, PhD`. DECIDED: a particle run and the one name word it attaches to are one word in all three counts, so a particle surname reads as a one-word surname does.
THE COMMA SETTLES P1's FORK, and that is the decision rather than a side effect. Standing alone, `Van Buren` is given 'Van', family 'Buren' with a `particle-or-given` report, so counting it as one surname makes the comma's reading disagree with the standalone one. Derek's call, after weighing the narrow alternative (count only a part P1 already reads as all surname, i.e. one led by a never-given particle like `de`): the listing form puts a surname before the comma, so the comma is evidence the standalone parse does not have, and the readings it gives are the ones a person takes — `Van Buren, Ed` given 'Ed', family 'Van Buren'; `van der Berg, MA` family 'van der Berg', suffix 'MA' (one name word, so C1 reads the case, and capitals in a mixed-case name make the credential). `Van Johnson, Dr.` moves with it, family 'Van Johnson', title 'Dr.', where 2.2.0 and 2.3.0 read given 'Van' and `tests/v2/pipeline/test_assign.py` had pinned the positional fork; that test now uses `Van Johnson Smith, Dr.`, which still has two name words.
THE PARTICLE REACHES ONE WORD, as P1's fold does, not to the end of the part as P2's chain does: `de Mesnil Jean, Dr.` under a family-first order is family 'de Mesnil', given 'Jean', two name words. The cost, and it is real for an ambiguous particle: a `van`-led part with a word after the surname counts two and keeps the positional read with P1's fork, so `van Buren John, Ed` reads given 'van', middle 'Buren', family 'John' and `van der Berg Smith, PhD` given 'van', family 'der Berg Smith', both unchanged from master. Under the default order a never-given particle loses nothing (P1's fold makes `de Mesnil Jean` all surname anyway). `_vocab.unit_ends` carries both reaches behind a `chain` flag, post_rules' fold reading the full chain as before.
ONLY PARTICLE CHAINS, not mechanisms.md#UNIT-PARTITION's full set — stated there as the Contract's one exception, and in P3's statement, whose "ONE name word wherever another rule counts them" now names C1's count as the exception. A bound given-name pair builds a GIVEN name, and P5 gives up a family word where the name has no other (`abdul Salam` alone is given 'abdul', family 'Salam'), so `abdul Salam, Ed` keeps the credential reading. Whether P3 joins a single-letter connective depends on the words of the whole name — `Carod i Rovira, Josep` joins as four words where `Ortega y Gasset` alone, three, does not — and the count is part of deciding what the whole name is, so segment could reach P3's answer only by copying P3; `Ortega y Gasset, Ed` keeps the credential reading. Approved as "particle chains and connective joins", narrowed to particles when P3's exception surfaced in implementation (the first draft joined `John e Smith, III` into one family name, caught by `tests/test_conjunctions.py`).
A WORD THAT IS ALSO A TITLE IS NO PARTICLE WHERE IT OPENS THE PART. `TITLES ∩ particles` is `{freiherr, st}` and `bound_given_names ∩ particles` is `{abu, أبو, ابو}` (recompute: `L = Parser().lexicon; L.particles & L.titles`, `L.particles & L.bound_given_names`). The first draft counted every particle, and the docs review found it reading `Freiherr von Berg, PhD` as family 'Freiherr von Berg' and `St John, PhD` as family 'St John', where every release reads the title. The first fix then excluded both overlaps in EVERY position, and the review of that fix found two errors in it. Position: inside a surname a title-particle chains as P1 and P2 chain it, and the blanket exclusion brought #575's own defect back for `de St Pierre, Ed`, `De St. Croix, Ed` and `de Abu Bakar, Ed` (family only, no given name, where 2.3.0 read given 'Ed'). And direction: 2.0 through 2.3 read `Abu Bakar, Ed` as given 'Ed', family 'Abu Bakar', so excluding the bound-given particle gave released behavior up rather than keeping it — and Abu Bakar is a common Malay surname. DECIDED: a title-particle is no particle only where it LEADS the part, and a bound given-name particle is a particle. So `Freiherr von Berg, PhD` and `St John, PhD` keep the title, `de St Pierre, Ed` reads given 'Ed', and `Abu Bakar, Ed` reads given 'Ed', family 'Abu Bakar' while `Abu Bakar, PhD` reads family 'Abu Bakar', suffix 'PhD' where 1.4.0 through 2.3.0 read given 'Abu'. The agreement test compares each token in both positions; it cannot see which position a caller passes, so case rows pin each count's flag: `de St Pierre, Ed` segment's, and `St John, Dr.` and `von St Johann, PhD` assign's, one per direction (the review of this round found assign's flag fixed to either constant passing every test, and both mutants now fail). `Freiherr von Berg, Ed` reads family 'Freiherr von Berg', given 'Ed' — 2.0.0's through 2.3.0's reading too, so no change against a release; only this cycle's count had read title 'Freiherr', family 'von Berg', suffix 'Ed'. The surname counts once and the title is no name word, so the count reads the listing form, which keeps a leading title in the family as it already did for `Prof. Cruz, Ed`. For `Prof.` that is a cost (rules.md#C1's Accepted line); for a German rank it is the right reading (Derek, 2026-10-01): since 1919 a former noble title is part of the legal surname, written between the given name and the particle (Karl-Theodor Freiherr von und zu Guttenberg), so `Freiherr von Berg` is the family name. That is also why `freiherr` is in PARTICLES at all — added with the German prefixes in 3e14ea20 (#18) without a stated reason, and the reason is this one. `Freiherr von Berg, PhD` keeps title 'Freiherr', the unambiguous credential's count taking the title as a word; both are readings of one name, and neither is changed here. A title in front is otherwise outside the fix: v1's count for an unambiguous credential counts the title as a word, so `Dr. van der Berg, PhD` reads as `Dr. van der Berg` does alone, given 'van'.
`De La Cruz, M.J. K.L.` reads given 'M.J.', suffix 'K.L.', reporting twice, as `Cruz, M.J. K.L.` does (Derek, 2026-10-01): with one name word before the comma the paired-initials count no longer reaches it, and C1's two-dotted-groups sentence now says "behind two or more name words". #563 had read it as a flipped credential run with no given name.
SEGMENT RUNS BEFORE CLASSIFY, so its count builds the two facts it reads (particle, suffix) from the vocabulary (`_vocab.surname_unit_tags`), while assign derives them from classify's tags (`_vocab.surname_unit_facts`) over every token of the part, so a suffix word stops a particle in both: the first draft's assign count dropped suffix pieces before walking and read `van Jr. Berg, Mr.` as family 'van Berg' (the code review). `test_classify.test_surname_unit_tags_agree_with_classify` sweeps every single-word vocabulary entry in three casings; its first run caught `JD.CPA` and `Msc.Ed.`, which classify tags as suffixes through the period-joined derivation, now mirrored, and that pair is the recorded negative control (`_SURNAME_UNIT_CONTROL`, asserted with the mirror patched out).
BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way.
COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it.

### T1 — separators, not joiners

Expand Down
2 changes: 1 addition & 1 deletion docs/design/mechanisms.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Problem shape. A rule needs to know how words were JOINED (chained titles, parti

## UNIT-PARTITION — count the units the joining rules built

Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all. How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_post_rules.py (`_unit_end`, `_units`); rules.md P1 is the counting rule, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember.
Problem shape. A rule counts "one name word", but the input holds words that another rule has already joined into one name — and the joining structure it would read has been merged away. Contract statement. Three rules build multi-word units: a particle chain (P2), a conjunction join (P3), and a bound given-name pair (P5). A rule that counts name words counts those units, and takes each whole or not at all — with one stated exception, rules.md#C1's count of the words before a comma, which counts particle surnames alone (decisions.md#C1, #575). How it works. group builds each join as a piece, but its own prefix chain then merges the joined piece into a longer one, so PIECES no longer carries the boundary — the units are rebuilt from the tags the vocabulary layer left (`particle`, `conjunction`, `vocab:bound-given`). The rebuild is RECURSIVE: what a conjunction or a bound word joins is the next UNIT, not the next word, and absorbing a single token instead strands a particle severed from the words it chains. Note the two joins arrive here for opposite reasons — the conjunction join was built and then swallowed, while the bound-given join was never built at all (P5 joins only at the first non-title piece), so restoring piece boundaries in group would fix the first and silently split the second. Lives in. nameparser/_pipeline/_vocab.py (`unit_ends`, the walk, shared since #575) and its two readers: _post_rules.py (`_units`, over classify's tags, for P1's fold) and rules.md#C1's counts before a comma (`_vocab.name_word_count`, `surname_unit_count`, and assign's positional-read test), which run partly before classify and so build their facts from the vocabulary (`surname_unit_tags`, held to classify's tags by an agreement test). C1 counts particle chains only, reaching one word as the fold does: a bound pair builds a given name, and whether P3 joins a connective depends on the whole name the count is helping to decide, so neither is one surname before a comma (`_vocab.SURNAME_UNIT_TAGS`, decisions.md#C1). rules.md P1 and C1 are the counting rules, P2/P3/P5 the joining ones. Reach for it when. A rule says "one name word" and the input can contain a join — enumerate the joining rules out of rules.md rather than the ones you remember.

## MARK-DONT-STRIP — record the decision, keep the fact

Expand Down
Loading
Loading