Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/design/decisions.md

Large diffs are not rendered by default.

9 changes: 7 additions & 2 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1702,7 +1702,9 @@ C1. Rationale: a credential run after the comma means the name is in
credential (S2), and the part reads as the credential run on that
evidence rather than on the count; a word of the class by shape
alone carries no such lean, so a part holding one is read by the
count. A decision either way at this comma
count, and so is a part holding two particles side by side,
which P2 joins into one particle run that S2 leaves to no word's
capitals. A decision either way at this comma
is reported; for a run of words the decision is the flip to the
credential run, reported once over the whole part. A flip in which
no listed word of this class takes part is the exception and is
Expand Down Expand Up @@ -1793,6 +1795,9 @@ C1. Rationale: a credential run after the comma means the name is in
"John Smith, Ed Ma" → suffix="Ed Ma"
"Jane Doe, MS LAc" → suffix="MS LAc"
"Smith, PhD MEng" → family="Smith" · boundary
"John Smith, PhD DO DO" → suffix="PhD DO DO"
"John Smith, PhD DO DO" → ambiguities=("suffix-or-name",)
"John Smith, PhD vd DO" → suffix="PhD vd DO"
"John Smith, X.Y.Z." → suffix="X.Y.Z."
"John Smith, X.Y.Z." → ambiguities=()
"John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z."
Expand Down Expand Up @@ -1849,7 +1854,7 @@ C1. Rationale: a credential run after the comma means the name is in
V` reads the suffix and `Smith, John PhD I.` continues the run,
while adding a suffix comma after either turns that same letter
into the middle initial.
history: decisions.md#C1 · interacts: H2, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py
history: decisions.md#C1 · interacts: H2, P2, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py

C2. Rationale: text beyond the recognized comma parts should be
taken in without silent guessing.
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,8 @@ Release Log

- **Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.** ``HumanName("John Smith, Ed Ma")`` gives first ``John``, last ``Smith``, suffix ``Ed Ma``, where 2.0 through 2.3 gave first ``Ed``, middle ``Ma``, last ``John Smith`` -- 1.4.0's reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (``John Smith, MA``), and ``parse()`` reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report, the capitals having decided it (``John Smith, PhD MA``, ``John Smith, MS MA``). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so ``john smith, v phd ma`` gives first ``john``, last ``smith``, suffix ``v phd ma``, where 2.3 gave first ``v``, middle ``ma``, last ``john smith``. ``john smith, md ma`` and ``John Smith, Ms Ma`` move the same way, where 2.0 through 2.3 gave title ``md``/``Ms`` -- the second is the accepted cost, ``Ms`` read as the suffix word it also is, as ``John Smith, Ms`` alone already reads it -- while one name word before the comma keeps the listing form (``Smith, Ms Ma`` gives title ``Ms``, first ``Ma``). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential's company whatever its case: ``John Smith PhD MEng`` and ``Doe, Jane PhD MEng`` give suffix ``PhD MEng``, the fields every release gave (``PhD, MEng`` through 2.2), now reported, and ``doe, jane v phd do`` gives suffix ``v phd do`` where 2.3.0 gave last ``do doe`` -- a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: ``Wang Ma PhD`` keeps last ``Ma``. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: ``Smith, PhD Ma`` gives last ``Smith``, suffix ``PhD Ma``, where 2.3 gave first ``PhD``, middle ``Ma``. A title that is also a credential (``MD``, ``Ms``) opening that part stays a title and nothing in the part speaks, so ``Smith, MD PhD Ma`` keeps title ``MD``, first ``PhD``, middle ``Ma`` and ``Smith, Ms MD Ma`` title ``Ms MD``, first ``Ma``, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so ``Wang M.Eng.`` gives suffix ``M.Eng.``, as ``Wang M.A.`` does and as 2.0 through 2.3 did. See the #544 entry under ``S2`` in ``docs/design/decisions.md`` (closes #544)

- **Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run.** ``HumanName("John Smith, PhD DO DO")`` gives first ``John``, last ``Smith``, suffix ``PhD DO DO``, where 2.2 and 2.3 gave first ``PhD``, last ``DO DO John Smith`` and 2.0 and 2.1 gave title ``PhD``, first ``DO DO`` -- 1.4.0's reading, restored. ``DO``, ``MC`` and ``VD`` are credentials and surname particles at once, and two particles next to each other join into one particle run whatever their capitals, so the capitals no longer settle such a run silently: it is read by the count of name words before the comma, and ``parse()`` reports the call. ``John Smith, PhD vd DO`` gives suffix ``PhD vd DO`` the same way, and ``John Smith, MD DO DO`` gives suffix ``MD DO DO`` where 2.3 gave title ``MD``, first ``DO``, middle ``DO``. A single ``DO`` is still left to its capitals (``John Smith, PhD DO`` gives suffix ``PhD DO`` with no report), and one name word before the comma keeps the listing form (``Smith, PhD DO DO`` gives first ``PhD``, last ``DO DO Smith``, as 2.3 did). See the #562 entry under ``C1`` in ``docs/design/decisions.md`` (closes #562)

- **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516)

- **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516)
Expand Down
16 changes: 16 additions & 0 deletions nameparser/_pipeline/_segment.py
Original file line number Diff line number Diff line change
Expand Up @@ -306,6 +306,7 @@ def class_run(seg: tuple[int, ...]) -> bool:
shaped = 0
pairs = 0
lexicon = state.lexicon
prev_particle = False
for i in groups[1]:
text = state.tokens[i].text
fold = run_word_fold(text, lexicon, state.policy)
Expand Down Expand Up @@ -337,6 +338,21 @@ def class_run(seg: tuple[int, ...]) -> bool:
break
else:
rest.append(text)
# #562, rules.md#C1: "and so is a part holding two
# particles side by side, which P2 joins into one particle
# run that S2 leaves to no word's capitals" -- group chains
# the pair, and the family-comma path can no longer promise
# to read the part whole: it read 'PhD DO DO', 'PhD vd DO'
# and 'MA vd vd' as name text. Where it did read the part
# whole ('VD DO', 'MD DO DO'), the count flips it to the
# same fields and reports the call, as any flip does.
# Asked only while the run is still settled -- nothing
# re-settles it -- with classify's own particle test, since
# group's chain is what the capitals lose to.
if settled:
particle = _normalize(text) in lexicon.particles
settled = not (particle and prev_particle)
prev_particle = particle
else:
# Every word is a member or left to the suffix predicate.
# A run whose every member the WRITING already settles as a
Expand Down
39 changes: 39 additions & 0 deletions tests/v2/cases.py
Original file line number Diff line number Diff line change
Expand Up @@ -835,6 +835,45 @@ def _check_cjk_shape_purity(self) -> None:
"make. Pinned because the reading now rests on the "
"family-comma path alone. 1.4.0 read the same; 2.3.0 "
"read given 'PhD', middle 'MA'"),
Case("comma_run_with_a_particle_chain_takes_the_count",
"John Smith, PhD DO DO",
{"given": "John", "family": "Smith", "suffix": "PhD DO DO"},
ambiguities=("suffix-or-name",),
notes="#562: the two 'DO's stand side by side, so S2 joins "
"them into one particle run rather than leaving either "
"to its capitals, and the row above's stand-down does "
"not apply: the run is the count's, flipped and "
"reported. 1.4.0 read the same; 2.0.0 read title "
"'PhD', given 'DO DO', and 2.3.0 given 'PhD', family "
"'DO DO John Smith'",
shape=3),
Case("comma_run_chained_behind_a_suffix_particle_takes_the_count",
"John Smith, PhD vd DO",
{"given": "John", "family": "Smith", "suffix": "PhD vd DO"},
ambiguities=("suffix-or-name",),
notes="#562: the particles need not be members of the class "
"-- 'vd' is particle and unambiguous suffix vocabulary, "
"and 'DO' beside it is chained all the same. 1.4.0 read the same; 2.3.0 read given 'PhD', "
"family 'vd DO John Smith'",
shape=3),
Case("comma_run_a_particle_pair_opens_is_flipped_and_reported",
"John Smith, vd DO",
{"given": "John", "family": "Smith", "suffix": "vd DO"},
ambiguities=("suffix-or-name",),
notes="#562's accepted cost: the family-comma path already "
"read this part whole, but the pair is a particle run "
"the capitals do not settle, so the count flips it to "
"the same fields and reports the call, as every flip "
"at this comma does. 2.3.0 read family 'John Smith vd "
"DO'"),
Case("comma_run_with_a_lone_trailing_particle_member_stays_settled",
"John Smith, PhD DO",
{"given": "John", "family": "Smith", "suffix": "PhD DO"},
notes="#562's boundary: one 'DO' beside no other particle "
"is left to its capitals (S2), so the run stays settled "
"and silent, as 'John Smith, PhD MA' is. Untagged: every "
"2.x release misread it, and its fix is #289's (#530), "
"not #562's"),
Case("comma_run_by_shape_member_takes_the_count",
"John Smith, X.Y.Z. MA",
{"given": "John", "family": "Smith", "suffix": "X.Y.Z. MA"},
Expand Down
Loading
Loading