Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 23 additions & 13 deletions docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -529,20 +529,30 @@ listed below.
``"John Smith 1.4.2"`` keeps family ``1.4.2``; that retirement
is not behind this switch).
* - ``unlisted_caps_suffixes``
- ``bool``
- Reads an unlisted all-caps word of two or more letters, with no
period in it, in a name written in more than one case as a
credential where the position allows it: ``"John Smith XYZ"``
gives suffix ``XYZ``, and since 2.4 so do the family-comma
form ``"Doe, John XYZ"`` and the word ending a maiden marker's
- ``CapsSuffixes``
- Where an unlisted all-caps word of two or more letters, with no
period in it, reads as a credential. The name must contrast it
with a word holding a capital whose last letter is lowercase
(``Smith``, ``DiCaprio``) that the vocabulary
does not claim as a title, particle or credential: a record
written wholly in capitals, or wholly in lowercase, keeps every
word a name word. ``CapsSuffixes.AFTER_COMMA``, the default, reads
it only in the part right after a comma with two or more name
words before it: ``"John Smith, XYZ"`` gives suffix ``XYZ``,
while ``"Smith, XYZ"`` keeps given ``XYZ`` and a lone two-letter
word, how initials are written, stays the given name
(``"García Márquez, MJ"``). The all-caps SURNAME convention
(``"Jean DUPONT"``, ``"DUPONT, Jean"``) never writes the
capitals there. ``CapsSuffixes.EVERYWHERE`` also reads the end
of a name, the given part's last word after a family comma
(``"Doe, John XYZ"``) and the word ending a maiden marker's
clause (``"Jane Doe nee Smith XYZ"`` gives maiden ``Smith``
with suffix ``XYZ``, where off it keeps maiden
``Smith XYZ``). Defaults to ``False``, and
deliberately:
an all-caps surname is a real writing convention that shape
cannot separate from a credential, so ``"Jean Pierre DUPONT"``
gives family ``Pierre``, suffix ``DUPONT`` with this on. Off,
nothing changes and nothing is reported.
with suffix ``XYZ``), where that convention does write them:
``"Jean Pierre DUPONT"`` then gives family ``Pierre``, suffix
``DUPONT``. ``CapsSuffixes.OFF`` reads none of them and reports
nothing, 2.3's reading -- and the way to keep a given name
written in capitals after a two-word surname, which the default
reads as a credential (``"García Márquez, GABRIEL"``).
* - ``strip_emoji``
- ``bool``
- Excludes emoji from tokenization — they appear in no field or
Expand Down
6 changes: 6 additions & 0 deletions docs/design/decisions.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion docs/design/mechanisms.md

Large diffs are not rendered by default.

69 changes: 50 additions & 19 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -1075,15 +1075,18 @@ S2. Rationale: generational suffixes and credentials are recognized
one such shape, admitted by default (S3); an unlisted all-caps
alphabetic word of two or more letters, standing in a suffix
position of a mixed-case name and belonging to no wordlist, is
the other, admitted only under the caller switch the example
lines below name. That second shape is OFF by default because
French and Korean records write the SURNAME in capitals, so it
is a surname as often as it is a credential and only the caller
knows which corpus this is; the cost of turning it on is that a
three-word name gives up its family name to the acronym, while a
two-word name keeps it — there are no words to spare there, so
the class is considered and declined and only the fork is
reported. The dotted shape does not take every promise of the
the other. By default that second shape is admitted in one
position only, the part after a comma with two or more name
words before it (C1): French and Korean records write the
SURNAME in capitals, and that convention puts the capitals at
the end of a name or before a comma, never after a comma behind
a full name. Everywhere else it is admitted only under the
caller setting the example lines below name, being a surname
there as often as a credential, which only the caller can know;
the cost of that setting is that a three-word name gives up its
family name to the acronym, while a two-word name keeps it —
there are no words to spare there, so the class is considered
and declined and only the fork is reported. The dotted shape does not take every promise of the
class with it: at the first slot after a comma, C1 reads paired
initials as the given name unless something speaks for them, and
makes a flip to the credential run in which no listed member takes
Expand Down Expand Up @@ -1129,10 +1132,10 @@ S2. Rationale: generational suffixes and credentials are recognized
"Doe, John DO Ed" → middle="DO Ed" · boundary
"Doe, John DO Ed" → ambiguities=() · boundary
"John Smith XYZ" → family="XYZ"
"John Smith XYZ" unlisted_caps_suffixes-on → suffix="XYZ"
"Jean DUPONT" unlisted_caps_suffixes-on → family="DUPONT"
"Jean Pierre DUPONT" unlisted_caps_suffixes-on → suffix="DUPONT"
"Jean Pierre DUPONT" unlisted_caps_suffixes-on → family="Pierre"
"John Smith XYZ" unlisted_caps_suffixes-everywhere → suffix="XYZ"
"Jean DUPONT" unlisted_caps_suffixes-everywhere → family="DUPONT"
"Jean Pierre DUPONT" unlisted_caps_suffixes-everywhere → suffix="DUPONT"
"Jean Pierre DUPONT" unlisted_caps_suffixes-everywhere → family="Pierre"
"Jack Ma." → family="Ma." · boundary
"Ph. D. Van Johnson" → family="Van Johnson"
"Ph. D. Van Johnson" → title="Ph."
Expand Down Expand Up @@ -1719,7 +1722,25 @@ C1. Rationale: a credential run after the comma means the name is in
there. Behind two or more name words, paired initials that are
not the only shape word in the part are a credential however
little else speaks for them, since no one writes a person's
initials as two dotted groups. The same count reads a part of two or
initials as two dotted groups. An unlisted all-caps word joins
the class in such a part only where the name carries the
contrast: one of its own words before the comma written the way a
name is written in mixed case, holding a capital with its last
cased letter lowercase — read after composing the word, passing
over every non-letter and caseless letter and any lowercase letter
whose capital is not a single letter, as ß's is not — with no
period and not claimed by the vocabulary
as a title, particle, connective, credential or generation. A
surname written in capitals ends in a capital whatever is glued in
front of it, and a word written wholly in lowercase holds none, so
such a record keeps its given name beside its lowercase particles,
titles and clauses ("GISCARD d'ESTAING, VALÉRY", 'LLOYD
FitzGERALD, RONALD'), and a name written wholly in lowercase reads
the listing form. Where the name carries it, an
unlisted word of two capitals after the comma is the paired
initials' shape undotted and is decided at this comma exactly as
they are, by the sentences above ('García Márquez, MJ' and 'García
Márquez, MJ PhD' keep given 'MJ'). The same count reads a part of two or
more words as the credential run when every word of it is a
suffix word or a word of this class, at least one of them of this
class, and none of them a single-letter roman numeral, in any
Expand Down Expand Up @@ -1850,6 +1871,16 @@ C1. Rationale: a credential run after the comma means the name is in
"John Smith, X.Y.Z." → suffix="X.Y.Z."
"John Smith, X.Y.Z." → ambiguities=()
"John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z."
"John Smith, XYZ" → suffix="XYZ"
"John Smith, XYZ" → ambiguities=("suffix-or-name",)
"John Smith, XYZ" unlisted_caps_suffixes-off → given="XYZ"
"John Smith, LEED AP" → suffix="LEED AP"
"Smith, XYZ" → given="XYZ" · boundary
"García Márquez, MJ" → given="MJ" · boundary
"MÜLLER WEIß, HANS" → given="HANS" · boundary
"García Márquez, MJ PhD" → given="MJ" · boundary
"García Márquez, MJ JK" → suffix="MJ JK"
"John Smith, PhD XYZ" → suffix="PhD XYZ"
"Smith, A.B." → given="A.B." · boundary
"García Márquez, G.J." → given="G.J."
"García Márquez, G.J." → family="García Márquez"
Expand Down Expand Up @@ -1911,7 +1942,6 @@ C1. Rationale: a credential run after the comma means the name is in
not structure — v1 applied the delimiter to the suffix-comma
form alone, and that limitation is kept as parity: "Smith, RN -
CRNA" reads given "RN" under the policy as without it.
"John Smith, LEED AP" → family="Smith" deviates: #291 (today: family="John Smith")
Accepted: the further-comma qualifier carries no example line of
its own. It discriminates PAIRS and spans both branches, so
exemplifying it means a with-comma partner for each — every one
Expand Down Expand Up @@ -1945,7 +1975,7 @@ C2. Rationale: text beyond the recognized comma parts should be
whether the name is written in one case so that nothing leans at
all, or the member is written the way a name is written. And
S2's other by-shape half, the unlisted all-caps word, does not
reach here under its switch either: the shape a tail segment is
reach here in any setting: the shape a tail segment is
recognized by is the dotted one alone.
A part the parse consumes wholly as suffixes raises no report
about reading a word of it as a name, whether it is the part
Expand All @@ -1963,7 +1993,7 @@ C2. Rationale: text beyond the recognized comma parts should be
"John Smith, MD, Ma" → ambiguities=("comma-structure",) · boundary
"Steven Hardman, MD, DO, DDS" → ambiguities=()
"STEVEN HARDMAN, MD, DO, DDS" → ambiguities=("comma-structure",) · boundary
"John Smith, MD, XYZ" unlisted_caps_suffixes-on → ambiguities=("comma-structure",)
"John Smith, MD, XYZ" unlisted_caps_suffixes-everywhere → ambiguities=("comma-structure",)
Accepted: the no-name-reading clause carries no example line of
its own. What it moves is a report with no field beside it, and
an example line would enter the rules corpus for that report
Expand Down Expand Up @@ -2533,8 +2563,9 @@ R4. Rationale: case repair is a display concern, applied only on
mark one, R5 defers to it. A credential acronym the exceptions map
does not carry is an initialism, so a single-case word the parse
put in the suffix role from the acronym vocabulary, or read as a
credential by its dotted shape alone (S3), repairs to its
all-caps spelling rather than a title-cased one, and that repair
credential by its dotted shape alone (S3) or by its capitals
(S2), repairs to its all-caps spelling rather than a title-cased
one, and that repair
outranks the Mac/Mc convention where a word fits both (MCSE, not
McSe). A roman numeral the parse put in the suffix role is
written in capitals the way a generation is written, whether or
Expand Down
3 changes: 3 additions & 0 deletions docs/modules.rst
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,9 @@ Configuration
.. autoclass:: nameparser.PatronymicRule
:members:

.. autoclass:: nameparser.CapsSuffixes
:members:

.. autoclass:: nameparser.Script
:members:

Expand Down
2 changes: 1 addition & 1 deletion docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ Release Log

- **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516)

- **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516)
- **New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request.** Its value is a ``CapsSuffixes``. The default, ``CapsSuffixes.AFTER_COMMA``, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: ``HumanName("John Smith, XYZ")`` gives first ``John``, last ``Smith``, suffix ``XYZ``, where 1.4.0 through 2.3.0 gave first ``XYZ``, last ``John Smith``; ``John Smith, LEED AP`` and ``John Smith, PhD XYZ`` give suffix ``LEED AP`` and ``PhD XYZ`` the same way, and ``John Smith, RAI`` gives suffix ``RAI`` again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (``Jean DUPONT``, ``DUPONT, Jean``) and never there. A word after a one-word surname stays the given name (``Smith, XYZ``), a two-letter word reads exactly as dotted initials do (``García Márquez, MJ`` and ``García Márquez, MJ PhD`` keep first ``MJ``), and the name has to contrast the capitals with a word of its own holding a capital whose last letter is lowercase (``Smith``, ``DiCaprio``); a name typed with decomposed accents reads as its composed spelling. A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (``GISCARD d'ESTAING, VALÉRY``, ``LLOYD FitzGERALD, RONALD``, ``LLOYD WEBBER, ANDREW PhD``), as does a name written wholly in lowercase. ``CapsSuffixes.EVERYWHERE`` also reads the end of a name, the given part's last word after a family comma and the word ending a maiden marker's clause: ``.parse("John Smith XYZ")`` gives suffix ``XYZ``, and ``Jean Pierre DUPONT`` gives last ``Pierre``, suffix ``DUPONT`` -- why it is not the default. ``CapsSuffixes.OFF`` reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (``García Márquez, GABRIEL`` gives suffix ``GABRIEL``). The field reaches the core parser only, through ``Parser(policy=Policy(unlisted_caps_suffixes=...))``; a ``HumanName`` tracks the parser's defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor ``unlisted_dotted_suffixes`` has a v1 ``Constants`` manager. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` (closes #516, closes #564)

- **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``John Smith, X.Y. P.Q.`` gives last ``Smith``, suffix ``X.Y. P.Q.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md``

Expand Down
Loading
Loading