Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -1433,6 +1433,7 @@ Every number below is a py3.11 measurement of 2026-08-31, recomputable with `uv
ONE REPAIR WAS TRIED AND REVERTED: collapsing `post_rules`' three separate role-index scans into one pass. It saves **6 calls per parse** on 3.11 (the first draft said one, which is not even reachable — `_idx` is a one-line list comprehension, so each removed call frees two frames). Reverted because it was measured against a noisy timing harness that showed nothing; at 6 calls it is 9% of the cycle's growth and worth reconsidering against the band, which can now see it.
RAISING OR LOWERING A ROW IS A DECISION, not a maintenance chore. Append here with the interpreter and the harness invocation, as above.

- 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one. One parse of the reference name makes 18 copies: six of the state, one each from tokenize, segment, classify, group, assign and post_rules, and twelve of single tokens, six each from classify and assign (`extract_delimited` returns the state unchanged when there is no delimiter, and `script_segment` returns early on ASCII input). That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). A field copy builds what `replace` builds for a dataclass that is decorated itself rather than inheriting the decoration, keeps the generated `__init__`, and has no `__post_init__` and no `init=False` field. `_copyable_fields` checks exactly those four, and `_COPY_FIELDS` runs it over `WorkToken`, `PendingAmbiguity` and `ParseState` at import, so a class that stops qualifying fails there and `copy_with` copies nothing else. `test_the_guard_refuses_a_class_a_field_copy_would_get_wrong` records, for each refused shape, what `replace` builds and what an unguarded copy would build instead. To mypy, `copy_with` is `from dataclasses import replace as copy_with`, which keeps the dataclass plugin's keyword and type checks at every call site; an assignment (`copy_with = dataclasses.replace`) would not, since the plugin keys on the callee's full name (measured: a misspelled field and a wrong-typed value both pass through the assignment and both fail through the import). Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. The rows drop by 40 and 58 rather than 36 and 54 because e0f1a2f already read 4 under every row, inside the band, and the new rows are set to what the harness reads now. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header.

### removed-v1-surface

Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,8 @@ Release Log

- **Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals.** ``HumanName("john smith x.y.z.").capitalize()`` gives ``John Smith X.Y.Z.`` where every release gave ``John Smith X.y.z.``, the dotted word being a suffix now (the ``unlisted_dotted_suffixes`` change above) and repaired as a listed acronym is; and ``john smith vi`` gives ``John Smith VI`` where every release gave ``John Smith Vi``, with ``vii``, ``viii`` and ``ix`` alike. Both are keyed on the suffix role: ``Jack X.Y.Z.``, which keeps its surname, still repairs as a name word (``Jack X.y.z.`` under ``force=True``), and ``john smith xi`` still gives ``John Smith Xi``, the parser reading ``xi`` as the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (``john smith B.Tech.`` gives ``John Smith B.Tech.``), and reads all capitals under ``force=True`` (``John Smith B.TECH.``, where every release gave ``John Smith B.tech.``), which a ``capitalization_exceptions`` mask such as ``{"btech": "BTech"}`` undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 under ``force=True``. The recipe is the ``R4`` entry's 2026-09-23 MEASURED bullet in ``docs/design/decisions.md`` (#459)

- **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546)

**Additions**

- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)
Expand Down
13 changes: 6 additions & 7 deletions nameparser/_pipeline/_assign.py
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,6 @@
"""
from __future__ import annotations

import dataclasses
from collections.abc import Sequence, Set
from typing import NamedTuple

Expand All @@ -78,15 +77,15 @@
)
from nameparser._pipeline._state import (
AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure,
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, _NEVER_FLIPPED,
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, _NEVER_FLIPPED, copy_with,
)
from nameparser._policy import Policy, Script
from nameparser._types import AmbiguityKind, Role

def _set_roles(tokens: list[WorkToken], piece: tuple[int, ...],
role: Role) -> None:
for i in piece:
tokens[i] = dataclasses.replace(tokens[i], role=role)
tokens[i] = copy_with(tokens[i], role=role)


#: Tags that say the word's own reading was claimed before position
Expand Down Expand Up @@ -672,7 +671,7 @@ def previous_kept(m: int, titled: tuple[int, ...]) -> int:
#: reads `.role`, verified by reading all three
#: (2026-09-19). The one thing this segment's code rewrites
#: between the two passes is the role, through `_set_roles`,
#: which is a `dataclasses.replace(role=...)` and leaves
#: which is a `copy_with(role=...)` and leaves
#: text and tags identical.
floors: dict[tuple[int, ...], tuple[int, bool]] = {}

Expand Down Expand Up @@ -1024,6 +1023,6 @@ def reads_as_a_suffix(m: int, titled: tuple[int, ...]) -> bool:
for seg_idx in range(tail, len(state.segments)):
for piece in state.pieces[seg_idx]:
_set_roles(tokens, piece, Role.SUFFIX)
return dataclasses.replace(state, tokens=tuple(tokens),
order=order,
ambiguities=tuple(ambiguities))
return copy_with(state, tokens=tuple(tokens),
order=order,
ambiguities=tuple(ambiguities))
11 changes: 5 additions & 6 deletions nameparser/_pipeline/_classify.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,12 +49,11 @@
"""
from __future__ import annotations

import dataclasses

from nameparser._lexicon import _normalize
from nameparser._pipeline._state import (
AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, ParseState, PendingAmbiguity,
WorkToken,
WorkToken, copy_with,
)
from nameparser._types import AmbiguityKind, Role
from nameparser._pipeline._vocab import (
Expand Down Expand Up @@ -264,7 +263,7 @@ def classify(state: ParseState) -> ParseState:
# class they consult. No extra frame -- it is one more boolean in a
# comprehension that already walks every token.
tokens = tuple(
dataclasses.replace(
copy_with(
t, tags=_tags_for(t, folded[i], state, marker_tags.get(i),
one_case_own=one_case and i < clause_at
and t.role is None, one_case=one_case))
Expand Down Expand Up @@ -324,6 +323,6 @@ def classify(state: ParseState) -> ParseState:
(i,)))
# The write rides the replace this stage already makes, so
# recording the fact costs no frame of its own.
return dataclasses.replace(state, tokens=tokens,
ambiguities=tuple(ambiguities),
one_case=one_case)
return copy_with(state, tokens=tokens,
ambiguities=tuple(ambiguities),
one_case=one_case)
5 changes: 2 additions & 3 deletions nameparser/_pipeline/_extract.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,12 +32,11 @@
from __future__ import annotations

import bisect
import dataclasses
import functools

from nameparser._lexicon import Lexicon, _normalize
from nameparser._pipeline._state import (
COMMA_CHARS, ParseState, PendingAmbiguity,
COMMA_CHARS, ParseState, PendingAmbiguity, copy_with,
)
from nameparser._pipeline._vocab import maiden_marker_run
from nameparser._types import AmbiguityKind, Role, Span
Expand Down Expand Up @@ -298,6 +297,6 @@ def extract_delimited(state: ParseState) -> ParseState:
continue
reported.add(j)
ambiguities.append(_unmatched(close, j)[1])
return dataclasses.replace(
return copy_with(
state, extracted=tuple(extracted), masked=tuple(masked),
ambiguities=state.ambiguities + tuple(ambiguities))
9 changes: 4 additions & 5 deletions nameparser/_pipeline/_group.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,6 @@
from __future__ import annotations

import bisect
import dataclasses
from collections.abc import Iterable, Sequence, Set
from enum import IntEnum
from typing import assert_never
Expand All @@ -52,7 +51,7 @@
)
from nameparser._pipeline._state import (
AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure,
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS,
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, copy_with,
)
from nameparser._pipeline._vocab import D, PH
from nameparser._pipeline._vocab import delimiter_cores
Expand Down Expand Up @@ -1687,7 +1686,7 @@ def group(state: ParseState) -> ParseState:
dropped.extend(marker_piece)
for piece in maiden_pieces:
for i in piece:
tokens[i] = dataclasses.replace(
tokens[i] = copy_with(
tokens[i], role=Role.MAIDEN)
# rules.md#C1: "a part that is nothing but suffix words is the
# credential run and reads as suffixes, whole" -- WHOLE is this
Expand Down Expand Up @@ -1737,7 +1736,7 @@ def group(state: ParseState) -> ParseState:
for piece, piece_tags_ in zip(pieces, ptags):
if "suffix" in piece_tags_ and len(piece) > 1:
for i in piece[1:]:
tokens[i] = dataclasses.replace(
tokens[i] = copy_with(
tokens[i], tags=tokens[i].tags | {"joined"})
all_pieces.append(tuple(tuple(p) for p in pieces))
all_ptags.append(tuple(frozenset(t) for t in ptags))
Expand Down Expand Up @@ -1816,7 +1815,7 @@ def group(state: ParseState) -> ParseState:
if (first + run < len(tokens)
and tokens[first + run].span.end <= clause.end):
dropped.extend(range(first, first + run))
return dataclasses.replace(
return copy_with(
state, tokens=tuple(tokens), pieces=tuple(all_pieces),
piece_tags=tuple(all_ptags), dropped=tuple(dropped),
ambiguities=tuple(ambiguities))
23 changes: 11 additions & 12 deletions nameparser/_pipeline/_post_rules.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,14 +20,13 @@
"""
from __future__ import annotations

import dataclasses
import re

from nameparser._lexicon import _run_addresses_by_given
from nameparser._pipeline._assign import _name_positions
from nameparser._pipeline._state import (
AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure,
WorkToken, _NEVER_FLIPPED, comma_bucket,
WorkToken, _NEVER_FLIPPED, comma_bucket, copy_with,
)
from nameparser._pipeline._vocab import delimiter_cores
from nameparser._policy import PatronymicRule
Expand Down Expand Up @@ -193,7 +192,7 @@ def _mark_suffix_entries(tokens: list[WorkToken], state: ParseState) -> None:
else tokens[between].role not in _RENDERS_ELSEWHERE
for between in range(previous + 1, current))
if same_part and not parted:
tokens[current] = dataclasses.replace(
tokens[current] = copy_with(
tokens[current], tags=tokens[current].tags | {"joined"})


Expand All @@ -214,7 +213,7 @@ def suffix_entries(state: ParseState) -> ParseState:
nothing else."""
tokens = list(state.tokens)
_mark_suffix_entries(tokens, state)
return dataclasses.replace(state, tokens=tuple(tokens))
return copy_with(state, tokens=tuple(tokens))


def _idx(tokens: list[WorkToken], role: Role) -> list[int]:
Expand Down Expand Up @@ -247,7 +246,7 @@ def _leading_name_piece(state: ParseState,


def _retag(tokens: list[WorkToken], i: int, role: Role) -> None:
tokens[i] = dataclasses.replace(tokens[i], role=role)
tokens[i] = copy_with(tokens[i], role=role)


# rules.md#P2: "a particle joins the words after it into one name
Expand Down Expand Up @@ -652,7 +651,7 @@ def post_rules(state: ParseState) -> ParseState:
f"of its own",
tuple(sorted(run))))
for j in run:
tokens[j] = dataclasses.replace(
tokens[j] = copy_with(
tokens[j], role=Role.FAMILY,
tags=tokens[j].tags | {FOLDED_TAG})
# recomputed for H1's reason, stated at H1: a stale index
Expand Down Expand Up @@ -811,7 +810,7 @@ def post_rules(state: ParseState) -> ParseState:
f"rather than standing as a name word of its own",
tuple(run)))
for i in run:
tokens[i] = dataclasses.replace(
tokens[i] = copy_with(
tokens[i], role=Role.FAMILY,
tags=tokens[i].tags | {FOLDED_TAG})

Expand All @@ -822,7 +821,7 @@ def post_rules(state: ParseState) -> ParseState:
# tags the token, and the rendering views consult the tag"
if state.policy.middle_as_family:
for i in _idx(tokens, Role.MIDDLE):
tokens[i] = dataclasses.replace(
tokens[i] = copy_with(
tokens[i], role=Role.FAMILY,
tags=tokens[i].tags | {FOLDED_TAG})
# rules.md#R2: "a name part whose every word is particle
Expand Down Expand Up @@ -874,7 +873,7 @@ def post_rules(state: ParseState) -> ParseState:
others += 1
if all_particle:
for i in part:
tokens[i] = dataclasses.replace(
tokens[i] = copy_with(
tokens[i], tags=tokens[i].tags | {UNJOINED_TAG})
elif conj and not others:
# #461: nothing here for the connective to join. The `elif`
Expand All @@ -883,9 +882,9 @@ def post_rules(state: ParseState) -> ParseState:
# and connective included, which is what keeps a caller's
# `add(particles={"y"})` readings unchanged.
for i in conj:
tokens[i] = dataclasses.replace(
tokens[i] = copy_with(
tokens[i],
tags=tokens[i].tags | {UNJOINED_CONJUNCTION_TAG})
_mark_suffix_entries(tokens, state)
return dataclasses.replace(state, tokens=tuple(tokens),
ambiguities=tuple(ambiguities))
return copy_with(state, tokens=tuple(tokens),
ambiguities=tuple(ambiguities))
13 changes: 6 additions & 7 deletions nameparser/_pipeline/_script_segment.py
Original file line number Diff line number Diff line change
Expand Up @@ -58,13 +58,12 @@
"""
from __future__ import annotations

import dataclasses
import functools
from collections.abc import Sequence

from nameparser._lexicon import FULL_STOPS
from nameparser._pipeline._state import (
ParseState, PendingAmbiguity, Structure, WorkToken,
ParseState, PendingAmbiguity, Structure, WorkToken, copy_with,
)
from nameparser._pipeline._vocab import (
effective_script, is_suffix_strict, is_wholly_suffix,
Expand Down Expand Up @@ -157,11 +156,11 @@ def _split(state: ParseState, i: int, splits: tuple[int, ...],
start = 0
for piece in _pieces(token.text, splits):
end = start + len(piece)
parts.append(dataclasses.replace(
parts.append(copy_with(
token, text=piece, span=Span(base + start, base + end)))
start = end
if tail_tag is not None:
parts[-1] = dataclasses.replace(
parts[-1] = copy_with(
parts[-1], tags=parts[-1].tags | {tail_tag})
added = len(splits)
tokens = state.tokens[:i] + tuple(parts) + state.tokens[i + 1:]
Expand All @@ -172,15 +171,15 @@ def _split(state: ParseState, i: int, splits: tuple[int, ...],
# pointing at the head.
segments = tuple(_remap(run, i, added) for run in state.segments)
ambiguities = tuple(
dataclasses.replace(a, indices=tuple(
copy_with(a, indices=tuple(
j + added if j > i else j for j in a.indices))
for a in state.ambiguities)
if detail is not None:
ambiguities += (PendingAmbiguity(
AmbiguityKind.SEGMENTATION, detail,
tuple(range(i, i + added + 1))),)
return dataclasses.replace(state, tokens=tokens, segments=segments,
ambiguities=ambiguities)
return copy_with(state, tokens=tokens, segments=segments,
ambiguities=ambiguities)


@functools.lru_cache(maxsize=16)
Expand Down
Loading
Loading