Skip to content

Commit 952ecf5

Browse files
authored
Merge pull request #546 from akamick86/perf/pipeline-state-copy
perf(pipeline): copy stage state without dataclasses.replace
2 parents f6a79ec + a2edd3f commit 952ecf5

13 files changed

Lines changed: 245 additions & 70 deletions

File tree

‎docs/design/decisions.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1433,6 +1433,7 @@ Every number below is a py3.11 measurement of 2026-08-31, recomputable with `uv
14331433
ONE REPAIR WAS TRIED AND REVERTED: collapsing `post_rules`' three separate role-index scans into one pass. It saves **6 calls per parse** on 3.11 (the first draft said one, which is not even reachable — `_idx` is a one-line list comprehension, so each removed call frees two frames). Reverted because it was measured against a noisy timing harness that showed nothing; at 6 calls it is 9% of the cycle's growth and worth reconsidering against the band, which can now see it.
14341434
RAISING OR LOWERING A ROW IS A DECISION, not a maintenance chore. Append here with the interpreter and the harness invocation, as above.
14351435

1436+
- 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one. One parse of the reference name makes 18 copies: six of the state, one each from tokenize, segment, classify, group, assign and post_rules, and twelve of single tokens, six each from classify and assign (`extract_delimited` returns the state unchanged when there is no delimiter, and `script_segment` returns early on ASCII input). That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). A field copy builds what `replace` builds for a dataclass that is decorated itself rather than inheriting the decoration, keeps the generated `__init__`, and has no `__post_init__` and no `init=False` field. `_copyable_fields` checks exactly those four, and `_COPY_FIELDS` runs it over `WorkToken`, `PendingAmbiguity` and `ParseState` at import, so a class that stops qualifying fails there and `copy_with` copies nothing else. `test_the_guard_refuses_a_class_a_field_copy_would_get_wrong` records, for each refused shape, what `replace` builds and what an unguarded copy would build instead. To mypy, `copy_with` is `from dataclasses import replace as copy_with`, which keeps the dataclass plugin's keyword and type checks at every call site; an assignment (`copy_with = dataclasses.replace`) would not, since the plugin keys on the callee's full name (measured: a misspelled field and a wrong-typed value both pass through the assignment and both fail through the import). Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. The rows drop by 40 and 58 rather than 36 and 54 because e0f1a2f already read 4 under every row, inside the band, and the new rows are set to what the harness reads now. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header.
14361437

14371438
### removed-v1-surface
14381439

‎docs/release_log.rst‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -36,6 +36,8 @@ Release Log
3636

3737
- **Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals.** ``HumanName("john smith x.y.z.").capitalize()`` gives ``John Smith X.Y.Z.`` where every release gave ``John Smith X.y.z.``, the dotted word being a suffix now (the ``unlisted_dotted_suffixes`` change above) and repaired as a listed acronym is; and ``john smith vi`` gives ``John Smith VI`` where every release gave ``John Smith Vi``, with ``vii``, ``viii`` and ``ix`` alike. Both are keyed on the suffix role: ``Jack X.Y.Z.``, which keeps its surname, still repairs as a name word (``Jack X.y.z.`` under ``force=True``), and ``john smith xi`` still gives ``John Smith Xi``, the parser reading ``xi`` as the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (``john smith B.Tech.`` gives ``John Smith B.Tech.``), and reads all capitals under ``force=True`` (``John Smith B.TECH.``, where every release gave ``John Smith B.tech.``), which a ``capitalization_exceptions`` mask such as ``{"btech": "BTech"}`` undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 under ``force=True``. The recipe is the ``R4`` entry's 2026-09-23 MEASURED bullet in ``docs/design/decisions.md`` (#459)
3838

39+
- **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546)
40+
3941
**Additions**
4042

4143
- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)

‎nameparser/_pipeline/_assign.py‎

Lines changed: 6 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -63,7 +63,6 @@
6363
"""
6464
from __future__ import annotations
6565

66-
import dataclasses
6766
from collections.abc import Sequence, Set
6867
from typing import NamedTuple
6968

@@ -78,15 +77,15 @@
7877
)
7978
from nameparser._pipeline._state import (
8079
AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure,
81-
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, _NEVER_FLIPPED,
80+
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, _NEVER_FLIPPED, copy_with,
8281
)
8382
from nameparser._policy import Policy, Script
8483
from nameparser._types import AmbiguityKind, Role
8584

8685
def _set_roles(tokens: list[WorkToken], piece: tuple[int, ...],
8786
role: Role) -> None:
8887
for i in piece:
89-
tokens[i] = dataclasses.replace(tokens[i], role=role)
88+
tokens[i] = copy_with(tokens[i], role=role)
9089

9190

9291
#: Tags that say the word's own reading was claimed before position
@@ -672,7 +671,7 @@ def previous_kept(m: int, titled: tuple[int, ...]) -> int:
672671
#: reads `.role`, verified by reading all three
673672
#: (2026-09-19). The one thing this segment's code rewrites
674673
#: between the two passes is the role, through `_set_roles`,
675-
#: which is a `dataclasses.replace(role=...)` and leaves
674+
#: which is a `copy_with(role=...)` and leaves
676675
#: text and tags identical.
677676
floors: dict[tuple[int, ...], tuple[int, bool]] = {}
678677

@@ -1024,6 +1023,6 @@ def reads_as_a_suffix(m: int, titled: tuple[int, ...]) -> bool:
10241023
for seg_idx in range(tail, len(state.segments)):
10251024
for piece in state.pieces[seg_idx]:
10261025
_set_roles(tokens, piece, Role.SUFFIX)
1027-
return dataclasses.replace(state, tokens=tuple(tokens),
1028-
order=order,
1029-
ambiguities=tuple(ambiguities))
1026+
return copy_with(state, tokens=tuple(tokens),
1027+
order=order,
1028+
ambiguities=tuple(ambiguities))

‎nameparser/_pipeline/_classify.py‎

Lines changed: 5 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -49,12 +49,11 @@
4949
"""
5050
from __future__ import annotations
5151

52-
import dataclasses
5352

5453
from nameparser._lexicon import _normalize
5554
from nameparser._pipeline._state import (
5655
AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, ParseState, PendingAmbiguity,
57-
WorkToken,
56+
WorkToken, copy_with,
5857
)
5958
from nameparser._types import AmbiguityKind, Role
6059
from nameparser._pipeline._vocab import (
@@ -264,7 +263,7 @@ def classify(state: ParseState) -> ParseState:
264263
# class they consult. No extra frame -- it is one more boolean in a
265264
# comprehension that already walks every token.
266265
tokens = tuple(
267-
dataclasses.replace(
266+
copy_with(
268267
t, tags=_tags_for(t, folded[i], state, marker_tags.get(i),
269268
one_case_own=one_case and i < clause_at
270269
and t.role is None, one_case=one_case))
@@ -324,6 +323,6 @@ def classify(state: ParseState) -> ParseState:
324323
(i,)))
325324
# The write rides the replace this stage already makes, so
326325
# recording the fact costs no frame of its own.
327-
return dataclasses.replace(state, tokens=tokens,
328-
ambiguities=tuple(ambiguities),
329-
one_case=one_case)
326+
return copy_with(state, tokens=tokens,
327+
ambiguities=tuple(ambiguities),
328+
one_case=one_case)

‎nameparser/_pipeline/_extract.py‎

Lines changed: 2 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -32,12 +32,11 @@
3232
from __future__ import annotations
3333

3434
import bisect
35-
import dataclasses
3635
import functools
3736

3837
from nameparser._lexicon import Lexicon, _normalize
3938
from nameparser._pipeline._state import (
40-
COMMA_CHARS, ParseState, PendingAmbiguity,
39+
COMMA_CHARS, ParseState, PendingAmbiguity, copy_with,
4140
)
4241
from nameparser._pipeline._vocab import maiden_marker_run
4342
from nameparser._types import AmbiguityKind, Role, Span
@@ -298,6 +297,6 @@ def extract_delimited(state: ParseState) -> ParseState:
298297
continue
299298
reported.add(j)
300299
ambiguities.append(_unmatched(close, j)[1])
301-
return dataclasses.replace(
300+
return copy_with(
302301
state, extracted=tuple(extracted), masked=tuple(masked),
303302
ambiguities=state.ambiguities + tuple(ambiguities))

‎nameparser/_pipeline/_group.py‎

Lines changed: 4 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -38,7 +38,6 @@
3838
from __future__ import annotations
3939

4040
import bisect
41-
import dataclasses
4241
from collections.abc import Iterable, Sequence, Set
4342
from enum import IntEnum
4443
from typing import assert_never
@@ -52,7 +51,7 @@
5251
)
5352
from nameparser._pipeline._state import (
5453
AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure,
55-
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS,
54+
WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, copy_with,
5655
)
5756
from nameparser._pipeline._vocab import D, PH
5857
from nameparser._pipeline._vocab import delimiter_cores
@@ -1687,7 +1686,7 @@ def group(state: ParseState) -> ParseState:
16871686
dropped.extend(marker_piece)
16881687
for piece in maiden_pieces:
16891688
for i in piece:
1690-
tokens[i] = dataclasses.replace(
1689+
tokens[i] = copy_with(
16911690
tokens[i], role=Role.MAIDEN)
16921691
# rules.md#C1: "a part that is nothing but suffix words is the
16931692
# credential run and reads as suffixes, whole" -- WHOLE is this
@@ -1737,7 +1736,7 @@ def group(state: ParseState) -> ParseState:
17371736
for piece, piece_tags_ in zip(pieces, ptags):
17381737
if "suffix" in piece_tags_ and len(piece) > 1:
17391738
for i in piece[1:]:
1740-
tokens[i] = dataclasses.replace(
1739+
tokens[i] = copy_with(
17411740
tokens[i], tags=tokens[i].tags | {"joined"})
17421741
all_pieces.append(tuple(tuple(p) for p in pieces))
17431742
all_ptags.append(tuple(frozenset(t) for t in ptags))
@@ -1816,7 +1815,7 @@ def group(state: ParseState) -> ParseState:
18161815
if (first + run < len(tokens)
18171816
and tokens[first + run].span.end <= clause.end):
18181817
dropped.extend(range(first, first + run))
1819-
return dataclasses.replace(
1818+
return copy_with(
18201819
state, tokens=tuple(tokens), pieces=tuple(all_pieces),
18211820
piece_tags=tuple(all_ptags), dropped=tuple(dropped),
18221821
ambiguities=tuple(ambiguities))

‎nameparser/_pipeline/_post_rules.py‎

Lines changed: 11 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -20,14 +20,13 @@
2020
"""
2121
from __future__ import annotations
2222

23-
import dataclasses
2423
import re
2524

2625
from nameparser._lexicon import _run_addresses_by_given
2726
from nameparser._pipeline._assign import _name_positions
2827
from nameparser._pipeline._state import (
2928
AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure,
30-
WorkToken, _NEVER_FLIPPED, comma_bucket,
29+
WorkToken, _NEVER_FLIPPED, comma_bucket, copy_with,
3130
)
3231
from nameparser._pipeline._vocab import delimiter_cores
3332
from nameparser._policy import PatronymicRule
@@ -193,7 +192,7 @@ def _mark_suffix_entries(tokens: list[WorkToken], state: ParseState) -> None:
193192
else tokens[between].role not in _RENDERS_ELSEWHERE
194193
for between in range(previous + 1, current))
195194
if same_part and not parted:
196-
tokens[current] = dataclasses.replace(
195+
tokens[current] = copy_with(
197196
tokens[current], tags=tokens[current].tags | {"joined"})
198197

199198

@@ -214,7 +213,7 @@ def suffix_entries(state: ParseState) -> ParseState:
214213
nothing else."""
215214
tokens = list(state.tokens)
216215
_mark_suffix_entries(tokens, state)
217-
return dataclasses.replace(state, tokens=tuple(tokens))
216+
return copy_with(state, tokens=tuple(tokens))
218217

219218

220219
def _idx(tokens: list[WorkToken], role: Role) -> list[int]:
@@ -247,7 +246,7 @@ def _leading_name_piece(state: ParseState,
247246

248247

249248
def _retag(tokens: list[WorkToken], i: int, role: Role) -> None:
250-
tokens[i] = dataclasses.replace(tokens[i], role=role)
249+
tokens[i] = copy_with(tokens[i], role=role)
251250

252251

253252
# rules.md#P2: "a particle joins the words after it into one name
@@ -652,7 +651,7 @@ def post_rules(state: ParseState) -> ParseState:
652651
f"of its own",
653652
tuple(sorted(run))))
654653
for j in run:
655-
tokens[j] = dataclasses.replace(
654+
tokens[j] = copy_with(
656655
tokens[j], role=Role.FAMILY,
657656
tags=tokens[j].tags | {FOLDED_TAG})
658657
# recomputed for H1's reason, stated at H1: a stale index
@@ -811,7 +810,7 @@ def post_rules(state: ParseState) -> ParseState:
811810
f"rather than standing as a name word of its own",
812811
tuple(run)))
813812
for i in run:
814-
tokens[i] = dataclasses.replace(
813+
tokens[i] = copy_with(
815814
tokens[i], role=Role.FAMILY,
816815
tags=tokens[i].tags | {FOLDED_TAG})
817816

@@ -822,7 +821,7 @@ def post_rules(state: ParseState) -> ParseState:
822821
# tags the token, and the rendering views consult the tag"
823822
if state.policy.middle_as_family:
824823
for i in _idx(tokens, Role.MIDDLE):
825-
tokens[i] = dataclasses.replace(
824+
tokens[i] = copy_with(
826825
tokens[i], role=Role.FAMILY,
827826
tags=tokens[i].tags | {FOLDED_TAG})
828827
# rules.md#R2: "a name part whose every word is particle
@@ -874,7 +873,7 @@ def post_rules(state: ParseState) -> ParseState:
874873
others += 1
875874
if all_particle:
876875
for i in part:
877-
tokens[i] = dataclasses.replace(
876+
tokens[i] = copy_with(
878877
tokens[i], tags=tokens[i].tags | {UNJOINED_TAG})
879878
elif conj and not others:
880879
# #461: nothing here for the connective to join. The `elif`
@@ -883,9 +882,9 @@ def post_rules(state: ParseState) -> ParseState:
883882
# and connective included, which is what keeps a caller's
884883
# `add(particles={"y"})` readings unchanged.
885884
for i in conj:
886-
tokens[i] = dataclasses.replace(
885+
tokens[i] = copy_with(
887886
tokens[i],
888887
tags=tokens[i].tags | {UNJOINED_CONJUNCTION_TAG})
889888
_mark_suffix_entries(tokens, state)
890-
return dataclasses.replace(state, tokens=tuple(tokens),
891-
ambiguities=tuple(ambiguities))
889+
return copy_with(state, tokens=tuple(tokens),
890+
ambiguities=tuple(ambiguities))

‎nameparser/_pipeline/_script_segment.py‎

Lines changed: 6 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -58,13 +58,12 @@
5858
"""
5959
from __future__ import annotations
6060

61-
import dataclasses
6261
import functools
6362
from collections.abc import Sequence
6463

6564
from nameparser._lexicon import FULL_STOPS
6665
from nameparser._pipeline._state import (
67-
ParseState, PendingAmbiguity, Structure, WorkToken,
66+
ParseState, PendingAmbiguity, Structure, WorkToken, copy_with,
6867
)
6968
from nameparser._pipeline._vocab import (
7069
effective_script, is_suffix_strict, is_wholly_suffix,
@@ -157,11 +156,11 @@ def _split(state: ParseState, i: int, splits: tuple[int, ...],
157156
start = 0
158157
for piece in _pieces(token.text, splits):
159158
end = start + len(piece)
160-
parts.append(dataclasses.replace(
159+
parts.append(copy_with(
161160
token, text=piece, span=Span(base + start, base + end)))
162161
start = end
163162
if tail_tag is not None:
164-
parts[-1] = dataclasses.replace(
163+
parts[-1] = copy_with(
165164
parts[-1], tags=parts[-1].tags | {tail_tag})
166165
added = len(splits)
167166
tokens = state.tokens[:i] + tuple(parts) + state.tokens[i + 1:]
@@ -172,15 +171,15 @@ def _split(state: ParseState, i: int, splits: tuple[int, ...],
172171
# pointing at the head.
173172
segments = tuple(_remap(run, i, added) for run in state.segments)
174173
ambiguities = tuple(
175-
dataclasses.replace(a, indices=tuple(
174+
copy_with(a, indices=tuple(
176175
j + added if j > i else j for j in a.indices))
177176
for a in state.ambiguities)
178177
if detail is not None:
179178
ambiguities += (PendingAmbiguity(
180179
AmbiguityKind.SEGMENTATION, detail,
181180
tuple(range(i, i + added + 1))),)
182-
return dataclasses.replace(state, tokens=tokens, segments=segments,
183-
ambiguities=ambiguities)
181+
return copy_with(state, tokens=tokens, segments=segments,
182+
ambiguities=ambiguities)
184183

185184

186185
@functools.lru_cache(maxsize=16)

0 commit comments

Comments
 (0)