Skip to content

Commit a2edd3f

Browse files
committed
perf(pipeline): address review -- typed copy_with, import-time table, exact guard
- mypy sees copy_with as `from dataclasses import replace as copy_with`, so the dataclass plugin checks its keywords again. An assignment (`copy_with = dataclasses.replace`) does not: the plugin keys on the callee's full name, and a misspelled or wrong-typed field passes. - _COPY_FIELDS is a read-only table built at import over the three pipeline classes, replacing the lazily filled dict. copy_with refuses any other class, and the zero-field falsy-cache wrinkle is gone. - _copyable_fields checks the class's own __dataclass_params__ and a generated __init__, so a validating __init__ and an undecorated subclass are refused. The redundant params.init test is dropped. - The refusal test records what replace builds and what an unguarded copy would build for each refused shape; each guard clause fails its own row when removed. - Release log states the saving in calls; the decisions entry gives the 6 + 12 copy breakdown and why the rows drop by 40 and 58. - Continuation lines realigned, _segment's double blank line removed.
1 parent d665457 commit a2edd3f

10 files changed

Lines changed: 152 additions & 69 deletions

File tree

‎docs/design/decisions.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1433,7 +1433,7 @@ Every number below is a py3.11 measurement of 2026-08-31, recomputable with `uv
14331433
ONE REPAIR WAS TRIED AND REVERTED: collapsing `post_rules`' three separate role-index scans into one pass. It saves **6 calls per parse** on 3.11 (the first draft said one, which is not even reachable — `_idx` is a one-line list comprehension, so each removed call frees two frames). Reverted because it was measured against a noisy timing harness that showed nothing; at 6 calls it is 9% of the cycle's growth and worth reconsidering against the band, which can now see it.
14341434
RAISING OR LOWERING A ROW IS A DECISION, not a maintenance chore. Append here with the interpreter and the harness invocation, as above.
14351435

1436-
- 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one, and one parse of the reference name makes 18 copies: every stage returns a copy of the state, and classify, assign, group and post_rules also copy tokens one at a time. That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). For a dataclass whose generated `__init__` only assigns its fields, which all three pipeline classes are (`WorkToken`, `ParseState`, `PendingAmbiguity`), a field copy builds the same object. `copy_with` checks that once per class, refuses a class that validates in `__post_init__` or carries an `init=False` field, and raises `TypeError` for an unknown field as `replace` does. Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header.
1436+
- 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one. One parse of the reference name makes 18 copies: six of the state, one each from tokenize, segment, classify, group, assign and post_rules, and twelve of single tokens, six each from classify and assign (`extract_delimited` returns the state unchanged when there is no delimiter, and `script_segment` returns early on ASCII input). That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). A field copy builds what `replace` builds for a dataclass that is decorated itself rather than inheriting the decoration, keeps the generated `__init__`, and has no `__post_init__` and no `init=False` field. `_copyable_fields` checks exactly those four, and `_COPY_FIELDS` runs it over `WorkToken`, `PendingAmbiguity` and `ParseState` at import, so a class that stops qualifying fails there and `copy_with` copies nothing else. `test_the_guard_refuses_a_class_a_field_copy_would_get_wrong` records, for each refused shape, what `replace` builds and what an unguarded copy would build instead. To mypy, `copy_with` is `from dataclasses import replace as copy_with`, which keeps the dataclass plugin's keyword and type checks at every call site; an assignment (`copy_with = dataclasses.replace`) would not, since the plugin keys on the callee's full name (measured: a misspelled field and a wrong-typed value both pass through the assignment and both fail through the import). Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. The rows drop by 40 and 58 rather than 36 and 54 because e0f1a2f already read 4 under every row, inside the band, and the new rows are set to what the harness reads now. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header.
14371437

14381438
### removed-v1-surface
14391439

‎docs/release_log.rst‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -36,7 +36,7 @@ Release Log
3636

3737
- **Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals.** ``HumanName("john smith x.y.z.").capitalize()`` gives ``John Smith X.Y.Z.`` where every release gave ``John Smith X.y.z.``, the dotted word being a suffix now (the ``unlisted_dotted_suffixes`` change above) and repaired as a listed acronym is; and ``john smith vi`` gives ``John Smith VI`` where every release gave ``John Smith Vi``, with ``vii``, ``viii`` and ``ix`` alike. Both are keyed on the suffix role: ``Jack X.Y.Z.``, which keeps its surname, still repairs as a name word (``Jack X.y.z.`` under ``force=True``), and ``john smith xi`` still gives ``John Smith Xi``, the parser reading ``xi`` as the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (``john smith B.Tech.`` gives ``John Smith B.Tech.``), and reads all capitals under ``force=True`` (``John Smith B.TECH.``, where every release gave ``John Smith B.tech.``), which a ``capitalization_exceptions`` mask such as ``{"btech": "BTech"}`` undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 under ``force=True``. The recipe is the ``R4`` entry's 2026-09-23 MEASURED bullet in ``docs/design/decisions.md`` (#459)
3838

39-
- **Change the parse pipeline to copy its state without dataclasses.replace, about 10% faster per name.** Every stage returns a copy of its frozen state, and several copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that refuses any class where that would differ from ``replace``. On py3.11 ``parse`` drops from 406 to 370 calls per parse of the benchmark's reference name and ``HumanName`` from 443 to 407, and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546)
39+
- **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546)
4040

4141
**Additions**
4242

‎nameparser/_pipeline/_assign.py‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1024,5 +1024,5 @@ def reads_as_a_suffix(m: int, titled: tuple[int, ...]) -> bool:
10241024
for piece in state.pieces[seg_idx]:
10251025
_set_roles(tokens, piece, Role.SUFFIX)
10261026
return copy_with(state, tokens=tuple(tokens),
1027-
order=order,
1028-
ambiguities=tuple(ambiguities))
1027+
order=order,
1028+
ambiguities=tuple(ambiguities))

‎nameparser/_pipeline/_classify.py‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -324,5 +324,5 @@ def classify(state: ParseState) -> ParseState:
324324
# The write rides the replace this stage already makes, so
325325
# recording the fact costs no frame of its own.
326326
return copy_with(state, tokens=tokens,
327-
ambiguities=tuple(ambiguities),
328-
one_case=one_case)
327+
ambiguities=tuple(ambiguities),
328+
one_case=one_case)

‎nameparser/_pipeline/_post_rules.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -887,4 +887,4 @@ def post_rules(state: ParseState) -> ParseState:
887887
tags=tokens[i].tags | {UNJOINED_CONJUNCTION_TAG})
888888
_mark_suffix_entries(tokens, state)
889889
return copy_with(state, tokens=tuple(tokens),
890-
ambiguities=tuple(ambiguities))
890+
ambiguities=tuple(ambiguities))

‎nameparser/_pipeline/_script_segment.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -179,7 +179,7 @@ def _split(state: ParseState, i: int, splits: tuple[int, ...],
179179
AmbiguityKind.SEGMENTATION, detail,
180180
tuple(range(i, i + added + 1))),)
181181
return copy_with(state, tokens=tokens, segments=segments,
182-
ambiguities=ambiguities)
182+
ambiguities=ambiguities)
183183

184184

185185
@functools.lru_cache(maxsize=16)

‎nameparser/_pipeline/_segment.py‎

Lines changed: 6 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -37,7 +37,6 @@
3737
"""
3838
from __future__ import annotations
3939

40-
4140
from nameparser._pipeline._pieces import own_words
4241
from nameparser._pipeline._state import (
4342
ParseState, PendingAmbiguity, Structure, comma_bucket, copy_with,
@@ -56,10 +55,10 @@ def segment(state: ParseState) -> ParseState:
5655
main = [i for i, t in enumerate(state.tokens) if t.role is None]
5756
if not main:
5857
return copy_with(state, segments=(),
59-
structure=Structure.NO_COMMA)
58+
structure=Structure.NO_COMMA)
6059
if not state.comma_offsets:
6160
return copy_with(state, segments=(tuple(main),),
62-
structure=Structure.NO_COMMA)
61+
structure=Structure.NO_COMMA)
6362
buckets: list[list[int]] = [[] for _ in range(len(state.comma_offsets) + 1)]
6463
for i in main:
6564
# _state.comma_bucket, not a local bisect: classify asks the
@@ -79,7 +78,7 @@ def segment(state: ParseState) -> ParseState:
7978
if len(groups) <= 1:
8079
segs = tuple(groups) if groups and groups[0] else (tuple(main),)
8180
return copy_with(state, segments=segs,
82-
structure=Structure.NO_COMMA)
81+
structure=Structure.NO_COMMA)
8382

8483
# The case fact, asked LAZILY: only a comma form can turn on it
8584
# here, and only where the part after the first comma is a single
@@ -306,6 +305,6 @@ def class_run(seg: tuple[int, ...]) -> bool:
306305
f"structures; consumed as suffix best-effort",
307306
tuple(seg)))
308307
return copy_with(state, segments=tuple(groups),
309-
structure=structure,
310-
ambiguities=tuple(ambiguities),
311-
one_case=one_case)
308+
structure=structure,
309+
ambiguities=tuple(ambiguities),
310+
one_case=one_case)

‎nameparser/_pipeline/_state.py‎

Lines changed: 57 additions & 43 deletions
Original file line numberDiff line numberDiff line change
@@ -12,58 +12,18 @@
1212

1313
import bisect
1414
import dataclasses
15-
from collections.abc import Sequence
15+
from collections.abc import Mapping, Sequence
1616
from dataclasses import dataclass
1717
from enum import Enum, auto
18-
from typing import TypeVar
18+
from types import MappingProxyType
19+
from typing import TYPE_CHECKING, TypeVar
1920

2021
from nameparser._lexicon import Lexicon
2122
from nameparser._policy import Policy
2223
from nameparser._types import (SHAPE_ACRONYM_TAG, AmbiguityKind,
2324
Role, Segmenter, Span)
2425

2526

26-
_T = TypeVar("_T")
27-
28-
#: Field names per class, recorded the first time `copy_with` copies one.
29-
_COPY_FIELDS: dict[type, tuple[str, ...]] = {}
30-
31-
32-
def copy_with(obj: _T, /, **changes: object) -> _T:
33-
"""`dataclasses.replace` for the pipeline's own dataclasses, without
34-
its per-call cost.
35-
36-
Every stage returns a copy of the state, and several copy tokens one
37-
at a time, so this runs many times per parse. The stdlib replace
38-
walks `fields()` and goes through `__init__` on every call; for a
39-
class whose `__init__` only assigns its fields, copying the fields
40-
directly gives the same object. `_copy_fields` checks that once per
41-
class and refuses any class where it would not hold
42-
(decisions.md#parse-cost has the measurement).
43-
"""
44-
cls = type(obj)
45-
names = _COPY_FIELDS.get(cls) or _copy_fields(cls)
46-
new = object.__new__(cls)
47-
for name in names:
48-
value = changes.pop(name) if name in changes else getattr(obj, name)
49-
object.__setattr__(new, name, value)
50-
if changes:
51-
raise TypeError(
52-
f"{cls.__name__} has no field named {', '.join(sorted(changes))}")
53-
return new
54-
55-
56-
def _copy_fields(cls: type) -> tuple[str, ...]:
57-
params = getattr(cls, "__dataclass_params__", None)
58-
if (params is None or not params.init or hasattr(cls, "__post_init__")
59-
or not all(f.init for f in dataclasses.fields(cls))):
60-
raise TypeError(
61-
f"copy_with cannot copy {cls.__name__}: it needs a dataclass "
62-
"whose generated __init__ only assigns its fields")
63-
names = _COPY_FIELDS[cls] = tuple(f.name for f in dataclasses.fields(cls))
64-
return names
65-
66-
6727
# The comma characters (ASCII/Arabic/fullwidth, #265). Shared here so
6828
# tokenize (separators/segmentation) and extract (close-quote
6929
# boundaries) cannot drift apart.
@@ -253,3 +213,57 @@ class ParseState:
253213
#: lean) rather than asking again.
254214
one_case: bool | None = None
255215
ambiguities: tuple[PendingAmbiguity, ...] = ()
216+
217+
218+
def _copyable_fields(cls: type) -> tuple[str, ...]:
219+
"""The fields `copy_with` carries for `cls`, or TypeError where a
220+
field copy would not build what `dataclasses.replace` builds: the
221+
class must be decorated itself (not inherit the decoration) and keep
222+
the generated `__init__`, with no `__post_init__` and no
223+
`init=False` field."""
224+
params = cls.__dict__.get("__dataclass_params__")
225+
init = cls.__dict__.get("__init__")
226+
# dataclasses compiles the __init__ it generates from a string; one
227+
# written in the class body carries its source file instead.
228+
generated = init is not None and init.__code__.co_filename == "<string>"
229+
if (params is None or not generated
230+
or hasattr(cls, "__post_init__")
231+
or not all(f.init for f in dataclasses.fields(cls))):
232+
raise TypeError(
233+
f"copy_with cannot copy {cls.__name__}: it needs a dataclass "
234+
"whose generated __init__ only assigns its fields")
235+
return tuple(f.name for f in dataclasses.fields(cls))
236+
237+
238+
#: Built once, at import: a pipeline class that stops qualifying fails
239+
#: here rather than in a parse, and `copy_with` copies nothing else.
240+
_COPY_FIELDS: Mapping[type, tuple[str, ...]] = MappingProxyType({
241+
cls: _copyable_fields(cls)
242+
for cls in (WorkToken, PendingAmbiguity, ParseState)})
243+
244+
_T = TypeVar("_T")
245+
246+
if TYPE_CHECKING:
247+
# mypy's dataclass plugin checks replace's keywords against the
248+
# class, which a `**changes: object` signature would not.
249+
from dataclasses import replace as copy_with
250+
else:
251+
def copy_with(obj: _T, /, **changes: object) -> _T:
252+
"""`dataclasses.replace` for the pipeline's own dataclasses,
253+
without its per-call cost: a direct field copy, which for the
254+
classes in `_COPY_FIELDS` builds the same object
255+
(decisions.md#parse-cost has the measurement)."""
256+
cls = type(obj)
257+
names = _COPY_FIELDS.get(cls)
258+
if names is None:
259+
raise TypeError(
260+
f"copy_with copies only the pipeline's own dataclasses, "
261+
f"not {cls.__name__}")
262+
new = object.__new__(cls)
263+
for name in names:
264+
value = changes.pop(name) if name in changes else getattr(obj, name)
265+
object.__setattr__(new, name, value)
266+
if changes:
267+
raise TypeError(f"{cls.__name__} has no field named "
268+
f"{', '.join(sorted(changes))}")
269+
return new

‎nameparser/_pipeline/_tokenize.py‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -191,6 +191,6 @@ def _containing(offset: int) -> tuple[int, ...]:
191191
else copy_with(a, indices=_containing(a.origin))
192192
for a in ambiguities)
193193
return copy_with(state, tokens=tuple(tokens),
194-
comma_offsets=tuple(sorted(commas)),
195-
interpunct_offsets=tuple(sorted(interpuncts)),
196-
ambiguities=ambiguities)
194+
comma_offsets=tuple(sorted(commas)),
195+
interpunct_offsets=tuple(sorted(interpuncts)),
196+
ambiguities=ambiguities)

0 commit comments

Comments
 (0)