You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/design/decisions.md
+1Lines changed: 1 addition & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1433,6 +1433,7 @@ Every number below is a py3.11 measurement of 2026-08-31, recomputable with `uv
1433
1433
ONE REPAIR WAS TRIED AND REVERTED: collapsing `post_rules`' three separate role-index scans into one pass. It saves **6 calls per parse** on 3.11 (the first draft said one, which is not even reachable — `_idx` is a one-line list comprehension, so each removed call frees two frames). Reverted because it was measured against a noisy timing harness that showed nothing; at 6 calls it is 9% of the cycle's growth and worth reconsidering against the band, which can now see it.
1434
1434
RAISING OR LOWERING A ROW IS A DECISION, not a maintenance chore. Append here with the interpreter and the harness invocation, as above.
1435
1435
1436
+
- 2026-09-26 #546 -- the stages copy their state with `_state.copy_with` instead of `dataclasses.replace`, and every row moves DOWN. `dataclasses.replace` walks `fields()` and calls `__init__` on each copy, three frames on 3.11 and 3.12 and four from 3.13 where `copy_with` is one. One parse of the reference name makes 18 copies: six of the state, one each from tokenize, segment, classify, group, assign and post_rules, and twelve of single tokens, six each from classify and assign (`extract_delimited` returns the state unchanged when there is no delimiter, and `script_segment` returns early on ASCII input). That is the whole of the drop, 18 x 2 = 36 frames on 3.11 and 3.12 and 18 x 3 = 54 from 3.13 (recompute: wrap `copy_with` with a counter in the eight stage modules and parse the reference name). A field copy builds what `replace` builds for a dataclass that is decorated itself rather than inheriting the decoration, keeps the generated `__init__`, and has no `__post_init__` and no `init=False` field. `_copyable_fields` checks exactly those four, and `_COPY_FIELDS` runs it over `WorkToken`, `PendingAmbiguity` and `ParseState` at import, so a class that stops qualifying fails there and `copy_with` copies nothing else. `test_the_guard_refuses_a_class_a_field_copy_would_get_wrong` records, for each refused shape, what `replace` builds and what an unguarded copy would build instead. To mypy, `copy_with` is `from dataclasses import replace as copy_with`, which keeps the dataclass plugin's keyword and type checks at every call site; an assignment (`copy_with = dataclasses.replace`) would not, since the plugin keys on the callee's full name (measured: a misspelled field and a wrong-typed value both pass through the assignment and both fail through the import). Measured 2026-09-26 with `uv run python tools/perf/call_count.py --against e0f1a2f`, each row on its own interpreter, parse/facade: 3.11 406/443 → 370/407, 3.12 384/421 → 348/385, 3.13, 3.14 and 3.15 402/439 → 348/385. The rows drop by 40 and 58 rather than 36 and 54 because e0f1a2f already read 4 under every row, inside the band, and the new rows are set to what the harness reads now. `_LINK_BASELINE`'s 64-link clause reads 2587 → 2301 on 3.11. By stage (`--stages`, py3.11, ms per 1000 parses of the reference name): group 19.6 → 17.3, classify 13.4 → 9.8, assign 11.1 → 8.0, tokenize 8.1 → 6.6, post_rules 7.9 → 6.5, segment 2.6 → 1.7. BEHAVIOR IDENTICAL: the differential gate's report at all five baselines matches e0f1a2f's line for line apart from the path header.
Copy file name to clipboardExpand all lines: docs/release_log.rst
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -36,6 +36,8 @@ Release Log
36
36
37
37
- **Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals.** ``HumanName("john smith x.y.z.").capitalize()`` gives ``John Smith X.Y.Z.`` where every release gave ``John Smith X.y.z.``, the dotted word being a suffix now (the ``unlisted_dotted_suffixes`` change above) and repaired as a listed acronym is; and ``john smith vi`` gives ``John Smith VI`` where every release gave ``John Smith Vi``, with ``vii``, ``viii`` and ``ix`` alike. Both are keyed on the suffix role: ``Jack X.Y.Z.``, which keeps its surname, still repairs as a name word (``Jack X.y.z.`` under ``force=True``), and ``john smith xi`` still gives ``John Smith Xi``, the parser reading ``xi`` as the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (``john smith B.Tech.`` gives ``John Smith B.Tech.``), and reads all capitals under ``force=True`` (``John Smith B.TECH.``, where every release gave ``John Smith B.tech.``), which a ``capitalization_exceptions`` mask such as ``{"btech": "BTech"}`` undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 under ``force=True``. The recipe is the ``R4`` entry's 2026-09-23 MEASURED bullet in ``docs/design/decisions.md`` (#459)
38
38
39
+
- **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546)
40
+
39
41
**Additions**
40
42
41
43
- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)
0 commit comments