Skip to content

Commit 60b040d

Browse files
derek73claude
andcommitted
Correct five claims the exclusion prose got wrong
A second review round over the corrections themselves. All five are prose; no behavior changes, and the gate is unmoved at 107/0. "Silencing all 47 would turn 47 legitimate classifications into UNEXPLAINED" was the biggest, and the first review round already raised it -- I rewrote the surrounding block and left the number. Measured by adding exactly that exclusion and running the gate: three names, all CJK, and the other 44 do not diff at all. The real cost is those three plus 44 shapes pre-silenced, which counting names both overstates and hides. 'Senator "Rick" Edmonds' was cited as a leading pair the entry fails to protect. It is medial and IS captured -- 'Senator' supplies the left flank. The corpus holds a separate '"Rick" Edmonds', which is the genuinely unguarded one; name that instead, and note the trap. "The 21 it gains are almost entirely that family" read as "credentials" against its own antecedent. 20 of 21 put the delimiter last, but only 8 are credentials and 6 are trailing NICKNAMES -- the reading this entry exists to protect. So the medial cut is about position, not about credentials: at the trailing position a nickname and a credential wear the same shape, and protection and over-reach arrive together. Whether `fields` already makes the wider shape safe is now flagged as an open question rather than answered by assertion. "Only 9 of those 47 are parenthesised credentials" is 8. Nine is the count a name_regex-carrying rule reaches -- two measurements conflated. 'Xyz. (Bud) Smith' was the example for a middle initial; it parses 'Xyz.' as a title. 'Cherice J. (Johnson) Williams' is the real one. Also: the _EXCLUSION_EFFECT preamble claimed in the present tense that dropping the Ph. D. anchor keeps the suite green -- true before this pin existed, self-refuting in the tree that contains it. And the fields-deletion note had the mechanism backwards: absorbed_by GROWS from () to ('fix(suffix-routing)',), it does not see a reading go unclaimed. Both now say which era they describe. Two limits named rather than left implicit: absorbed_by records only the first rule matching each subset, so a fully shadowed rule does not move it; and an entry carrying `fields` is asked about only the subsets those cover, which is where the 387 comes from. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent 91a5f03 commit 60b040d

4 files changed

Lines changed: 72 additions & 32 deletions

File tree

‎AGENTS.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -125,8 +125,9 @@ uv run sphinx-build -b html docs dist/docs
125125
# tests/v2/test_ledger_guards.py records what each one silences --
126126
# which rules would claim it with exclusions off, and how much corpus it
127127
# captures -- in _EXCLUSION_EFFECT. That IS an enrollment: an entry
128-
# added or edited moves the record and must be re-recorded deliberately,
129-
# the same forcing function _CORPUS_CLAIMS applies to rules.
128+
# added, or its name_regex/fields/examples edited, moves the record and
129+
# must be re-recorded deliberately, the same forcing function
130+
# _CORPUS_CLAIMS applies to rules. (A `why`-only edit does not move it.)
130131
# 9. Open the next cycle's VERSION: bump VERSION in nameparser/_version.py to
131132
# the minor now being worked, and set PRE_RELEASE = 'dev'. The tree then says
132133
# what it is building rather than what it last shipped -- docs/conf.py reads

‎tests/v2/test_ledger_guards.py‎

Lines changed: 17 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -1149,10 +1149,19 @@ class _Excluded(NamedTuple):
11491149
#: harness cannot report because the exclusion (correctly) hides it.
11501150
#:
11511151
#: `captures` and `digest` cover the opposite direction, which nothing
1152-
#: else watches: an exclusion widened by name_regex silences real
1153-
#: classifications, and CI stays green while the release gate breaks.
1154-
#: Measured: dropping the Ph. D. entry's ASCII anchor keeps the whole
1155-
#: suite passing and takes a bare run to unexplained: 1.
1152+
#: else watched: an exclusion widened by name_regex silences real
1153+
#: classifications, and BEFORE this record CI stayed green while the
1154+
#: release gate broke. Measured then: dropping the Ph. D. entry's
1155+
#: Latin anchor left the whole suite passing and took a bare run to
1156+
#: unexplained: 1. Measured now: the same edit fails this pin -- the
1157+
#: record is keyed by name_regex, so editing one is exactly what it
1158+
#: catches -- while the gate still goes to unexplained: 1.
1159+
#:
1160+
#: One limit worth naming: `absorbed_by` records only the FIRST rule
1161+
#: matching each subset, so a rule appended behind one that already
1162+
#: answers those subsets is invisible here. A rule reaching a
1163+
#: protected READING is caught; a rule shadowed by an existing one on
1164+
#: every subset it claims is not.
11561165
_EXCLUSION_EFFECT: dict[str, _Excluded] = {
11571166
"(?i)^[\\u0000-\\u024f]*\\bph\\.\\s*d\\.\\s*$":
11581167
_Excluded(3, "5a12a8117651",
@@ -1255,7 +1264,10 @@ def test_a_fields_narrowing_actually_narrows_something() -> None:
12551264
a deleted key drops the entry from the loop entirely. That works
12561265
only while one entry carries `fields`. A second one would leave the
12571266
deletion green here -- caught instead by `absorbed_by` in the
1258-
recorded pin, which sees the reading go unclaimed.
1267+
recorded pin, which is then asked about every reading rather than
1268+
the three the key covers, and sees rules claim them. Measured:
1269+
deleting this entry's `fields` grows its `absorbed_by` from () to
1270+
('fix(suffix-routing)',).
12591271
12601272
Measured: deleting `fields = ["nickname", "middle"]` from the
12611273
ASCII-pairs entry passes every other check in this tree. The entry

‎tools/differential/README.md‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -325,7 +325,10 @@ before the rules are reached, so the answer is `None` however the rules
325325
change. The pin therefore asks with exclusions switched OFF and records
326326
which rules WOULD claim each protected reading. A rule widened to reach
327327
one changes that record and fails in CI, rather than being invisible
328-
until someone reasons about it. The same record carries the number and
328+
until someone reasons about it. It records the FIRST rule matching each
329+
subset, so a rule shadowed on every subset it claims by one already
330+
sitting ahead of it does not move the record; a rule reaching a
331+
protected reading does. The same record carries the number and
329332
digest of the corpus names an entry captures, which is the opposite
330333
drift: an over-wide exclusion silences real classifications, and that
331334
is loud at release but otherwise silent on a push.

‎tools/differential/expected_since_1.4.0.toml‎

Lines changed: 48 additions & 24 deletions
Original file line numberDiff line numberDiff line change
@@ -424,9 +424,17 @@ fields = ["title", "given", "middle", "suffix"]
424424
# Shapes that must never be explained. A [[change]] rule says "this
425425
# diff is intended, and here is what changed"; there was no rule
426426
# meaning "whatever happens here is a regression", so the two promises
427-
# below could only be written as prose -- and both were false from the
428-
# day they were written (the first shipped with the harness itself,
429-
# the second three days later) until #328 found them.
427+
# below could only be written as prose -- and both were false as
428+
# written from the day they were written (the first shipped with the
429+
# harness itself, the second three days later) until #328 found them.
430+
#
431+
# The two are not false alike, and the difference matters now that
432+
# this file scopes the second promise to one reading. Both said
433+
# "unclassified", unqualified, and both were claimed on SOME reading
434+
# from day one. But only the Ph. D. promise was claimed on the reading
435+
# it now protects; no rule has ever claimed the ASCII pairs' NICKNAME
436+
# reading, then or today. That entry is a tripwire under no present
437+
# tension, which is the point of setting it before the tension exists.
430438
#
431439
# classify() consults these before any rule, so a matching name
432440
# reports UNEXPLAINED however many rules would claim it.
@@ -435,7 +443,10 @@ fields = ["title", "given", "middle", "suffix"]
435443
# any corpus -- 'John "Jack" Kennedy' does not -- so the entry carries
436444
# its own test data, which tests/v2/test_ledger_guards.py runs against
437445
# every non-empty subset of the fields a rule may name: the seven
438-
# roles and the `_ambiguities` pseudo-field.
446+
# roles and the `_ambiguities` pseudo-field. An entry carrying
447+
# `fields` is asked only about the subsets those cover -- 3 for the
448+
# ASCII pairs, not 255 -- which is where the 387 in that file comes
449+
# from.
439450

440451
[[never]]
441452
why = "trailing 'Ph. D.' split-token healing is PARITY, not a 2.0 change: v1 healed the adjacent pair too, so a diff here is a regression"
@@ -461,8 +472,8 @@ why = "trailing 'Ph. D.' split-token healing is PARITY, not a 2.0 change: v1 hea
461472
# '田中さん, Ph. D.', whose diff is a real fix(cjk-comma-compound)
462473
# classification -- measured, it took the gate to unexplained: 1. The
463474
# cut is at U+0250, the same threshold _is_latin_only uses in
464-
# compare.py, so it covers Latin-1 and Latin Extended-A rather than
465-
# bare ASCII: the split-token healing is script-independent, and
475+
# compare.py, so it covers Latin-1 and Latin Extended-A and -B rather
476+
# than bare ASCII: the split-token healing is script-independent, and
466477
# 'José Smith, Ph. D.' is as much the protected shape as 'Jose Smith,
467478
# Ph. D.' is. An ASCII-only class protected one and not the other.
468479
name_regex = "(?i)^[\\u0000-\\u024f]*\\bph\\.\\s*d\\.\\s*$"
@@ -473,26 +484,39 @@ why = "feat(#273) recognizes TYPOGRAPHIC nickname delimiters; the ASCII pairs we
473484
# Narrowed twice, and both narrowings are load-bearing.
474485
#
475486
# By shape: a delimited run MEDIAL to the name -- flanked on the left
476-
# by a word character or a period (so a middle initial counts: 'Xyz.
477-
# (Bud) Smith') and on the right by a word character. A bare [("']
478-
# class reaches 47 corpus names, all of them claimed by a rule today,
479-
# so silencing them would turn 47 legitimate classifications into
480-
# UNEXPLAINED. Only 9 of those 47 are parenthesised credentials;
481-
# 11 match on a bare apostrophe -- "Brian O'connor",
482-
# "Harietta Keopuolani Nahi'ena'ena" -- which no delimiter regex can
483-
# tell from a quote.
487+
# by a word character or a period (so a middle initial counts:
488+
# 'Cherice J. (Johnson) Williams') and on the right by a word
489+
# character. A bare [("'] class reaches 47 corpus names. Measured by
490+
# adding exactly that exclusion and running the gate, silencing all 47
491+
# costs THREE classifications -- '山田 太郎 (マイケル・ジャクソン)',
492+
# '김, 민준씨 (Jimmy)', '김민준씨 (Jimmy)' -- because only those three
493+
# diff against 1.4.0 at all. The other 44 would be pre-silenced: no
494+
# diff today, and no way to report one tomorrow. Counting names
495+
# overstates the first cost and hides the second.
496+
#
497+
# Of the 47, 8 are parenthesised credentials and 11 match on a bare
498+
# apostrophe -- "Brian O'connor", "Harietta Keopuolani Nahi'ena'ena" --
499+
# which no delimiter regex can tell from a quote.
484500
#
485-
# The medial cut is what keeps the credentials out: 'Andrew Perkins
486-
# (JD)' and its kin put the parens LAST. Measured, widening the shape
487-
# to leading and trailing runs takes it from 13 corpus names to 34,
488-
# and the 21 it gains are almost entirely that family.
501+
# The medial cut is about POSITION, not about credentials. Widening to
502+
# leading and trailing runs takes the shape from 13 corpus names to
503+
# 34, and 20 of the 21 it gains put the delimited run last. Only 8 of
504+
# those are credentials; 6 are trailing NICKNAMES -- 'Franklin,
505+
# Benjamin (Ben)', 'Rev John A. Kenneth Doe III (Kenny)',
506+
# '김민준씨 (Jimmy)' -- which is the reading this entry exists to
507+
# protect. The rest are trailing maiden names and junk
508+
# ('Bridge (1.4)', 'John Jones (Google Docs)').
489509
#
490-
# So the promise is kept for medial pairs only. A leading or trailing
491-
# ASCII pair -- 'Senator "Rick" Edmonds' reads as protected and is
492-
# not, 'Jenny Baker (Johnson)' likewise -- is still unguarded prose,
493-
# for want of a regex that separates a trailing nickname from a
494-
# trailing credential. Narrowing a rule would be the way in, not
495-
# widening this entry.
510+
# So the trailing position is where protection and over-reach arrive
511+
# TOGETHER: a nickname and a credential wear the identical shape
512+
# there, and the delimiter alone cannot separate them. The promise is
513+
# kept for medial pairs only, and a leading or trailing ASCII pair --
514+
# '"Rick" Edmonds' reads as protected and is not, 'Jenny Baker
515+
# (Johnson)' likewise -- is still unguarded prose. (The corpus also
516+
# holds 'Senator "Rick" Edmonds', which IS medial and IS captured;
517+
# the two are easy to confuse.) Whether the `fields` narrowing already
518+
# makes the wider shape safe is an open question, unmeasured here;
519+
# narrowing a rule is the other way in.
496520
#
497521
# By role: ASCII parens mark nicknames, maiden names, suffixes and
498522
# credentials alike, and no regex tells them apart. The shape reaches

0 commit comments

Comments
 (0)