Skip to content

An ideographic zero reads with the letter after it, and a lookalike on a template comment is refused with its reason - #335

Merged
HackingGate merged 4 commits into
mainfrom
unicode-followups-334
Oct 9, 2026
Merged

HackingGate merged 4 commits into
mainfrom
unicode-followups-334

Conversation

@HackingGate

@HackingGate HackingGate commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

Part of #334: items 1, 2 and 3. Item 4 (harness hook titles) follows after #333 merges, because #333 rewrites the same text-subject plumbing in src/text.rs.

Special characters are spelled as codepoints here.

Ruling on items 1 and 2 (final, commit 04e8d62)

This went through two reviews:

  • First commit: read U+3007 with the letter after it. That let a Latin word ending in U+3007 pass before any East Asian letter.
  • aa6b9ec: added a closed COUNTERS list. The list held the commonest first kanji of Japanese compounds, so DEM + U+3007 + U+65E5 U+672C U+8A9E, REP + U+3007 + U+540D U+524D and nine more disguises passed. main refused them.
  • 04e8d62: removes COUNTERS entirely.

reading_scripts (src/guard/unicode.rs) decides each run of U+3007 as one unit. It looks at the nearest letter after the run and the nearest letter before it in the same word. The search skips marks and letters no script owns, such as U+30FC.

  1. The next letter is in the script U+3007 is drawn as (Latin): the run joins that word. This applies even after a kanji numeral, because U+3007 + K is the disguise itself.
  2. Otherwise: the run is read with the letter before it, if U+3007 is drawn as that letter's script. After a kanji, a kana, a Hangul letter, a digit or nothing, it stays Han.

lookalikes also passes each letter's reading script to lookalikes_in_word and script_of_word, and each letter votes for its reading script. A U+3007 read as Latin is therefore a Latin vote. Without that, a Cyrillic or Greek half of the word could outvote it.

Refused (each names the letter drawn as Latin):

  • g + U+3007 + od, C + U+3007 + DE, U+3007 + K, HELL + U+3007
  • the three item-1 lines: kana + U+3007 + K + kana, U+4EF6 + U+3007 + K, U+4E8C + U+3007 + K
  • TOD / DEM / REP + U+3007 + a kanji compound, kana, or a katakana word
  • HELL + U+3007 + Hangul, a + U+3007 + U+30A2
  • HELL + U+3007 U+3007, and Z + U+3007 U+3007 (names both zeros)
  • g + U+3007 + U+043E, A + U+3007 + U+03B1 U+0441
  • U+0421 U+041E + U+3007 + L, U+0430 U+043E + U+3007 + d, U+03BF + U+3007 + d + U+03B1 (all judged as Latin words)
  • all eleven compound cases from the second review, and TOD + U+3007 + U+4EF6

Passing:

False positives, refused and documented: U+3007 + GB, the U+3007 U+3007 placeholder before a Latin word, and a Latin word + U+3007 + a counter, such as API + U+3007 + U+4EF6 and PR + U+3007 + U+56DE. The report names U+3007, and allow = ["U+3007"] admits it. A U+3007 numeral glued to a Latin word is rare in real text, so this cost is accepted to refuse the disguise everywhere it sits.

docs/REFERENCE.md states these two rules, the skipped characters, the Cyrillic/Greek fallback, and the false positives with their escape.

Item 3

  • comment_line_note now says "The first non-blank line opens with the comment character". It takes the calling rule's own finding test as a predicate, so it is no longer tied to non_ascii_in_subject.
  • prevent_unusual_unicode at the git hooks goes through the new unusual_unicode_in_message. That function attaches the note when a lookalike sits on the comment-opened candidate.
  • An invisible character on that line does not attach the note, because an invisible is refused on every line anyway.
  • The title seam (unusual_unicode_in) never attaches the note, because every line of a title is a headline there.
  • The lookalike pass is split out of unusual_findings into lookalike_findings so both callers can use it.

Done when

  • 1. an_ideographic_zero_inside_a_latin_word_is_a_lookalike refuses the three item-1 lines, each exactly once and naming IDEOGRAPHIC NUMBER ZERO. In the same test, the kanji numeral U+4E8C U+3007 U+4E8C U+516D U+5E74 and API run into a kanji compound still pass.
  • 2. Ruled refused, accepted and documented. an_ideographic_zero_before_latin_is_refused_as_written refuses all three issue cases: API + U+3007 + U+4EF6, U+3007 U+3007 + API + kana, and U+3007 + GB.
  • 3.
    • the_note_names_the_first_non_blank_line covers "\n\n# " + Japanese + "\nFix the parser\n".
    • a_lookalike_on_a_comment_opened_first_line_carries_the_note covers three cases: a lookalike on the comment-opened first line carries the note, an ordinary subject and a subject under an ASCII comment do not, and an invisible on the comment line does not.
    • The existing a_non_ascii_comment_opening_the_message_is_judged_as_a_subject still pins the note's absence on ordinary subjects for ascii-only-commit-subject.
  • 4. Not in this PR; follows The shim judges what a forge holds of the text: the merge message it composes, the lines an edit adds, and a title set through api or glab mr merge #333.

Removed during review: the tests an_ideographic_zero_reads_with_the_letter_after_it and an_ideographic_zero_before_a_counter_is_a_numeral, and the COUNTERS const added in aa6b9ec. No function from main is removed.

https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4

…n a template comment is refused with its reason

Part of #334, items 1, 2 and 3. Item 4 follows after #333.

1 and 2, ruled together. U+3007 IDEOGRAPHIC NUMBER ZERO is Han with a
Latin O for its UTS #39 skeleton. lookalikes in src/guard/unicode.rs
kept it in a Latin word only when every letter so far was drawn as Latin,
so a kanji or kana run before U+3007 + `K` split the word after U+3007
and the disguise passed. It now reads such a letter with the letter
AFTER it (new reading_scripts, resolved from the right): before a Latin
letter it joins that Latin word wherever it sits, so kana + U+3007 + `K`
+ kana, U+4EF6 + U+3007 + `K` and U+4E8C + U+3007 + `K` are each refused;
before a kanji it is the numeral zero and stays in the kanji run, so
`API` + U+3007 + U+4EF6 and U+5168 U+3007 U+4EF6 pass; with no letter
after it, it reads with the letter before it (`HELL` + U+3007 is
refused). Residual: U+3007 + `GB` and the U+3007 U+3007 placeholder
before a product name are refused, and `allow = ["U+3007"]` admits them;
a Latin word ending in U+3007 run straight into a kanji passes.
docs/REFERENCE.md says so in place of "stays one word".

3. comment_line_note says "The first non-blank line", and takes the
rule's own finding test as a predicate. The commit-msg path of
prevent_unusual_unicode now goes through the new
unusual_unicode_in_message, which attaches the note when a lookalike is
on the comment-opened candidate; the title seam does not, since every
line of a title is a headline. The lookalike pass is split out of
unusual_findings into lookalike_findings.

Tests: an_ideographic_zero_inside_a_latin_word_is_a_lookalike (three
disguise lines, `HELL` + U+3007, and three more kanji numerals passing),
an_ideographic_zero_reads_with_the_letter_after_it,
the_note_names_the_first_non_blank_line,
a_lookalike_on_a_comment_opened_first_line_carries_the_note, and in
tests/guard_cli.rs
an_ideographic_zero_before_latin_is_refused_and_its_allowance_admits_it.
No function or test is removed.

Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 2 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 000791d2-c849-4fe2-9825-9ba732ae6318

📥 Commits

Reviewing files that changed from the base of the PR and between ffd19e1 and d5a9e9e.


📒 Files selected for processing (4)
  • docs/REFERENCE.md
  • src/guard/message.rs
  • src/guard/unicode.rs
  • tests/guard_cli.rs


  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov-commenter

codecov-commenter commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.62846% with 6 lines in your changes missing coverage. Please review.
✅ Project coverage is 94.30%. Comparing base (ffd19e1) to head (d5a9e9e).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/guard/message.rs 96.15% 3 Missing ⚠️
src/guard/unicode.rs 98.28% 3 Missing ⚠️

❌ Your patch status has failed because the patch coverage (97.62%) is below the target coverage (100.00%). You can increase the patch coverage or adjust the target coverage.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #335      +/-   ##
==========================================
+ Coverage   94.26%   94.30%   +0.04%     
==========================================
  Files          46       46              
  Lines       21985    22758     +773     
==========================================
+ Hits        20724    21463     +739     
- Misses       1261     1295      +34     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

…nter follows

Review of #335. Reading U+3007 with the letter after it let a Latin word
ending in U+3007 pass before any East Asian letter (`TOD` + U+3007 + kana,
`HELL` + U+3007 + Hangul, `a` + U+3007 + U+30A2), and before a kanji
(`TOD` + U+3007 + U+4E00 U+89A7). The fallback to the letter before also
failed on runs (`HELL` + U+3007 U+3007) and before a Cyrillic or Greek
letter (`g` + U+3007 + U+043E). main refused all of these.

reading_scripts in src/guard/unicode.rs is rewritten to decide each run
of letters drawn across the East Asian boundary whole, left to right,
against the first letter after the run and the last one before it:
(1) before a letter of the script it is drawn as, it joins that word,
so U+3007 + `K` after kana, a counter, or a kanji numeral (U+4E8C +
U+3007 + `K`) is still refused; (2) before a counter kanji in the new
closed COUNTERS list it stays Han, so `API` + U+3007 + U+4EF6 passes;
(3) otherwise it reads with the letter before it when it is drawn as that
letter's script, and stays Han after kanji, kana or nothing. New helpers:
own_script, drawn_across, all_drawn_as.

script_of_word no longer counts a letter drawn across the boundary
toward the word's majority, so `Z` + U+3007 U+3007 names both zeros
instead of naming `Z` as a Latin letter in a Han word.

Residual, documented in docs/REFERENCE.md: a Latin word ending in U+3007
right before a counter (`TOD` + U+3007 + U+4EF6) passes; U+3007 + `GB`
and the U+3007 U+3007 placeholder are refused, and `allow = ["U+3007"]`
admits them.

Tests: an_ideographic_zero_inside_a_latin_word_is_a_lookalike extended
with every case above; an_ideographic_zero_before_a_counter_is_a_numeral
and an_ideographic_zero_before_latin_is_refused_as_written replace
an_ideographic_zero_reads_with_the_letter_after_it, which is removed.

Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
…anji, and votes as Latin

Second review of #335. COUNTERS held the commonest first kanji of
Japanese compounds (U+65E5, U+672C, U+884C, U+4EBA, U+5206, U+540D,
U+756A, U+5EA6, U+6642, U+5E74, U+5341), so reading U+3007 as a numeral
before one let `DEM` + U+3007 + U+65E5 U+672C U+8A9E and ten more
disguises through that main refused. The COUNTERS const and its rule
are removed. reading_scripts now has two rules: before a letter of the
script it is drawn as, a run of U+3007 joins that word; otherwise it is
read with the letter before it when drawn as that letter's script, and
stays Han after kanji, kana, Hangul, a digit or nothing. `API` + U+3007
+ U+4EF6 and `PR` + U+3007 + U+56DE are now refused, documented false
positives that `allow = ["U+3007"]` admits.

script_of_word stopped counting U+3007 at all in the previous commit,
so a Cyrillic or Greek half could outvote it: U+0421 U+041E + U+3007 +
`L` passed. lookalikes now hands lookalikes_in_word and script_of_word
the script each letter is read as, and each letter votes for that
script, so U+3007 read as Latin is a Latin vote and the tie-breaker
names the Cyrillic letters.

docs/REFERENCE.md states the two rules, that the search skips marks and
letters no script owns, the Cyrillic and Greek fallback, and the false
positives; the residual sentence about a counter is removed.

Tests: an_ideographic_zero_before_a_counter_is_a_numeral is removed.
Added an_ideographic_zero_after_a_latin_word_is_read_with_it_before_any_compound
and an_ideographic_zero_read_as_latin_votes_latin;
an_ideographic_zero_before_latin_is_refused_as_written now refuses
`API` + U+3007 + U+4EF6 and `PR` + U+3007 + U+56DE.

Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
…s vote, and the docs say so

Third review. docs/REFERENCE.md and the doc comment on lookalikes said a
U+3007 joining a Latin word is refused, and a Cyrillic or Greek letter
beside it named, without the condition that Latin must win or tie the
word's script vote. In a Cyrillic- or Greek-majority word, such as a
Cyrillic-spelled `exec` + U+3007 + `t`, U+3007 is not named: the same
majority-vote limit that lets an all-Cyrillic lookalike of `execot`
pass. Both now state it.

Rule 2's list of letters after which U+3007 stays Han now includes a
Cyrillic or Greek letter, in both places.

Tests: the comment in an_ideographic_zero_inside_a_latin_word_is_a_lookalike
described a third rule before a counter, which the previous commit
removed; it now describes rule 2, and rule 1's comment no longer names
a counter kanji. No code changes.

Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants