Repository navigation
An ideographic zero reads with the letter after it, and a lookalike on a template comment is refused with its reason - #335
Conversation
…n a template comment is refused with its reason Part of #334, items 1, 2 and 3. Item 4 follows after #333. 1 and 2, ruled together. U+3007 IDEOGRAPHIC NUMBER ZERO is Han with a Latin O for its UTS #39 skeleton. lookalikes in src/guard/unicode.rs kept it in a Latin word only when every letter so far was drawn as Latin, so a kanji or kana run before U+3007 + `K` split the word after U+3007 and the disguise passed. It now reads such a letter with the letter AFTER it (new reading_scripts, resolved from the right): before a Latin letter it joins that Latin word wherever it sits, so kana + U+3007 + `K` + kana, U+4EF6 + U+3007 + `K` and U+4E8C + U+3007 + `K` are each refused; before a kanji it is the numeral zero and stays in the kanji run, so `API` + U+3007 + U+4EF6 and U+5168 U+3007 U+4EF6 pass; with no letter after it, it reads with the letter before it (`HELL` + U+3007 is refused). Residual: U+3007 + `GB` and the U+3007 U+3007 placeholder before a product name are refused, and `allow = ["U+3007"]` admits them; a Latin word ending in U+3007 run straight into a kanji passes. docs/REFERENCE.md says so in place of "stays one word". 3. comment_line_note says "The first non-blank line", and takes the rule's own finding test as a predicate. The commit-msg path of prevent_unusual_unicode now goes through the new unusual_unicode_in_message, which attaches the note when a lookalike is on the comment-opened candidate; the title seam does not, since every line of a title is a headline. The lookalike pass is split out of unusual_findings into lookalike_findings. Tests: an_ideographic_zero_inside_a_latin_word_is_a_lookalike (three disguise lines, `HELL` + U+3007, and three more kanji numerals passing), an_ideographic_zero_reads_with_the_letter_after_it, the_note_names_the_first_non_blank_line, a_lookalike_on_a_comment_opened_first_line_carries_the_note, and in tests/guard_cli.rs an_ideographic_zero_before_latin_is_refused_and_its_allowance_admits_it. No function or test is removed. Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 2 minutes. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codecov Report❌ Patch coverage is
❌ Your patch status has failed because the patch coverage (97.62%) is below the target coverage (100.00%). You can increase the patch coverage or adjust the target coverage. Additional details and impacted files@@ Coverage Diff @@
## main #335 +/- ##
==========================================
+ Coverage 94.26% 94.30% +0.04%
==========================================
Files 46 46
Lines 21985 22758 +773
==========================================
+ Hits 20724 21463 +739
- Misses 1261 1295 +34 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…nter follows Review of #335. Reading U+3007 with the letter after it let a Latin word ending in U+3007 pass before any East Asian letter (`TOD` + U+3007 + kana, `HELL` + U+3007 + Hangul, `a` + U+3007 + U+30A2), and before a kanji (`TOD` + U+3007 + U+4E00 U+89A7). The fallback to the letter before also failed on runs (`HELL` + U+3007 U+3007) and before a Cyrillic or Greek letter (`g` + U+3007 + U+043E). main refused all of these. reading_scripts in src/guard/unicode.rs is rewritten to decide each run of letters drawn across the East Asian boundary whole, left to right, against the first letter after the run and the last one before it: (1) before a letter of the script it is drawn as, it joins that word, so U+3007 + `K` after kana, a counter, or a kanji numeral (U+4E8C + U+3007 + `K`) is still refused; (2) before a counter kanji in the new closed COUNTERS list it stays Han, so `API` + U+3007 + U+4EF6 passes; (3) otherwise it reads with the letter before it when it is drawn as that letter's script, and stays Han after kanji, kana or nothing. New helpers: own_script, drawn_across, all_drawn_as. script_of_word no longer counts a letter drawn across the boundary toward the word's majority, so `Z` + U+3007 U+3007 names both zeros instead of naming `Z` as a Latin letter in a Han word. Residual, documented in docs/REFERENCE.md: a Latin word ending in U+3007 right before a counter (`TOD` + U+3007 + U+4EF6) passes; U+3007 + `GB` and the U+3007 U+3007 placeholder are refused, and `allow = ["U+3007"]` admits them. Tests: an_ideographic_zero_inside_a_latin_word_is_a_lookalike extended with every case above; an_ideographic_zero_before_a_counter_is_a_numeral and an_ideographic_zero_before_latin_is_refused_as_written replace an_ideographic_zero_reads_with_the_letter_after_it, which is removed. Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
…anji, and votes as Latin Second review of #335. COUNTERS held the commonest first kanji of Japanese compounds (U+65E5, U+672C, U+884C, U+4EBA, U+5206, U+540D, U+756A, U+5EA6, U+6642, U+5E74, U+5341), so reading U+3007 as a numeral before one let `DEM` + U+3007 + U+65E5 U+672C U+8A9E and ten more disguises through that main refused. The COUNTERS const and its rule are removed. reading_scripts now has two rules: before a letter of the script it is drawn as, a run of U+3007 joins that word; otherwise it is read with the letter before it when drawn as that letter's script, and stays Han after kanji, kana, Hangul, a digit or nothing. `API` + U+3007 + U+4EF6 and `PR` + U+3007 + U+56DE are now refused, documented false positives that `allow = ["U+3007"]` admits. script_of_word stopped counting U+3007 at all in the previous commit, so a Cyrillic or Greek half could outvote it: U+0421 U+041E + U+3007 + `L` passed. lookalikes now hands lookalikes_in_word and script_of_word the script each letter is read as, and each letter votes for that script, so U+3007 read as Latin is a Latin vote and the tie-breaker names the Cyrillic letters. docs/REFERENCE.md states the two rules, that the search skips marks and letters no script owns, the Cyrillic and Greek fallback, and the false positives; the residual sentence about a counter is removed. Tests: an_ideographic_zero_before_a_counter_is_a_numeral is removed. Added an_ideographic_zero_after_a_latin_word_is_read_with_it_before_any_compound and an_ideographic_zero_read_as_latin_votes_latin; an_ideographic_zero_before_latin_is_refused_as_written now refuses `API` + U+3007 + U+4EF6 and `PR` + U+3007 + U+56DE. Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
…s vote, and the docs say so Third review. docs/REFERENCE.md and the doc comment on lookalikes said a U+3007 joining a Latin word is refused, and a Cyrillic or Greek letter beside it named, without the condition that Latin must win or tie the word's script vote. In a Cyrillic- or Greek-majority word, such as a Cyrillic-spelled `exec` + U+3007 + `t`, U+3007 is not named: the same majority-vote limit that lets an all-Cyrillic lookalike of `execot` pass. Both now state it. Rule 2's list of letters after which U+3007 stays Han now includes a Cyrillic or Greek letter, in both places. Tests: the comment in an_ideographic_zero_inside_a_latin_word_is_a_lookalike described a third rule before a counter, which the previous commit removed; it now describes rule 2, and rule 1's comment no longer names a counter kanji. No code changes. Claude-Session: https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4
Part of #334: items 1, 2 and 3. Item 4 (harness hook titles) follows after #333 merges, because #333 rewrites the same text-subject plumbing in
src/text.rs.Special characters are spelled as codepoints here.
Ruling on items 1 and 2 (final, commit 04e8d62)
This went through two reviews:
COUNTERSlist. The list held the commonest first kanji of Japanese compounds, soDEM+ U+3007 + U+65E5 U+672C U+8A9E,REP+ U+3007 + U+540D U+524D and nine more disguises passed. main refused them.COUNTERSentirely.reading_scripts(src/guard/unicode.rs) decides each run of U+3007 as one unit. It looks at the nearest letter after the run and the nearest letter before it in the same word. The search skips marks and letters no script owns, such as U+30FC.Kis the disguise itself.lookalikesalso passes each letter's reading script tolookalikes_in_wordandscript_of_word, and each letter votes for its reading script. A U+3007 read as Latin is therefore a Latin vote. Without that, a Cyrillic or Greek half of the word could outvote it.Refused (each names the letter drawn as Latin):
g+ U+3007 +od,C+ U+3007 +DE, U+3007 +K,HELL+ U+3007K+ kana, U+4EF6 + U+3007 +K, U+4E8C + U+3007 +KTOD/DEM/REP+ U+3007 + a kanji compound, kana, or a katakana wordHELL+ U+3007 + Hangul,a+ U+3007 + U+30A2HELL+ U+3007 U+3007, andZ+ U+3007 U+3007 (names both zeros)g+ U+3007 + U+043E,A+ U+3007 + U+03B1 U+0441L, U+0430 U+043E + U+3007 +d, U+03BF + U+3007 +d+ U+03B1 (all judged as Latin words)TOD+ U+3007 + U+4EF6Passing:
2026+ U+5E74 U+4E00 U+3007 U+6708, U+5168 U+3007 U+4EF6, U+4E8C U+3007 U+3007 U+30071+ U+3007 + U+4EF6, U+7B2C U+3007 U+7248, U+3007 U+6642 U+3007 U+5206, U+5E73 U+6210 U+3007 U+5E74v1+ U+3007 U+3007False positives, refused and documented: U+3007 +
GB, the U+3007 U+3007 placeholder before a Latin word, and a Latin word + U+3007 + a counter, such asAPI+ U+3007 + U+4EF6 andPR+ U+3007 + U+56DE. The report namesU+3007, andallow = ["U+3007"]admits it. A U+3007 numeral glued to a Latin word is rare in real text, so this cost is accepted to refuse the disguise everywhere it sits.docs/REFERENCE.mdstates these two rules, the skipped characters, the Cyrillic/Greek fallback, and the false positives with their escape.Item 3
comment_line_notenow says "The first non-blank line opens with the comment character". It takes the calling rule's own finding test as a predicate, so it is no longer tied tonon_ascii_in_subject.prevent_unusual_unicodeat the git hooks goes through the newunusual_unicode_in_message. That function attaches the note when a lookalike sits on the comment-opened candidate.unusual_unicode_in) never attaches the note, because every line of a title is a headline there.unusual_findingsintolookalike_findingsso both callers can use it.Done when
an_ideographic_zero_inside_a_latin_word_is_a_lookalikerefuses the three item-1 lines, each exactly once and namingIDEOGRAPHIC NUMBER ZERO. In the same test, the kanji numeral U+4E8C U+3007 U+4E8C U+516D U+5E74 andAPIrun into a kanji compound still pass.an_ideographic_zero_before_latin_is_refused_as_writtenrefuses all three issue cases:API+ U+3007 + U+4EF6, U+3007 U+3007 +API+ kana, and U+3007 +GB.tests/guard_cli.rs::an_ideographic_zero_before_latin_is_refused_and_its_allowance_admits_it, nottests/text_cli.rs. The text seam is prose-only and cannot carry a title-kind subject without changes tosrc/text.rs(The shim judges what a forge holds of the text: the merge message it composes, the lines an edit adds, and a title set through api or glab mr merge #333's territory). The new test uses the commit-msg subject line, which gets the same lookalike pass. U+3007 +GB+ kana is refused, the report namesU+3007, and the same run withallow = ["U+3007"]passes.the_note_names_the_first_non_blank_linecovers"\n\n# " + Japanese + "\nFix the parser\n".a_lookalike_on_a_comment_opened_first_line_carries_the_notecovers three cases: a lookalike on the comment-opened first line carries the note, an ordinary subject and a subject under an ASCII comment do not, and an invisible on the comment line does not.a_non_ascii_comment_opening_the_message_is_judged_as_a_subjectstill pins the note's absence on ordinary subjects forascii-only-commit-subject.Removed during review: the tests
an_ideographic_zero_reads_with_the_letter_after_itandan_ideographic_zero_before_a_counter_is_a_numeral, and theCOUNTERSconst added in aa6b9ec. No function from main is removed.https://claude.ai/code/session_01QJkWXa2HxJ6q9TMNAGSNv4