Skip to content

The publication audit reads each reachable blob through its encoding, and reports one that does not decode as unread - #340

Merged
HackingGate merged 1 commit into
mainfrom
audit-reads-blobs-through-their-encoding
Oct 9, 2026
Merged

HackingGate merged 1 commit into
mainfrom
audit-reads-blobs-through-their-encoding

Conversation

@HackingGate

Copy link
Copy Markdown
Owner

audit --for-publication decoded every reachable blob with String::from_utf8_lossy. A UTF-16 file that named a private owner was therefore searched as replacement characters, so the name matched nothing, and the blob still counted among the surfaces read. When the forge was reachable, the run ended with "every one of them is clean" and exit 0, just before the private-to-public flip, which cannot be undone.

Blobs now go through guard::scope::decode, the decoder the guards already use. A UTF-16 file with a byte-order mark is decoded and searched as the text it holds. A binary blob is still searched lossily, as before, because a name written into a binary is still published by the flip. A blob that does not decode (for example Latin-1) is listed among the surfaces the run could not read and is not counted as read, and the run exits 2. It is still searched for what its bytes spell in ASCII. If anything is found, exit 1 takes precedence and both are printed.

Tests added in tests/audit_publication_cli.rs:

  • a_name_in_a_utf16_file_is_found
  • a_blob_that_does_not_decode_is_reported_unread
  • a_name_in_a_blob_that_does_not_decode_is_still_found
  • a_name_in_a_binary_blob_is_still_found

docs/REFERENCE.md now documents how the audit decodes blobs.

https://claude.ai/code/session_01TG5hbR4Re6gcEZTpBCBymc

… and reports one that does not decode as unread

audit --for-publication read every reachable blob with
String::from_utf8_lossy. A UTF-16 file became replacement characters
with NULs between them, so a private name inside it matched nothing,
and the blob was still counted among the surfaces read. With a forge
that answered, the run ended on "every surface a flip would republish
was read, and every one of them is clean" and exit 0, over a name the
same text in UTF-8 is refused for, immediately before the flip that
cannot be taken back. A blob that was not text at all, such as a
Latin-1 file, was counted as read the same way.

The blobs now go through the decoder the guards already use, which
honours a byte-order mark. Text is searched as the text it holds. A
blob that is neither text nor binary is listed among the surfaces this
run could not read, is left out of the count of surfaces read, and
makes the exit 2. It is still searched lossily, as before, because
that keeps every ASCII byte: a private name written in ASCII inside a
Latin-1 file is still a finding, and the finding outranks the unread
surface, so that run exits 1 with both printed. A binary blob is still
searched as its bytes, as it always was here: the guards skip one for
want of lines to point at, but a name written into a binary is
published by the flip all the same.

The CLI tests cover a private name in a UTF-16 file, a Latin-1 blob
reaching exit 2 without being counted as read, a private name in a
Latin-1 blob found and the blob still listed as unread, and a name in a
binary blob still being found. The reference documents how blobs are
decoded.

Claude-Session: https://claude.ai/code/session_01TG5hbR4Re6gcEZTpBCBymc
@coderabbitai

coderabbitai Bot commented Oct 9, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 38 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 55be756f-f1a0-4d34-8812-65c0c03b4127

📥 Commits

Reviewing files that changed from the base of the PR and between c192382 and 89531dd.


📒 Files selected for processing (3)
  • docs/REFERENCE.md
  • src/audit.rs
  • tests/audit_publication_cli.rs


  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.32%. Comparing base (c192382) to head (89531dd).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #340   +/-   ##
=======================================
  Coverage   94.32%   94.32%           
=======================================
  Files          46       46           
  Lines       22820    22835   +15     
=======================================
+ Hits        21524    21539   +15     
  Misses       1296     1296           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@HackingGate
HackingGate merged commit 4b19d94 into main Oct 9, 2026
12 checks passed
@HackingGate
HackingGate deleted the audit-reads-blobs-through-their-encoding branch October 9, 2026 16:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants