Repository navigation
v0.85 lane outcomes: the measured results, two refuted premises, and a blocker priced against the wrong operation - #1518
Merged
Merged
Conversation
…a blocker priced against the wrong operation Every outcome below was MEASURED before it was written, and each lane's refuting command was RUN first. Two lanes refute a premise this plan shipped with, which is the lane working rather than failing. THE BLOCKER THAT WAS PRICED AGAINST THE WRONG OPERATION. Three MUST lanes have recorded UNPROVEN or UNRUN across multiple releases "for want of a full build" against a multi-gigabyte disk budget. Building only the CLI crate yields a binary reporting the workspace version in under twenty seconds, well under a gigabyte of target, and it compiles a real module to a valid object. That figure prices a full TEST build — all targets plus dev-dependencies — while every one of those lanes needed a BINARY. The cost and the requirement were two different populations, and a cost measured for one build shape was used to decline a different shape, in the constraint this programme repeated most often. (The binary on PATH reports a long-superseded version, so refusing it was correct; the answer was to build one, not to use the wrong one or give up.) FALCON2 — the live figure HELD. Using the reporter's own command line, taken from the issue body: his optimized module skips exactly one of seventeen functions and the single skip is the frame-home export, named rather than left as a count; his fused module is clean. The residual is NOT a float decline — the class oracle passes every cell, and within it the f32 value-carrying branch on ARM now compiles while the f64 one still declines. The module was fetched on the maintainer's explicit approval, listed and traversal-checked before extraction, and treated as untrusted input: compiled, never executed. WIDEFILE8 — the root-cause step ran for the first time in five releases, and it explains the stall. Driving the recovery-statistics flag shows this repository's pressure fixture rescued by the lower rungs so the SUSPECT RUNG IS NEVER TRIED; the reporter's module reaches it once and rescues nothing. That is the first confirmed input on the suspect path. AND THE PROBE IS GETTING RARER: against the reading recorded in the rung's own test file at an earlier release, its try count has halved while a lower rung now rescues a function it did not before. If a later release rescues that last function earlier, the defect becomes UNOBSERVABLE while remaining live — an independent measured argument for the middle step of the fixed order. The retry STAYS OUT; the order was not jumped. MEMLAYER2 — the compile half is PROVEN and three carried claims are corrected. The refusal reproduces with two controls. "The refusal lives in the selector, not the CLI, so a CLI-level probe misses it" is FALSE: there are TWO refusals and a CLI pre-flight bails first, which is what a probe hits immediately. The bounding design assertion is intact but was cited at the wrong path — and a wrong path looks exactly like a deleted assertion, so reading the abbreviation literally would have rewritten the scope on a typo. FIRSTTOKEN4 — REFUTED as stated. There is no dead code branch: every red branch is unreached over the live population, which is what a green tree MEANS, so "never reached" is true of every working guard and discriminates nothing. Synthetic reachability shows every branch reachable by a legal artifact and behaving correctly, with two passing controls. What is actually dead is a VOCABULARY — the permissive gate admits two non-delivery dispositions the strict gate forbids and the tree uses zero of them — which makes this lane and DISPVOCAB2 one defect seen from two sides. DISPVOCAB2 — the asymmetry is demonstrated on one tree with a paired control: one disposition value makes the strict gate fail while the permissive gate reports no disagreement; a shared value passes both. The demonstration target was chosen mechanically by a listing and selected the previous release's artifact ABOUT this defect — the same shape as the release before, where the one artifact whose subject was check-run states was the only one an audit misread. SELFTEST — confirmed live with a matrix in both directions. With the wiring removed the self-test still exits zero while its sibling unittest reds, and the self-test's printed output is UNCHANGED between runs: it tests a neighbouring thing, not a weakened one. PROVENANCE — confirmed on the shipped record, with a positive control proving the check is not passing everything. The measurement also exposed a FALSE-POSITIVE CLASS in the check itself: a workflow run identifier is all decimal digits, decimal digits are valid hexadecimal, so a hex-shaped pattern reports a correct run reference as a broken commit citation. The obvious narrowing would make the check silent on a mistyped identifier, so out-of-population must be reported rather than dropped. NOTESPOP2 — all three parts confirmed. The decisive one is gate potency: the only occurrence of the generator's name in any workflow is its UNIT TEST, with a positive control showing the same search does find a wired script when one exists. The external sync does run in CI, so the generator's only venue is the operator's synced checkout — the single place where the guard cannot fire. ENVIRONMENT STATED WITH EVERY POTENCY VERDICT, because the previous release proved potency can depend on untracked local state. The self-test, branch reachability and disposition demonstrations are pure-interpreter work in detached worktrees with no build, so those verdicts are environment-independent. Gates at this commit: rivet validate PASS; status_evidence rc=0; claim_check 75/75; verdict-prose 0 disagreements with the population risen by exactly the number of outcomes added, which is the evidence that no lead word escaped classification rather than merely passing. Refs #1318, #1145, #1439, #1259, #1458, #1476, #1515, #1516. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
The v0.85 lane outcomes measured so far. Every one was measured before it was
written, and each lane's refuting command was run first. Two lanes refute a
premise this plan shipped with — which is the lane working, not failing.
The blocker that was priced against the wrong operation
Three MUST lanes have recorded UNPROVEN or UNRUN across multiple releases "for
want of a full build" against a multi-gigabyte disk budget.
Building only the CLI crate yields a binary reporting the workspace version
in under twenty seconds, well under a gigabyte of target, and it compiles a real
module to a valid object. That figure prices a full test build — all targets
plus dev-dependencies — while every one of those lanes needed a binary.
The cost and the requirement were two different populations, and a cost measured
for one build shape was used to decline a different shape — in the constraint
this programme repeated most often. (The binary on
PATHreports along-superseded version, so refusing it was correct; the answer was to build one.)
Per lane
The two refutations
FIRSTTOKEN4's premise does not hold. Every red branch is unreached over the
live population — which is what a green tree means. A reached red branch is
a defect, so "never reached" is true of every working guard and discriminates
nothing. Synthetic reachability shows every branch reachable by a legal artifact
and behaving correctly, with two passing controls. What is actually dead is a
vocabulary: the permissive gate admits two non-delivery dispositions the
strict gate forbids, and the tree uses zero of them. That makes FIRSTTOKEN4 and
DISPVOCAB2 one defect seen from two sides.
WIDEFILE8's five-release stall is explained. This repository's pressure
fixture is rescued by the lower rungs, so the suspect rung is never tried —
the fixture cannot exercise the path whose emitted code is the suspect. The
reporter's module reaches it once and rescues nothing: the first confirmed input
on that path. And the probe is getting rarer — against the reading in the
rung's own test file at an earlier release, its try count has halved while a
lower rung now rescues a function it did not before. If a later release rescues
that last function earlier, the defect becomes unobservable while remaining
live. That is an independent, measured argument for the middle step of the
fixed order. The retry stays out; the order was not jumped.
Three things caught before they became corrections
design assertion was cited at an abbreviated path. Reading it literally would
have concluded the constraint was gone and rewritten the scope on a typo.
returns more lines than call sites, because the population includes the
definition and two comments. The carried figure was right.
workflow run identifier as a broken commit citation. The obvious narrowing
would make it silent on a genuinely mistyped sha, so out-of-population must be
reported, not dropped.
Environment stated with every potency verdict
The previous release proved a guard's potency can depend on untracked local
state. The self-test, branch-reachability and disposition demonstrations are
pure-interpreter work in detached worktrees with no build, so those verdicts are
environment-independent.
Gates at this commit
rivet validate— PASSscripts/status_evidence_check.py— rc=0scripts/claim_check.py claims.yaml— 75/75scripts/verdict_prose_check.py— 0 disagreements, and the population rose byexactly the number of outcomes added. That is the real check: an unclassified
lead word would remove an artifact from the population rather than fail, so a
smaller rise would have meant an escape.
Refs #1318, #1145, #1439, #1259, #1458, #1476, #1515, #1516.
🤖 Generated with Claude Code
https://claude.ai/code/session_01YJK5LZZEkV5smCY1jKn18L