Repository navigation
Conversation
Plan controlled interventions on known hard examples and add a paired CPU smoke runner using the public model-editing API. [force ci]
Add meeting decisions and checks for valid controls, seeds and budgets. [force ci]
Run and replay a controlled Waterbirds editing pilot; add an offline demo. [force ci]
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This continues the hard-example diagnostics work from #307 on the requested
srini/hard-example-diagnosticsbranch. The model-editing API is the foundation; this draft adds the comparison evidence needed to tell whether an intervention helped.Contribution
wl.compare_predictions(control, intervention)API. It validates supplied checkpoint/split/preprocessing/budget/optimizer provenance, joins stable sample IDs, rejects changed cohorts or labels, and reports per-group deltas plus corrected/regressed cases. It does not silently intersect different populations or independently certify the truth of supplied metadata.Actual experiment
Three paired seeds, one baseline checkpoint per seed, 500 baseline steps and 250 steps in each continuation branch. All 5,794 official test images are retained. Target selection used baseline validation only.
Mean waterbird-on-land accuracy:
Balanced sampling also regresses pooled common cases by 2.63 points, so it does not meet the proposed one-point regression budget. This is a diagnostic pilot, not a claim that editing reliably improves performance or evidence of a capacity bottleneck.
Known dataset label problems and failed IG completeness checks remain visible. Frozen-backbone attention cannot change after a head-only edit. Preview images are explicitly post-hoc illustrations; full metrics use the complete held-out split.
Validation
Large datasets, checkpoints, raw reports and generated image-bearing HTML remain outside Git. The branch commits the recipe, verifier, compact results and demo source.
Read measured results and demo script and reproduction instructions.
Remaining work: live Studio transport/synchronization, full training-state replay and structural transformer editing. This draft follows #287 and relates to #267; it does not close that issue. Earlier planning PR #307 is preserved.