Skip to content

Add paired prediction diagnostics and a controlled editing demo - #311

Draft
Vi-Sri wants to merge 4 commits into
GrayboxTech:mainfrom
Vi-Sri:srini/hard-example-diagnostics
Draft

Vi-Sri wants to merge 4 commits into
GrayboxTech:mainfrom
Vi-Sri:srini/hard-example-diagnostics

Conversation

@Vi-Sri

@Vi-Sri Vi-Sri commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

This continues the hard-example diagnostics work from #307 on the requested srini/hard-example-diagnostics branch. The model-editing API is the foundation; this draft adds the comparison evidence needed to tell whether an intervention helped.

Contribution

  • New public wl.compare_predictions(control, intervention) API. It validates supplied checkpoint/split/preprocessing/budget/optimizer provenance, joins stable sample IDs, rejects changed cohorts or labels, and reports per-group deltas plus corrected/regressed cases. It does not silently intersect different populations or independently certify the truth of supplied metadata.
  • Runnable official Waterbirds adapter, pinned frozen ViT-B/16 feature cache, four-arm paired runner, saved-head replay verifier and class-conditioned image attribution.
  • Self-contained offline demo with group results, learning/gradient traces, architecture changes, corrected/regressed images, attribution warnings and provenance. No new Studio RPC or live synchronization is claimed.
  • API reference, experiment contracts, meeting proposal, reproduction commands and a measured-results presentation script.

Actual experiment

Three paired seeds, one baseline checkpoint per seed, 500 baseline steps and 250 steps in each continuation branch. All 5,794 official test images are retained. Target selection used baseline validation only.

Mean waterbird-on-land accuracy:

Arm Accuracy Change vs continued training
Continue 59.45% —
Widen head by 16 57.11% -2.34 pp
Balanced sampling 82.87% +23.42 pp
Widen + balanced sampling 74.71% +15.26 pp

Balanced sampling also regresses pooled common cases by 2.63 points, so it does not meet the proposed one-point regression budget. This is a diagnostic pilot, not a claim that editing reliably improves performance or evidence of a capacity bottleneck.

Known dataset label problems and failed IG completeness checks remain visible. Frozen-backbone attention cannot change after a head-only edit. Preview images are explicitly post-hoc illustrations; full metrics use the complete held-out split.

Validation

  • 26 new comparison tests and 4 existing model-graph tests passed (plus 2 subtests).
  • 10 planner tests and 4 sampler/attribution numerical tests passed.
  • Two full GPU experiments reproduced identical final predictions and sampled histories.
  • Independently replayed all 12 saved final heads on 5,794 cases each; predicted labels matched exactly.
  • Ruff and whitespace checks passed. Desktop browser filters, architecture view, image selection and attribution warnings were exercised with no JavaScript errors or broken images.

Large datasets, checkpoints, raw reports and generated image-bearing HTML remain outside Git. The branch commits the recipe, verifier, compact results and demo source.

Read measured results and demo script and reproduction instructions.

Remaining work: live Studio transport/synchronization, full training-state replay and structural transformer editing. This draft follows #287 and relates to #267; it does not close that issue. Earlier planning PR #307 is preserved.

Plan controlled interventions on known hard examples and add a paired CPU smoke runner using the public model-editing API.

[force ci]
Add meeting decisions and checks for valid controls, seeds and budgets.

[force ci]
Run and replay a controlled Waterbirds editing pilot; add an offline demo.

[force ci]

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant