Skip to content

Plan hard-example diagnostics and scaffold paired editing experiments - #307

Draft
Vi-Sri wants to merge 2 commits into
GrayboxTech:mainfrom
Vi-Sri:codex/hard-example-diagnostics
Draft

Vi-Sri wants to merge 2 commits into
GrayboxTech:mainfrom
Vi-Sri:codex/hard-example-diagnostics

Conversation

@Vi-Sri

@Vi-Sri Vi-Sri commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Engineers can edit a model through the merged API, but need a reproducible way to inspect a recurring failure and compare an intervention against continued training from the same checkpoint.

This draft starts that workflow with an 8 October meeting handout, a detailed proposal, proposed backend/Studio contracts, and runnable scaffolding under wl-model-editing/hard-example-diagnostics.

Proposed first milestone

Use a pretrained frozen ViT-B/16 and editable head on known Waterbirds failure groups. Compare continued training, head widening, balanced sampling, and widening plus balanced sampling across three paired seeds. Record subgroup improvement, ordinary-case regressions, and the intervention history. Follow with last-block fine-tuning and a natural-variation experiment on Oxford Pets.

The handout covers today's five decisions, ownership, a workflow diagram, the four-arm comparison, acceptance criteria and implementation gates. The contracts cover cases, snapshots, layer identity, bounded diagnostic requests, interventions and comparisons, including stale UI responses after editing. The representation and API are the architectural contribution; the prototype UI will help validate them.

Implemented in this draft

  • Standard-library planner that generates the 12-run experiment matrix without launching training.
  • Planner guardrails reject edited controls, duplicate or invalid seeds/arms, invalid budgets, unsupported sampling and mismatched optimizer-reset policies. Ten standard-library tests exercise the plan and its safeguards.
  • CPU smoke runner restores the same learned head parameters into two branches, applies a real add_neurons operation through the public API, checks dependency propagation and optimizer references, then exports predictions and corrected/regressed case IDs.
  • Single wrapped hyperparameter configuration for the smoke; existing public model-signal tracking is reused.

Local validation, 8 October

  • Ten planner tests passed; generated 12 jobs with one shared baseline identifier per seed.
  • Synthetic smoke passed again: checkpoint prediction parity, widening propagation, optimizer binding, finite training, two branches with eight continuation steps each, 40 fixed evaluation cases per branch, JSON report export.
  • Ruff passed for the planner, tests and smoke runner; whitespace check passed.

These are targeted local checks, not a claim that the full repository CI suite passes.

Remaining work

Real dataset adapters, pretrained feature extraction, attribution, persistent replay, new RPCs and Studio screens are planned, not implemented. Synthetic smoke metrics are not evidence of real-data improvement. Frozen-backbone attention must remain unchanged after head-only edits; output-conditioned attribution is a separate measurement.

Follows #287. Related to #267; does not close it.

Plan controlled interventions on known hard examples and add a paired CPU smoke runner using the public model-editing API.

[force ci]
Add meeting decisions and checks for valid controls, seeds and budgets.

[force ci]

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant