Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
176 changes: 176 additions & 0 deletions docs/proposals/hard-example-diagnostics.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,176 @@
# Make one failure understandable — and test whether we can fix it

**WeightsLab · project brief · updated 8 October 2026**

**Status:** proposal and runnable scaffolding; real-data results are pending.

For the short presentation and decision checklist, open the
[8 October meeting handout](hard-example-meeting-2026-10-08.md).

**Foundation:** [merged model-editing API, PR #287](https://github.com/GrayboxTech/weightslab/pull/287).

## The outcome we want

An engineer selects a meaningful failure, compares it with examples the model
gets right, inspects the relevant layers, and tries a controlled intervention.
The result shows what improved, what regressed, and exactly what changed.

Our first demonstration should answer: **Can an intervention improve one
coherent failure mode beyond simply training longer, while preserving ordinary
cases?** Improvement is a hypothesis to test, not an outcome we assume.

The transcript's rare-tiger example is the product story: explain a recurring
failure and test a remedy. An attribution heatmap alone cannot establish that
new neurons learned a particular concept or that insufficient capacity caused
the original failure.

## Decisions and todos for today's meeting

Suggested 25-minute agenda. Owners below are proposed, not assigned commitments.

| Time | Decision / todo | Proposed owner | Leave with |
|---|---|---|---|
| 0–5 min | Choose one dataset and failure group | Vi-Sri + ML reviewer | Waterbirds first; explicit group definition |
| 5–10 min | Agree the comparison and success criteria | Vi-Sri + ML reviewer | Equal training budget; ordinary-case regression bound |
| 10–15 min | Agree the first visual journey | Vi-Sri + Studio maintainer | Case selection → layer inspection → branch → comparison |
| 15–20 min | Review API and history ownership | Vi-Sri + backend maintainer | Layer identity, sampled signals, checkpoint/event contract |
| 20–25 min | Choose the next demo checkpoint | Team | GPU time, reviewer, next biweekly meeting and fallback day |

- [ ] Confirm Waterbirds as the first controlled experiment; Oxford Pets next.
- [ ] Pick the target group using training/validation evidence, then freeze the choice.
- [ ] Confirm frozen ViT-B/16 + editable MLP head as the first editing scope.
- [ ] Agree whether the proposed +5 percentage-point target and ≤1-point common-case
regression budget are meaningful for this demo.
- [ ] Confirm where longer experiments belong and what small checks belong in CI.
- [ ] Request the existing Studio demo/reference so the prototype can follow its interaction style.
- [ ] Confirm the team member who will review model diagnostics and the Studio API contract.

## Three experiments, in order

| Experiment | Concrete question | Setup and intervention | Evidence / stop condition |
|---|---|---|---|
| **E1 · Waterbirds: unusual backgrounds** | Does head capacity or sample exposure explain failure on an atypical group? | Pretrained frozen ViT-B/16; fork one trained head checkpoint into continued training, wider head, group-balanced sampling, and wider head + balanced sampling. | Accuracy for every class/background group, margin, loss, corrected and newly broken examples. If the baseline has no recurring group failure, stop and report that before changing the protocol. |
| **E2 · Representation bottleneck** | Is the frozen representation limiting adaptation? | On the same selected cases, compare continued training with an unchanged head versus unfreezing the last ViT block. Keep architecture fixed and use a conventional PyTorch fine-tuning path until full-model editing is capability-tested. | Per-layer update/gradient traces and class-conditioned input attribution. Treat this as a separate representation experiment; it does not validate structural transformer edits. |
| **E3 · Oxford Pets: natural appearance variation** | Does the workflow transfer to a visually recognizable, naturally occurring subgroup? | Breed classification; manually review a coherent pose/occlusion subgroup and its confusable breed, then repeat the selected controlled comparison. | Fixed, labelled subgroup across train/validation/test. If subgroup support is too small, report an exploratory case study rather than a population-level result. |

Waterbirds provides bird-class/background groups and a reproducible setting for
atypical-context failures. It uses **composited images**, so it is a controlled
first experiment, not evidence about natural wildlife rarity. Preserve the
published splits and account for the known label notes in the authors' repository.
Source: [Waterbirds authors' dataset and protocol](https://github.com/kohpangwei/group_DRO#waterbirds).

Oxford-IIIT Pets provides 37 breeds, head boxes, and foreground trimaps. Pose or
occlusion subgroup labels would be our annotations, not supplied dataset labels.
Source: [Oxford-IIIT Pet dataset](https://www.robots.ox.ac.uk/~vgg/data/pets/).

Pin `ViT_B_16_Weights.IMAGENET1K_V1` and its preprocessing rather than a moving
default. Model reference: [Torchvision ViT-B/16](https://docs.pytorch.org/vision/stable/models/generated/torchvision.models.vit_b_16.html).

## What the engineer sees

```mermaid
flowchart LR
A[Select known hard cases] --> B[Compare with correct examples]
B --> C[Inspect predictions and layer signals]
C --> D[Record a hypothesis]
D --> E[Fork the same checkpoint]
E --> F[Continue training: control]
E --> G[Edit model or training sample policy]
F --> H[Compare on fixed held-out cases]
G --> H
H --> I[Keep, revise, or reject intervention]
I --> J[Replayable experiment history]
classDef proposed fill:#e8f1ff,stroke:#3064ae,color:#173153;
class A,B,C,D,E,H,I,J proposed;
```

Blue nodes describe the new diagnostics and comparison workflow. Model edits,
the ledger, and step-level model signals already have backend foundations.

The first screen should have four linked areas:

1. **Cases:** original image, label, prediction, confidence, margin, group,
and human notes. A hard positive is a same-class example distant from an
anchor; a hard negative is a different-class example close to it under a
stated representation. Store the anchor and distance definition.
2. **Model inspector:** graph with selected layer, dimensions, frozen state,
activation summaries, gradient norms, and weight-change history. Curves
suggest hypotheses; they do not automatically diagnose capacity or label errors.
3. **Intervention:** checkpoint, target layer, edit preview, sample policy,
and engineer's reason. The engineer chooses whether a valid difficult
sample stays, is emphasized, or is excluded from a training branch.
4. **Comparison and history:** baseline/control/intervention side by side;
subgroup results, ordinary-case regressions, selected images and attributions,
plus a timeline of the checkpoint and every edit.

Start from a small known case set. Automated embedding-based discovery follows
after we establish which comparisons actually help engineers make decisions.

## Make the result defensible

- Select the failure mode on train/validation; lock the test manifest before
intervention selection. Do not remove evaluation cases because they are hard.
- Fork identical model weights. Match seeds, update counts, batch schedule,
preprocessing, and optimizer policy between paired arms. Log changed sampling
separately from changed capacity. Reinitialize optimizers equally when a
shape-changing edit cannot preserve optimizer state.
- Compare **post-training edit versus post-training control**, not only edited
versus pre-training checkpoint. Also evaluate immediately after editing to
separate the edit's instantaneous effect from subsequent learning.
- Run three paired seeds initially. Show group counts, per-seed results, and
uncertainty. Three seeds and a small subgroup are preliminary evidence.
- Proposed pilot target: ≥5 percentage points on the selected subgroup versus
continued training, ≤1 point common-case regression, and consistent direction
across seeds. Agree these thresholds before looking at final test results;
meeting them alone does not establish statistical significance.
- Show both corrections and regressions. A null result is useful: record whether
more ordinary training, balanced exposure, or changing representation helped.

**Attribution boundary:** a frozen, deterministic backbone produces the same
attention for the same input after a head-only edit. Class-conditioned input
gradients may change because the head changed. Compute attribution through the
full image→backbone→head graph, with input gradients enabled; cached embeddings
alone cannot yield pixel attribution. Attention is a view of token interactions,
not a literal picture of everything the model "sees". Pin target class, baseline,
preprocessing and color scale across comparisons; record approximation error
for [Integrated Gradients](https://captum.ai/docs/extension/integrated_gradients).

If widening helps, a later ablation can disable only the newly added units to
test their contribution. Even that supports a contribution claim, not a claim
that an individual neuron represents "stripes".

## Delivery plan

These are work packages, not promised calendar dates. Start the first GPU pilot
after today's dataset and scope decisions.

| Package | Deliverable | Acceptance / handoff |
|---|---|---|
| **Now · planning draft** | This brief, experiment matrix generator, local synthetic smoke runner, proposed data contracts | Runnable scaffolding; synthetic results clearly labelled |
| **1 · baseline and cases** | Pretrained ViT checkpoint, dataset/split hashes, selected validation cases, locked test set | Repeatable subgroup failure; enough examples to evaluate it |
| **2 · controlled interventions** | Four E1 arms × three seeds, immediate-edit and post-training results | Checkpoint parity, optimizer rebinding, every-group metrics and regressions |
| **3 · diagnostic prototype** | Image comparison, layer histories, attribution, edit/history timeline | Same case and layer refer to the same snapshot across views |
| **4 · Studio integration** | Versioned requests, bounded signal fetches, edit acknowledgements and refreshed graph | Browser/server training pause, edit, resume and synchronization verified end to end |
| **Later · discovery** | Candidate hard-pair mining using the task model; optional external embeddings | Discovery quality evaluated separately; human review remains available |

Vi-Sri's proposed scope is the experiment, model representation, API/data
contracts and diagnostic prototype. Backend and Studio maintainers review how
those contracts join their existing systems. Keep each package reviewable in
its own PR; agree test placement with maintainers before adding long GPU jobs.

## Five-minute demonstration script

1. Show several examples of the same validation failure mode and a correct reference.
2. Point to the prediction margin and one informative layer comparison.
3. State the hypothesis, show the common checkpoint, and preview the intervention.
4. Replay recorded control/intervention runs; live training is optional.
5. Show held-out improvements **and** regressions, then the history needed to reproduce them.

Opening line for today's meeting:

> The editing API is merged. Next I want to make one recurring failure
> inspectable and test whether editing actually helps beyond training longer.
> We can start with known hard examples, keep the engineer in control, and build
> the API and visual comparison needed to bring that workflow into Studio.

Implementation entry point: [diagnostics scaffold](../../weightslab/examples/PyTorch/wl-model-editing/hard-example-diagnostics/README.md).
125 changes: 125 additions & 0 deletions docs/proposals/hard-example-meeting-2026-10-08.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
# From model editing to explainable experiments

WeightsLab · 8 October 2026 · meeting handout

**Decision requested:** agree one controlled failure-mode experiment and the
smallest useful inspection/comparison workflow. This is a proposal, not a
report of measured improvement.

> The editing API is merged. Next we want to show a real failure, explain what
> we suspect, try a model or data change, and see whether it helped more than
> simply training longer. The engineer should be able to inspect both the
> improvements and the damage, and replay exactly what changed.

## Today's five decisions

- [ ] **First dataset:** approve Waterbirds as a controlled starting point.
Keep the published splits; identify the failure group on validation, not test.
- [ ] **Editing scope:** frozen pretrained ViT-B/16 plus an editable MLP head.
Structural changes inside attention blocks are a separate capability gate.
- [ ] **Fair comparison:** one baseline checkpoint per seed, four arms, equal
continuation budgets, and an identical optimizer-reset policy in every arm.
- [ ] **Demo contract:** show cases, layer evidence, the proposed intervention,
and before/control/after comparison with history. Review the data contract
before extending the shared proto.
- [ ] **Ownership and next checkpoint:** confirm an ML reviewer and a
backend/Studio reviewer, GPU availability, and the next demo date. Suggested
owner of the experiment, representation and prototype: Vi-Sri.

## First experiment: capacity or sample exposure?

Hypothesis: a recurring atypical-background failure may respond to more head
capacity, more exposure to underrepresented groups, both, or neither. Do not
assume a failure proves a capacity bottleneck.

| Arm | Hidden head width | Training sampling | What it isolates |
|---|---|---|---|
| A · Continue | 64 | Original | Additional training alone |
| B · Widen | 80 (+16) | Original | Added capacity versus A |
| C · Resample | 64 | Group-balanced | Changed sample exposure versus A |
| D · Both | 80 (+16) | Group-balanced | Capacity versus C; sampling versus B |

Proposed recipe: `ViT_B_16_Weights.IMAGENET1K_V1`, seeds **17, 29, 43**,
**500 baseline steps**, then **250 steps per arm**. Cache frozen features for
fast head experiments. These budgets are pilot settings, not a claim that the
baseline has converged. Check validation learning curves before locking the
protocol; record any revision before final test evaluation.

**Evidence required:** identical starting predictions; immediate-edit versus
post-training effect; every group's accuracy and count; worst-group accuracy;
corrected and newly broken examples; per-seed paired differences; checkpoint,
split and configuration hashes. Report empirical test accuracy separately from
the training-frequency-weighted benchmark average.

**Proposed demo target, to agree today:** at least +5 percentage points on the
selected group versus A, no more than 1 point common-case regression, and a
consistent direction across the three seeds. Define the common groups before
evaluation. These are practical pilot criteria, not statistical significance.
A null result is a valid outcome; do not keep changing the protocol until an
edit wins.

Waterbirds deliberately combines bird foregrounds with backgrounds. That makes
the groups reproducible, but it is not a natural rare-animal benchmark.
[Dataset and evaluation protocol](https://github.com/kohpangwei/group_DRO#waterbirds).

## What the demo should show

```mermaid
flowchart LR
A[Surface known hard cases] --> B[Inspect predictions and layers]
B --> C[Record hypothesis and fork checkpoint]
C --> D[Continue: control]
C --> E[Edit model or sampling]
D --> F[Compare held-out cases]
E --> F
F --> G[Keep or reject with explainable history]
```

One selected image stays selected across four panels:

1. **Cases:** image, label, prediction, margin, group and human notes; also a
correct reference case. Begin with known cases, not automated mining.
2. **Model:** layer path/identity, shape, activation summary, gradient norm and
weight updates. Treat these as evidence for a hypothesis, not its proof.
3. **Decision:** show affected layers, sampling choice, parent checkpoint,
optimizer policy and the engineer's reason before applying the edit.
4. **Comparison:** control and intervention results, corrections, regressions,
and a replayable event timeline. Use a read-only recorded-run prototype first.

**Important visual boundary:** head-only edits cannot change attention in a
frozen deterministic backbone. Class-conditioned input attribution can change;
compute it through image → backbone → head with a fixed target class and
display scale. Do not draw a spatial heatmap from a cached CLS vector or call
attention a complete explanation of what the model sees.

## Work after the meeting, in order

| Gate | Concrete deliverable | Done when |
|---|---|---|
| 1 · Establish failure | Dataset adapter, pinned feature cache, baseline and validation case manifest | Failure is coherent and reproducible; test selection is locked |
| 2 · Test intervention | Four arms × three seeds with immediate and final snapshots | Same-parent checks, shape propagation, optimizer references and matched budgets pass |
| 3 · Make it inspectable | Case comparison, bounded layer diagnostics, attribution and history | Every view identifies the same case and model snapshot; missing data is explicit |
| 4 · Integrate Studio | Versioned requests and edit acknowledgement with architecture revision | Pause/edit/resume works; stale responses are rejected and the graph refreshes |

The architectural contribution is the **model representation and diagnostic
API/data contract**: stable case/snapshot/layer references, bounded measurement
requests, and explainable intervention history. The prototype UI tests whether
that information helps an engineer; it can later be ported into Studio.

Follow-ups: **E2**, keep the architecture fixed and unfreeze the last ViT block
to test a representation bottleneck; **E3**, repeat the workflow on a reviewed
natural-variation subgroup in Oxford Pets. Neither belongs in the first demo's
completion claim. [Full proposal](hard-example-diagnostics.md).

## What is available versus planned

- **Available:** a 12-job plan generator, proposed contracts, and a CPU
synthetic smoke runner exercising public neuron addition, dependency
propagation, optimizer binding and same-checkpoint comparisons.
- **Planned:** Waterbirds execution, pretrained feature extraction, image
attribution, persistent replay, new RPCs and the Studio comparison screen.
- **Review boundary:** draft PR [#307](https://github.com/GrayboxTech/weightslab/pull/307)
remains a planning/scaffolding PR. It does not claim to close issue #267.

Start with the [runnable scaffold](../../weightslab/examples/PyTorch/wl-model-editing/hard-example-diagnostics/README.md)
and its [proposed contracts](../../weightslab/examples/PyTorch/wl-model-editing/hard-example-diagnostics/CONTRACTS.md).
Loading
Loading