Skip to content

Use unchanged upstream GPT-2 with MuonClip and AdamW - #165

Merged
charlesmartin14 merged 36 commits into
mainfrom
codex/continuous-muonclip-8
Oct 6, 2026
Merged

charlesmartin14 merged 36 commits into
mainfrom
codex/continuous-muonclip-8

Conversation

@charlesmartin14

@charlesmartin14 charlesmartin14 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

The earlier experiment used a modified six-head model; the first correction still rewrote GPT-2 with separate Q/K/V parameters. This revision meets the stricter architecture requirement by importing the unchanged upstream GPT class directly.

Model: byte-identical karpathy/llm.c train_gpt2.py at 7ecd8906afe6ed7a2b2cdb731c042f26d525b820, checked against upstream and guarded by SHA256 757d0cea0d48cbc4c7d7d70371f955d49cdf3a7cfb4c87701a720c2fe0905c34. No replacement model layers or rewritten forwards. GPT-2 Small has 12 blocks, 12 heads, width 768, packed QKV, learned positions, LayerNorm, original GELU, biases and tied embeddings: 124,439,808 parameters. Default single-run seed is upstream's 42.

Optimizers: MuonClip (default) and AdamW operate on the same unchanged model. MuonClip uses LR 0.02, Nesterov momentum 0.95, five NS iterations, RMS scale 0.2, matrix decay 0.1, per-head QK clip threshold 100 and balance 0.5. Hooks observe exact causal QK maxima and return None; Q/K weights and biases are rescaled only in the optimizer update. Auxiliary tensors use AdamW. Plain Muon remains explicitly selectable.

Configuration: original GPT-2/FineWeb reference horizon, 19,560 updates, 524,288 tokens/update, 700 warmup updates, cosine decay, gradient clipping 1.0 and 10,485,760-token validation every 250 updates. MuonClip is an explicit optimizer substitution, not a published GPT-2 speed record. TPU execution and Python sequential shard order differ from the CUDA C record and are documented.

Launch: require a live lease, matching guest, attention parity and a disposable full-size model/optimizer update at the actual batch/accumulation before fresh training. A failed or mismatched preflight prevents training. Deadline-interrupted full-budget runs cannot claim completion. A single run has a 12-hour cap; the six-run paired suite requires 72h45m.

Tracking: 48 stored block matrices, with Q/K/V slices extracted only from saved CPU weights for 72 spectral traces. Plans/reports carry new architecture/protocol identities and reject old/split-QKV results. STOCK_ARCHITECTURE.md and stock_weight_matrices.csv list every stored matrix.

Validation: exact upstream initialization/logits/gradients; independent clipped AdamW update; MuonClip update math, exact QK observations, forced per-head weight/bias clipping and V preservation; ownership, checkpoints, spectral slices, preflight failure gates, schedules, data, legacy references and reporting. Locally 72 passed, 1 optional WW skipped and 1 real-WW test deselected; CI now runs the full relevant suite with WW installed. Source syntax: 2 passed. Prior stale notebook/path CI failures were repaired and 13 affected tests passed.

This branch also carries the GPT-2/TPU experiment infrastructure developed since main. Historical modified-model experiments remain identifiable.

No TPU training has been launched. Live TPU memory fit/throughput/convergence remain unverified; the recorded allocation expired October 6 at 01:47:39 UTC.

Add four-chip data parallelism, global QK clipping, full-state resume checks,
long-run and smoke configs, and a GitHub-based TPU runbook. Preserve tied
embedding weights through XLA device conversion.
Apply the native XLA 2.6 precision control, use highest precision in both
SPMD configurations, and validate HLO precision before gradient comparisons.
Keep existing numerical tolerances. Record precision in runtime identity.

Validation: 79 tests passed, including the real four-device CPU-XLA check.
Physical TPU rerun remains pending.
Preserve full optimizer, RNG and sampler state across bounded training segments.
Track cumulative steps and test probes; add pause, recovery and checkpoint retention.
Include TPU runbook and CPU/XLA regression coverage.
Require the exact configured snapshot grid for continuation segments while
preserving the historical campaign minimum. Exercise real raw/clipped-alpha
training, pause/resume and fresh-worker segments in the integration test.

Validation: 53 continuation, completion and clip-Xmax campaign tests passed.
Prepare shared 5B-token corpus on Cloud Shell before requesting TPUs. Preserve
fixed probes and continuous MuonClip training; force a full checkpoint at the
wall-time stop, record measured throughput projections, and cap allocation with
Flex-start expiry. Add independent seed support and budget-ordering tests.

Validation: 24 CPU tests passed; one hardware-only check skipped. TPU preflight
must pass on allocation before scientific training starts.
Submit the bounded TPU request from Cloud Shell without a long /tmp preparation
job. Use the VM CPU and persistent disk for ordered parallel tokenization,
validate and archive the corpus, then train in one continuous process. Preserve
the six-hour allocation cap and no-restart guard. Save launch phases/errors in
Cloud Shell HOME and a transcript via run.sh; report empty status explicitly.

Validation: 27 CPU tests passed, one hardware check skipped. Parallel and serial
split files match byte for byte; tested budget flags and failure records.
Use a persistent pip cache, 300-second timeouts and bounded install retries.
Install pinned CPU PyTorch with XLA to avoid unnecessary CUDA downloads.
Preserve the allocation deadline; refuse concurrent workers or any scientific
run restart before changing dependencies. Keep local failure status when the
cloud reporter is unavailable during installation.

Validation: 18 local tests passed, including interrupted downloads, exhausted
retries and expired allocation deadlines. Shell syntax and diff checks passed.
TPU execution remains subject to the existing on-machine preflight.
Add a Cloud Shell replacement command that stops the old worker, removes TPU
requests and VMs in the two used zones, verifies deletion, and reuses the existing
data disk for one fresh v5e-8 request. Preserve old results, share installed
packages and data caches, and block duplicate allocation on repeated invocation.

Use a separate 24-hour configuration with paired token-error/alpha measurements
every 500 updates and a 23.5-hour stop deadline from VM boot. Archive all
checkpoints in GCS; retain three verified epoch files locally plus rolling
checkpoints. Add benchmark stage messages and periodic waiting heartbeats.

Validation: 35 CPU tests passed, one hardware-only test skipped. Tests cover
reallocation ordering/failure/duplicate guards, preserved disk, the 24-hour
budget, unchanged optimizer/probes, aligned measurements and verified pruning.
TPU numerical and shape preflight must pass again on the new allocation.
…ied rolling backups

Preserve the existing TPU, FineWeb corpus and all crash evidence. Launch one fresh process with scalar finite guards, paired token error/alpha monitoring, a stage stall watchdog, and allocation cutoff. Retain all metrics with bounded local/cloud checkpoints.

Validation: 43 CPU tests passed; TPU stability remains unverified.
…easurements

Reuse the successful replay's scalar finite checks and synchronized optimizer stages in a fresh uninterrupted run on the existing TPU. Save a full checkpoint before evaluation and complete pending measurements on explicit resume. Retain token-error measurements every 25 updates and all layer spectra at step 25 and every 100 updates.

Keep the current PyTorch/XLA environment and existing data/evidence. Verify numerical equivalence, interrupted measurement recovery, invalid-update handling, and mutable diagnostic backups with 57 passing CPU tests. Continuous TPU stability still requires the live run.
Stop the current MuonClip service and launch one bounded reference job with the pinned upstream GPT-2 model, canonical tokenized FineWeb data, AdamW, the published batch/schedule, and fixed benchmark validation tokens. Remove WeightWatcher, per-tensor diagnostics, preflight, and periodic checkpoint uploads from this runner.

Limit the whole new service to 30 minutes including compilation, data, final evaluation/save, and backup. Preserve old artifacts and the current allocation; refuse concurrent jobs. Report partial training/evaluation and hardware/data-order differences honestly, with the published reference curve at matching token counts.

CPU tests verify accumulated updates against upstream AdamW, exact sequential data windows, corrupted download rejection, external timeout and concurrent-launch prevention. Existing suite passed; TPU execution remains to be tested.
…ud saves

Add --kill-current to terminate the entire old service cgroup with SIGKILL rather than requesting a final checkpoint/backup. Add --no-save to skip the new run's model checkpoint and final cloud upload, retaining ordinary loss/throughput logs and the hard 30-minute deadline.

Nine focused CPU tests pass, including immediate-kill behavior and proof that no-save supervision never launches backup.
…d Muon recipe

Keep training settings and the 125-update evaluation cadence unchanged. Compute token error from the same validation logits; analyze immutable transformer-weight snapshots in a separate CPU process. Preserve per-layer raw/clipped fits, paired tables, and tracking status on disk and back up the tables to cloud. Add an explicit stop-and-fresh-start option that is idempotent across SSH retries.

Validation: 27 focused CPU tests passed, including source-model/optimizer parity, exact token counts, immutable snapshot pairing and real WeightWatcher analysis. Separate-process WeightWatcher smoke test passed. TPU verification remains live.
…ing and direct TPU replacement

Keep the successful 3000-update Muon model, data, batch and schedule. Use AdamW for the control, with 0.1 decoupled decay on hidden matrices and unchanged auxiliary updates. Record optimizer groups and label the published curve as Muon. Add direct TPU launch and explicit stopping of the recorded long-run service without deleting files.

Validation: PyTorch 2.6 CPU speedrun suite: 25 passed, 1 optional real-WeightWatcher test skipped. Includes decoupled-decay math, auxiliary update parity, mixed-precision learning/checkpointing and stop/launch guards. TPU convergence/performance not yet validated.
… workflow

Add a main speedrun entry point and fixed six-run protocol, three matched initialization seeds each for Muon and AdamW. Preserve the model, data, optimizer recipes and 3000-update schedule; complete the full budget instead of early-stopping at target. Keep paired raw/clipped WeightWatcher and token-error tracking every 125 updates, checkpoints and cloud backups.

Run sequentially under a bounded systemd service on an existing leased eight-chip TPU; require enough live allocation time, refuse concurrent training, and stop on incomplete jobs without automatic retry. Report seed-level mean/SD and paired final differences with explicit sample counts, plus per-seed elapsed times and target crossing outcomes. Document recipe/weight-decay confounding and partial-result handling.

Validation: PyTorch 2.6 CPU suites: 35 passed, one optional real-WeightWatcher test skipped. Tests exercise full-budget stopping, seed metadata, AdamW math and learning, paired statistics, lease rejection and failure-stop behavior. Live six-run TPU performance remains to be measured.
Replace the active six-head modified-speedrun model with standard GPT-2 Small: 12 blocks, 12 heads, width 768, learned positions, affine LayerNorm, GPT-2 GELU, biases, and tied embeddings/output (124,439,808 parameters). Preserve separate Q/K/V storage for 72 spectral traces and verify logits, loss and all gradients against the pinned packed-QKV GPT-2 reference.

Route both current single and paired-seed launch paths through the stock model. Use one optimizer owner for tied weights, FP32 parameters/states and BF16 activations; auxiliary AdamW uses 6e-4 instead of incompatible legacy embedding/head rates. Preserve the historical 25k model and its explicit six-head attention preflight. Tag plans/manifests/results, reject mixed-architecture pooling, and label old loss targets as historical.

Add a full architecture audit and generated 75-entry matrix inventory (74 unique matrices). Validation: 53 CPU tests passed, one optional WeightWatcher test skipped and one real-WeightWatcher test excluded because the local environment lacks WeightWatcher; final stock-model and suite checks: 14 passed. Live TPU training was not launched.
Use the published 19,560-update GPT-2/FineWeb horizon, 700 warmup updates, cosine decay and global gradient clipping for the stock Muon/AdamW experiments. Tag plans and manifests with the pinned protocol, verify the complete data corpus, and reject mismatched results.

Require a fresh lease for single-run launches and preserve incomplete status for interrupted full-budget runs. Document the original-baseline choice and explicit TPU/Muon differences.
Exercise the current multiseed notebook's result cell instead of asserting removed wrapper source. Preserve all returned tables and metadata and verify displayed matrix filtering. Place launcher transcripts in the working directory by default with an explicit override.

The affected notebook and path checks pass locally (13 tests).
Verify and import the byte-identical pinned llm.c GPT model directly, retaining packed QKV, original layers and forward definitions. Add packed-matrix MuonClip with observation-only causal QK hooks and optimizer-only Q/K weight/bias clipping. Use auxiliary AdamW and the pinned full GPT-2/FineWeb schedule.

Require a disposable full-model accumulated TPU optimizer preflight before fresh training. Preserve 72 spectral projection traces by slicing saved CPU QKV weights. Update architecture/protocol identities, defaults, matrix inventory, documentation and CI.

Validation: 72 focused CPU tests passed, one optional WeightWatcher test skipped, one real-WW test deselected locally; GitHub CI runs the full reference/optimizer tests with WeightWatcher installed. Two source syntax checks passed. Live TPU preflight remains required.
@charlesmartin14 charlesmartin14 changed the title Use stock GPT-2/FineWeb baseline for Muon and AdamW experiments Use unchanged upstream GPT-2 with MuonClip and AdamW Oct 6, 2026
@charlesmartin14
charlesmartin14 merged commit bb3114a into main Oct 6, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant