diff --git a/.github/workflows/baseline-tests.yml b/.github/workflows/baseline-tests.yml index e19b8176..c829e2c9 100644 --- a/.github/workflows/baseline-tests.yml +++ b/.github/workflows/baseline-tests.yml @@ -71,3 +71,14 @@ jobs: PYTHONPATH: baseline/nanogpt_one_head/src MPLBACKEND: Agg run: pytest -q baseline/nanogpt_one_head/tests + + - name: Verify upstream GPT-2, MuonClip and AdamW integration + env: + PYTHONPATH: baseline/nanogpt_one_head/src + MPLBACKEND: Agg + run: >- + pytest -q baseline/gpt2_small/tests/test_stock_gpt2.py + baseline/gpt2_small/tests/test_repeated_speedrun.py + baseline/gpt2_small/tests/test_muon_speedrun.py + baseline/gpt2_small/tests/test_speedrun30.py + baseline/gpt2_small/tests/test_muon_longrun.py diff --git a/README.md b/README.md index bcab5cdb..6b769abf 100644 --- a/README.md +++ b/README.md @@ -5,6 +5,18 @@ spectral renormalization-group program. The repository keeps the unmodified reference baselines separate from every RG intervention so optimizer claims can be tested against strong, restartable, statistically controlled experiments. +## nanoGPT speedrun: repeated-seed comparison + +Use `python3 baseline/gpt2_small/speedrun.py plan` for the fixed-recipe +MuonClip/AdamW comparison using the **unchanged upstream GPT-2 Small (124,439,808 parameters)**: +12 blocks, 12 heads, width 768, context 1024. Three matched seeds per optimizer, 19,560 updates each, +raw/clipped WeightWatcher spectra and validation token error every 250 updates. +The [pinned original GPT-2/FineWeb baseline](baseline/gpt2_small/muon_speedrun/BENCHMARK.md) uses 700 warmup updates, cosine decay and global gradient clipping at 1.0; MuonClip is an explicit optimizer substitution. +[Protocol, launch instructions and seed-level statistics](baseline/gpt2_small/muon_speedrun/REPEATED_SEEDS.md). +[Architecture audit and every matrix dimension](baseline/gpt2_small/STOCK_ARCHITECTURE.md). +The current paired runs use the standard architecture; the historical modified +six-head speedrun and one-head experiments remain separately identified. + ## Baseline status The baseline suite has completed a recipe audit, an executable audit, and a diff --git a/baseline/README.md b/baseline/README.md index 0dcfec5c..64bb39a1 100644 --- a/baseline/README.md +++ b/baseline/README.md @@ -382,3 +382,8 @@ GitHub Actions additionally compiles all Python sources, parses notebook code cells, runs a real pinned WeightWatcher integration, and executes a pinned nanochat CPU model/optimizer preflight. These bounded checks do not replace the full three-seed long-horizon campaigns or the required target-MPS preflight. + +## GPT-2 Small + +See [gpt2_small/README.md](gpt2_small/README.md) for the 124M-parameter, +12-layer, 12-head FineWeb TPU validation workflow. diff --git a/baseline/gpt2_small/README.md b/baseline/gpt2_small/README.md new file mode 100644 index 00000000..35290a12 --- /dev/null +++ b/baseline/gpt2_small/README.md @@ -0,0 +1,389 @@ +# Upstream GPT-2 Small / FineWeb: MuonClip and AdamW + +The current runner uses the **unmodified pinned upstream GPT-2 model directly**: +12 blocks, 12 heads, width 768, MLP 3072, context 1024, vocabulary 50257, +packed QKV, learned positions, LayerNorm, GELU, biases and tied output embeddings. +There are 124,439,808 unique parameters. Source SHA256 is verified before import. + +[Single-run instructions and required TPU preflight](muon_speedrun/README.md) · +[All matrix sizes](STOCK_ARCHITECTURE.md) · +[Pinned benchmark configuration](muon_speedrun/BENCHMARK.md). + +Use `--optimizer muon_clip` (default) or `--optimizer adamw` with the single-run +launcher and a live node. The original GPT-2/FineWeb baseline uses 19,560 updates, +700 warmup updates, cosine decay, 524,288 tokens per update and gradient clipping +at 1.0. Full validation and spectral observations occur every 250 updates. + +`python3 baseline/gpt2_small/speedrun.py plan` shows the +[three-seed MuonClip/AdamW comparison](muon_speedrun/REPEATED_SEEDS.md). +Its six 12-hour caps require 72h45m of lease; use a single run on a shorter allocation. +CPU verification is complete only when the pinned commit's CI passes. A live TPU +preflight is required to establish device memory fit and a successful optimizer update. + +The [historical 25k experiment](muon_longrun/README.md) uses the modified six-head +model and remains separately identified; it is not this new upstream-model benchmark. + +For the **30-minute GPT-2/FineWeb reference run**, see +[speedrun30/README.md](speedrun30/README.md). Its launcher stops the current +MuonClip service, preserves prior data, and uses the published benchmark's +tokenized FineWeb validation set and upstream GPT-2 model. It does not run +WeightWatcher or per-tensor diagnostics. The hard time cap permits a partial +reference run; it is not a promise of reaching the final published loss. + +This is the GPT-2 Small experiment, under `baseline/gpt2_small`. Its training module +is `rg_gpt2_small.experiment`; configurations and launchers live here. It reuses the +shared GPT implementation, optimizers, SPMD, corpus validation and WeightWatcher +adapter from the older `rg_nanogpt_one_head` infrastructure package. That dependency +does not constrain the head count. Existing one-head and continuous8 workflows stay +in their original directory. Existing disk and bucket paths are retained for data reuse. + +The TPU launch scripts expose both source packages through `PYTHONPATH`. For local +development, from the repository root, install both with +`pip install -e baseline/nanogpt_one_head -e baseline/gpt2_small`. + +## Continuous run using the successful replay's numerical checks + +The instrumented TPU replay `muonclip-replay-20261004-150734` passed update 2 and +all three evaluation splits. Both optimizer stages and gradient checks passed. +Training NLL decreased from 10.998178 to 10.982910 and test NLL from 10.984691 to +10.969128. Its cloud backup was verified. This is one-update evidence; the original +failure's cause and long-run stability remain unresolved. + +The user authorized a continuous run on the installed PyTorch/XLA environment while +waiting for access to Google's TorchTPU. The launcher now enables the shared +`execution_checks` path used by the passing replay: scalar finite reductions before +and after clipping, separate synchronized MuonClip and auxiliary AdamW steps, checks +of updated weights/moments, and synchronized evaluation with scalar host averaging. +The older `stack`/extrema per-tensor diagnostic stays disabled. Model, corpus, input +windows, optimizer formulas and learning-rate settings are unchanged. + +Before every scheduled evaluation, a full local checkpoint is committed with +`measurement_pending=true`. Successful evaluation/spectra replace it with a full +checkpoint containing the pending immutable result records, then publish verified +cloud uploads. If evaluation or spectra fails, the pre-evaluation checkpoint remains +on disk and is included in the worker's exit backup. An explicit recovery completes +that step's measurement before training further; there is no automatic restart. +Three rolling local/cloud checkpoints are retained plus initialization, and every +scalar/spectral result is retained. Routine finite-check reports retain early and +measurement steps plus current rolling reports; failures always retain their report. + +The checks add overhead and are retained throughout this run. A 25-update tiny-model +CPU test matches the original CPU optimizer trajectory exactly, and injected +evaluation/spectral failures verify checkpoint recovery and unchanged old metrics. +The full 124M continuous path still needs the live TPU run to establish stability. + +## Replaying the earlier failed update + +The `muonclip-night-20261004-053212` run stopped during evaluation after update 2. +Its step-1 checkpoint and cloud backup are preserved. Do not interpret that run as +passed validation or restart a long run on the strength of its finite initial loss. + +```bash +python3 baseline/gpt2_small/scripts/replay_muonclip.py start +python3 baseline/gpt2_small/scripts/replay_muonclip.py status +``` + +This uses the existing TPU and FineWeb, reads the saved step-1 checkpoint, and +replays exactly one next update into a fresh `muonclip-replay-` directory. +It checks the original config/data/software fingerprint, restores model, both +optimizers, sampler and RNG states, and retains the original LR schedule. It checks +the saved/restored tensors and fixed evaluation probes, gradients before/after +clipping, and weights/moments after primary MuonClip and auxiliary AdamW separately. +Only scalar finite flags cross to CPU; the diagnostic avoids the `stack`/extrema +expression implicated in the earlier diagnostic abort. `FIRST_INVALID.json` names +the first detected invalid stage/tensors. Native aborts and timeouts also get a +separate `TPU_PORT_FAILURE.json`, stage log, environment and available XLA metrics. + +The updated model/optimizer state is saved **before** post-update evaluation to +`diagnostic/update_state.pt`. This is diagnostic evidence, explicitly marked +non-resumable. Existing checkpoints, crash logs, disk and cloud objects are untouched. +The diagnostic is bounded to 30 minutes, plus up to 10 minutes for verified backup, +within the current allocation. No long run, allocation, restart, package install, +data download or cleanup follows automatically. Other launchers reject an active replay. + +CPU tests compare the replay with uninterrupted update 2, including weights, +optimizer states, sampler state and evaluation metrics. TPU equivalence is still +unproven: the extra synchronization changes graph boundaries. A passing replay +requires a subsequent check of the original execution path; it does not establish +that the original numerical failure is fixed. The replay module also accepts +`--device cpu` for a separately requested comparison of the same saved state. + +## Launch on the existing TPU + +The user authorized a fresh continuous MuonClip run after the successful replay. +Use the current 48-hour TPU, installed environment and preserved FineWeb: + +```bash +python3 baseline/gpt2_small/scripts/run_muonclip.py start +python3 baseline/gpt2_small/scripts/run_muonclip.py status +``` + +This starts one fresh process under `rg-gpt2-muonclip-.service` with +`Restart=no`, in `/mnt/disks/rg-data/gpt2small/muonclip-continuous-`. +It never invokes the cleanup/reallocation scripts, changes the installed packages, +downloads the corpus, or writes into previous experiment directories/cloud prefixes. +The saved `port-check-20261004-045633` crash evidence remains intact. A shared launch +lock and checks for existing services/trainers prevent simultaneous TPU jobs. + +The replay-style finite checks described above remain enabled throughout training. +The scalar loss/gradient-norm guard also remains before each update. The 124M GPT-2 +model, data, MuonClip/auxiliary AdamW settings, batch size and long-run LR schedule +are retained. The original numerical failure is not yet diagnosed. + +Training stays in one process from initialization until the existing allocation +cutoff (20 minutes before its recorded expiry), manual STOP, token budget or error. +There is no AdamW gate, stop/resume transition or automatic restart. A watchdog stops +a phase with no progress for 30 minutes and saves a failure report. First completed +updates are printed explicitly; a watch heartbeat alone is not completion evidence. + +Token error, NLL and full checkpoints are recorded at steps 0, 1, 2, 4 and every 25 +updates. All 72 matrix raw/clipped alphas are paired with the same-step token metrics +at step 25 and every 100 updates. Evaluation windows stay fixed. Synchronization, +evaluation, spectra and checkpoint upload overhead contribute to elapsed runtime. + +Three rolling local full checkpoints and an initialization milestone limit disk +usage. During training, verified uploads keep three rotating cloud slots plus +initialization; every scalar/spectral JSON record is retained. The cloud prefix is +`gs://tpu-builders-504820-ww-continuous8/gpt2small/`. +`muonclip/checkpoints/LATEST_VERIFIED.json` is published only after checkpoint and +metrics uploads, with object generation and CRC32C. It names a cloud slot, not a +local filename; validate generation/checksum when downloading for recovery. A failed +upload stops the trainer, retains local evidence, and does not advance this pointer. +Exit backup additionally archives the remaining local files using object permissions. +Old-run checkpoints and cloud objects are never deleted by this workflow. + +## Current validation failure + +Observed failures and their attribution are tracked in [TPU_PORT_BUGS.md](TPU_PORT_BUGS.md). +Diagnosing bugs in the PyTorch/TPU port is an explicit experiment objective. + +From an updated Cloud Shell checkout, use the existing 48-hour TPU for a bounded +four-update AdamW diagnostic (no allocation or data download): + +```bash +python3 baseline/gpt2_small/scripts/retry_adamw.py start +python3 baseline/gpt2_small/scripts/retry_adamw.py status +``` + +It runs independently of Cloud Shell under a new service and fresh output directory, +pins the checked-in source, and refuses concurrent trainers. Training is limited to +20 minutes with up to 10 more for backup. It records per-layer gradients, clipped +gradients, weights and optimizer moments as finite flags/extrema; only small reduced +tables move to CPU. Full checkpoints are saved at each completed update. Source and +runtime versions, input offsets, XLA metrics, stage times and failures accompany the +run. `TPU_PORT_FAILURE.json` records observed failure without assuming upstream fault; +`PROBE_STATUS.json` reports completion or failure. Existing cloud upload verification +runs on exit. This is diagnostic instrumentation, not a throughput measurement. + +The saved 2026-10-04 04:19 UTC stack from commit `4631a3d` is inside the +nonfinite-gradient diagnostic's per-parameter CPU copy. That branch is reached only +after detecting a nonfinite loss or aggregate gradient norm. The latest checkpoint +pointer was step 0. This is a failed numerical validation, not evidence of healthy +training. The origin of the nonfinite result is still unconfirmed. + +The failure handler now writes the scalar failure and raises immediately, avoiding +the full gradient-copy loop. Validation synchronizes XLA before reading diagnostic +scalars, prints its pre-update stage, and emits Python stacks every five minutes. +Each validation phase has a 30-minute limit (also bounded by allocation time); +on expiry its child is terminated and the report says validation is incomplete. +These changes improve failure reporting and bound waits; TPU numerical stability +has not yet been demonstrated. They do not change the model, data or learning rates. + +## Model and data + +124,439,808 trainable parameters: 12 independent blocks, 12 heads, width 768, +head dimension 64, MLP width 3072, context 1024, vocabulary 50257. GELU, LayerNorm, +causal attention and tied embedding/output weights; biases enabled. The six separate +projection matrices per block yield exactly 72 named trajectories. No QKV fusion or +layer sharing is introduced. + +The existing corpus is `/mnt/disks/rg-data/continuous8/data`: 5 billion training +GPT-2 tokens and 10 million tokens each for validation/test, stored as three uint16 +files with document indexes. Every launch validates metadata, sizes, SHA256 identities +and the document-disjoint split contract. This workflow never prepares/downloads data. +Evaluation uses deterministic windows within each split; train/validation/test probes +are fixed across optimizers and resumes. They are not the speedrun benchmark protocol. + +## Configurations + +- `configs/gpt2_small_fineweb_adamw_baseline.yaml`: AdamW 6e-4, betas .9/.95, + weight decay .1, gradient clipping 1.0. +- `configs/gpt2_small_fineweb_muonclip_baseline.yaml`: existing advanced MuonClip, + matrix LR .02, momentum .95, Nesterov, five NS iterations, QK threshold 100; + auxiliary AdamW LR 6e-4 independently configurable. +- `configs/gpt2_small_fineweb_muonclip_long_ww.yaml`: 100B token presentations; + refuses CLI launch without `--allow-long-run`. No loss-based stopping. +- `configs/gpt2_small_cpu_smoke.yaml`: tiny synthetic-corpus CPU integration tests. + +**LR caveat:** advanced MuonClip scales the orthogonal update by +`update_rms_scale * sqrt(max(matrix.shape))`, here .2 times that square root. +The requested .02 is retained explicitly, but is not equivalent to the plain Muon +parameterization. Its stability is NOT established by this configuration. Do not +call this a reproduced reference benchmark or launch a long run before validating it. + +Global microbatch 8, accumulation 4, context 1024 = 32,768 tokens/update, split over +eight chips. Batch size, accumulation, warmup, cosine horizon, floor, max_steps and +max_tokens are configurable. Stop-after is an interruption point, not an LR-horizon +change. max_tokens is rounded up to the next complete optimizer update. + +## Run only the short validation on the existing TPU + +From a clean Cloud Shell checkout of this commit: + +```bash +python3 baseline/gpt2_small/scripts/cloudshell.py +``` + +This verifies the existing active queue/node, gracefully stops the old trainer, +archives compact scientific results and verifies the archive, then deletes only +recognized model checkpoint files in the known old run directory. It does not +format/delete disks, remove the corpus/caches/indexes, or allocate a replacement TPU. +Old scalar records are retained as well as archived; ambiguous output is not removed. + +The exact commit is checked out under a fresh persistent-disk validation directory. +The sequence is AdamW through step 4, MuonClip through step 4, then fresh-process +resume of each through step 25. The baseline horizon stays fixed at 1000 steps. +MuonClip WW runs at steps 4 and 25. All earlier metric/WW file hashes must survive +resume unchanged. Short validation stops on nonfinite metrics, increasing training +probe NLL, incomplete matrix inventory or null controls. Failed alpha fits are retained +as NaN with status/reason, never replaced with clipped alpha. + +The cutoff is read from the previous allocation record; five minutes are reserved +for checkpoint/backup. Setup/compilation/WW may consume the remaining window: an +incomplete validation is reported honestly, and no new machine is requested. +Outputs are copied to `gs://tpu-builders-504820-ww-continuous8/gpt2small/` on exit; +a failed cloud copy is an error and leaves persistent-disk outputs intact. + +25 updates prove plumbing and initial learning direction only. They do not establish +convergence or reproduce CE 3.28. A longer matched benchmark is needed before claiming +normal GPT-2 training quality, and before the long scientific experiment. + +## Replace an expired allocation while retaining FineWeb + +From the updated, clean Cloud Shell checkout: + +```bash +python3 baseline/gpt2_small/scripts/reallocate_validation.py launch +python3 baseline/gpt2_small/scripts/reallocate_validation.py status +``` + +The launcher replaces only `ww-continuous8-24h-20261003-s1337` and its node. +It retains the existing `ww-continuous8-pilot-20261002-s1337-data` disk and all +cloud objects. It waits for disk detachment and attaches that same disk to one +new v5litepod-8 with a server-enforced four-hour allocation limit, including setup. +The startup script mounts the existing ext4 filesystem; it never formats a disk. +The corpus and Python environment are reused without downloading or reinstalling. +The fixed replacement queue name prevents duplicate launches on repeated commands. + +Short validation runs independently of Cloud Shell in `rg-gpt2-validation.service`. +The new root is `/mnt/disks/rg-data/gpt2small/ww-gpt2-validation-20261004-s1337`. +Cloud uploads are tested before training using object permissions, CRC32C and +size verification. Exit backup uses this same uploader rather than bucket-metadata +operations. The complete logs and outputs remain on the persistent disk on failure. + +This is a fresh validation from initialization, not continuation of the failed run. +Before each optimizer update, the runner checks loss and gradient norm. A failure +writes `nonfinite_diagnostics.json` with the step, loss values and aggregate gradient +norm, then stops before applying the invalid update. The earlier nonfinite +gradient's cause is still unconfirmed; these checks do not claim to fix it. +Only if the short AdamW checks pass does MuonClip validation proceed, followed by +the existing resume checks. No long run starts automatically. The service does +not restart automatically after failures or reboot. A finished service does not +delete its TPU: the allocation limit remains four hours unless stopped earlier. + +### Allocate 48 hours instead + +```bash +python3 baseline/gpt2_small/scripts/reallocate_validation.py launch --hours 48 +python3 baseline/gpt2_small/scripts/reallocate_validation.py status --hours 48 +``` + +The queued-resource API exposes no update operation for extending the requested +lifetime. This replaces only the four-hour validation request with +`ww-gpt2-validation-48h-20261004-s1337`. If its validation service is already +running, it is stopped before deletion; all existing files remain on the same +data disk. The new VM mounts that disk, reuses FineWeb and the environment, and +runs fresh short validation in a separate directory. Repeating this command +does not replace an existing 48-hour request or create a second machine. + +Both the server-enforced allocation limit and the startup deadline use 48 hours; +the worker deadline reserves ten minutes, and validation reserves another five. +The queue can wait up to four hours for capacity, independently of the 48-hour +allocation lifetime. At the published $0.60/chip-hour Flex-start rate, eight +chips for 48 hours cost $230.40 before storage and any earlier allocation usage. +Short validation still stops after its checks; it does not automatically start +the long experiment. The allocation remains available until deletion or expiry. + +## Records, checkpoints and timing + +### Capture a stalled validation + +From Cloud Shell on the updated branch: + +```bash +python3 baseline/gpt2_small/scripts/capture_stall.py --stop +``` + +This targets the 48-hour validation run above. It captures process CPU/memory, +logs, saved XLA metrics and live Python/native stacks before stopping only +`rg-gpt2-validation.service`. A bounded, optional py-spy installation goes in a +temporary directory, leaving the training environment and checkout unchanged. +If profiling is unavailable, process/log diagnostics are still retained. +Reports remain under the run's `diagnostics/stall-*` directory on the persistent +disk; a compact summary is printed for sharing. This command does not upload them. +Omit `--stop` for a read-only capture. A blocked update may be lost on stop; no new +checkpoint is promised. Existing checkpoints, FineWeb and the TPU allocation +remain. The allocation continues to incur compute usage until deletion/expiry. +Use `--show-last` to print the saved capture, native stack and current service state +without another profiler installation or stop. Both old and new package paths are +recognized when diagnosing the already allocated machine. + +The validator's `WAIT` lines only indicate a live process, not a completed update. +Likewise `before_update: 2` confirms the gradient check before update 2, not its +completion. An extended wait requires inspection, not an assumed compilation ETA. + +### Scientific records + +Per-step immutable JSON scalar and WW records are written incrementally. NLL, +perplexity, top-1 accuracy, error (fraction), steps, token presentations, wall time, +LR and pre-clipping gradient norm are included. All WW library columns are retained, +including raw/clipped fits, randomized distance, bounds, KS/D, spectral measures and +fit status. `xmin/xmax/D` are library-returned clipped-fit fields, not invented raw-fit +bounds. Raw failure is explicit; alpha is never constrained toward 2. + +Three rolling full checkpoints contain model, both optimizer states where applicable, +step-derived cosine scheduler state/config, token count, data sampler RNG, all training +RNG states and run identity. Checkpoint writes are atomic; old checkpoints are pruned +only after publishing the new pointer. A checkpoint includes pending metrics/WW records: +resume completes missing writes, appends later records, and refuses histories ahead of +the checkpoint rather than erasing data. Config/data/software version mismatch fails +closed. Optional milestone copies are independent of the rolling set. A single writer +lock prevents concurrent mutation. CPU exact-resume tests compare model and optimizer +states bit for bit. Numerical equivalence on TPU replacement still requires TPU testing. + +The first two update durations include compilation and execution and are reported +separately; they are **not pure compiler timing**. Native XLA CompileTime/ExecuteTime +metrics are also saved in `logs/xla_compile_metrics_after_step_*.txt`. Later synchronized update durations +measure training throughput, excluding evaluation/WW/checkpoint overhead. End-to-end +throughput is reported separately. The validation runner intentionally synchronizes +once per update to measure completed TPU work; the long configuration only synchronizes +at measurement boundaries after the first two updates. This is not a reproduced speedrun. + +WW supports fixed `interval`, explicit `steps`, and `logarithmic` 1/2/5-per-decade +schedules. A full pass reports seconds and recommends at least 9x that duration of +training between passes for at most 10% WW-only overhead. Include other overhead when +selecting the eventual long-run schedule. Scalar checkpoints are bounded in count; +no thousands of full checkpoint copies are retained. + +## Analysis and tests + +```bash +python baseline/gpt2_small/scripts/analyze.py /mnt/disks/rg-data/gpt2small/VALIDATION/muonclip +PYTHONPATH=baseline/gpt2_small/src:baseline/nanogpt_one_head/src python -m pytest baseline/gpt2_small/tests -q +``` + +Analysis creates CSVs and loss/perplexity/error-vs-token plots; mean/min raw and clipped +alpha trajectories; raw-alpha/error regression plots; all-layer raw/clipped plots +by matrix type; and per-matrix initial/latest/min/delta/recent slope/recent variance. +Correlations along one trajectory do not establish causation or independent-sample +significance. Tokens are presentations, not necessarily unique training tokens. diff --git a/baseline/gpt2_small/STOCK_ARCHITECTURE.md b/baseline/gpt2_small/STOCK_ARCHITECTURE.md new file mode 100644 index 00000000..162c2c4c --- /dev/null +++ b/baseline/gpt2_small/STOCK_ARCHITECTURE.md @@ -0,0 +1,72 @@ +# Unmodified upstream GPT-2 Small architecture + +Both MuonClip and AdamW instantiate `GPT` directly from the unchanged +`speedrun30/vendor/llmc_train_gpt2.py`, pinned to llm.c commit +`7ecd8906afe6ed7a2b2cdb731c042f26d525b820`. The adapter checks the complete file's +SHA256 before import. No model class, layer or forward method is rewritten. +The architecture identifier is `gpt2-small-upstream-packed-v2`. + +| Property | Value | +|---|---| +| Blocks / heads / hidden width | 12 / 12 / 768 | +| Head width / MLP width | 64 / 3072 | +| Context / vocabulary | 1024 / 50257 | +| Positions | Learned absolute embeddings | +| Normalization | Pre-LayerNorm, affine scales and biases, epsilon 1e-5 | +| Activation | Original GPT-2 tanh GELU | +| Linear biases | Enabled; vocabulary head has no bias | +| Output head | Tied to token embedding | +| Dropout | 0, as in the upstream pretraining model | +| Unique trainable parameters | **124,439,808** | +| Single-run initialization | Upstream initializer, seed 42 | + +## Stored layer weight matrices + +Shapes use PyTorch `[output features, input features]` storage. **QKV is packed** +in the upstream `attn.c_attn.weight` parameter. The following four matrices occur +in each of the twelve blocks, indexed 0–11: + +| Block | QKV `attn.c_attn` | O `attn.c_proj` | MLP IN `mlp.c_fc` | MLP OUT `mlp.c_proj` | +|---|---|---|---|---| +| L00 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L01 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L02 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L03 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L04 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L05 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L06 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L07 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L08 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L09 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L10 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | +| L11 | 2304 × 768 | 768 × 768 | 3072 × 768 | 768 × 3072 | + +| Other weight | Shape | +|---|---| +| Token embedding `transformer.wte.weight` | 50257 × 768 | +| Position embedding `transformer.wpe.weight` | 1024 × 768 | +| Vocabulary output `lm_head.weight` | 50257 × 768; alias of token embedding | + +There are **48 stored block matrices**, two embedding matrices and one named tied +output alias: **50 unique matrix parameters, 51 named entries**. The full machine +readable list is [stock_weight_matrices.csv](stock_weight_matrices.csv). +Each block also has LayerNorm scales/biases of length 768, packed QKV bias 2304, +attention output bias 768, MLP input bias 3072 and output bias 768. Final LayerNorm +has scale and bias vectors of length 768. Each block contains 7,087,872 parameters. +The upstream causal-mask buffers are not learned parameters. + +Q, K and V are each a 768 × 768 **slice** of the packed 2304 × 768 matrix; a single +head occupies 64 × 768 rows. WeightWatcher extracts these slices only from saved +CPU weights, retaining 72 projection traces without changing trainable storage. + +MuonClip observes QK logits without modifying the forward output. Its parameter +and bias rescaling occurs in the optimizer update. TPU attention selection and +BF16 autocast live outside the upstream model. The worker requires attention +parity plus a full-sized accumulated optimizer preflight before fresh training. + +[Configuration and provenance](muon_speedrun/BENCHMARK.md) · +[Launch instructions](muon_speedrun/README.md) · +[Paired-seed protocol](muon_speedrun/REPEATED_SEEDS.md) + +Historical six-head modified-model runs and the earlier split-QKV port have +different architecture/protocol identifiers and are excluded from new comparisons. diff --git a/baseline/gpt2_small/TPU_PORT_BUGS.md b/baseline/gpt2_small/TPU_PORT_BUGS.md new file mode 100644 index 00000000..b63cde16 --- /dev/null +++ b/baseline/gpt2_small/TPU_PORT_BUGS.md @@ -0,0 +1,183 @@ +# TPU port bug log + +Finding and isolating correctness/performance defects in the PyTorch-to-TPU port is +an explicit objective of this project. Preserve failures as evidence. A passing +CPU smoke test does not establish TPU correctness. Do not silently lower learning +rates, change data, or skip nonfinite checks to make a run appear successful. + +## 2026-10-04: Muon speedrun missing Pallas dependency and HBM exhaustion + +Status: two configuration failures confirmed; corrected retry requires live TPU validation. +These observations do not establish an upstream PyTorch/XLA or hardware bug. + +- Run `muon-speedrun-muon-20261004-225635`, commit `1b4548d`, v5litepod-8. +- Flash check failed with `ModuleNotFoundError: No module named 'jax'`. + `torch_xla[tpu]` does not include the Pallas extras; XLA 2.6 setup.py pins + both JAX and jaxlib to 0.4.38 for that optional dependency group. +- Automatic math-attention fallback kept global microbatch 128 (16/chip). + First forward/backward compilation failed with `RESOURCE_EXHAUSTED`: + 16.89G required versus 15.75G HBM, exceeding capacity by 1.14G. +- No completed training update. Step-zero checkpoint and final cloud backup + were saved. Disk evidence: + `/mnt/disks/rg-data/gpt2small/muon-speedrun-muon-20261004-225635`; + cloud prefix: `gs://tpu-builders-504820-ww-continuous8/gpt2small/muon-speedrun-muon-20261004-225635`. +- Correction: per-run pinned Pallas overlay; flash forward/backward check required; + no implicit math fallback; microbatch 64, eight accumulation passes, unchanged + global batch 524,288; early checkpoints after updates 1 and 5. Failure status now + replaces stale `training` status and includes the actual exception. +- The smaller microbatch and kernel must still pass a live full-model run. Do not + report the OOM fixed solely because CPU tests or attention-only checks pass. + +## 2026-10-04: GPT-2 AdamW nonfinite result, followed by stalled failure reporting + +Status: numerical cause open; blocking diagnostic implementation replaced. +Upstream attribution: **unconfirmed**. No upstream issue has been submitted. + +- Run: `ww-gpt2-validation-48h-20261004-s1337`, original commit `4631a3d`. +- Machine: one v5litepod-8, eight chips, SPMD; project `tpu-builders-504820`, + zone `us-west4-a`. Reported runtime: Python 3.10.12, PyTorch 2.6.0+cpu, + PyTorch/XLA 2.6.0. The library's `+cpu` build string does not identify where + the model computations executed; the run explicitly selected XLA/TPU. +- Model: GPT-2 Small, 124,439,808 parameters; 12 layers, 12 heads, width 768, + context 1024, tied embeddings. FineWeb reused from the persistent disk. +- Before updates 1 and 2, reported losses were finite (approximately 11.01 and + 10.24) and aggregate gradient norms were 16.177856 and 7.746079. +- After approximately 2.5 hours, the four-update check had not completed. + Latest checkpoint pointer: step 0. Process 6802 had roughly 212 GiB RSS. +- Saved Python and native stacks identify `require_finite_update`, line 141, + at `p.grad.detach().float().cpu()`. This branch executes only after detecting + a nonfinite loss or aggregate gradient norm. It does not reveal which scalar + failed, the first affected matrix, or the numerical root cause. +- Native frames include `THPVariable_cpu` and tensor conversion. They confirm + waiting in the host transfer path; they do not prove a hardware failure, + compiler bug, deadlock, or that all training ran on CPU. +- Service subsequently confirmed `MainPID=0`, `ActiveState=inactive`, + `SubState=dead`. TPU allocation and data remain available. +- Evidence on disk: + `/mnt/disks/rg-data/gpt2small/ww-gpt2-validation-48h-20261004-s1337/diagnostics/stall-20261004-041944-714308`. + +### Reporting defect and correction + +The old failure handler copied full gradients to CPU, serially, before writing +the failure report. That code could stall and hide the already detected failure. +Commit `f887593` removed these copies, added an execution barrier before host +scalar reads, saved the scalar failure immediately, and bounded validation phases. +These changes have local test coverage; they have not established numerical +correctness on TPU. GPT-2 now has its own `baseline/gpt2_small` package. + +### Next diagnostic and attribution criteria + +Run four fresh AdamW updates on the same allocation, with the original model, +corpus, seed and optimizer hyperparameters. Save each completed update. Record: + +- Source commit, Python/torch/torch_xla/libtpu versions, relevant XLA settings. +- Exact input-window offsets, corpus identities, initialization/rolling full states. +- Per-tensor finite flags and extrema before clipping, after clipping, and after + the optimizer update; parameter names plus Adam moment names. +- XLA compilation/execution counters, fallback counters, stage timestamps, + Python tracebacks and a structured failure report. + +Checks reduce tensors on device and transfer a small summary table, never full +gradients. Additional synchronization is recorded as instrumentation: it can change +fusion/compilation behavior. A pass under instrumentation does not by itself clear +the original execution path. The diagnostic stops after four updates or 20 minutes; +up to 10 additional minutes are reserved for verified cloud backup. Disk evidence +remains if upload fails. No long experiment starts automatically. + +To attribute an upstream bug, isolate the first failing operation and compare a +matched CPU/TPU replay with the same inputs, weights and optimizer state. Preserve +both results and a minimal reproducer before claiming a PyTorch/XLA defect. + +## 2026-10-04: instrumented diagnostic aborts at its first gradient check + +Status: open; fatal native message still required for diagnosis. + +Further stack evidence localizes this abort to `port_debug.py:68`: the new +diagnostic's `torch.stack((isfinite(value).all().float(), value.amin(), value.amax()))`. +Native frames include `torch_xla::Stack::Stack`, `XlaNode::GetOpShape`, and +`XLANativeFunctions::stack`. The abort occurs while assembling diagnostic summaries, +before the first optimizer update. It therefore does not reproduce or explain the +earlier nonfinite-result failure. CPU tests passed this operation, but TPU behavior +has not passed validation. The preceding native assertion/status text is still +needed; the Python abort trace alone is insufficient to identify its cause. + +- Run: `port-check-20261004-045633`, commit `19e2bb0`. +- Supervisor report: child exit code `-6` (SIGABRT), with last recorded stage + `before_clipping_started`, update 1, Unix time `1791089838.8617651`. +- No per-tensor results from this check were reported. This abort does not by + itself establish a nonfinite gradient, a particular failing operator, or an + upstream runtime/hardware bug. It is a separate observed failure from the + earlier numerical check and stalled host copy. +- Cloud backup was explicitly verified for this run. Logs, initialization, + exact input-window offsets and environment metadata remain on disk and in + its cloud prefix. The next evidence to inspect is the fatal native message + immediately preceding the abort in `run.log`. +- A later launch was blocked by an untracked repository-root `FETCH_HEAD` + file in Cloud Shell. This local checkout issue is separate from the TPU + abort; moving that file outside the checkout preserves it and clears this + particular cleanliness check. The real Git metadata is under `.git`. + +### User-authorized MuonClip bypass, 2026-10-04 + +`run_muonclip.py` starts a separate continuous MuonClip experiment on the remaining +48-hour allocation. Its config sets `validation_tensor_checks=false` and +`validation_gradient_checks=false`, bypassing the crashing per-tensor stack and its +verbose validation path. `finite_update_guard=true` still checks scalar losses and +aggregate norm before the optimizer update; the existing numerical guard is not +removed. Model/data/optimizer and long-run LR settings are unchanged. Additional +synchronization, reporting and denser checkpoint/spectral measurements are explicit. + +The workaround has CPU integration coverage; TPU success is not claimed. It is not +a repair or root-cause diagnosis for either SIGABRT or the earlier nonfinite result. +The earlier source checkouts, diagnostics, initialization checkpoints and verified +cloud archives are retained. New failures produce separate evidence in the new run. +No automatic restart, cleanup, disk formatting or TPU reallocation is performed. + +### MuonClip update-2 evaluation failure, 2026-10-04 + +Run `muonclip-night-20261004-053212`, commit +`6f8d59214f51be00212a9956182a0c697225f086`, failed about ten minutes after launch; +it did not train overnight. The first update and its evaluation completed. Step 1 +was saved locally and its cloud checkpoint upload was verified. The final cloud +backup was also explicitly verified. + +- Before update 2, four microbatch losses were finite (approximately 10.99) and + the aggregate gradient norm was `14.34039306640625`. +- The log then reported `completed_update: 2`, followed by `RuntimeError: + Nonfinite train NLL` in evaluation. The saved progress stage was `evaluating`, + completed step 2. The latest saved checkpoint was step 1. +- The recorder evaluated before saving the checkpoint, so the failed update's + weights and optimizer state were not preserved. Neither a corrupt update nor + an evaluation/runtime fault has yet been isolated. Finite pre-update losses + and norm do not prove finite updated parameters or optimizer moments. +- Evidence remains under `/mnt/disks/rg-data/gpt2small/muonclip-night-20261004-053212` + and `gs://tpu-builders-504820-ww-continuous8/gpt2small/muonclip-night-20261004-053212`. + +The new `replay_muonclip.py` diagnostic restores the saved step-1 state and exact +next input windows without changing optimizer settings. It checks saved/restored +state, fixed evaluation probes, pre/post-clipping gradients, and state after each +of primary MuonClip and auxiliary AdamW. It saves the resulting diagnostic state +before evaluation. Per-tensor finite reductions use scalar host transfers without +the earlier stack/extrema diagnostic. No attribution to upstream PyTorch/XLA or +TPU hardware is justified yet. These synchronization changes are recorded; a pass +does not by itself reproduce or fix the original continuous execution path. + +### Saved-state TPU replay passed, 2026-10-04 + +The user supplied the final results for `muonclip-replay-20261004-150734` at commit +`1b039b39761b13184f1cb09585b2e1e4343f8474`: `one_update_passed`, child exit 0. +All recorded checks passed and the cloud backup was explicitly verified. +Post-update NLL: train `10.98291015625`, validation `10.991303443908691`, +test `10.969128131866455`. Pre-update train/test NLL reproduced the prior saved +checkpoint's evaluation. This localizes neither the original fault nor a fix: +additional synchronization, reductions, process state and checkpoint restoration +differ from the failing run. + +The next user-authorized continuous run retains these synchronization/reduction +boundaries on the existing PyTorch/XLA environment, starting from initialization. +The shared `execution_checks.py` implementation is used by both replay and training. +Checkpointing now precedes evaluation and spectra, with a pending-measurement flag +for explicit recovery. CPU tests cover 25 continuous updates and exact state +recovery after injected measurement failures. No continuous TPU success is claimed +until its output is inspected. TorchTPU migration awaits repository/package access; +the installed environment and all previous evidence remain intact. diff --git a/baseline/gpt2_small/configs/gpt2_small_cpu_smoke.yaml b/baseline/gpt2_small/configs/gpt2_small_cpu_smoke.yaml new file mode 100644 index 00000000..1936b785 --- /dev/null +++ b/baseline/gpt2_small/configs/gpt2_small_cpu_smoke.yaml @@ -0,0 +1,56 @@ +run_id: tiny_cpu_smoke +seed: 1337 +model: + vocab_size: 64 + block_size: 8 + n_layer: 2 + n_head: 2 + n_embd: 32 + dropout: 0.0 + bias: true + tie_weights: true +dataset: + name: HuggingFaceFW/fineweb-edu + config: sample-10BT + split: train + revision: 593b3a867298afb8ce42625a270ef20ddcad28f9 + tokenizer: gpt2 + encoding_workers: 16 + encoding_batch_size: 256 + train_tokens: 512 + val_tokens: 128 + test_tokens: 128 +runtime: + matmul_precision: highest + mps_fallback: true + deterministic_algorithms: false + empty_mps_cache_after_weightwatcher: true + tpu_spmd: false + tpu_expected_chips: 8 +training: + batch_size: 2 + grad_accum_steps: 1 + max_steps: 4 + max_tokens: 64 + warmup_steps: 1 + schedule_steps: 4 + grad_clip: 1.0 +eval_batches: 1 +metrics_interval: 2 +ww: + enabled: false + interval: 0 + steps: [] + logarithmic: false + min_evals: 20 +optimizer: + display_name: AdamW + family: adamw + learning_rate: 0.0006 + min_learning_rate: 6.0e-05 + warmup_fraction: 0.01 + schedule: warmup_cosine + beta1: 0.9 + beta2: 0.95 + epsilon: 1.0e-08 + weight_decay: 0.1 diff --git a/baseline/gpt2_small/configs/gpt2_small_fineweb_adamw_baseline.yaml b/baseline/gpt2_small/configs/gpt2_small_fineweb_adamw_baseline.yaml new file mode 100644 index 00000000..9bc415c9 --- /dev/null +++ b/baseline/gpt2_small/configs/gpt2_small_fineweb_adamw_baseline.yaml @@ -0,0 +1,56 @@ +run_id: gpt2_small_fineweb_adamw_baseline_s1337 +seed: 1337 +model: + vocab_size: 50257 + block_size: 1024 + n_layer: 12 + n_head: 12 + n_embd: 768 + dropout: 0.0 + bias: true + tie_weights: true +dataset: + name: HuggingFaceFW/fineweb-edu + config: sample-10BT + split: train + revision: 593b3a867298afb8ce42625a270ef20ddcad28f9 + tokenizer: gpt2 + encoding_workers: 16 + encoding_batch_size: 256 + train_tokens: 5000000000 + val_tokens: 10000000 + test_tokens: 10000000 +runtime: + matmul_precision: highest + mps_fallback: true + deterministic_algorithms: false + empty_mps_cache_after_weightwatcher: true + tpu_spmd: true + tpu_expected_chips: 8 +training: + batch_size: 8 + grad_accum_steps: 4 + max_steps: 1000 + max_tokens: 32768000 + warmup_steps: 20 + schedule_steps: 1000 + grad_clip: 1.0 +eval_batches: 2 +metrics_interval: 25 +ww: + enabled: false + interval: 0 + steps: [] + logarithmic: false + min_evals: 20 +optimizer: + display_name: AdamW + family: adamw + learning_rate: 0.0006 + min_learning_rate: 6.0e-05 + warmup_fraction: 0.01 + schedule: warmup_cosine + beta1: 0.9 + beta2: 0.95 + epsilon: 1.0e-08 + weight_decay: 0.1 diff --git a/baseline/gpt2_small/configs/gpt2_small_fineweb_muonclip_baseline.yaml b/baseline/gpt2_small/configs/gpt2_small_fineweb_muonclip_baseline.yaml new file mode 100644 index 00000000..1472b7b9 --- /dev/null +++ b/baseline/gpt2_small/configs/gpt2_small_fineweb_muonclip_baseline.yaml @@ -0,0 +1,68 @@ +run_id: gpt2_small_fineweb_muonclip_baseline_s1337 +seed: 1337 +model: + vocab_size: 50257 + block_size: 1024 + n_layer: 12 + n_head: 12 + n_embd: 768 + dropout: 0.0 + bias: true + tie_weights: true +dataset: + name: HuggingFaceFW/fineweb-edu + config: sample-10BT + split: train + revision: 593b3a867298afb8ce42625a270ef20ddcad28f9 + tokenizer: gpt2 + encoding_workers: 16 + encoding_batch_size: 256 + train_tokens: 5000000000 + val_tokens: 10000000 + test_tokens: 10000000 +runtime: + matmul_precision: highest + mps_fallback: true + deterministic_algorithms: false + empty_mps_cache_after_weightwatcher: true + tpu_spmd: true + tpu_expected_chips: 8 +training: + batch_size: 8 + grad_accum_steps: 4 + max_steps: 1000 + max_tokens: 32768000 + warmup_steps: 20 + schedule_steps: 1000 + grad_clip: 1.0 +eval_batches: 2 +metrics_interval: 25 +ww: + enabled: true + interval: 0 + steps: + - 25 + - 1000 + logarithmic: false + min_evals: 20 +optimizer: + display_name: MuonClip + RMS-matched updates + auxiliary AdamW + family: muon_clip + learning_rate: 0.02 + min_learning_rate: 0.002 + warmup_fraction: 0.02 + schedule: warmup_cosine + momentum: 0.95 + nesterov: true + newton_schulz_steps: 5 + muon_epsilon: 1.0e-07 + weight_decay: 0.1 + update_rms_scale: 0.2 + qk_clip_threshold: 100.0 + qk_clip_balance: 0.5 + qk_diagnostics_interval: 25 + beta1: 0.9 + beta2: 0.95 + epsilon: 1.0e-08 + aux_learning_rate: 0.0006 + aux_min_learning_rate: 6.0e-05 diff --git a/baseline/gpt2_small/configs/gpt2_small_fineweb_muonclip_long_ww.yaml b/baseline/gpt2_small/configs/gpt2_small_fineweb_muonclip_long_ww.yaml new file mode 100644 index 00000000..24ee84ff --- /dev/null +++ b/baseline/gpt2_small/configs/gpt2_small_fineweb_muonclip_long_ww.yaml @@ -0,0 +1,69 @@ +run_id: gpt2_small_fineweb_muonclip_long_ww_s1337 +seed: 1337 +model: + vocab_size: 50257 + block_size: 1024 + n_layer: 12 + n_head: 12 + n_embd: 768 + dropout: 0.0 + bias: true + tie_weights: true +dataset: + name: HuggingFaceFW/fineweb-edu + config: sample-10BT + split: train + revision: 593b3a867298afb8ce42625a270ef20ddcad28f9 + tokenizer: gpt2 + encoding_workers: 16 + encoding_batch_size: 256 + train_tokens: 5000000000 + val_tokens: 10000000 + test_tokens: 10000000 +runtime: + matmul_precision: highest + mps_fallback: true + deterministic_algorithms: false + empty_mps_cache_after_weightwatcher: true + tpu_spmd: true + tpu_expected_chips: 8 +training: + batch_size: 8 + grad_accum_steps: 4 + max_tokens: 100000000000 + warmup_steps: 2000 + schedule_steps: 3051758 + grad_clip: 1.0 +eval_batches: 2 +metrics_interval: 500 +ww: + enabled: true + interval: 10000 + steps: + - 100 + - 500 + - 1000 + logarithmic: false + min_evals: 20 +optimizer: + display_name: MuonClip + RMS-matched updates + auxiliary AdamW + family: muon_clip + learning_rate: 0.02 + min_learning_rate: 0.002 + warmup_fraction: 0.02 + schedule: warmup_cosine + momentum: 0.95 + nesterov: true + newton_schulz_steps: 5 + muon_epsilon: 1.0e-07 + weight_decay: 0.1 + update_rms_scale: 0.2 + qk_clip_threshold: 100.0 + qk_clip_balance: 0.5 + qk_diagnostics_interval: 25 + beta1: 0.9 + beta2: 0.95 + epsilon: 1.0e-08 + aux_learning_rate: 0.0006 + aux_min_learning_rate: 6.0e-05 +long_run: true diff --git a/baseline/gpt2_small/muon_longrun/README.md b/baseline/gpt2_small/muon_longrun/README.md new file mode 100644 index 00000000..831d8647 --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/README.md @@ -0,0 +1,185 @@ +# Fresh 25,000-update Muon trajectory + +This extends the successful `muon-speedrun-muon-20261005-030026` recipe from a +new seed-1337 initialization. It never loads that speedrun's trained checkpoint. +The original run directory remains the read-only reference. The experiment +does **not** stop when validation NLL reaches 3.28. + +## Exact reference and intentional changes + +The reference is the modified `2024-11-10_UNetDoubleLr` transformer, **not stock +12-head GPT-2 and not MuonClip**: 162,201,642 parameters, 12 blocks, 6 heads, +width 768, context 1,024, vocab 50,304. QK normalization is present; QK clipping +and gradient clipping are absent. The benchmark dataset is pinned GPT-2-tokenized +FineWeb (`kjj0/fineweb10B-gpt2`, revision +`889765ea1f903759787add96995d81171b632d0c`), not FineWeb-Edu. + +The launcher compares the actual reference manifest and hashes of the model, +optimizer, runtime, data implementation/manifest and Pallas installer. Those +files are imported unchanged. Numerical settings remain: + +| Setting | Value | +|---|---| +| Hardware | One v5litepod-8 host, eight-chip SPMD | +| Attention | Verified TPU flash; no math fallback | +| Global microbatch | 64 sequences, eight per chip | +| Gradient accumulation | 8 microbatches | +| Effective batch | 524,288 tokens/update | +| Precision | BF16 activations, embedding/scalars; FP32 linear weights | +| Muon peak LR | 0.04 | +| Muon momentum | 0.85 to 0.95 over first 500 updates, then 0.95 | +| Newton–Schulz | 5 iterations; original coefficients and BF16 operation order | +| Auxiliary Adam peak LRs | Embedding 0.6; head 0.008; scalars 0.04 | +| Auxiliary Adam | betas (0.9, 0.95), eps 1e-8; foreach/fused false; TPU capturable | +| Weight decay / clipping | 0 / none | +| Initialization seed | 1337 | + +Only the run horizon/schedule, measurement cadence, gradient-norm logging, +full-state recovery, and orchestration change. The no-warmup scheduler is +`min(1, max(0, (25000 - update_index) / 7500))`, where `update_index` is zero +based. Completed update 17,500 is immediately before cooldown; the next update +uses index 17,500. The last applied factor is `1/7500`; the next is zero. +The separate 500-update momentum ramp is unchanged. At step 3,000 the LR +factor remains **1**, whereas the short reference had finished cooldown. + +The old full validation at step 3,000 was NLL **3.28082160949707**, perplexity +**26.5976**. That narrowly missed a strict `<=3.28` threshold. The new step-3,000 +result is compared with it in `COMPARISON_3000.json`; the schedules intentionally +differ after index 2,100, so identical validation loss is not expected. + +## Corpus and duration + +25,000 updates process **13,107,200,000 token presentations**. All 103 pinned +training shards contain **10,255,324,043 tokens**, slightly fewer usable after +the unchanged per-shard batch truncation. The stream makes one full sequential +pass and repeats about 28%; this is not 13.1B distinct tokens. Epoch is +`tokens_seen / usable_tokens_per_full_corpus_pass`, recorded with its denominator. + +The existing benchmark cache is reused. Missing pinned shards are downloaded +before training; the older FineWeb/Edu corpus is never deleted or reformatted. +At the measured reference rate, expect roughly **10–11 hours**, plus any unusual +setup/measurement overhead. This is an estimate, not a completion guarantee. +The service has a 12-hour cap, with the final 30 minutes reserved for tracker +drain/backup; the trainer saves at its earlier deadline. + +## Start, inspect, stop + +From a clean checkout on your authenticated **Mac terminal or Cloud Shell**: + +```bash +python3 baseline/gpt2_small/muon_longrun/launch.py start +python3 baseline/gpt2_small/muon_longrun/launch.py status +python3 baseline/gpt2_small/muon_longrun/launch.py metrics +python3 baseline/gpt2_small/muon_longrun/launch.py stop +``` + +`start` describes the live node and its linked queued resource and prints queue +creation, node creation, maximum duration, explicit termination timestamp and +remaining hours. It requires **more than 12.5 hours**, an already-mounted data +disk with 45 GiB free, a healthy reference manifest, and an idle TPU. It never +creates/deletes an allocation or stops another running experiment. It refuses +to guess expiration from queue submission time. SSH retry is idempotent. + +The launcher prints the commit, directory, service, training deadline and cloud +prefix. The record is `/mnt/disks/rg-data/gpt2small/MUON_LONG25K_LATEST.json`. +Output is under `/mnt/disks/rg-data/gpt2small/muon-long25k-s1337-`. +`status` reports the systemd PID, latest metrics, checkpoint, and tracker state; +`metrics` also tails the scalar file. `stop` requests a checkpoint after the +current update, then drains tracking and performs backup. There is no automatic +restart. Do not delete the allocation before the stop/backup finishes. + +If your prompt is `charles@t1v-...`, you are already inside the TPU, not on your +Mac. Prefer running the launch command from the Mac session where you logged +into Google Cloud. Cloud CLI errors are now printed rather than hidden behind +`CalledProcessError`. No training starts when the lease lookup fails. + +For direct execution on the TPU, add `--here` (also supported for status/metrics/stop). +Start still queries the live cloud API using the invoking account, verifies the +guest's metadata IP against the checked node, and then uses local `sudo` to create +the service. This avoids SSHing back into the same machine. It does not grant +the VM service account additional permissions or bypass lease verification. + +## Measurement plan + +| Output | Cadence | +|---|---| +| Train NLL, global FP32 gradient L2 norm, all LRs, phase, tokens/s, time, epoch | Every 10 updates (also first 5) | +| Full validation NLL, perplexity, top-1 token error and accuracy | Every 500 updates and every spectral snapshot | +| WeightWatcher | 0, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 3000; then every 1000; plus 17500/final | +| Full checkpoints | Every 2500; also initial, 3000, 17500, final/clean stop | + +Every validation uses the **same 10,485,760 benchmark tokens** and the unchanged +evaluator. These are validation measurements, not a separate held-out test set. +Step-0 validation and CPU WeightWatcher must complete before the first update. + +All 72 hidden matrices are retained: Q, K, V, O, MLP_IN, MLP_OUT for each block. +`alpha_raw` comes only from `raw_alpha`; `alpha_clip_xmax` comes only from +WeightWatcher's clipped `alpha`. Fitting imposes no alpha=2 constraint. All +native scalar outputs are retained, including randomized/null and ERG statistics, +`alpha_weighted`, `log_alpha_norm`, matrix rank and fit diagnostics. The weighted +metrics retain WW's native clipped-alpha definition and are labelled accordingly. + +Zero-initialized matrices have rank 0 and explicit unavailable/degenerate fits; +they are not discarded or assigned invented alphas. Thus 72 rows at step 0 +does not imply 72 successful power-law fits. Across-matrix standard deviations +are not error bars across independent seeds. + +After step 0, the CPU tracker operates on immutable weight snapshots, independently +of training RNG and the data stream. Completed JSON/CSV measurements and training +metrics are flushed to the persistent disk. Transfer/enqueue overhead and CPU WW +time are recorded separately. Above 10% measured foreground overhead, or a >10% +median-update slowdown while the CPU tracker is active, later sampling drops to +every 2,000 updates. This second criterion is a conservative contention proxy, +**not proof** that WW caused the slowdown. Early/milestone/final measurements +remain, and scientific quantities do not change. The tracker drains at exit. + +## Startup and recovery + +The existing eight-chip flash forward/backward gate runs first. The trainer then +verifies its initial full validation and 72-row WW measurement. It checks finite +loss and sampled gradient norm while progressing. At update 2 it captures full +state, restores **its own step-0 initialization**, replays its first two updates, +and requires exact tensor/state equality. This is an in-process recovery test, +not a new training process and not a restore from the short reference. It stops +if the comparison fails and writes `RESUME_PARITY.json` only on success. + +After 100 updates it records steady timing relative to the reference, refuses a +>50% slowdown, and continues the **same process**. `STARTUP.json` is the evidence; +an allocated/active service alone does not mean those checks passed. + +Checkpoints atomically contain model, both optimizer states, scheduler, completed +step, token count, CPU/Python/NumPy/XLA RNG, exact shard/offset/cycle, corpus identity +and numerical source hashes. Keep the latest two rolling checkpoints plus +permanent 0, 3000, 10000, 17500, 25000 milestones. Final cloud backup copies these +retained checkpoint files and the small scientific outputs with CRC verification. +Intermediate checkpoints/results are immediately safe on the mounted disk; +cloud backup is an exit operation, not a per-update claim. + +For an **explicit recovery**, use the same commit and a compatible existing TPU: + +```bash +python3 baseline/gpt2_small/muon_longrun/launch.py recover \ + --checkpoint /mnt/disks/rg-data/gpt2small//checkpoints/step_0010000.pt +``` + +Recovery writes a new directory and rejects changed scheduler, data order, +numerical source, precision/runtime identity or model shape. A 3,000-step speedrun +checkpoint has a different schema and is rejected. No recovery happens implicitly. +CPU tests establish exact state replay across a shard boundary. The live TPU +startup gate establishes initial two-update replay only if it passes; it does +not prove that every later interruption or software/hardware change is bitwise +reproducible. Retain `RESUME_PARITY.json` and the pinned environment with the data. + +## Local verification + +```bash +OMP_NUM_THREADS=1 OPENBLAS_NUM_THREADS=1 \ +PYTHONPATH=baseline/gpt2_small/src:baseline/nanogpt_one_head/src \ +python3 -m pytest baseline/gpt2_small/tests/test_speedrun30.py \ + baseline/gpt2_small/tests/test_muon_speedrun.py \ + baseline/gpt2_small/tests/test_muon_longrun.py -q +``` + +Tests cover upstream model/optimizer parity, schedule boundaries, unchanged +sequential sampling, corpus budget, exact full-state CPU recovery, retention, +lease rejection, and real WeightWatcher raw/clipped/null fields with zero matrices. diff --git a/baseline/gpt2_small/muon_longrun/checkpoint.py b/baseline/gpt2_small/muon_longrun/checkpoint.py new file mode 100644 index 00000000..dcb55f1a --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/checkpoint.py @@ -0,0 +1,81 @@ +"""Atomic full-state saves, bounded local retention, explicit deterministic restore.""" +from dataclasses import asdict +import os +import random +import numpy as np +import torch +from common import BATCH_TOKENS, PERMANENT, atomic_json, sync_dir + + +def cpu(value): + if isinstance(value,torch.Tensor): return value.detach().cpu().clone() + if isinstance(value,dict): return {k:cpu(v) for k,v in value.items()} + if isinstance(value,list): return [cpu(v) for v in value] + if isinstance(value,tuple): return tuple(cpu(v) for v in value) + return value + + +def state(model,muon,adam,stream,rt,step,schedule,identity): + rt.step(wait=True) + rng={'torch':torch.get_rng_state(),'numpy':np.random.get_state(),'python':random.getstate()} + if rt.tpu: rng['xla']=int(rt.xm.get_rng_state(device=rt.device)) + return dict(schema=2,step=step,tokens_seen=step*BATCH_TOKENS,config=asdict(model.config), + model=cpu(model.state_dict()),muon=cpu(muon.state_dict()),adam=cpu(adam.state_dict()), + data_cursor=stream.state_dict(),rng=rng,scheduler=asdict(schedule), + next_update=schedule.values(step),identity=identity) + + +def save(root,payload): + step=payload['step']; folder=root/'checkpoints'; folder.mkdir(exist_ok=True) + path=folder/f'step_{step:07d}.pt'; temporary=path.with_suffix('.tmp') + with temporary.open('wb') as f: + torch.save(payload,f); f.flush(); os.fsync(f.fileno()) + temporary.replace(path) + sync_dir(folder) + link=root/'checkpoint_latest.pt.tmp'; link.unlink(missing_ok=True) + os.link(path,link); link.replace(root/'checkpoint_latest.pt') + atomic_json(root/'checkpoint_latest.json',{'file':str(path.relative_to(root)),'step':step, + 'tokens_seen':payload['tokens_seen'],'next_update':payload['next_update'], + 'scheduler':payload['scheduler'],'schema':2}) + rolling=sorted(p for p in folder.glob('step_*.pt') if int(p.stem.split('_')[1]) not in PERMANENT) + for old in rolling[:-2]: old.unlink() + atomic_json(root/'checkpoint_inventory.json',{'files':[str(p.relative_to(root)) for p in sorted(folder.glob('*.pt'))], + 'permanent_steps':sorted(PERMANENT),'rolling_keep':2}) + return path + + +def restore(payload,model,muon,adam,stream,rt,schedule,identity): + if payload.get('schema')!=2 or payload['scheduler']!=asdict(schedule) or payload['identity']!=identity: + raise RuntimeError('Resume source, scheduler, corpus or numerical settings differ') + if payload['config']!=asdict(model.config): raise RuntimeError('Model configuration mismatch') + model.load_state_dict(payload['model'],strict=True) + for p in (*model.parameters(),*model.buffers()): rt.replicate(p) + muon.load_state_dict(payload['muon']); adam.load_state_dict(payload['adam']) + for group in muon.groups: rt.shard_matrices(group['buffer']) + # Adam counters/moments must follow the live parameter device, including capturable steps. + for p,values in adam.state.items(): + for key,value in values.items(): + if isinstance(value,torch.Tensor): + values[key]=value.to(rt.device); rt.replicate(values[key]) + stream.load_state_dict(payload['data_cursor']) + torch.set_rng_state(payload['rng']['torch']); np.random.set_state(payload['rng']['numpy']) + random.setstate(payload['rng']['python']) + rt.step(wait=True) + # Restore the seed AFTER materializing restored state; mark_step may advance it. + if rt.tpu: rt.xm.set_rng_state(payload['rng']['xla'],device=rt.device) + return int(payload['step']) + + +def assert_same(a,b,path='state'): + """Exact CPU comparison of full replay state; used only by the startup gate.""" + if isinstance(a,torch.Tensor): + if not torch.equal(a,b): raise RuntimeError('TPU replay differs at '+path) + elif isinstance(a,np.ndarray): + if not np.array_equal(a,b): raise RuntimeError('TPU replay differs at '+path) + elif isinstance(a,dict): + if a.keys()!=b.keys(): raise RuntimeError('TPU replay keys differ at '+path) + for k in a: assert_same(a[k],b[k],path+'.'+str(k)) + elif isinstance(a,(tuple,list)): + if len(a)!=len(b): raise RuntimeError('TPU replay lengths differ at '+path) + for i,(x,y) in enumerate(zip(a,b)): assert_same(x,y,path+'.'+str(i)) + elif a!=b: raise RuntimeError('TPU replay differs at '+path) diff --git a/baseline/gpt2_small/muon_longrun/common.py b/baseline/gpt2_small/muon_longrun/common.py new file mode 100644 index 00000000..7abf83a6 --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/common.py @@ -0,0 +1,111 @@ +"""Fixed long-run plan; reuse the validated speedrun implementation.""" +from dataclasses import dataclass, asdict +import hashlib +import json +import os +from pathlib import Path +import sys + +HERE = Path(__file__).resolve().parent +SHORT = HERE.parent/'muon_speedrun' +sys.path.append(str(SHORT)) +STEPS = 25000 +BATCH_TOKENS = 524288 +VAL_TOKENS = 10485760 +MICROBATCH = 64 +CONTEXT = 1024 +REFERENCE_NAME = 'muon-speedrun-muon-20261005-030026' +CACHE = Path('/mnt/disks/rg-data/benchmark-fineweb10B-889765ea') +PERMANENT = {0, 3000, 10000, 17500, STEPS} +EARLY_WW = {0, 100, 250, 500, 750, 1000, 1500, 2000, 2500, 3000} + + +@dataclass(frozen=True) +class Schedule: + total_steps: int = STEPS + warmdown_steps: int = 7500 + warmup_steps: int = 0 + + @property + def cooldown_start(self): + return self.total_steps-self.warmdown_steps + + def factor(self, update_index): + return min(1., max(0., (self.total_steps-update_index)/self.warmdown_steps)) + + def phase(self, update_index): + return 'cooldown' if update_index >= self.cooldown_start else 'main' + + def values(self, update_index): + factor = self.factor(update_index) + return dict(lr_factor=factor, scheduler_phase=self.phase(update_index), + muon_lr=.04*factor, adam_embedding_lr=.6*factor, + adam_head_lr=.008*factor, adam_scalar_lr=.04*factor) + + +def ww_due(step, sparse=False): + interval = 2000 if sparse else 1000 + return step in EARLY_WW or step in PERMANENT or (step > 3000 and step % interval == 0) + + +def val_due(step, sparse=False): + return step % 500 == 0 or ww_due(step, sparse) + + +def adapt_tracking(step, foreground_fraction, median_seconds, reference_seconds, tracker_state): + """Conservative cadence reduction; observed host slowdown is NOT causal proof.""" + if step < 3000: return False + return foreground_fraction > .10 or ( + median_seconds > reference_seconds*1.10 and tracker_state.get('status')=='measuring') + + +def atomic_json(path, value): + path.parent.mkdir(parents=True, exist_ok=True) + temporary = path.with_suffix(path.suffix+'.tmp') + with temporary.open('w') as f: + json.dump(value, f, indent=2, allow_nan=False); f.write('\n') + f.flush(); os.fsync(f.fileno()) + temporary.replace(path) + sync_dir(path.parent) + + +def sync_dir(path): + fd = os.open(path, os.O_RDONLY) + try: os.fsync(fd) + finally: os.close(fd) + + +def sha(path): + h=hashlib.sha256() + with path.open('rb') as f: + for chunk in iter(lambda:f.read(8*1024*1024),b''): + h.update(chunk) + return h.hexdigest() + + +def verify_reference(reference): + manifest=json.loads((reference/'manifest.json').read_text()) + val=json.loads((reference/'latest_validation.json').read_text()) + expected={'recipe':'2024-11-10_UNetDoubleLr','optimizer':'muon','seed':1337, + 'config':{'vocab_size':50304,'n_layer':12,'n_head':6,'n_embd':768}, + 'batch_tokens':BATCH_TOKENS,'global_microbatch_sequences':64,'accumulation':8, + 'muon_lr':.04,'adam_embedding_lr':.6,'adam_head_lr':.008,'adam_scalar_lr':.04, + 'weight_decay':0,'gradient_clipping':False,'attention':'flash','warmup_updates':0, + 'data_repo':'kjj0/fineweb10B-gpt2','data_revision':'889765ea1f903759787add96995d81171b632d0c'} + for key,value in expected.items(): + if manifest.get(key)!=value: + raise RuntimeError(f'Reference setting differs: {key}: {manifest.get(key)!r}') + if not (val['step']==3000 and val['full_benchmark_evaluation'] + and val['evaluation_tokens']==VAL_TOKENS and 3 < val['val_nll'] < 3.3): + raise RuntimeError('Reference full 3,000-step validation is not available/healthy') + hashes={} + for relative in ('muon_speedrun/model.py','muon_speedrun/optim.py','muon_speedrun/runtime.py', + 'muon_speedrun/data.py','muon_speedrun/pallas_dependencies.py', + 'speedrun30/train.py','speedrun30/data_manifest.json'): + old=reference/'repo/baseline/gpt2_small'/relative + new=HERE.parent/relative + if sha(old)!=sha(new): + raise RuntimeError('Validated implementation changed: '+relative) + hashes[relative]=sha(new) + return {'reference_run':str(reference),'manifest':manifest,'final_validation':val, + 'unchanged_source_sha256':hashes} diff --git a/baseline/gpt2_small/muon_longrun/launch.py b/baseline/gpt2_small/muon_longrun/launch.py new file mode 100644 index 00000000..d806a1f8 --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/launch.py @@ -0,0 +1,224 @@ +"""Start/status/stop the fixed 25k plan on the existing TPU; never allocate a node.""" +import argparse +import datetime as dt +import fcntl +import importlib.util +import json +import os +from pathlib import Path +import re +import shlex +import shutil +import subprocess +import sys +import time +import uuid +import urllib.request + +PROJECT='tpu-builders-504820'; ZONE='us-west4-a' +NODE='ww-gpt2-validation-48h-20261004-s1337-node' +BASE=Path('/mnt/disks/rg-data/gpt2small'); LATEST=BASE/'MUON_LONG25K_LATEST.json' +MIN_REMAINING=12.5*3600 + + +def run(command,**kwargs): + try: + return subprocess.run(command,check=True,text=True,**kwargs) + except subprocess.CalledProcessError as exc: + # capture_output previously swallowed the reason gcloud refused the query. + if exc.stderr: print(exc.stderr.rstrip(),file=sys.stderr,flush=True) + raise + + +def timestamp(value): + value=re.sub(r'(\.\d{6})\d+',r'\1',value) + return dt.datetime.fromisoformat(value.replace('Z','+00:00')).timestamp() + + +def lease_from(node,queue,now): + if node.get('state')!='READY' or queue.get('state',{}).get('state')!='ACTIVE': + raise RuntimeError('Existing TPU/queue is not READY/ACTIVE') + if node.get('acceleratorType')!='v5litepod-8': + raise RuntimeError('Expected the existing single-host v5litepod-8') + candidates=[node.get('schedulingConfig',{}).get('terminationTimestamp'), + queue.get('runDuration',{}).get('terminationTime')] + for spec in queue.get('tpu',{}).get('nodeSpec',[]): + if spec.get('nodeId')==NODE: + candidates.append(spec.get('node',{}).get('schedulingConfig',{}).get('terminationTimestamp')) + expiries=[timestamp(x) for x in candidates if x] + if not expiries: + raise RuntimeError('API did not return an explicit termination time; refusing to infer it from queue creation') + expiry=min(expiries) + result={'node':NODE,'queue':queue['name'].rsplit('/',1)[-1], + 'queue_created':queue.get('createTime'),'node_created':node.get('createTime'), + 'max_run_duration':queue.get('runDuration',{}).get('maxRunDuration'), + 'node_internal_ips':[item['ipAddress'] for item in node.get('networkEndpoints',[]) if item.get('ipAddress')], + 'termination_unix':expiry,'termination_utc':dt.datetime.fromtimestamp(expiry,dt.timezone.utc).isoformat(), + 'checked_unix':now,'remaining_hours':(expiry-now)/3600} + print(json.dumps(result,indent=2),flush=True) + if expiry-now<=MIN_REMAINING: + raise RuntimeError('Need more than 12.5 hours remaining; no run started and no allocation changed') + return result + + +def live_lease(): + flags=['--project='+PROJECT,'--zone='+ZONE,'--format=json'] + try: + node=json.loads(run(['gcloud','alpha','compute','tpus','tpu-vm','describe',NODE,*flags],capture_output=True).stdout) + except subprocess.CalledProcessError as exc: + raise RuntimeError('Cloud lease lookup failed before launch. See the gcloud error above. ' + 'Use the Mac/Cloud Shell session authenticated as your project user; ' + 'the TPU guest may be using a different account. No trainer started.') from exc + queue_name=node.get('queuedResource','').rsplit('/',1)[-1] + if not queue_name: raise RuntimeError('Could not determine the node\'s queued resource') + queue=json.loads(run(['gcloud','alpha','compute','tpus','queued-resources','describe',queue_name,*flags],capture_output=True).stdout) + return lease_from(node,queue,time.time()) + + +def verify_local_host(lease): + """Do not launch on a different VM merely because a data mount exists there.""" + req=urllib.request.Request('http://metadata.google.internal/computeMetadata/v1/instance/network-interfaces/0/ip', + headers={'Metadata-Flavor':'Google'}) + with urllib.request.urlopen(req,timeout=5) as response: + address=response.read().decode().strip() + if address not in lease.get('node_internal_ips',[]): + raise RuntimeError('This VM does not match the checked TPU endpoint; no local launch') + + +def dispatch(command,here): + # --here keeps the cloud lookup in the invoking account, then elevates locally. + # Credentials, IAM policies and the lease requirement are never changed. + if here: + return subprocess.run(command if os.geteuid()!=0 else command[1:]).returncode + return subprocess.run(['gcloud','compute','tpus','tpu-vm','ssh',NODE,'--project='+PROJECT, + '--zone='+ZONE,'--worker=0','--command='+shlex.join(command)]).returncode + + +def active(unit): + r=subprocess.run(['systemctl','show',unit,'--property=ActiveState','--value'],capture_output=True,text=True) + return r.stdout.strip() in ('active','activating','deactivating','reloading') + + +def status_remote(tail_metrics=False): + if not LATEST.exists(): print('No long run launched.'); return + record=json.loads(LATEST.read_text()); root=Path(record['root']) + print(json.dumps(record,indent=2),flush=True) + subprocess.run(['systemctl','--no-pager','--full','status',record['unit']]) + for name in ('RUN_STATUS.json','STARTUP.json','status.json','latest_validation.json', + 'checkpoint_latest.json','TRACKING_STATUS.json','RESUME_PARITY.json'): + if (root/name).exists(): print(name+'\n'+(root/name).read_text(),flush=True) + subprocess.run(['tail','-n','12',str(root/('metrics.jsonl' if tail_metrics else 'run.log'))]) + + +def stop_remote(): + record=json.loads(LATEST.read_text()); root=Path(record['root']) + (root/'STOP').touch() + print('Safe stop requested. Trainer will finish the current update, save full state, then drain tracking/backup.') + print('Watch:',root/'run.log') + + +def start_remote(commit,lease,request_id,resume=None): + if os.geteuid()!=0 or not os.path.ismount('/mnt/disks/rg-data'): + raise RuntimeError('The persistent data disk must already be mounted; root required') + if not re.fullmatch('[0-9a-f]{40}',commit): raise ValueError('Pinned Git commit required') + if resume: + resume=Path(resume).resolve() + if BASE.resolve() not in resume.parents or resume.suffix!='.pt' or not resume.is_file(): + raise RuntimeError('Recovery checkpoint must exist under the persistent experiment directory') + with (BASE/'port-check-launch.lock').open('a') as lock: + fcntl.flock(lock,fcntl.LOCK_EX|fcntl.LOCK_NB) + if LATEST.exists(): + previous=json.loads(LATEST.read_text()) + if previous.get('request_id')==request_id or active(previous['unit']): + print('Existing launch retained; no duplicate or restart.'); status_remote(); return + if not -30 <= time.time()-lease['checked_unix'] <= 600: + raise RuntimeError('Lease check is stale; rerun the launcher') + if lease['node']!=NODE or lease['termination_unix']-time.time()<=MIN_REMAINING: + raise RuntimeError('Need more than 12.5 hours remaining on the checked node') + if shutil.disk_usage(BASE).free < 45*1024**3: + raise RuntimeError('Need 45 GiB free for remaining shards, spectra and retained checkpoints; nothing deleted') + stamp=dt.datetime.now(dt.timezone.utc).strftime('%Y%m%d-%H%M%S') + root=BASE/('muon-long25k-s1337-'+stamp); root.mkdir(); repo=root/'repo'; repo.mkdir() + run(['git','-C',str(repo),'init','-q']) + run(['git','-C',str(repo),'remote','add','origin','https://github.com/CalculatedContent/rg_optimizers.git']) + run(['git','-C',str(repo),'fetch','--depth','1','origin',commit],timeout=180) + run(['git','-C',str(repo),'checkout','--detach',commit]) + scripts=repo/'baseline/gpt2_small' + spec=importlib.util.spec_from_file_location('training_guard',scripts/'scripts/run_muonclip.py') + guard=importlib.util.module_from_spec(spec); spec.loader.exec_module(guard); guard.assert_idle() + # Verify reference BEFORE launching a service; never change that directory. + sys.path.insert(0,str(scripts/'muon_longrun')) + from common import verify_reference, REFERENCE_NAME, atomic_json + approved=verify_reference(BASE/REFERENCE_NAME) + deadline=time.time()+12*3600 + if lease['termination_unix']-deadline<1800: + raise RuntimeError('Less than 30 minutes lease margin after checkout; no service started') + unit='rg-muon-long25k-'+stamp+'.service' + record={'root':str(root),'unit':unit,'commit':commit,'request_id':request_id, + 'node':NODE,'started_unix':time.time(),'service_deadline_unix':deadline, + 'training_deadline_unix':deadline-1800,'steps':25000,'tokens':13107200000, + 'optimizer':'muon','config':approved['manifest']['config'], + 'batch_tokens':524288,'global_microbatch_sequences':64,'accumulation':8, + 'peak_lrs':{'muon':.04,'adam_embedding':.6,'adam_head':.008,'adam_scalar':.04}, + 'muon_momentum':{'initial':.85,'final':.95,'ramp_updates':500}, + 'newton_schulz':{'steps':5,'coefficients':[3.4445,-4.7750,2.0315]}, + 'adam':{'betas':[.9,.95],'eps':1e-8,'weight_decay':0}, + 'gradient_clipping':False,'qk_clipping':False, + 'scheduler':{'warmup_updates':0,'cooldown_start':17500,'warmdown_updates':7500,'final_step':25000}, + 'reference':approved['reference_run'],'lease':lease,'fresh_initialization':resume is None, + 'resume_source':str(resume) if resume else None, + 'automatic_restart':False,'cloud_uri':'gs://tpu-builders-504820-ww-continuous8/gpt2small/'+root.name} + atomic_json(root/'launch.json',record); atomic_json(root/'LEASE.json',lease) + (root/'commit.txt').write_text(commit+'\n') + env={'PYTHONPATH':str(scripts/'src')+':'+str(scripts.parent/'nanogpt_one_head/src'), + 'PJRT_DEVICE':'TPU','TPU_ACCELERATOR_TYPE':'v5litepod-8', + 'OMP_NUM_THREADS':'4','OPENBLAS_NUM_THREADS':'4','MKL_NUM_THREADS':'4', + 'TOKENIZERS_PARALLELISM':'false'} + command=['systemd-run','--unit='+unit,'--property=Type=exec','--property=Restart=no', + '--property=RuntimeMaxSec=43200','--property=TimeoutStopSec=15', + '--property=KillMode=control-group','--property=StandardOutput=append:'+str(root/'run.log'), + '--property=StandardError=append:'+str(root/'run.log')] + command+=['--setenv='+k+'='+v for k,v in env.items()] + command+=['/mnt/disks/rg-data/continuous8/venv/bin/python','-u', + str(scripts/'muon_longrun/long_worker.py'),str(root),str(deadline)] + if resume: command+=['--resume',str(resume)] + atomic_json(LATEST,record) # Publish request ID before systemd to make SSH retry idempotent. + run(command) + print(json.dumps(record,indent=2),flush=True) + run(['systemctl','show',unit,'--property=MainPID','--property=ActiveState']) + print('Muon run: 25,000 total steps, 13.1072B tokens; no target-loss stop.',flush=True) + print('Peak LR through 17,500; linear warmdown over final 7,500; no LR warmup.',flush=True) + print('Scalars every 10; full validation every 500 + spectral steps.',flush=True) + print('WW: 0,100,250,500,750,1000,1500,2000,2500,3000; then 1000, plus 17500/final.',flush=True) + print('Checkpoints every 2500; latest two rolling + permanent 0/3000/10000/17500/25000.',flush=True) + print('Estimated ~10–11 hours; 12-hour service cap includes preparation and final backup.',flush=True) + for action in ('status','metrics','stop'): + print(f'From your local checkout: python3 baseline/gpt2_small/muon_longrun/launch.py {action}',flush=True) + + +def main(): + p=argparse.ArgumentParser(); p.add_argument('action',choices=('start','recover','status','metrics','stop')) + p.add_argument('--checkpoint',type=Path,help='Explicit recovery from this long plan only; never used by start') + p.add_argument('--here',action='store_true',help='Run directly on this TPU; start still requires cloud read access') + p.add_argument('--on-tpu',action='store_true',help=argparse.SUPPRESS) + p.add_argument('--commit',help=argparse.SUPPRESS); p.add_argument('--lease',help=argparse.SUPPRESS) + p.add_argument('--request-id',help=argparse.SUPPRESS); a=p.parse_args() + if (a.action=='recover') != bool(a.checkpoint): p.error('Only recover requires --checkpoint') + if a.on_tpu: + if a.action in ('start','recover'): start_remote(a.commit,json.loads(a.lease),a.request_id,a.checkpoint) + elif a.action=='stop': stop_remote() + else: status_remote(a.action=='metrics') + return 0 + command=['sudo','python3','-c',Path(__file__).read_text(),a.action,'--on-tpu'] + if a.checkpoint: command+=['--checkpoint',str(a.checkpoint)] + if a.action in ('start','recover'): + repo=Path(__file__).resolve().parents[3] + if run(['git','-C',str(repo),'status','--porcelain'],capture_output=True).stdout.strip(): + raise RuntimeError('Use a clean checkout of the pushed commit') + commit=run(['git','-C',str(repo),'rev-parse','HEAD'],capture_output=True).stdout.strip() + lease=live_lease() + if a.here: verify_local_host(lease) + command+=['--commit',commit,'--lease',json.dumps(lease),'--request-id',uuid.uuid4().hex] + return dispatch(command,a.here) + +if __name__=='__main__': raise SystemExit(main()) diff --git a/baseline/gpt2_small/muon_longrun/long_data.py b/baseline/gpt2_small/muon_longrun/long_data.py new file mode 100644 index 00000000..197b611c --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/long_data.py @@ -0,0 +1,56 @@ +"""The same sequential stream, with explicit cycles and recoverable cursor.""" +from concurrent.futures import ThreadPoolExecutor +import json +from pathlib import Path +from common import SHORT, CACHE, MICROBATCH, CONTEXT, atomic_json, sha +from data import FineWeb, reference + + +class Stream(reference.TrainStream): + def __init__(self, source, batch=MICROBATCH, context=CONTEXT): + super().__init__(source,batch,context) + self.cycles=0 + + def next_batch(self): + wrap = (self.shard == len(self.names)-1 and + self.position+self.batch*self.context+1 > len(self.tokens)) + result=super().next_batch() + if wrap: + self.cycles+=1 + return result + + def state_dict(self): + return dict(shard=self.shard,position=self.position,cycles=self.cycles, + batch=self.batch,context=self.context,names=self.names) + + def load_state_dict(self, state): + if state['names']!=self.names or state['batch']!=self.batch or state['context']!=self.context: + raise RuntimeError('Resume data ordering/batch mismatch') + shard,position=int(state['shard']),int(state['position']) + if not 0 <= shard < len(self.names): raise ValueError('Invalid shard cursor') + tokens=self.source.array(self.names[shard]) + if not 0 <= position < len(tokens) or position % (self.batch*self.context): + raise ValueError('Invalid within-shard cursor') + self.shard,self.position,self.cycles,self.tokens=shard,position,int(state['cycles']),tokens + + +def corpus_metadata(source): + files={k:v for k,v in source.manifest['files'].items() if '_train_' in k} + count=MICROBATCH*CONTEXT + unique=sum((v['size']-1024)//2 for v in files.values()) + usable=sum((((v['size']-1024)//2-1)//count)*count for v in files.values()) + return dict(repo=source.manifest['repo'],revision=source.manifest['revision'], + train_shards=len(files),corpus_tokens=unique,usable_tokens_per_epoch=usable, + epoch_definition='tokens_seen / usable_tokens_per_full_sequential_corpus_pass', + ordering='unchanged sequential shard traversal; wraps to first shard after a full pass', + manifest_sha256=sha(SHORT.parent/'speedrun30/data_manifest.json')) + + +def prepare(root, deadline): + source=FineWeb(CACHE,deadline) + names=sorted(source.manifest['files']) + with ThreadPoolExecutor(max_workers=4) as pool: + for name,_ in zip(names,pool.map(source.array,names)): + print('Verified benchmark shard:',name,flush=True) + atomic_json(root/'data_receipts.json',source.receipts) + atomic_json(root/'corpus.json',corpus_metadata(source)) diff --git a/baseline/gpt2_small/muon_longrun/long_worker.py b/baseline/gpt2_small/muon_longrun/long_worker.py new file mode 100644 index 00000000..77fb4079 --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/long_worker.py @@ -0,0 +1,81 @@ +"""Bounded setup, one continuous trainer, async CPU WW, persistent scientific output.""" +import argparse +import json +import os +from pathlib import Path +import signal +import subprocess +import sys +import time +from common import HERE, SHORT, REFERENCE_NAME, atomic_json, verify_reference +from worker import bounded, finish_tracking + + +def cloud_backup(root): + from rg_nanogpt_one_head.continuous_support import CloudPublisher + publisher=CloudPublisher('gs://tpu-builders-504820-ww-continuous8/gpt2small/'+root.name) + receipts=[] + for folder in (root,root/'tracking'): + files=folder.iterdir() if folder==root else folder.rglob('*') + for path in sorted(files): + if path.is_file() and path.suffix in ('.json','.jsonl','.csv','.log','.txt'): + publisher.snapshot_text_file(path,path.relative_to(root).as_posix()) + for path in sorted((root/'checkpoints').glob('step_*.pt')): + receipts.append(publisher.file(path,path.relative_to(root).as_posix())) + result={'status':'verified','checkpoints':receipts,'unix_time':time.time()} + publisher.json(result,'CLOUD_BACKUP_VERIFIED.json'); atomic_json(root/'CLOUD_BACKUP_VERIFIED.json',result) + + +def main(): + p=argparse.ArgumentParser(); p.add_argument('root',type=Path); p.add_argument('deadline',type=float) + p.add_argument('--backup-only',action='store_true'); p.add_argument('--resume',type=Path) + a=p.parse_args(); root=a.root + if a.backup_only: cloud_backup(root); return 0 + train_deadline=a.deadline-1800 + state={'status':'preparing','automatic_restart':False,'deadline_unix':a.deadline, + 'training_deadline_unix':train_deadline,'steps':25000} + atomic_json(root/'RUN_STATUS.json',state); tracker=None + def phase(command,seconds,label): + if (root/'STOP').exists(): raise RuntimeError('Stop requested during setup') + result=bounded(command,min(seconds,train_deadline-time.time()),root,label) + if result['exit_code']!=0: raise RuntimeError(label+' failed: '+str(result)) + try: + approved=verify_reference(root.parent/REFERENCE_NAME) + logs=root.parent/REFERENCE_NAME/'metrics.jsonl' + if logs.exists(): + import statistics + durations=[r['seconds'] for line in logs.read_text().splitlines() + if (r:=json.loads(line)).get('kind')=='train' and r.get('step',0)>=100] + if durations: approved['median_update_seconds']=statistics.median(durations) + atomic_json(root/'REFERENCE.json',approved) + phase([sys.executable,str(SHORT/'pallas_dependencies.py'),str(root)],600,'pinned Pallas overlay') + os.environ['PYTHONPATH']=str(root/'pallas-deps')+os.pathsep+os.environ.get('PYTHONPATH','') + phase([sys.executable,'-u',str(HERE/'train_long.py'),'prepare','--root',str(root), + '--deadline',str(min(time.time()+1800,train_deadline-600))],1800,'verify full benchmark corpus') + phase([sys.executable,'-u',str(SHORT/'run.py'),'attention-check','--legacy-attention-check','--root',str(root), + '--microbatch','64','--deadline',str(time.time()+300)],300,'8-chip flash attention check') + env={**os.environ,'PJRT_DEVICE':'CPU','CUDA_VISIBLE_DEVICES':'', + 'OMP_NUM_THREADS':'1','OPENBLAS_NUM_THREADS':'1','MKL_NUM_THREADS':'1'} + tracker=subprocess.Popen([sys.executable,'-u',str(HERE/'track_long.py'),str(root), + str(a.deadline-1200)],env=env,start_new_session=True) + state['status']='training'; atomic_json(root/'RUN_STATUS.json',state) + command=[sys.executable,'-u',str(HERE/'train_long.py'),'train','--root',str(root), + '--deadline',str(train_deadline)] + if a.resume: command+=['--resume',str(a.resume)] + result=bounded(command,train_deadline-time.time(),root,'25,000 Muon updates',watch=True) + state.update(result) + if result['exit_code']!=0: raise RuntimeError('Continuous trainer failed; see FAILURE.json/run.log') + state.update(json.loads((root/'status.json').read_text())) + except Exception as exc: + state.update(status='failed',error=repr(exc)) + finally: + if tracker is not None: + state['tracking']=finish_tracking(tracker,root,min(time.time()+600,a.deadline-1200)) + atomic_json(root/'RUN_STATUS.json',state) + backup=bounded([sys.executable,'-u',__file__,str(root),str(a.deadline),'--backup-only'], + max(0,a.deadline-time.time()-20),root,'retained results cloud backup') + state['backup']=backup; atomic_json(root/'RUN_STATUS.json',state) + print(json.dumps(state),flush=True) + return 0 if state.get('status')=='schedule_complete' and backup['exit_code']==0 and state.get('tracking',{}).get('status')=='complete' else 1 + +if __name__=='__main__': raise SystemExit(main()) diff --git a/baseline/gpt2_small/muon_longrun/track_long.py b/baseline/gpt2_small/muon_longrun/track_long.py new file mode 100644 index 00000000..124df4f0 --- /dev/null +++ b/baseline/gpt2_small/muon_longrun/track_long.py @@ -0,0 +1,92 @@ +"""CPU spectral measurements; preserve explicit zero/unfit matrices at initialization.""" +import argparse +import importlib.metadata +import json +from pathlib import Path +import random +import time +from common import atomic_json, sha, EARLY_WW, PERMANENT +import tracking + + +def measure(path): + import numpy as np + import torch + import weightwatcher as ww + torch.set_num_threads(1) + began=time.monotonic(); payload=torch.load(path,map_location='cpu',weights_only=False) + seed=1001340+payload['step']; random.seed(seed); np.random.seed(seed); torch.manual_seed(seed) + names=list(payload['matrices']) + nonzero={n:v for n,v in payload['matrices'].items() if bool(torch.count_nonzero(v))} + identity={k:payload['validation'].get(k) for k in + ('evaluation_tokens','full_benchmark_evaluation','val_nll','val_perplexity', + 'val_token_error','val_accuracy','val_error_count','epoch','lr_factor', + 'scheduler_phase','muon_lr','adam_embedding_lr','adam_head_lr','adam_scalar_lr')} + identity.update(step=payload['step'],tokens_seen=payload['tokens_seen'],run_id=payload['run_id'], + snapshot_sha256=sha(path),diagnostic_seed=seed, + weightwatcher_version=importlib.metadata.version('weightwatcher')) + frame=ww.WeightWatcher(model=tracking.holder_from(nonzero)).analyze(**tracking.WW_OPTIONS) + if not {'alpha','raw_alpha'}.issubset(frame.columns): + raise RuntimeError('WeightWatcher raw/clipped fields missing') + rows=tracking.normalize_rows(frame,names,identity) + for row in rows: + name=row['matrix_name'] + if name not in nonzero: + row.update(status='zero_matrix',matrix_rank=0, + randomized_status='degenerate_zero_matrix') + else: + row['randomized_status']='available' if row.get('max_rand_eval') is not None else 'unavailable' + # Keep all WW scalar fields, plus explicit missing values where a fit is undefined. + for key in ('matrix_rank','alpha_weighted','log_alpha_norm','max_rand_eval', + 'rand_distance','rand_mp_softrank','rand_num_spikes','num_fingers'): + row.setdefault(key,None) + row['weighted_metric_alpha_source']='clipped alpha (WeightWatcher native definition)' + result=tracking.summary(rows,identity) + result.update(weightwatcher_seconds=time.monotonic()-began, + zero_matrix_count=len(names)-len(nonzero)) + return {'layers':rows,'summary':result,'options':tracking.WW_OPTIONS} + + +def configure(root): + version=importlib.metadata.version('weightwatcher') + if version!='0.7.7': raise RuntimeError('Expected the existing weightwatcher==0.7.7') + atomic_json(root/'TRACKING_CONFIG.json',dict(weightwatcher_version=version, + options=tracking.WW_OPTIONS,matrix_count=72,matrix_roles=list(tracking.ROLES.values()), + early_steps=sorted(EARLY_WW),later_interval=1000,always_steps=sorted(PERMANENT), + overhead_fallback_interval=2000,execution='separate CPU process, one BLAS thread', + raw_alpha_source='raw_alpha',clipped_alpha_source='alpha', + zero_matrix_policy='Keep row, rank=0, undefined alpha/null fits explicitly unavailable', + uncertainty='Across-matrix standard deviation is not uncertainty across seeds', + token_error='Validation teacher-forced top-1 error, same tokens and weights as NLL', + randomization='WeightWatcher randomize=True, all returned null/ERG fields retained')) + + +def watch(root,deadline): + configure(root); folder=root/'tracking' + while time.time()