Skip to content

docs(rocm): explain the granite-4.0-h-tiny w1 perplexity gap - #2240

Merged
inureyes merged 1 commit into
mainfrom
fix/issue-2154-granite-w1-ppl
Oct 8, 2026
Merged

inureyes merged 1 commit into
mainfrom
fix/issue-2154-granite-w1-ppl

Conversation

@inureyes

@inureyes inureyes commented Oct 8, 2026

Copy link
Copy Markdown
Member

Summary

Closes the granite-4.0-h-tiny w1 investigation with a documented conclusion: the +4.345% perplexity against Metal is not a ROCm defect. No code changes.

  • The gap still exists at a18d3d7: the w1 trace is byte-identical to c5fe9a16 and 96cbce84, with the HIP SSM update kernel on or off (a w1 chunk never has SSM state).
  • No Metal host here, and MLXCEL_DEVICE=cpu is unusable as a bf16 reference (top-1 logit 14.50 vs 18.875 on the GPU, about 400 s per chunk). The reference is the same forward with f32 activations; GPU and CPU stream agree to 0.0007 nats.
  • Per op, ROCm bf16 against f32 on the same inputs: every stage is 0.0014 to 0.0038 relative L2 (one bf16 rounding is 0.0017). None stands out.
  • Both backends sharpen the logits against f32 by a similar amount in the 128 positions Metal covers (spread +0.38 ROCm, +0.31 Metal). A mixed-precision sweep isolates that to bf16 rounding of the residual stream; no op family comes close.
  • Over 1024 positions ROCm's NLL is within noise of f32 (+0.011 nats, 0.6 clustered SE). The first 128-position window is the largest of eight (window SD 0.051 against a within-window SE of 0.025), and only 90 of its 128 inputs are distinct, so the 2.5 SE overstated the evidence.

Adds benchmarks/logit_traces/rocm_gfx1151_a18d3d76/ (traces, per-op table, probe source, METADATA, RUNS, SHA256SUMS) and a section in docs/benchmark_results/rocm-correctness-gfx1151-2026-09-30.md.

Verification

On gfx1151, every GPU run holding the host GPU lock: logit_trace w1 (128 and 1024 chunks), w1 with MLXCEL_SSM_KERNEL=0, w8, w256; the probe's f32, per-op and mixed-precision arms. w8 and w256 at a18d3d7 still pass against the Metal traces at --decided 2.0. Fast gates (verify-versions, verify-kernel-dtype-keys, verify-kernel-port-dispatch, verify-llama-compat, verify-fmt), dead_doc_pointers and check_binary_assets.py pass. make verify-rocm was not run: the change is docs and trace data only.

Not verified: Metal and CUDA (not available on this host). No 1024-position Metal trace or Metal per-op numbers could be produced; the remaining 0.042-nat difference between the backends is not attributed to any one op.

Closes #2154

The 2026-09-30 correctness matrix left granite-4.0-h-tiny at w1 (+4.345% perplexity against Metal, 2.5 SE) as weak evidence of a ROCm bias. Re-measured at a18d3d7 the w1 trace is byte-identical to the earlier ones, with the HIP SSM update kernel on or off, so the gap still exists.

Measured against an f32-activation reference (device-independent: GPU and CPU stream agree to 0.0007 nats), every op on ROCm's single-token path is within one to two and a half bf16 roundings of f32, and over 1024 positions ROCm's NLL is within noise of f32 (+0.011 nats, 0.6 clustered SE). Both backends sharpen the logits against f32 by a similar amount (+0.38 and +0.31 spread in the Metal window), and a mixed-precision sweep isolates that to rounding the residual stream to bf16. The 128-position window the Metal trace covers is the one where that rounding costs the most NLL (window SD 0.051 against a within-window SE of 0.025), and only 90 of its 128 inputs are distinct.

Adds benchmarks/logit_traces/rocm_gfx1151_a18d3d76/ (traces, per-op table, probe source, METADATA, RUNS, SHA256SUMS) and a results-page section. The CPU stream is recorded as unusable as a bf16 reference on this host. No code change.

Closes #2154
@inureyes inureyes added status:done Completed type:docs Documentation improvements or additions priority:low Low priority area:models Model architectures, weights, loading, metadata platform:linux Linux (CUDA / packaging) specific labels Oct 8, 2026
@inureyes
inureyes merged commit 1b4e365 into main Oct 8, 2026
27 checks passed
@inureyes
inureyes deleted the fix/issue-2154-granite-w1-ppl branch October 8, 2026 20:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:models Model architectures, weights, loading, metadata platform:linux Linux (CUDA / packaging) specific priority:low Low priority status:done Completed type:docs Documentation improvements or additions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(rocm): explain the granite-4.0-h-tiny w1 perplexity gap against Metal

1 participant