Skip to content

test(rocm): verify the unified engine on gfx1151 after the engine epic - #2250

Merged
inureyes merged 1 commit into
mainfrom
test/issue-2192-rocm-unified-engine
Oct 8, 2026
Merged

inureyes merged 1 commit into
mainfrom
test/issue-2192-rocm-unified-engine

Conversation

@inureyes

@inureyes inureyes commented Oct 8, 2026

Copy link
Copy Markdown
Member

Summary

ROCm verification of the unified engine (Phase 7 of epic #2166) on the gfx1151 host: the gate, every named ROCm test, where each ROCm behavior now lives on the engine path, the ROCm parity measurement and definition, and the engine's throughput against the pre-epic CxxGenerator on the same host, same day, interleaved with a null arm. Results page: docs/benchmark_results/rocm-unified-engine-gfx1151-2026-10-08.md, with the drivers and raw records under docs/benchmark_results/data/rocm-unified-engine-gfx1151-2026-10-08/.

Verification

On gfx1151 (ROCm 10.0.0, HIP 7.15), branch rebased onto main at 1b4e3657: make verify-rocm with MLXCEL_ROCM_SMOKE_MODEL=models/mlx/Qwen3-0.6B-4bit: 12194 passed, 0 failed, 403 ignored across 163 test binaries, [verify-rocm] OK (prerequisites: versions, dtype keys, kernel-port dispatch, llama-compat, overlay records, Python tooling, fmt, clippy with -D warnings, the GPU smoke). The named tests, the parity sessions and the benchmark rounds are logged in the results page and its data directory (every GPU run under scripts/rocm_gpu_guard.sh, all benchmark runs CLEAN).

Not verified: Metal and CUDA (not available on this host). Nothing on those paths changes; the only code change is a comment in src/models/llama3.rs.

Closes #2192

…ve engine epic

Phase 7 of epic #2166, run on the Radeon 8060S (gfx1151) host after Phases 0 to 6 merged on a CUDA machine. Every ROCm-only test and every test that skips without a kernel port ran here for the first time against the merged engine, and the results page records where each ROCm behavior now lives on the engine path, the ROCm parity measurement, and the engine's throughput against the pre-epic CxxGenerator on the same host and day.

- docs/benchmark_results/rocm-unified-engine-gfx1151-2026-10-08.md: the gate, the named tests (all ran, none skipped, including the 20 GiB u32 pool-offset test under --ignored), the post-epic location of each ROCm behavior, parity (every CLI-vs-server pair identical on Qwen3-0.6B, Llama-3.1-8B and granite-4.0-h-tiny; dense vs paged identical for greedy over 64 tokens and 0 decided-position mismatches over 128 teacher-forced decode steps with the HIP paged v2 kernel in the loop; the seeded stream diverges at token 6 and 18 with the paged kernel), and 11 interleaved decode cells against 4c44e31 with a null arm: ten within ADR 0007's 1.0 percent threshold, gemma-3-4b-it-4bit at 2048 tokens 1.2 percent below, isolated to the B=1 lookahead pipeline on the rotating cache (filed as #2239).
- Found on both builds and filed separately: affine MoE prefill lost the expert-batched kernel when #2137 cleared right_sorted on the sorted path (granite 573 vs 919 tok/s, Mixtral 26 vs 126), cross-backend (#2240).
- ADR 0007 and docs/CONTINUOUS_BATCHING.md record the ROCm parity definition applied and the measured result; docs/installation.md points at the page.
- The drivers and records are under docs/benchmark_results/data/rocm-unified-engine-gfx1151-2026-10-08/ (per-run guard logs, one lock hold for the whole session, a teacher-forced dense-vs-paged trace through the server's /completion n_probs, and a rounds summarizer).
- src/models/llama3.rs: the fused RoPE-append comment named update_and_fetch as the consumer; since #2171 it is cache.attend on either storage.

Verification: the named tests and the full gate on gfx1151 are listed in the page and the PR. Metal and CUDA are not available on this host; nothing on those paths changes (one comment in llama3.rs).

Refs #2192, #2166.
@inureyes inureyes added type:test Test related changes priority:high High priority area:inference Generation, sampling, decoding (incl. speculative, DRY) platform:linux Linux (CUDA / packaging) specific status:done Completed labels Oct 8, 2026
@inureyes
inureyes merged commit 7a3fcc4 into main Oct 8, 2026
26 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:inference Generation, sampling, decoding (incl. speculative, DRY) platform:linux Linux (CUDA / packaging) specific priority:high High priority status:done Completed type:test Test related changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

test(rocm): verify the unified engine on ROCm after the batch-native engine epic

1 participant