Repository navigation
test(rocm): verify the unified engine on gfx1151 after the engine epic - #2250
Merged
Merged
Conversation
…ve engine epic Phase 7 of epic #2166, run on the Radeon 8060S (gfx1151) host after Phases 0 to 6 merged on a CUDA machine. Every ROCm-only test and every test that skips without a kernel port ran here for the first time against the merged engine, and the results page records where each ROCm behavior now lives on the engine path, the ROCm parity measurement, and the engine's throughput against the pre-epic CxxGenerator on the same host and day. - docs/benchmark_results/rocm-unified-engine-gfx1151-2026-10-08.md: the gate, the named tests (all ran, none skipped, including the 20 GiB u32 pool-offset test under --ignored), the post-epic location of each ROCm behavior, parity (every CLI-vs-server pair identical on Qwen3-0.6B, Llama-3.1-8B and granite-4.0-h-tiny; dense vs paged identical for greedy over 64 tokens and 0 decided-position mismatches over 128 teacher-forced decode steps with the HIP paged v2 kernel in the loop; the seeded stream diverges at token 6 and 18 with the paged kernel), and 11 interleaved decode cells against 4c44e31 with a null arm: ten within ADR 0007's 1.0 percent threshold, gemma-3-4b-it-4bit at 2048 tokens 1.2 percent below, isolated to the B=1 lookahead pipeline on the rotating cache (filed as #2239). - Found on both builds and filed separately: affine MoE prefill lost the expert-batched kernel when #2137 cleared right_sorted on the sorted path (granite 573 vs 919 tok/s, Mixtral 26 vs 126), cross-backend (#2240). - ADR 0007 and docs/CONTINUOUS_BATCHING.md record the ROCm parity definition applied and the measured result; docs/installation.md points at the page. - The drivers and records are under docs/benchmark_results/data/rocm-unified-engine-gfx1151-2026-10-08/ (per-run guard logs, one lock hold for the whole session, a teacher-forced dense-vs-paged trace through the server's /completion n_probs, and a rounds summarizer). - src/models/llama3.rs: the fused RoPE-append comment named update_and_fetch as the consumer; since #2171 it is cache.attend on either storage. Verification: the named tests and the full gate on gfx1151 are listed in the page and the PR. Metal and CUDA are not available on this host; nothing on those paths changes (one comment in llama3.rs). Refs #2192, #2166.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ROCm verification of the unified engine (Phase 7 of epic #2166) on the gfx1151 host: the gate, every named ROCm test, where each ROCm behavior now lives on the engine path, the ROCm parity measurement and definition, and the engine's throughput against the pre-epic
CxxGeneratoron the same host, same day, interleaved with a null arm. Results page:docs/benchmark_results/rocm-unified-engine-gfx1151-2026-10-08.md, with the drivers and raw records underdocs/benchmark_results/data/rocm-unified-engine-gfx1151-2026-10-08/.paged_pool_past_u32_elements_matches_gatherunder--ignoredin a guarded idle window. The two CUDA-failingssm_update_parity_testscases pass on gfx1151.docs/CONTINUOUS_BATCHING.md.mlxcel run, chat) shows no gap on ROCm.right_sortedon the sorted path (granite 573 vs 919 tok/s, Mixtral 26 vs 126; cross-backend by the same flag on Metal and CUDA): filed as fix(moe): sorted gather_qmm prefill lost right_sorted since #2137, MoE prefill 1.7x to 4.8x slower #2241.src/models/llama3.rs: one stale comment corrected (the fused RoPE-append output is consumed bycache.attendsince refactor: one KV storage abstraction with dense and paged backends #2171).Verification
On gfx1151 (ROCm 10.0.0, HIP 7.15), branch rebased onto
mainat1b4e3657:make verify-rocmwithMLXCEL_ROCM_SMOKE_MODEL=models/mlx/Qwen3-0.6B-4bit: 12194 passed, 0 failed, 403 ignored across 163 test binaries,[verify-rocm] OK(prerequisites: versions, dtype keys, kernel-port dispatch, llama-compat, overlay records, Python tooling, fmt, clippy with-D warnings, the GPU smoke). The named tests, the parity sessions and the benchmark rounds are logged in the results page and its data directory (every GPU run underscripts/rocm_gpu_guard.sh, all benchmark runs CLEAN).Not verified: Metal and CUDA (not available on this host). Nothing on those paths changes; the only code change is a comment in
src/models/llama3.rs.Closes #2192