Repository navigation
perf(rocm): add host-gap accounting per port unit to the decode profile - #2238
Merged
Merged
Conversation
A port unit's share of decode GPU time implies a ceiling of 1 / (1 - share) that leaves out the host launch gaps between dispatches, so the profile underrated ports that replace many small dispatches: the #2067 SSM port measured 1.46x and 1.45x against GPU-share ceilings of 1.42x and 1.25x. scripts/rocm_decode_gaps.py charges each idle stretch of the decode window to the dispatch that ends it. summarize() now writes role_host_gap_ms_per_token and, per port unit (fallback and reached-default roles), host gap per token and per dispatch, wall share, ceiling_gpu_share, ceiling_wall and plain_ceiling_wall_est (gaps scaled to the plain run). A residual check pins the per-role gaps to the window's idle time, and report() prints a ceiling table. Existing fields are unchanged. Validated on gfx1151 at 6dbfe7d with MLXCEL_SSM_KERNEL=0: plain wall ceilings 1.60x (granite) and 1.56x (Nemotron-H) against measured medians of 1.49x and 1.44x; GPU-share ceilings 1.45x and 1.26x. Results in benchmarks/rocm_profiles/gfx1151_6dbfe7d7_ssm-graph/, a dated note on the 2026-09-30 profile page, and docs/benchmarks.md describes the new fields. Closes #2148
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The decode profile ranked ports by their share of decode GPU time, whose ceiling
1 / (1 - share)ignores the host launch gaps between dispatches. That underrated dispatch-heavy ports: #2067 measured 1.46x and 1.45x against GPU-share ceilings of 1.42x and 1.25x. The profile now charges each idle stretch of the decode window to the dispatch that ends it and gives every port unit a wall-time share and ceiling that count its launch gaps.scripts/rocm_decode_gaps.py(new, split out becauserocm_decode_profile.pywas already over 500 lines):attribute_gaps, the ceilings, and the scaling of traced gaps to the plain run.summarize()addsrole_host_gap_ms_per_tokenand, per unit for the fallback and reached-default roles, host gap per token and per dispatch, wall share,ceiling_gpu_share,ceiling_wallandplain_ceiling_wall_est. Ahost_gap_attribution_residual_nscheck (expected 0) pins the per-role gaps to the window's idle time. Existing fields are unchanged;report()adds a ceiling table and shows-for summaries written before this change.docs/benchmarks.mddescribes the fields and namesplain_ceiling_wall_estas the bound to compare a measured speedup against; the 2026-09-30 profile page gets a dated note with the new ceilings.mlxcel-bench-decodeoutput and phase marks (both profiled runs cut cleanly: 0 kernels straddle the decode start), so no parser fix was needed.Validation of the method (gfx1151,
6dbfe7d7,MLXCEL_SSM_KERNEL=0)Report output (
python3 scripts/rocm_decode_profile.py report benchmarks/rocm_profiles/gfx1151_6dbfe7d7_ssm-graph), ceiling table:NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy(plain)granite-4.0-h-tiny-4bit_greedy(plain)Measured #2067 speedup on the same binary, median of three alternating runs per arm in one guard session per model: granite 59.88 to 89.40 tok/s, 1.49x; Nemotron-H 52.29 to 75.26 tok/s, 1.44x. Spread under 2% except one granite kernel-off run at 52.18 (median unaffected). Both
plain_ceiling_wall_estvalues (1.60x, 1.56x) are at or above the measured gains, and both GPU-share ceilings (1.45x, 1.26x) are below them. Today's measured gains differ slightly from PR #2099's 1.46x and 1.45x because main has moved since; the comparison uses gains and ceilings from the same commit. Raw results,ab_decode.txtand both guard logs (every attempt CLEAN) are inbenchmarks/rocm_profiles/gfx1151_6dbfe7d7_ssm-graph/.Verification
a18d3d76:make verify-rocmOK, 163 cargo test suites, 12194 passed, 0 failed, 403 ignored, plus 100 Python tooling tests (including the new ones);cargo test --features rocm --test dead_doc_pointerspassed.make verify-versions verify-kernel-dtype-keys verify-kernel-port-dispatch verify-llama-compat verify-fmt verify-python-toolingpassed.Not verified: Metal and CUDA (not available on this host). The change touches only the ROCm profiling scripts, their tests and docs; no Rust or backend code.
Closes #2148