Skip to content

perf(rocm): add host-gap accounting per port unit to the decode profile - #2238

Merged
inureyes merged 1 commit into
mainfrom
perf/issue-2148-profile-host-gaps
Oct 8, 2026
Merged

inureyes merged 1 commit into
mainfrom
perf/issue-2148-profile-host-gaps

Conversation

@inureyes

@inureyes inureyes commented Oct 8, 2026

Copy link
Copy Markdown
Member

Summary

The decode profile ranked ports by their share of decode GPU time, whose ceiling 1 / (1 - share) ignores the host launch gaps between dispatches. That underrated dispatch-heavy ports: #2067 measured 1.46x and 1.45x against GPU-share ceilings of 1.42x and 1.25x. The profile now charges each idle stretch of the decode window to the dispatch that ends it and gives every port unit a wall-time share and ceiling that count its launch gaps.

  • scripts/rocm_decode_gaps.py (new, split out because rocm_decode_profile.py was already over 500 lines): attribute_gaps, the ceilings, and the scaling of traced gaps to the plain run.
  • summarize() adds role_host_gap_ms_per_token and, per unit for the fallback and reached-default roles, host gap per token and per dispatch, wall share, ceiling_gpu_share, ceiling_wall and plain_ceiling_wall_est. A host_gap_attribution_residual_ns check (expected 0) pins the per-role gaps to the window's idle time. Existing fields are unchanged; report() adds a ceiling table and shows - for summaries written before this change.
  • Docs: docs/benchmarks.md describes the fields and names plain_ceiling_wall_est as the bound to compare a measured speedup against; the 2026-09-30 profile page gets a dated note with the new ceilings.
  • The parser still reads the post-epic: one batch-native engine for CLI and server decode #2166 mlxcel-bench-decode output and phase marks (both profiled runs cut cleanly: 0 kernels straddle the decode start), so no parser fix was needed.

Validation of the method (gfx1151, 6dbfe7d7, MLXCEL_SSM_KERNEL=0)

Report output (python3 scripts/rocm_decode_profile.py report benchmarks/rocm_profiles/gfx1151_6dbfe7d7_ssm-graph), ceiling table:

Run: ceiling wall (GPU share) #2063 #2064 #2065 #2067 #2068
NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy (plain) 1.00 (1.00) 1.00 (1.00) 1.26 (1.37) 1.56 (1.26) 1.00 (1.00)
granite-4.0-h-tiny-4bit_greedy (plain) 1.00 (1.00) 1.00 (1.00) 1.25 (1.29) 1.60 (1.45) 1.00 (1.00)

Measured #2067 speedup on the same binary, median of three alternating runs per arm in one guard session per model: granite 59.88 to 89.40 tok/s, 1.49x; Nemotron-H 52.29 to 75.26 tok/s, 1.44x. Spread under 2% except one granite kernel-off run at 52.18 (median unaffected). Both plain_ceiling_wall_est values (1.60x, 1.56x) are at or above the measured gains, and both GPU-share ceilings (1.45x, 1.26x) are below them. Today's measured gains differ slightly from PR #2099's 1.46x and 1.45x because main has moved since; the comparison uses gains and ceilings from the same commit. Raw results, ab_decode.txt and both guard logs (every attempt CLEAN) are in benchmarks/rocm_profiles/gfx1151_6dbfe7d7_ssm-graph/.

Verification

  • On gfx1151 after rebasing onto a18d3d76: make verify-rocm OK, 163 cargo test suites, 12194 passed, 0 failed, 403 ignored, plus 100 Python tooling tests (including the new ones); cargo test --features rocm --test dead_doc_pointers passed.
  • make verify-versions verify-kernel-dtype-keys verify-kernel-port-dispatch verify-llama-compat verify-fmt verify-python-tooling passed.

Not verified: Metal and CUDA (not available on this host). The change touches only the ROCm profiling scripts, their tests and docs; no Rust or backend code.

Closes #2148

A port unit's share of decode GPU time implies a ceiling of 1 / (1 - share) that leaves out the host launch gaps between dispatches, so the profile underrated ports that replace many small dispatches: the #2067 SSM port measured 1.46x and 1.45x against GPU-share ceilings of 1.42x and 1.25x.

scripts/rocm_decode_gaps.py charges each idle stretch of the decode window to the dispatch that ends it. summarize() now writes role_host_gap_ms_per_token and, per port unit (fallback and reached-default roles), host gap per token and per dispatch, wall share, ceiling_gpu_share, ceiling_wall and plain_ceiling_wall_est (gaps scaled to the plain run). A residual check pins the per-role gaps to the window's idle time, and report() prints a ceiling table. Existing fields are unchanged.

Validated on gfx1151 at 6dbfe7d with MLXCEL_SSM_KERNEL=0: plain wall ceilings 1.60x (granite) and 1.56x (Nemotron-H) against measured medians of 1.49x and 1.44x; GPU-share ceilings 1.45x and 1.26x. Results in benchmarks/rocm_profiles/gfx1151_6dbfe7d7_ssm-graph/, a dated note on the 2026-09-30 profile page, and docs/benchmarks.md describes the new fields.

Closes #2148
@inureyes inureyes added status:done Completed type:enhancement New features, capabilities, or significant additions priority:low Low priority area:benchmark Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks) platform:linux Linux (CUDA / packaging) specific labels Oct 8, 2026
@inureyes
inureyes merged commit f715fe0 into main Oct 8, 2026
27 checks passed
@inureyes
inureyes deleted the perf/issue-2148-profile-host-gaps branch October 8, 2026 16:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:benchmark Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks) platform:linux Linux (CUDA / packaging) specific priority:low Low priority status:done Completed type:enhancement New features, capabilities, or significant additions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(rocm): add host-gap accounting per port unit to the decode profile

1 participant