Skip to content

[TEST] Cover Gluon Blackwell tcgen05 ops - #459

Merged
Jokeren merged 10 commits into
mainfrom
test/gluon-blackwell-tcgen05-ops
Jun 26, 2026
Merged

Jokeren merged 10 commits into
mainfrom
test/gluon-blackwell-tcgen05-ops

Conversation

@Jokeren

@Jokeren Jokeren commented Jun 20, 2026

Copy link
Copy Markdown
Member

Summary

Add CPU-result coverage for Blackwell tcgen05_mma simulation through Gluon.

This adds a focused end-to-end test file that:

  • loads A/B/C tiles through TMA into shared memory
  • seeds tensor memory with the accumulator tile
  • runs blackwell.tcgen05_mma using both tcgen05_commit and explicit mbarrier signaling paths
  • checks the simulated CPU result against a.float() @ b.float() + c

Validation

  • PYTHONPATH=/mnt/keren/triton-viz pytest tests/end_to_end/test_gluon_blackwell_tcgen05_ops.py -q
  • PYTHONPATH=/mnt/keren/triton-viz pytest tests/end_to_end/test_gluon_blackwell_tcgen05_ops.py tests/end_to_end/test_gluon_blackwell_tensor_memory_ops.py tests/end_to_end/test_gluon_wgmma_ops.py tests/end_to_end/test_gluon_tma_im2col_ops.py tests/end_to_end/test_gluon_blackwell_tma_ops.py tests/end_to_end/test_gluon_tma_ops.py tests/end_to_end/test_gluon_async_copy_ops.py tests/end_to_end/test_gluon_core_ops.py tests/end_to_end/test_gluon.py tests/end_to_end/test_core.py tests/unit/test_adapters.py -q
  • pre-commit run --files tests/end_to_end/test_gluon_blackwell_tcgen05_ops.py

@Jokeren
Jokeren force-pushed the test/gluon-blackwell-tensor-memory-ops branch 9 times, most recently from 4d83822 to 60c2748 Compare June 26, 2026 13:21
Base automatically changed from test/gluon-blackwell-tensor-memory-ops to main June 26, 2026 14:36
@Jokeren
Jokeren force-pushed the test/gluon-blackwell-tcgen05-ops branch from 6049438 to b1422ef Compare June 26, 2026 14:48
@github-actions

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.105s 0.106s +0.4% 20 / 20
gemm_oob 0.118s 0.117s -0.3% 20 / 20
indirect_load 0.022s 0.022s +0.1% 20 / 20
nested_loop 0.235s 0.238s +1.1% 20 / 20
block_pointer_loop_advance 0.206s 0.213s +3.3% 20 / 20
liger_jsd 0.137s 0.138s +1.0% 20 / 20
flaggems_layernorm 0.397s 0.398s +0.4% 20 / 20
swiglu 0.171s 0.172s +1.0% 20 / 20
cross_entropy 0.980s 0.983s +0.3% 20 / 20
fused_linear_jsd 0.208s 0.209s +0.3% 20 / 20
Total 2.579s 2.597s +0.7% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

@Jokeren
Jokeren marked this pull request as ready for review June 26, 2026 16:14
@Jokeren
Jokeren merged commit 22731c9 into main Jun 26, 2026
4 checks passed
@Jokeren
Jokeren deleted the test/gluon-blackwell-tcgen05-ops branch June 26, 2026 16:14
mark14wu added a commit that referenced this pull request Jul 10, 2026
Resolves the PR #361 merge conflict (CI could not run on the merge ref).
Main brings the merged #435 LoopSite refactor and the Gluon Blackwell
test coverage (#456-#459).

Resolutions:
- race_detector.py _process_pending_check: keep the branch's
  _unsupported_capture early-return guard on top of main's version.
- core/frontend/base.py: take main's simplified _LangPatchScope wholesale
  (3-tuple changes, no set_item/mark_removed) — the branch's 4-tuple
  machinery has no remaining callers on either side, and the auto-merge
  hybrid would have crashed restore().
- Adapt the race detector's loop hooks to #435's LoopSite keying: hook
  params and loop_stack matching use LoopSite (a hashable NamedTuple)
  instead of bare linenos; _finished_loop_iter_subs keys become LoopSites
  (two loops at the same function-relative line in different files no
  longer share bookkeeping).
- simulation/gluon.py: broaden the optional gfx1250 import guard — a
  triton at the declared floor (>= 3.6.0) lacks the newer submodules
  (cluster/tdm/...) and the narrow is_hip_gfx1250-only guard killed every
  consumer at collection time on such installs.

Full suite: 717 passed; the only delta vs the pre-merge baseline is
inside the environment-dependent gluon family (4 old failures fixed by
main, 3 new tests hit the local triton-3.6.0 x numpy-2 interpreter
boundary that CI's environment does not).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant