[TEST] Cover Gluon Blackwell tensor memory ops - #458
Conversation
2d4a108 to
afd6898
Compare
d199fc5 to
8a915c5
Compare
Performance Benchmark
Iterations: 1 warmup + 20 measured |
2754b31 to
1c0b764
Compare
d9be516 to
389a784
Compare
bbf9c2a to
4d83822
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4d83822c4a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| return gluon_blackwell.TensorMemoryLayout( | ||
| tuple(block), | ||
| int(col_stride), | ||
| [] if cga_layout is None else [list(basis) for basis in cga_layout], | ||
| bool(two_ctas), | ||
| bool(fp4_padded), | ||
| ) |
There was a problem hiding this comment.
Preserve FP4 padding when rebuilding tensor-memory layouts
When a Gluon kernel uses a Blackwell tensor-memory layout with fp4_padded=True, _to_ir passes that flag into this builder, but this constructor call now drops it and always returns a layout constructed without the FP4 padding state. That makes the interpreter treat FP4-padded tensor-memory descriptors as ordinary layouts, so any later layout-dependent register conversion or descriptor behavior can diverge from real Gluon semantics. Please propagate fp4_padded here, using a version-aware constructor path if needed for older Triton builds.
Useful? React with 👍 / 👎.
4d83822 to
60c2748
Compare
Summary
Adds CPU regression coverage for Blackwell tensor-memory roundtrip operations, stacked on the Hopper WGMMA PR.
The test covers:
blackwell.allocate_tensor_memoryTensorMemoryLayoutstoreloadget_reg_layout(num_warps=...)The simulated output is compared against the original torch CPU input tile.
This also makes Gluon builtin patching avoid injecting
_semanticinto helper functions that are marked builtin but do not accept that keyword, which is needed by Blackwell tensor-memory layout helpers.Validation
PYTHONPATH=/mnt/keren/triton-viz pytest tests/end_to_end/test_gluon_blackwell_tensor_memory_ops.py -qPYTHONPATH=/mnt/keren/triton-viz pytest tests/end_to_end/test_gluon_blackwell_tensor_memory_ops.py tests/end_to_end/test_gluon_wgmma_ops.py tests/end_to_end/test_gluon_tma_im2col_ops.py tests/end_to_end/test_gluon_blackwell_tma_ops.py tests/end_to_end/test_gluon_tma_ops.py tests/end_to_end/test_gluon_async_copy_ops.py tests/end_to_end/test_gluon_core_ops.py tests/end_to_end/test_gluon.py tests/end_to_end/test_core.py tests/unit/test_adapters.py -qpre-commit run --files triton_viz/core/simulation/gluon.py tests/end_to_end/test_gluon_blackwell_tensor_memory_ops.py