Skip to content

[TEST] Cover Gluon core simulation ops - #451

Merged
Jokeren merged 3 commits into
mainfrom
test/gluon-core-simulation-ops
Jun 22, 2026
Merged

Jokeren merged 3 commits into
mainfrom
test/gluon-core-simulation-ops

Conversation

@Jokeren

@Jokeren Jokeren commented Jun 20, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds focused CPU regression coverage for Gluon core simulation ops and fixes the simulator layout propagation needed by those kernels. The new tests cover scalar range loops, program_id, arange, masked load/store, 2D broadcasted offsets, SliceLayout, convert_layout, and elementwise add without requiring a CUDA device.

Refactors the Gluon simulator builder so inherited Triton interpreter operations can attach Gluon layouts through a shared _WRAPPED_LAYOUT_OPS path while keeping special layout/signature cases as explicit builder methods. create_broadcast, create_expand_dims, and create_histogram now delegate to the parent interpreter implementation for data behavior and attach the Gluon layout locally. Tensor-producing paths use _tensor_result so layout metadata stays attached to returned handles.

Validation

  • PYTHONPATH=/mnt/keren/triton-viz pytest tests/end_to_end/test_gluon_core_ops.py tests/end_to_end/test_gluon.py tests/end_to_end/test_core.py tests/unit/test_adapters.py -q
  • pre-commit run --files triton_viz/core/simulation/gluon.py

@github-actions

github-actions Bot commented Jun 20, 2026 •

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.096s 0.097s +1.1% 20 / 20
gemm_oob 0.108s 0.108s +0.5% 20 / 20
indirect_load 0.019s 0.019s +0.5% 20 / 20
nested_loop 0.201s 0.202s +0.4% 20 / 20
block_pointer_loop_advance 0.195s 0.196s +0.7% 20 / 20
liger_jsd 0.133s 0.133s -0.2% 20 / 20
flaggems_layernorm 0.352s 0.353s +0.2% 20 / 20
swiglu 0.164s 0.164s +0.4% 20 / 20
cross_entropy 0.920s 0.914s -0.6% 20 / 20
fused_linear_jsd 0.202s 0.203s +0.7% 20 / 20
Total 2.390s 2.390s +0.0% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

@Jokeren
Jokeren force-pushed the test/gluon-core-simulation-ops branch from 773718e to 8928cab Compare June 22, 2026 14:31
@Jokeren
Jokeren marked this pull request as ready for review June 22, 2026 14:36
@Jokeren
Jokeren merged commit 2b2818d into main Jun 22, 2026
4 checks passed
@Jokeren
Jokeren deleted the test/gluon-core-simulation-ops branch June 22, 2026 14:36

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8928cab66b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

elem_type,
),
elem_type,
_get_handle_layout(scale),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the data layout for FP4 upcasts

When an FP4 value is upcast with a separate scale tensor, this tags the unpacked result with the scale tensor's layout rather than the packed data/output layout. In MXFP4-style kernels the scale tile often has a different layout/shape from the values it scales, so the next Gluon operation infers the wrong distributed layout and can reject otherwise valid combinations with data-layout tensors as a layout mismatch. The FP8 path just above keeps the source layout; the FP4 path should not derive the result layout from scale.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant