Skip to content

[FIX] Concretize symbolic ints for tuple indexing on Triton 3.6 - #486

Closed
mark14wu wants to merge 1 commit into
mainfrom
claude/eager-sanitizer-failing-tests-8d5619
Closed

mark14wu wants to merge 1 commit into
mainfrom
claude/eager-sanitizer-failing-tests-8d5619

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

On Triton 3.6, the sanitizer crashes when a kernel indexes a pointer tuple with a tl.static_range loop variable (peer_ptrs[i]). The crash message is:

ValueError: SymbolicExprDataWrapper cannot coerce type <class 'z3.z3.ArithRef'> to int

This makes tests/end_to_end/test_sanitizer.py::test_tuple_pointer_item_selection_uses_registered_tuple_ranges fail on Triton 3.6. CI only tests the latest Triton pre-release, so it never showed up there. pyproject.toml still declares triton>=3.6.0.

Cause: Triton 3.6 implements tensor.__index__ as int(self.handle.data), which calls SymbolicExprDataWrapper.__int__. That method converted the expression to Z3, where a loop index or program id is a free variable rather than a number. Newer Triton releases use int(self.handle.data.squeeze()) instead. That path goes through _scalar_data(), which returns the current value.

Fix: __int__ now returns the current value through _scalar_data(), like squeeze() and __bool__ already do. Both Triton versions now take the same path. In the interpreter, only Triton 3.6's __index__ calls int() on the handle data.

Test Plan

  • New unit test tests/unit/test_symbolic_client.py::test_loop_index_int_is_the_current_iteration checks that int() on a loop index gives the current iteration. It does not depend on the Triton version. It fails before this change and passes after it.
  • Triton 3.8: uv run pytest tests/ → 365 passed, 6 skipped.
  • Triton 3.6 (installed with uv pip install --target <dir> --no-deps triton==3.6.0, run with PYTHONPATH=<dir> uv run --no-sync pytest ...):
    • tests/unit/test_sanitizer.py, tests/end_to_end/test_sanitizer.py, tests/unit/test_symbolic_client.py: 140 passed. Without the fix, the tuple test and the new unit test fail.
    • Rest of tests/: the only failures are the existing Gluon ones described below. They also fail without this change.

Related Issues

Two other sanitizer tests, test_gemm_oob_call_stack and test_cli_code_context_points_to_kernel, also failed locally alongside this one. They are not code bugs. The venv used had a stale pre-rename triton-viz install, so their subprocesses could not import tilelens or find tile-sanitizer. Both pass in a correctly installed env on Triton 3.6 and 3.8.

Not addressed here: on Triton 3.6 the Gluon frontend fails to import (cannot import name 'cluster' from 'triton.experimental.gluon.language.amd.gfx1250'). As a result tests/end_to_end/test_gluon.py cannot be collected and 6 Gluon tests fail. That is a separate issue.

Breaking Changes

None. int() on a symbolic scalar now returns its current value where it used to raise on free Z3 variables. For expressions that reduce to a constant, the result is the same.

Checklist

  • I added tests to all new functionality I added/bugs I fixed.
  • I verified that a human has reviewed all code in this PR.
  • I ran npm run build:frontend if the PR modified any TypeScript code.
  • I made sure that my code is well documented (comments explaining strange code, docstrings for functions, website modified if new functionality added).

Triton 3.6 implements tensor.__index__ as int(handle.data), which reaches
SymbolicExprDataWrapper.__int__. That method evaluated the expression to
Z3, where a loop index or program id is a free variable, so indexing a
pointer tuple with a static_range loop variable raised "cannot coerce
ArithRef to int". Newer Triton calls handle.data.squeeze() instead, which
already concretizes to the current value.

Make __int__ concretize the same way so both Triton versions behave alike,
and add a version-independent unit test for int() on a loop index.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.056s 0.054s -2.9% 20 / 20
gemm_oob 0.063s 0.060s -4.3% 20 / 20
indirect_load 0.011s 0.011s -3.1% 20 / 20
nested_loop 0.118s 0.116s -1.9% 20 / 20
block_pointer_loop_advance 0.063s 0.060s -5.6% 20 / 20
liger_jsd 0.078s 0.081s +4.6% 20 / 20
flaggems_layernorm 0.206s 0.201s -2.3% 20 / 20
swiglu 0.091s 0.096s +5.5% ⚠️ 20 / 20
cross_entropy 0.528s 0.545s +3.2% 20 / 20
fused_linear_jsd 0.120s 0.118s -1.8% 20 / 20
Total 1.334s 1.342s +0.6% N/A

Threshold: >5% regression flagged with ⚠️
Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

@mark14wu mark14wu closed this Oct 3, 2026
@mark14wu

mark14wu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Closed due to deprecation of triton 3.6.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant