Conversation
…ng memory Triton's interpreter runs loads, stores and atomics on raw host memory. An unmasked out-of-bounds access in a traced kernel therefore wrote past the end of the tensor's allocation, corrupted glibc heap metadata, and the process aborted or segfaulted later at exit (rc 134/139). This showed up on CPU-only machines but happens with CUDA tensors too, since the interpreter overruns its host copies the same way. The tracer now records the storage ranges of its tensor arguments and, in its own before-callbacks, raises IndexError before a Triton load, store or atomic (or a Gluon gl.load/gl.store) whose pointer block touches a tensor argument's storage while an active lane falls outside every known storage. Accesses that never touch a known storage and programs that cast integers to pointers are not judged, so the guard is best-effort; the Sanitizer remains the complete check. Shared code only gains argument-normalising adapters for AtomicRMW and AtomicCas.
Performance Benchmark
Iterations: 1 warmup + 20 measured |
mark14wu
added this pull request to stack #493
October 4, 2026 02:32
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Tracing a kernel with an out-of-bounds access, for example a load or store that forgot its mask, made the process abort or segfault at exit (rc 134
corrupted size vs. prev_size/ rc 139). The trigger is the out-of-bounds access, not the CPU:Triton's interpreter runs loads, stores and atomics on raw host memory. For CPU tensors it uses the caller's own storage, so the overrun writes past the end of the torch allocation and corrupts glibc heap metadata. The crash only surfaces when the allocator frees at exit. It was first noticed on CPU, but it also happens with CUDA tensors, because the interpreter overruns its host copies the same way. Plain
TRITON_INTERPRET=1without TileLens crashes the same way: this is UB in the kernel, and the interpreter is not patched here.What the tracer does now (
tilelens/clients/tracer/tracer.py):arg_callbackrecords the storage byte range behind every tensor argument. It recurses into tuples and uses.basefortriton.reinterpretwrappers and hostTensorDescriptors.gl.load/gl.store, the tracer's own before-callbacks run_check_in_bounds. If any lane of the pointer block, masked-off lanes included, touches a tensor argument's storage, every active lane must lie fully inside some known storage. Otherwise it raisesIndexErrorbefore the access runs. The message names the program, the kernel line, the argument, and the offending byte range, and points to the Sanitizer. Masked-off lanes only attribute the access to a tensor. That matters because Triton issues a floatatomic_max/atomic_minas two calls split by sign, and one of them can hold only the out-of-bounds lanes.TRITON_ADAPTERSgains argument-normalising adapters forAtomicRMW(ptr, mask) andAtomicCas(ptr), mirroring the Load/Store ones. All policy stays in the tracer. Symbolic clients use op overriders, which do not go through these adapters.Known gaps, addressed in the stacked follow-up PR: under Gluon, the integer-to-pointer exemption never fires, Gluon atomics are not checked, and Gluon async copies are not checked.
Test Plan
New tests:
tests/end_to_end/test_tracer.py: refusal of out-of-bounds loads (with and withoutgrid_idxsampling), stores (num_sms1 and 4),atomic_add,atomic_cas, a sign-split floatatomic_max, an int32 word overlapping the end of a byte tensor, atriton.reinterpretargument, and a tuple argument. There are also allow-tests for a view reading its base storage and for a pointer-table gather that mixes an argument row with a non-argument row.tests/end_to_end/test_gluon.py: an unmasked Gluongl.storeis refused (as anInterpreterErrorcaused byIndexError).tests/unit/test_tracer.py: a Pythonboolmask;tests/unit/test_adapters.py: the two new adapters.torch.from_numpyon a slice of a larger numpy buffer, so out-of-bounds writes land in sentinels rather than the heap. The tests then assert the sentinels are untouched, which is deterministic.Runs (CPU, Triton 3.8):
IndexError.mainall new refusal tests fail. Each of the reinterpret, tuple, sign-split and word-overlap tests fails when its specific code path is removed. The pointer-table test fails without the integer-to-pointer exemption.pytest tests/ -n 8: 378 passed, 3 skipped.