Conversation
… and async copies The tracer's guard missed three Gluon paths. Gluon builds atomics and int->pointer casts on TileLens's simulation Builder, which no namespace patched, so straddling gl.atomic_* / buffer_atomic_* calls still corrupted memory, and the int->pointer exemption never fired, which refused valid pointer-table gathers. Async copies are Load/Store subclasses, which the tracer's exact-type callback lookup ignored. The Gluon frontend now dispatches create_atomic_rmw, create_atomic_cas, create_buffer_atomic_rmw and create_int_to_ptr on the simulation Builder as Gluon-specific subclasses with argument-normalising adapters, so clients that match op types exactly keep running them concretely. It also dispatches the Ampere async copies and accepts CDNA4's keyword spelling in the async-copy adapters. The tracer resolves callbacks along the op's MRO and broadcasts pointer and mask lanes together, so async copies are checked and recorded as loads and stores.
mark14wu
added this pull request to stack #493
October 4, 2026 02:32
Performance Benchmark
Iterations: 1 warmup + 20 measured |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #491. This closes the three Gluon gaps that #491 left in the tracer's out-of-bounds guard:
gl.atomic_*and AMDbuffer_atomic_*were never checked, so a straddling atomic still wrote past the tensor.shared_to_globaloverrun corrupted memory, and the Ampere and AMD global-to-shared copies read past the tensor.Changes:
tilelens/core/frontend/gluon.py(dispatch only).Builder, not on Triton's interpreter builder, so nothing patched them. A newGLUON_NAMESPACESentry, keyed by thatBuilderclass, dispatchescreate_atomic_rmw,create_atomic_cas,create_buffer_atomic_rmwandcreate_int_to_ptr. Each gets a Gluon-specific subclass (GluonAtomicRMW,GluonAtomicCas,GluonBufferAtomicRMW,GluonIntToPtr) and an argument-normalising adapter. The buffer atomic adapter hands clients the per-lane pointers.main. Dispatching the plainAtomicRMW/AtomicCas/IntToPtrtypes was measured to crash the Sanitizer and RaceDetector on Gluon spin-lock and pointer-cast kernels that pass onmain.async_copy.async_load/async_copy_global_to_shared, which were not intercepted at all. The async-copy adapters now accept CDNA4's(dest, ptr)keyword spelling. The TMA descriptor adapter now applies only to the TMA modules, so the Ampere functions, which share the TMA functions' names, no longer get the descriptor adapter.tilelens/clients/tracer/tracer.py(policy).register_op_callbackresolves callbacks along the op type's MRO. The Gluon subclasses above and the existingGluonAsyncCopyLoad/GluonAsyncCopyStore/GluonBufferLoadToSharednow reach the tracer's existing Load/Store/atomic/IntToPtr callbacks.Visualizer-visible change: async copies are now recorded. A sampled program's async copy appears as a Load tile (
shared_to_globalas a Store tile) with the copy's offsets and mask. Before, these copies were invisible. For example, an Ampereasync_loadof a 4×8 tile with a K-tail maskcols[None, :] < 6now records a Load with offsets and mask both of shape (4, 8), andget_visualization_data()renders it.Still not checked, as the README says:
buffer_load/buffer_store. Their adapters hand clients the base pointer, and fixing that changes the symbolic clients' buffer-op behaviour, so it needs its own PR.Inside
gl.warp_specializepartitions, the tracer's per-program state (sampling, the int->pointer exemption, the program id in messages) is not inherited by the partition threads. That predates this PR and also affectsgl.load/gl.store.Test Plan
New tests:
tests/end_to_end/test_gluon.py:gl.atomic_add,gl.atomic_casandcdna3.buffer_atomic_addare refused, and the sentinels are untouched. In-bounds and masked atomics still run.shared_to_globalis refused with the sentinels intact.tests/unit/test_adapters.py: the builder atomic adapters, the buffer atomic adapter (per-lane pointers; their.value()no-mask placeholder becomesNone), the CDNA4 keyword spelling, and the Ampere async copies not getting the TMA adapter.tests/unit/test_tracer.py: op subclasses get their core op's callbacks.Runs (CPU, Triton 3.8):
pytest tests/ -n 8: 397 passed, 3 skipped.