Skip to content

[FIX] Extend the tracer's out-of-bounds guard to Gluon atomics, casts and async copies - #492

Open
mark14wu wants to merge 1 commit into
claude/tracer-cpu-exit-crash-0d5953from
claude/tracer-gluon-oob-guard
Open

mark14wu wants to merge 1 commit into
claude/tracer-cpu-exit-crash-0d5953from
claude/tracer-gluon-oob-guard

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator

Summary

Stacked on #491. This closes the three Gluon gaps that #491 left in the tracer's out-of-bounds guard:

Changes:

  • tilelens/core/frontend/gluon.py (dispatch only).
    • Gluon builds atomics and int->pointer casts on TileLens's own simulation Builder, not on Triton's interpreter builder, so nothing patched them. A new GLUON_NAMESPACES entry, keyed by that Builder class, dispatches create_atomic_rmw, create_atomic_cas, create_buffer_atomic_rmw and create_int_to_ptr. Each gets a Gluon-specific subclass (GluonAtomicRMW, GluonAtomicCas, GluonBufferAtomicRMW, GluonIntToPtr) and an argument-normalising adapter. The buffer atomic adapter hands clients the per-lane pointers.
    • The subclasses are deliberate. The Sanitizer, RaceDetector and Profiler match op types exactly, so they keep running these ops concretely, as on main. Dispatching the plain AtomicRMW/AtomicCas/IntToPtr types was measured to crash the Sanitizer and RaceDetector on Gluon spin-lock and pointer-cast kernels that pass on main.
    • It dispatches Ampere async_copy.async_load / async_copy_global_to_shared, which were not intercepted at all. The async-copy adapters now accept CDNA4's (dest, ptr) keyword spelling. The TMA descriptor adapter now applies only to the TMA modules, so the Ampere functions, which share the TMA functions' names, no longer get the descriptor adapter.
  • tilelens/clients/tracer/tracer.py (policy).
    • register_op_callback resolves callbacks along the op type's MRO. The Gluon subclasses above and the existing GluonAsyncCopyLoad / GluonAsyncCopyStore / GluonBufferLoadToShared now reach the tracer's existing Load/Store/atomic/IntToPtr callbacks.
    • Pointer and mask lanes are broadcast together before checking and recording. Gluon adapters see a builtin's operands before it broadcasts them, so an async copy can pass a row or scalar mask for a full tile.

Visualizer-visible change: async copies are now recorded. A sampled program's async copy appears as a Load tile (shared_to_global as a Store tile) with the copy's offsets and mask. Before, these copies were invisible. For example, an Ampere async_load of a 4×8 tile with a K-tail mask cols[None, :] < 6 now records a Load with offsets and mask both of shape (4, 8), and get_visualization_data() renders it.

Still not checked, as the README says:

  • AMD buffer_load/buffer_store. Their adapters hand clients the base pointer, and fixing that changes the symbolic clients' buffer-op behaviour, so it needs its own PR.
  • TMA/TDM descriptor copies, which the simulation bounds by the descriptor shape.
  • Accesses that never touch a tensor argument, and programs after an int->pointer cast.

Inside gl.warp_specialize partitions, the tracer's per-program state (sampling, the int->pointer exemption, the program id in messages) is not inherited by the partition threads. That predates this PR and also affects gl.load/gl.store.

Test Plan

New tests:

  • tests/end_to_end/test_gluon.py:
    • Out-of-bounds gl.atomic_add, gl.atomic_cas and cdna3.buffer_atomic_add are refused, and the sentinels are untouched. In-bounds and masked atomics still run.
    • All five global-to-shared async copies (Ampere ×2, CDNA4 ×2, gfx1250) run, record a Load, and refuse an overrun read. An overrunning gfx1250 shared_to_global is refused with the sentinels intact.
    • An async copy with a broadcast K-tail mask records shape-consistent Loads and visualizes.
    • The Gluon pointer-table gather is allowed.
    • Non-regression for the Sanitizer: a Gluon spin-lock kernel and a pointer bitcast after an int->pointer cast still run cleanly.
  • tests/unit/test_adapters.py: the builder atomic adapters, the buffer atomic adapter (per-lane pointers; the ir.value() no-mask placeholder becomes None), the CDNA4 keyword spelling, and the Ampere async copies not getting the TMA adapter.
  • tests/unit/test_tracer.py: op subclasses get their core op's callbacks.

Runs (CPU, Triton 3.8):

… and async copies

The tracer's guard missed three Gluon paths. Gluon builds atomics and
int->pointer casts on TileLens's simulation Builder, which no namespace
patched, so straddling gl.atomic_* / buffer_atomic_* calls still corrupted
memory, and the int->pointer exemption never fired, which refused valid
pointer-table gathers. Async copies are Load/Store subclasses, which the
tracer's exact-type callback lookup ignored.

The Gluon frontend now dispatches create_atomic_rmw, create_atomic_cas,
create_buffer_atomic_rmw and create_int_to_ptr on the simulation Builder as
Gluon-specific subclasses with argument-normalising adapters, so clients that
match op types exactly keep running them concretely. It also dispatches the
Ampere async copies and accepts CDNA4's keyword spelling in the async-copy
adapters. The tracer resolves callbacks along the op's MRO and broadcasts
pointer and mask lanes together, so async copies are checked and recorded as
loads and stores.
@mark14wu
mark14wu added this pull request to stack #493 October 4, 2026 02:32
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.056s 0.055s -1.0% 20 / 20
gemm_oob 0.061s 0.062s +0.6% 20 / 20
indirect_load 0.011s 0.011s +0.5% 20 / 20
nested_loop 0.116s 0.115s -0.8% 20 / 20
block_pointer_loop_advance 0.061s 0.061s -0.3% 20 / 20
liger_jsd 0.079s 0.078s -1.0% 20 / 20
flaggems_layernorm 0.199s 0.196s -1.8% 20 / 20
swiglu 0.094s 0.092s -1.4% 20 / 20
cross_entropy 0.523s 0.504s -3.7% 20 / 20
fused_linear_jsd 0.118s 0.119s +0.8% 20 / 20
Total 1.318s 1.293s -1.9% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant