Conversation
Lets a client analyze a kernel's compiled IR instead of interpreting it, without a GPU. - core: clients declare NEEDS_INTERPRETER, IR_STAGES and LAUNCH; IR clients get launch events from a capture around jit_fn.run (every autotune/heuristics config, deduplicated by binding), take no part in op/loop patching or the pre_run vote, and conflicting launch preferences are refused at registration. Each launch gets its own Launch, client finalize is isolated, and runner chains are rebuilt instead of mutating the user's Autotuner/Heuristics. - host compile: TTIR through the JIT's own binder and specialization, compiled on the host for a configurable target (default cuda:89, TILELENS_IR_TARGET); no driver or device access on the IR path. - tilelens.ir: a TTIR reader that walks the MLIR bindings and reads the attributes they cannot expose from the aligned text; the AccessGraph term model with bit widths and width obligations; capture, launch binding, IRClient base and IRVerdict records (saved by tilelens.save). - Tested on Triton 3.6 and 3.8; other releases are refused unless TILELENS_IR_ALLOW_UNTESTED_TRITON is set, and IR tests skip there. - Tests: reader conformance suite (static footprint vs Triton's interpreter), golden TTIR per release, lifecycle and host-compile tests, and tools/ir_bulk_conformance.py.
Performance Benchmark
Iterations: 1 warmup + 20 measured |
mark14wu
added this pull request to stack #481
September 29, 2026 20:19
mark14wu
marked this pull request as ready for review
September 29, 2026 20:20
The golden regeneration test took the generator's path out of the locs but not Triton's: a golden whose kernel calls into Triton's own sources (tl.cdiv, tl.zeros, ...) names the directory Triton is installed at, so it never regenerated byte for byte on another machine (CI failed on golden_matmul_tma_s1_sm90 under Triton 3.8). Take both paths out before comparing.
The TTIR goldens' locs named the absolute paths of the machine that printed them: the checkout, Triton's installation, and local directories outside the repository. They now name a file of this repository from tests/, one of Triton's own sources from triton/, and a kernel kept outside the repository by its file name only. - Both generators write portable locs, so regenerating a golden does not bring the paths back. - The regeneration test compares the generator's portable output with the golden byte for byte instead of masking paths at comparison time. - Only loc strings change; line and column numbers are untouched.
This was referenced Oct 3, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds the mechanism that lets a TileLens client analyze a kernel's compiled IR (TTIR) instead of interpreting it, and do so without a GPU. This PR has no user-facing client of its own; the compiled sanitizer in the stacked #480 (base: this branch) is its first consumer.
What's included
tilelens/core/client.py,trace.py)NEEDS_INTERPRETER,IR_STAGESandLAUNCH(skip/run/indifferent); conflicting launch preferences are refused when clients are registered.jit_fn.run, so every autotune / heuristics config is seen (deduplicated by specialization and call binding). They take no part in op/loop patching or thepre_runvote, so they cannot starve or clobber interpreting clients.Launchobject per launch, isolatedfinalizeper client,begin_launch/abort_launchhooks.tilelens/core/host_compile.py): TTIR through the JIT's own binder and specialization, compiled on the host for a configurable target (defaultcuda:89,TILELENS_IR_TARGET), only up to the stages clients ask for. No driver or device access on the IR path;tl.target_infoanswers for the configured target.tilelens.irttir_reader.py+_mlir_walk.py: walks Triton's MLIR bindings for structure and reads the attributes the bindings cannot expose from the aligned text; any misalignment is a refusal, never a silent misread.AccessGraphterm model with bit widths andwidth_obligations(); refusals are a typedUnsupportedTTIR(kind, ...)that clients decide about.capture.py(artifact log, parse cache),launch.py(launch binding, tensor facts),client.py(IRClientbase without analysis defaults),verdict.py(IRVerdict, saved bytilelens.save()).TILELENS_IR_ALLOW_UNTESTED_TRITON=1, and IR tests skip there.Behaviour changes for existing clients
tilelens.launchesholds one object per launch. Before, every launch of a trace appended the sameLaunchobject, so later launches overwrote earlier records.Autotuner/Heuristics. This also fixes three crashes on main: a plain@triton.heuristicskernel (duplicatewarmupkeyword),@autotuneover@heuristics("missing BLOCK"), and a Profiler-traced kernel calling a traced device function ("Unsupported function referenced").patch_warmupasks every client to vote (theall(not ...)generator stopped at the firstTrue).finalizeno longer prevents the others from finalizing.Testing
tests/conformance/): for 124 kernels, the reader's static footprint is compared with the footprint Triton's own interpreter touches (an independent numpy evaluator on one side). 102 compared, all conforming; the rest are refusals pinned by kind. A mutation run reverting reader fixes is caught 27/28 (the survivor is equivalent at the graph interface).tools/ir_bulk_conformance.py): 0 misaligned of 14,336 Triton 3.6 TTIR texts and 0 of 1,127 Triton 3.8 texts.tests/end_to_end/test_sanitizer.py::{test_tuple_pointer_item_selection_uses_registered_tuple_ranges, test_gemm_oob_call_stack, test_cli_code_context_points_to_kernel}on 3.6 (the last two on 3.8; they need an installedtile-sanitizer/ an importabletilelensin subprocesses).