Skip to content

[FEAT] Add a host-compiled IR layer for clients - #479

Open
mark14wu wants to merge 3 commits into
mainfrom
ir-mode-core
Open

mark14wu wants to merge 3 commits into
mainfrom
ir-mode-core

Conversation

@mark14wu

@mark14wu mark14wu commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds the mechanism that lets a TileLens client analyze a kernel's compiled IR (TTIR) instead of interpreting it, and do so without a GPU. This PR has no user-facing client of its own; the compiled sanitizer in the stacked #480 (base: this branch) is its first consumer.

What's included

  • Core IR lifecycle (tilelens/core/client.py, trace.py)
    • Clients declare NEEDS_INTERPRETER, IR_STAGES and LAUNCH (skip / run / indifferent); conflicting launch preferences are refused when clients are registered.
    • IR clients receive launch events from a capture around jit_fn.run, so every autotune / heuristics config is seen (deduplicated by specialization and call binding). They take no part in op/loop patching or the pre_run vote, so they cannot starve or clobber interpreting clients.
    • One Launch object per launch, isolated finalize per client, begin_launch / abort_launch hooks.
  • Host compile (tilelens/core/host_compile.py): TTIR through the JIT's own binder and specialization, compiled on the host for a configurable target (default cuda:89, TILELENS_IR_TARGET), only up to the stages clients ask for. No driver or device access on the IR path; tl.target_info answers for the configured target.
  • tilelens.ir
    • ttir_reader.py + _mlir_walk.py: walks Triton's MLIR bindings for structure and reads the attributes the bindings cannot expose from the aligned text; any misalignment is a refusal, never a silent misread.
    • AccessGraph term model with bit widths and width_obligations(); refusals are a typed UnsupportedTTIR(kind, ...) that clients decide about.
    • capture.py (artifact log, parse cache), launch.py (launch binding, tensor facts), client.py (IRClient base without analysis defaults), verdict.py (IRVerdict, saved by tilelens.save()).
  • Triton versions: tested on 3.6 and 3.8 (per-release printer/vocabulary/runtime tables); other releases are refused unless TILELENS_IR_ALLOW_UNTESTED_TRITON=1, and IR tests skip there.

Behaviour changes for existing clients

  • tilelens.launches holds one object per launch. Before, every launch of a trace appended the same Launch object, so later launches overwrote earlier records.
  • Tracing no longer mutates the user's Autotuner / Heuristics. This also fixes three crashes on main: a plain @triton.heuristics kernel (duplicate warmup keyword), @autotune over @heuristics ("missing BLOCK"), and a Profiler-traced kernel calling a traced device function ("Unsupported function referenced").
  • patch_warmup asks every client to vote (the all(not ...) generator stopped at the first True).
  • One client raising in finalize no longer prevents the others from finalizing.

Testing

  • Reader conformance suite (tests/conformance/): for 124 kernels, the reader's static footprint is compared with the footprint Triton's own interpreter touches (an independent numpy evaluator on one side). 102 compared, all conforming; the rest are refusals pinned by kind. A mutation run reverting reader fixes is caught 27/28 (the survivor is equivalent at the graph interface).
  • Alignment over real kernels (tools/ir_bulk_conformance.py): 0 misaligned of 14,336 Triton 3.6 TTIR texts and 0 of 1,127 Triton 3.8 texts.
  • This commit alone, CPU only: Triton 3.6 1056 passed / 9 skipped; Triton 3.8 1072 passed / 1 skipped. Pre-existing failures, identical on a clean HEAD copy: tests/end_to_end/test_sanitizer.py::{test_tuple_pointer_item_selection_uses_registered_tuple_ranges, test_gemm_oob_call_stack, test_cli_code_context_points_to_kernel} on 3.6 (the last two on 3.8; they need an installed tile-sanitizer / an importable tilelens in subprocesses).

Lets a client analyze a kernel's compiled IR instead of interpreting it,
without a GPU.

- core: clients declare NEEDS_INTERPRETER, IR_STAGES and LAUNCH; IR
  clients get launch events from a capture around jit_fn.run (every
  autotune/heuristics config, deduplicated by binding), take no part in
  op/loop patching or the pre_run vote, and conflicting launch
  preferences are refused at registration. Each launch gets its own
  Launch, client finalize is isolated, and runner chains are rebuilt
  instead of mutating the user's Autotuner/Heuristics.
- host compile: TTIR through the JIT's own binder and specialization,
  compiled on the host for a configurable target (default cuda:89,
  TILELENS_IR_TARGET); no driver or device access on the IR path.
- tilelens.ir: a TTIR reader that walks the MLIR bindings and reads the
  attributes they cannot expose from the aligned text; the AccessGraph
  term model with bit widths and width obligations; capture, launch
  binding, IRClient base and IRVerdict records (saved by tilelens.save).
- Tested on Triton 3.6 and 3.8; other releases are refused unless
  TILELENS_IR_ALLOW_UNTESTED_TRITON is set, and IR tests skip there.
- Tests: reader conformance suite (static footprint vs Triton's
  interpreter), golden TTIR per release, lifecycle and host-compile
  tests, and tools/ir_bulk_conformance.py.
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.053s 0.054s +0.6% 20 / 20
gemm_oob 0.060s 0.060s +0.3% 20 / 20
indirect_load 0.011s 0.011s -0.5% 20 / 20
nested_loop 0.113s 0.112s -0.6% 20 / 20
block_pointer_loop_advance 0.060s 0.062s +3.4% 20 / 20
liger_jsd 0.076s 0.079s +4.1% 20 / 20
flaggems_layernorm 0.197s 0.201s +2.1% 20 / 20
swiglu 0.091s 0.094s +3.0% 20 / 20
cross_entropy 0.518s 0.536s +3.3% 20 / 20
fused_linear_jsd 0.115s 0.116s +0.7% 20 / 20
Total 1.294s 1.324s +2.3% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

@mark14wu
mark14wu added this pull request to stack #481 September 29, 2026 20:19
@mark14wu
mark14wu marked this pull request as ready for review September 29, 2026 20:20
The golden regeneration test took the generator's path out of the locs
but not Triton's: a golden whose kernel calls into Triton's own sources
(tl.cdiv, tl.zeros, ...) names the directory Triton is installed at, so
it never regenerated byte for byte on another machine (CI failed on
golden_matmul_tma_s1_sm90 under Triton 3.8). Take both paths out before
comparing.
The TTIR goldens' locs named the absolute paths of the machine that
printed them: the checkout, Triton's installation, and local directories
outside the repository. They now name a file of this repository from
tests/, one of Triton's own sources from triton/, and a kernel kept
outside the repository by its file name only.

- Both generators write portable locs, so regenerating a golden does not
  bring the paths back.
- The regeneration test compares the generator's portable output with
  the golden byte for byte instead of masking paths at comparison time.
- Only loc strings change; line and column numbers are untouched.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant