Conversation
Triton 3.8's Autotuner walks its .fn chain down to a JITFunction before calling knobs.autotuning.listener, and check_disk_cache does the same walk. The interpreted path replaced autotuner.fn with an InterpretedFunction, whose .fn is the plain Python function, so a traced autotuned launch raised AttributeError: 'function' object has no attribute 'fn' when a listener was set (and with cache_results=True or TRITON_CACHE_AUTOTUNING=1 even without one). Put an _InterpretedLeaf under traced Autotuner/Heuristics layers instead: it runs the interpreted kernel, forwards other attributes to it, and exposes the JITFunction as .fn. Also turn the traced autotuner's disk cache off so dummy interpreter timings are never persisted, shallow-copy the warmup runner (a JITFunction holds an RLock), and unpack a leaf left by an earlier trace so re-tracing keeps jit_fn. GluonTrace gets the same leaf and now records jit_fn.
Performance Benchmark
Threshold: >5% regression flagged with |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
On Triton 3.8, launching a traced autotuned kernel on the interpreted path raises:
This happens whenever
knobs.autotuning.listeneris set. Before calling the listener,Autotuner.runwalksself.fndown to theJITFunctionit tunes. The interpreted path replacesautotuner.fnwith anInterpretedFunction, whose.fnis the plain Python function, so the walk fails.check_disk_cachedoes the same walk, socache_results=TrueandTRITON_CACHE_AUTOTUNING=1hit the same error even without a listener. That part already failed before the 3.8 requirement.Changes in
tilelens/core/trace.py:_InterpretedLeaf. Traced Autotuner/Heuristics layers now sit on an_InterpretedLeafinstead of a bare interpreted function. Itsruninterprets the kernel, other attributes forward to the interpreted function, andfnis the kernel'sJITFunction. The listener therefore receives the user's ownJITFunction, the same object an untraced launch reports. The leaf mirrors_IRLeaffrom the IR-mode work.cache_results = False. Triton only turns its disk cache off for autotuners decorated in interpreter mode, but TileLens enables interpreter mode at launch time, after that decision. Without this, the now-successful walk would either fail ondriver.active(CPU) or write the dummy(1.0, 1.0, 1.0)timings into Triton's cache (GPU), where a later real run would read them._warmup_runnerusescopyinstead ofdeepcopy, because the chain now reaches aJITFunction, which holds an RLock.Autotuner.warmuponly rebindsnargs, so the shared state is safe.unpack_kernelin bothTritonTraceandGluonTracerecognizes a leaf left by an earlier trace of the same autotuner. Re-tracing therefore keepsjit_fn; before, the second trace lost it.GluonTracehit the same crash; it now builds the same leaf and recordsjit_fn.Not changed: stock Triton 3.8 with
TRITON_INTERPRET=1and a listener set crashes the same way without TileLens, because noJITFunctionexists anywhere in that chain. That is an upstream interpreter issue, so sources that are alreadyInterpretedFunctions keep the old behaviour.Test Plan
New tests in
tests/end_to_end/test_core.py:test_autotune_listener_sees_jit_functionandtest_gluon_autotune_listener_sees_jit_function: the listener fires once with the tracedjit_fn, and the kernel output is correct.test_autotune_retrace_keeps_jit_functionandtest_gluon_autotune_retrace_keeps_jit_function: tracing the same autotuner twice keepsjit_fnand still launches.test_autotune_interpreter_skips_disk_cache: acache_results=Trueautotuner launches and writes no*.autotune.json.The Triton test kernels are decorated inside
knobs.runtime.scope()withinterpret = False. Otherwise a worker that also collectedtests/unit/test_multithreading.py(which setsTRITON_INTERPRET=1at import) would give anInterpretedFunctionand silently test the upstream case instead.Runs (CPU, Triton 3.8):
AttributeErrorabove.pytest tests/ -n autohas no new failures. The only failures are 6 CLI tests that also fail onmainin this environment because thetile-*entry points are not installed.tests/unit/test_multithreading.py.test_core.py/test_tracer.pypass.Related Issues
None filed. Stock
TRITON_INTERPRET=1+knobs.autotuning.listenerfails in Triton itself and should be reported upstream separately.Breaking Changes
None intended. After tracing, the user's Autotuner/Heuristics object now holds an
_InterpretedLeafat.fnrather than anInterpretedFunction. Attribute access andautotuner.warmup(...)on that object behave as before, and.fn.fnis now theJITFunction. Traced autotuners no longer use Triton's autotune disk cache.Checklist
npm run build:frontendif the PR modified any TypeScript code. (no TypeScript changes)