Skip to content

[FIX] Skip the profiler buffer load check when Triton has no active driver - #484

Open
mark14wu wants to merge 1 commit into
mainfrom
claude/eager-profiler-cpu-drivers-b21f9f
Open

mark14wu wants to merge 1 commit into
mainfrom
claude/eager-profiler-cpu-drivers-b21f9f

Conversation

@mark14wu

@mark14wu mark14wu commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

On CPU-only hosts, every launch traced with the Profiler failed before the kernel ran:

RuntimeError: 0 active drivers ([]). There should only be one.

The buffer load check is enabled by default (PROFILER_DISABLE_BUFFER_LOAD_CHECK=0), so Profiler.pre_warmup_callback asks TritonTrace.run to call the real jit_fn.warmup before each launch. Warmup compiles the kernel through triton.runtime.driver.active, which raises when no backend driver is available. The Profiler is the only client that requests warmup. The existing profiler tests all set profiler_disable_buffer_load_check = True, so CPU CI never ran this path.

Changes

  • Profiler.pre_warmup_callback checks driver.active first. With no driver, it skips warmup and turns off the buffer load check for that profiler, so the launch runs normally.
  • The finalize report now says Skipped: no active Triton driver to compile the kernel (e.g. CPU-only host). instead of dropping the section or flagging a buffer load issue without any ASM to check.
  • GPU hosts are unchanged: warmup still runs and the check reports the same as before.

Tests

  • Added test_buffer_load_check_without_active_driver in tests/end_to_end/test_profiler.py. It monkeypatches Triton's driver factory to simulate a driverless host, so it also covers this path on GPU machines. Without the fix it fails with the same 0 active drivers error.
  • pytest tests/end_to_end/test_profiler.py tests/unit/test_profiler.py: 24 passed.
  • Manual repro with CUDA_VISIBLE_DEVICES="" and a masked add kernel under @tilelens.trace("profiler"): it used to fail with 0 active drivers and now completes with the skip note. With the GPU visible (RTX 4090), the output is the same as before.
  • pre-commit (ruff, ruff-format, mypy, codespell, …) passes on the changed files.

…river

The buffer load check is on by default and makes every profiled launch
call the real jit_fn.warmup first. On CPU-only hosts Triton has no
backend driver, so warmup fails with "0 active drivers ([]). There
should only be one." and the launch never runs.

Probe triton's driver.active in Profiler.pre_warmup_callback; when no
driver is available, skip warmup, disable the buffer load check for
that profiler, and say so in the finalize report instead of failing or
reporting a spurious buffer load issue.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Performance Benchmark

Benchmark main (min) PR (min) Change Samples
gemm 0.104s 0.103s -0.9% 20 / 20
gemm_oob 0.116s 0.116s -0.0% 20 / 20
indirect_load 0.021s 0.021s +0.4% 20 / 20
nested_loop 0.233s 0.232s -0.1% 20 / 20
block_pointer_loop_advance 0.124s 0.123s -0.2% 20 / 20
liger_jsd 0.138s 0.138s +0.3% 20 / 20
flaggems_layernorm 0.392s 0.391s -0.3% 20 / 20
swiglu 0.169s 0.169s +0.0% 20 / 20
cross_entropy 0.965s 0.967s +0.1% 20 / 20
fused_linear_jsd 0.209s 0.208s -0.2% 20 / 20
Total 2.470s 2.470s -0.0% N/A

Iterations: 1 warmup + 20 measured
Samples are shown as main / PR; long pytest benchmarks may use fewer samples.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant