Conversation
Small `KumoTabular` forwards leave the GPU waiting for kernel launches. Replay CUDA graphs for repeated input layouts, e.g., across folds of a table, with the graph cache from #971, but decide per layout from measurements instead of a fixed input size limit. `GraphCache` times eager calls with CUDA events and, once two of them have finished, captures a graph into its own memory pool. Replays after the first one are timed as well. The graph is kept if two of them save a larger share of the fastest eager call's time than the share of device memory taken by its pool, and dropped together with its pool as soon as one does not, so graphs of large, compute-bound inputs do not hold on to memory. Layouts whose eager calls take at least as long as one whose graph was dropped are not captured. Timings are read from completed events on later calls, so no synchronization is added. Signed-off-by: Jingang Qu <jqu@nvidia.com>
JingangQu
requested review from
RBendias,
akihironitta,
aw471 and
rusty1s
as code owners
September 27, 2026 09:13
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #994. Small
KumoTabularforwards leave the GPU waiting for kernel launches. This PR replays CUDA graphs for repeated input layouts, e.g., across folds of a table, with the graph cache from #971, but decides per layout from measurements instead of a fixed input size limit.Changes
Replay CUDA graphs of repeated forwards
_KumoTabular.forwardruns inference throughGraphCacheon CUDA, without gradients or the key/value cache and outside oftorch.compile. Graphs are keyed by input layout, autocast state and parameter storage, and replays are bit-identical to eager execution.Decide per layout from measurements
GraphCachetimes eager calls with CUDA events. Once two of them have finished, it captures a graph into its owntorch.cuda.MemPool. The first replay, which also uploads the graph, is not timed.Results
TabArena, outer protocol, 816 splits, Kumo-Tabular-S (8 estimators, fp16, no KV cache), 8× NVIDIA RTX PRO 6000 Blackwell with one task per GPU; each branch was run on its own:
diabetes).seismic-bumps0.86x,wine_quality0.87x).GiveMeSomeCredit,APSFailure,kddcup09_appetency), replays save at most 0.5% of the time while their pools would take 3–8% of device memory, so no graph is kept and no memory is held.