Skip to content

Serialize TensorRT engine builds across GPUs - #1225

Open
zsqdx wants to merge 1 commit into
lightvector:masterfrom
zsqdx:agent/fix-trt-multigpu
Open

Serialize TensorRT engine builds across GPUs#1225
zsqdx wants to merge 1 commit into
lightvector:masterfrom
zsqdx:agent/fix-trt-multigpu

Conversation

@zsqdx

@zsqdx zsqdx commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Keep buildSerializedNetwork() under the existing process-wide TensorRT tuning mutex, including timing-cache hits.
  • Use RAII for the mutex so a build/cache exception cannot leave later NN server threads deadlocked.

Why

With TensorRT 10.16.1, concurrent warm-cache engine builds on different GPUs can fail nondeterministically. The existing cache-hit path released tuneMutex immediately before buildSerializedNetwork(), allowing all GPU threads to enter TensorRT's builder at once. On an 8x RTX 5060 Ti machine this reproduced as two SIGSEGVs and one hang, always after both threads loaded the timing cache and before either engine finished initialization.

This only serializes engine construction during startup. Engine deserialization, execution contexts, and inference remain per-GPU and concurrent. It does not change the graph, precision, or inference hot path.

Validation

Tested with KataGo v1.17.1, TensorRT 10.16.1, CUDA 13.2, and b11c768h12nbt3tflrs-fson-silu:

  • dual-GPU warm-cache stress: 12/12 passes after the fix
  • 8-GPU cold- and warm-cache benchmark: both pass, with every GPU processing rows
  • USE_CACHE_TENSORRT_PLAN=1: builds and passes dual-GPU cold/warm smoke tests
  • katago runtests: all tests pass

@zsqdx
zsqdx marked this pull request as ready for review August 3, 2026 20:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant