Summary
_grow_allocation_slow_path (cuda_core/cuda/core/_memory/_virtual_memory_resource.py:362) frees the old buffer's VA range manually (cuMemAddressFree(int(buf.handle), aligned_prev_size), line 457), then calls buf._clear() (line 460) with a comment saying this stops the old buffer's destructor from freeing it again.
Buffer._clear() (_buffer.pyx:262) does self._h_ptr.reset(). For a shared_ptr, reset() isn't a no-op invalidation — if buf holds the last reference, it runs the deleter immediately. This buffer's handle was created with mr=self (Buffer.from_handle(..., mr=self) in allocate()), so its deleter calls back into mr.deallocate() with the old, already-freed pointer and size. deallocate() then calls cuMemRetainAllocationHandle on a VA range that no longer exists, which fails with CUDA_ERROR_INVALID_VALUE. Under the error-handling policy landed in #2759, that failure is caught in the deleter and reported as a CUDAWarning instead of being silent; previously it was silently swallowed.
Evidence
Seen in CI, test_vmm_allocator_grow_allocation:
_virtual_memory_resource.py:292: CUDAWarning: mr.deallocate() failed during Buffer
destruction; the allocation may have leaked: CUDA_ERROR_INVALID_VALUE: This indicates
that one or more of the parameters passed to the API call is not within an acceptable
range of values.
Hypothesis, not yet confirmed by reproduction
Read from the code during review of #2759, not verified by running the test with added instrumentation.
Suggested fix direction
_clear()'s intent here seems to be "drop bookkeeping for a range this code already freed by hand," not "release the handle normally." Those need to be different operations: either give Buffer a way to drop its handle without invoking the resource's deallocate() (e.g. release the underlying pointer from the shared_ptr without running the deleter), or have the grow path swap in a handle whose deleter is already a no-op.
Refs: found during review of #2759.
Summary
_grow_allocation_slow_path(cuda_core/cuda/core/_memory/_virtual_memory_resource.py:362) frees the old buffer's VA range manually (cuMemAddressFree(int(buf.handle), aligned_prev_size), line 457), then callsbuf._clear()(line 460) with a comment saying this stops the old buffer's destructor from freeing it again.Buffer._clear()(_buffer.pyx:262) doesself._h_ptr.reset(). For ashared_ptr,reset()isn't a no-op invalidation — ifbufholds the last reference, it runs the deleter immediately. This buffer's handle was created withmr=self(Buffer.from_handle(..., mr=self)inallocate()), so its deleter calls back intomr.deallocate()with the old, already-freed pointer and size.deallocate()then callscuMemRetainAllocationHandleon a VA range that no longer exists, which fails withCUDA_ERROR_INVALID_VALUE. Under the error-handling policy landed in #2759, that failure is caught in the deleter and reported as aCUDAWarninginstead of being silent; previously it was silently swallowed.Evidence
Seen in CI,
test_vmm_allocator_grow_allocation:Hypothesis, not yet confirmed by reproduction
Read from the code during review of #2759, not verified by running the test with added instrumentation.
Suggested fix direction
_clear()'s intent here seems to be "drop bookkeeping for a range this code already freed by hand," not "release the handle normally." Those need to be different operations: either giveBuffera way to drop its handle without invoking the resource'sdeallocate()(e.g. release the underlying pointer from theshared_ptrwithout running the deleter), or have the grow path swap in a handle whose deleter is already a no-op.Refs: found during review of #2759.