Repository navigation
Encode prompts in BF16 without stalling on references, and cut references to the generated length (0.2.4) - #72
Merged
Merged
Conversation
…nces to the generated length (0.2.4) ComfyUI's text-encoder base ran the H3 Qwen3-VL encoder in FP32, so the NVFP4 weights were dequantized to FP32 matmuls. The official MiniMax pipeline runs it in BF16, and so does FreeVideo now. Encoding memory is planned from the sequence the language model really sees, checked against the dedicated memory actually free after loading (on Windows also the DXGI budget), and encoder weights move back to host memory until the forward fits. Inputs too long for one forward run the language model in blocks without a dense attention mask. DeepStack features are joined once, and each vision block returns its cache. On Windows a forward that starts to spill into shared GPU memory is stopped within seconds and retried with more room, and the room each kind of input needed is remembered per machine. References are cut to the generated length, at most 15 s, as the official pipeline and ComfyUI's native H3 node do, instead of refusing clips over 15 s and using shorter ones whole. The media card, the generation progress and the result say which clip was cut and how much was used. On an RTX 5060 Ti 16 GB (Windows, 832x480, 5 s), prompt encoding with one to three reference images takes 22-24 s instead of 35-41 s at a 12.3-12.6 GiB instead of a 15.0-15.4 GiB peak; a 14.5 s reference video with three images, which stalled for over ten minutes in shared memory, encodes in 30 s; two 15 s videos and three images at 1344x768, which failed out of memory, now complete.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Prompts with reference images or videos could stop responding during encoding on 16 GB cards. ComfyUI's text-encoder base ran the H3 Qwen3-VL encoder in FP32, and on Windows a long reference pushed the encode into shared GPU memory, where it ran many times slower. This runs the encoder in BF16 like the official pipeline, keeps it in dedicated GPU memory, and cuts references to the generated length as the official pipeline does. Release 0.2.4.
BF16, as in the official pipeline
Memory planned from the real input
Never in shared GPU memory
References cut to the generated length, and said so
Measured on an RTX 5060 Ti 16 GB (Windows, 832×480, 5 s, the same long prompt; prompt encoding and the encoder's PyTorch peak)