Skip to content

Encode prompts in BF16 without stalling on references, and cut references to the generated length (0.2.4) - #72

Merged
KBRASK merged 2 commits into
mainfrom
fix/encoder-bf16-references
Oct 6, 2026
Merged

KBRASK merged 2 commits into
mainfrom
fix/encoder-bf16-references

Conversation

@KBRASK

@KBRASK KBRASK commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

Prompts with reference images or videos could stop responding during encoding on 16 GB cards. ComfyUI's text-encoder base ran the H3 Qwen3-VL encoder in FP32, and on Windows a long reference pushed the encode into shared GPU memory, where it ran many times slower. This runs the encoder in BF16 like the official pipeline, keeps it in dedicated GPU memory, and cuts references to the generated length as the official pipeline does. Release 0.2.4.

BF16, as in the official pipeline

  • The language model gets BF16 embeddings, DeepStack features and masks instead of FP32, so the NVFP4 weights no longer run as FP32 matmuls.
  • The BF16 output is kept as it is; FreeVideo stores conditioning in BF16 anyway.

Memory planned from the real input

  • Room is planned from the sequence the language model really sees: each image or video block expands to its patches / 4.
  • It is checked against the dedicated memory actually free after loading. On Windows this is also bounded by the DXGI budget, which cudaMemGetInfo overstates.
  • Encoder weights move back to host memory until the forward fits.
  • Inputs too long for one forward run the language model in blocks, without a dense T×T mask.
  • DeepStack features are joined once instead of once per image, and each vision block returns its cache. Without expandable segments, 33 blocks had grown reserved memory from 12.7 to 26.5 GiB.

Never in shared GPU memory

  • On Windows a watcher stops a forward whose allocations outgrow the dedicated room, within seconds.
  • The encode retries by itself with more room or in blocks.
  • What each forward used is remembered per machine and per kind of input, from the latest forwards only.

References cut to the generated length, and said so

  • A reference video or audio clip contributes as much as the video being generated, at most 15 s, as in the official pipeline and ComfyUI's native H3 node.
  • Before, clips over 15 s were refused and shorter ones were used whole. A 15 s reference on a 5 s video gave the model 2.5× the vision tokens and 2.8× the VAE frames and transformer reference rows the official pipeline would.
  • The media card, the generation progress and the result say which clip was cut and how much was used. A reused input cache reports it too.

Measured on an RTX 5060 Ti 16 GB (Windows, 832×480, 5 s, the same long prompt; prompt encoding and the encoder's PyTorch peak)

Input 0.2.3 This PR
Text only 22.7 s, 13.9 GiB 20.0–21.6 s, 13.4 GiB
1 reference image 34.5 s, 15.0 GiB 22.8 s, 12.6 GiB
2 reference images 37.0 s, 15.2 GiB 22.3 s, 12.4 GiB
3 reference images 41.2 s, 15.4 GiB 24.0 s, 12.3 GiB
14.5 s reference video + 3 images Stopped responding for over 10 min 30.2 s, 11.5 GiB
Two 15 s videos + 3 images, 1344×768 Out of memory 3 times Completes

KBRASK added 2 commits October 6, 2026 08:29
…nces to the generated length (0.2.4)

ComfyUI's text-encoder base ran the H3 Qwen3-VL encoder in FP32, so the NVFP4
weights were dequantized to FP32 matmuls. The official MiniMax pipeline runs it
in BF16, and so does FreeVideo now.

Encoding memory is planned from the sequence the language model really sees,
checked against the dedicated memory actually free after loading (on Windows
also the DXGI budget), and encoder weights move back to host memory until the
forward fits. Inputs too long for one forward run the language model in blocks
without a dense attention mask. DeepStack features are joined once, and each
vision block returns its cache. On Windows a forward that starts to spill into
shared GPU memory is stopped within seconds and retried with more room, and
the room each kind of input needed is remembered per machine.

References are cut to the generated length, at most 15 s, as the official
pipeline and ComfyUI's native H3 node do, instead of refusing clips over 15 s
and using shorter ones whole. The media card, the generation progress and the
result say which clip was cut and how much was used.

On an RTX 5060 Ti 16 GB (Windows, 832x480, 5 s), prompt encoding with one to
three reference images takes 22-24 s instead of 35-41 s at a 12.3-12.6 GiB
instead of a 15.0-15.4 GiB peak; a 14.5 s reference video with three images,
which stalled for over ten minutes in shared memory, encodes in 30 s; two 15 s
videos and three images at 1344x768, which failed out of memory, now complete.
@KBRASK
KBRASK merged commit f0960b8 into main Oct 6, 2026
4 checks passed
@KBRASK
KBRASK deleted the fix/encoder-bf16-references branch October 6, 2026 08:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant