Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ https://github.com/user-attachments/assets/ecda7d0d-7fbe-4e0c-8c29-8f3315bafc15

## News

- **2026-10-07** · **[v0.3.0](https://github.com/FlashML-org/FreeVideo/releases/tag/v0.3.0): a faster int8 model and prompt enhancement.** GeForce RTX 40/50 and RTX 30-series cards switch to an int8 model that is faster and closer to the original quality; an RTX 3090 makes a 5-second video about 2.3x faster. Existing installations upgrade in one click, and an optional local prompt enhancer rewrites prompts for MiniMax H3.
- **2026-10-06** · **[Gallery](https://freevideo-community.pages.dev/#gallery) is live.** Watch 20-second clips made with FreeVideo, and the four quality levels side by side.
- **2026-10-06** · **[v0.2.3](https://github.com/FlashML-org/FreeVideo/releases/tag/v0.2.3): videos carry their workflow.** Drop a FreeVideo video onto the ComfyUI canvas to restore its prompt, seed and settings. Earlier videos can get theirs too.
- **2026-10-05** · **[v0.2.0](https://github.com/FlashML-org/FreeVideo/releases/tag/v0.2.0): four quality levels.** Choose Light, Medium, High or Max for each video; higher levels give higher quality but take longer. Results can be exported as sharing images or videos with the generation time and GPU.
Expand All @@ -32,7 +33,7 @@ FreeVideo is a local inference engine for MiniMax H3 on consumer GPUs, built on

It coordinates VRAM, system memory and disk, adapting weight placement, compute precision and attention kernels to the available hardware. FreeVideo runs as a ComfyUI plugin, with a Windows launcher for setup and command-line support on Linux. Its core features include:

- **Hardware-adaptive execution**: Chooses the FP8 compute path for each GPU architecture, either native FP8 or FP8 storage with BF16 compute, and automatically probes the available attention kernels.
- **Hardware-adaptive execution**: Chooses the weight format for each GPU, int8 on GeForce and RTX 30-series cards and FP8 on workstation and datacenter cards, and automatically probes the available attention kernels.
- **Low-memory inference**: Weight streaming, asynchronous prefetching and chunked computation keep peak memory low, enabling inference with as little as 8 GB of VRAM and 16 GB of RAM.
- **Multimodal inputs**: Text prompts, first and last frames, and image, video and audio references.
- **Community LoRAs**: Use MiniMax H3 LoRAs in your workflow. See [examples](docs/LoRA.md).
Expand Down
3 changes: 2 additions & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ https://github.com/user-attachments/assets/ecda7d0d-7fbe-4e0c-8c29-8f3315bafc15

## 更新动态

- **2026-10-07** · **[v0.3.0](https://github.com/FlashML-org/FreeVideo/releases/tag/v0.3.0):更快的 int8 模型和提示词增强。** GeForce RTX 40/50 系和 RTX 30 系显卡改用更快、画质更接近原版的 int8 模型,RTX 3090 生成 5 秒视频快约 2.3 倍;已安装的用户可一键升级。新增可选的本地提示词增强,把提示词改写成适合 MiniMax H3 的写法。
- **2026-10-06** · **[作品展示](https://freevideo-community.pages.dev/#gallery)上线。** 观看 FreeVideo 生成的 20 秒视频,以及四档质量的并排对比。
- **2026-10-06** · **[v0.2.3](https://github.com/FlashML-org/FreeVideo/releases/tag/v0.2.3):视频自带工作流。** 把 FreeVideo 生成的视频拖到 ComfyUI 画布上,即可还原提示词、种子和全部参数,之前的视频也能补上。
- **2026-10-05** · **[v0.2.0](https://github.com/FlashML-org/FreeVideo/releases/tag/v0.2.0):四档质量。** 每次生成可选择轻量、标准、精细或极致,档位越高,生成质量越高,但耗时更长;生成结果可导出为带生成耗时和显卡信息的分享图片或视频。
Expand All @@ -32,7 +33,7 @@ FreeVideo 是面向消费级显卡的 MiniMax H3 本地推理引擎,基于 [Op

它统一调度显存、内存与磁盘,并根据硬件条件调整权重放置、计算精度和注意力内核。FreeVideo 以 ComfyUI 插件形式提供,配备 Windows 启动器,也支持 Linux 命令行。主要特性包括:

- **硬件自适应**:针对不同显卡架构选择 FP8 计算路径(原生 FP8,或 FP8 存储配合 BF16 计算),并自动探测可用的注意力内核,无需手动配置。
- **硬件自适应**:针对不同显卡选择权重格式(GeForce 和 RTX 30 系用 int8,专业卡和数据中心卡用 FP8),并自动探测可用的注意力内核,无需手动配置。
- **低显存推理**:通过权重流式加载、异步预取与分块计算降低峰值显存,最低只需 8GB 显存和 16GB 内存。
- **多模态输入**:支持文本、首帧、尾帧,以及图像、视频、音频参考输入。
- **社区 LoRA**:支持在工作流中使用 MiniMax H3 社区 LoRA。[查看效果对比](docs/LoRA.zh-CN.md)。
Expand Down
6 changes: 3 additions & 3 deletions docs/execution-planning.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
# FreeVideo Adaptive Execution Planner

The FreeVideo Adaptive Execution Planner computes an execution plan for every request from live device and host measurements. The plan fixes the residency of the 50 transformer blocks across VRAM, pinned host memory and disk, the transfer schedule, the FP8 GEMM path, the attention backend and head grouping, activation staging and the VAE decoder placement, within the measured VRAM and host-memory budgets.
The FreeVideo Adaptive Execution Planner computes an execution plan for every request from live device and host measurements. The plan fixes the residency of the 50 transformer blocks across VRAM, pinned host memory and disk, the transfer schedule, the int8 or FP8 GEMM path, the attention backend and head grouping, activation staging and the VAE decoder placement, within the measured VRAM and host-memory budgets.

## Inputs

| Input | Source | Determines |
| --- | --- | --- |
| Architecture and compute capability | CUDA device properties | FP8 GEMM path and kernel set |
| Architecture and compute capability | CUDA device properties | Weight format, GEMM path and kernel set |
| Free VRAM | `torch.cuda.mem_get_info` | VRAM budget |
| Available host memory | OS memory counters, cgroup v1/v2 limit, Windows commit headroom | Host-memory budget |
| Attention backends | On-device kernel probes | Backend selection |
Expand All @@ -23,7 +23,7 @@ VRAM budget: free VRAM minus a reserve of 2.5% of free VRAM, clamped to 0.5–1
| Transfer schedule | Two transfer slots (prefetch) at a VRAM budget of 14 GiB or more, with lower thresholds on Windows Blackwell and Ampere; one slot otherwise. |
| Attention | The first backend that passes its probe, in the order SageAttention 2, PyTorch flash attention, cuDNN, FlashAttention 2, FlashAttention 4. Head group: the widest of 16, 8 and 4 heads whose measured activation footprint at the request's token count fits the VRAM budget. |
| Activation staging | Below a 10 GiB VRAM budget, the residual stream and attention outputs are staged in host buffers when the on-device activation path does not fit. |
| FP8 GEMM | Per-tensor scales on Blackwell (SM120), per-channel scales on Ada (SM89) and Hopper (SM90), FP8 weights with BF16 compute on Ampere (SM80, SM86). |
| GEMM | int8 on GeForce Ada and Blackwell and on every Ampere card: activations get the cache's ConvRot rotation and one int8 scale per row. Other cards use FP8: per-tensor scales on Blackwell (SM120), per-channel scales on Ada (SM89) and Hopper (SM90). |
| Chunking | Feed-forward chunk 2048, projection chunk 1024, window batch 4 (1 on Ampere). |
| Text encoder | Runs in a separate process that exits before the transformer is loaded. |
| VAE decoder | Below a 20 GiB VRAM budget, ⌊(VRAM budget − 3.25 GiB) / 268.6 MB⌋ of its 36 blocks stay resident and the rest are streamed; temporal clips are decoded in sequence with the original tiles and blending. |
Expand Down
6 changes: 3 additions & 3 deletions docs/zh-CN/execution-planning.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
# FreeVideo Adaptive Execution Planner

FreeVideo Adaptive Execution Planner(自适应执行规划器)根据设备与主机的实时测量结果,为每个请求计算执行计划。执行计划在实测的显存与内存预算内确定:50 个 Transformer 块在显存、锁页内存与磁盘之间的驻留位置,传输调度,FP8 GEMM 路径,注意力后端与头分组,激活暂存,以及 VAE 解码器的放置。
FreeVideo Adaptive Execution Planner(自适应执行规划器)根据设备与主机的实时测量结果,为每个请求计算执行计划。执行计划在实测的显存与内存预算内确定:50 个 Transformer 块在显存、锁页内存与磁盘之间的驻留位置,传输调度,int8 或 FP8 GEMM 路径,注意力后端与头分组,激活暂存,以及 VAE 解码器的放置。

## 输入

| 输入 | 来源 | 决定 |
| --- | --- | --- |
| 架构与计算能力 | CUDA 设备属性 | FP8 GEMM 路径与内核集 |
| 架构与计算能力 | CUDA 设备属性 | 权重格式、GEMM 路径与内核集 |
| 空闲显存 | `torch.cuda.mem_get_info` | 显存预算 |
| 可用内存 | 系统内存计数、cgroup v1/v2 上限、Windows 提交余量 | 内存预算 |
| 注意力后端 | 设备上的内核探测 | 后端选择 |
Expand All @@ -23,7 +23,7 @@ FreeVideo Adaptive Execution Planner(自适应执行规划器)根据设备
| 传输调度 | 显存预算不低于 14 GiB 时使用两个传输槽(预取),Windows Blackwell 与 Ampere 的门槛更低;否则使用一个传输槽。 |
| 注意力 | 按 SageAttention 2、PyTorch flash attention、cuDNN、FlashAttention 2、FlashAttention 4 的顺序选择第一个通过探测的后端。头分组:在 16、8、4 头中选择该请求 token 数下实测激活占用不超过显存预算的最宽分组。 |
| 激活暂存 | 显存预算低于 10 GiB 且设备上的激活路径放不下时,把残差流和注意力输出暂存在主机缓冲区。 |
| FP8 GEMM | Blackwell(SM120)逐张量缩放,Ada(SM89)与 Hopper(SM90)逐通道缩放,Ampere(SM80、SM86)以 FP8 存储权重、BF16 计算。 |
| GEMM | GeForce Ada、Blackwell 与所有 Ampere 显卡使用 int8:激活按缓存的 ConvRot 旋转后逐行 int8 量化。其他显卡使用 FP8:Blackwell(SM120)逐张量缩放,Ada(SM89)与 Hopper(SM90)逐通道缩放。 |
| 分块 | 前馈分块 2048,投影分块 1024,窗口批次 4(Ampere 为 1)。 |
| 文本编码器 | 在独立进程中运行,Transformer 加载前退出。 |
| VAE 解码器 | 显存预算低于 20 GiB 时,36 个块中 ⌊(显存预算 − 3.25 GiB) / 268.6 MB⌋ 个常驻,其余流式加载;按时间片段依次解码,分块与融合方式与原解码器相同。 |
Expand Down
7 changes: 6 additions & 1 deletion freevideo_engine/backends/cuda.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ class CUDABackend(DeviceBackend):
device = 'cuda'
capabilities = BackendCapabilities(
name='cuda', memory_model='dedicated',
linear_policies=('native-fp8', 'bf16-weight-only'),
linear_policies=('native-fp8', 'bf16-weight-only', 'int8'),
attention_candidates=('cudnn', 'torch-flash', 'sage2', 'fa2', 'fa4'),
pinned_host_weights=True, streamed_weights=True)

Expand Down Expand Up @@ -85,6 +85,11 @@ def prepare_linears(self, model, manifest, linear_compute, fp8_gemm):
if linear_compute == 'native-fp8':
from ..fp8_gemm import install
actual_fp8_gemm = install(model, manifest['scale_granularity'], requested=fp8_gemm)
elif precision == 'int8':
if linear_compute != 'int8':
raise ValueError('Int8 weights on CUDA run with the int8 linear policy')
from ..int8_ops import install_cached_linears as install_int8_linears
install_int8_linears(model, manifest['linears'], manifest.get('rotation'))
elif precision != 'bf16':
raise ValueError('Unknown prepared weight precision')
return actual_fp8_gemm
Expand Down
Loading