Skip to content

Use a faster int8 model on GeForce and RTX 30-series cards, with a one-click upgrade (0.3.0) - #80

Merged
KBRASK merged 1 commit into
mainfrom
release/int8-030
Oct 6, 2026
Merged

KBRASK merged 1 commit into
mainfrom
release/int8-030

Conversation

@KBRASK

@KBRASK KBRASK commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

GeForce RTX 40/50 and RTX 30-series cards now run the prepared ConvRot int8 model, which is faster on these cards and closer to the original weights than FP8. Existing installations keep their model until the user upgrades from the launcher in one click: the new model downloads beside the current one, the switch happens between videos, and only then are the old model's own files removed. Release 0.3.0, which also ships the prompt enhancement (#75) and the installer retries and download cache cleanup (#73).

int8 on CUDA

  • Activations get the cache's group-256 ConvRot rotation and one scale per row; a Triton GEMM multiplies them with the int8 weight rows, and the int32 products are exact before the scale epilogue.
  • Chosen for GeForce Ada and Blackwell, where FP8 with FP32 accumulation runs at half rate, and for every Ampere card, which has no FP8 tensor cores. Workstation and datacenter cards keep FP8, which measured 5% faster on an RTX PRO 6000.
  • int8 rows carry their own scales, so the planner reserves no FP8 activation stash and spends the room on larger attention batches.
  • Setup checks the int8 path on the GPU before it finishes an int8 installation. Low-rank LoRAs run beside the int8 projections; adapters that must be merged keep the FP8 model.

Measured (Windows, 1344×768, Light 8+3, warm requests)

GPU Video FP8 int8
RTX 3090 24 GB 5 s 433.7 s 187.5 s (2.31×; sampling 329 → 113 s)
RTX 4070 12 GB 10 s 558 s (one out-of-memory retry) 433 s
RTX 5060 Ti 16 GB 15 s 945 s (FP8 staged to host memory) 655 s

On matched noise against BF16, int8 was closer than per-tensor FP8 in 7 of 7 scenes, row-wise FP8 in 6 of 7 and BF16 weight-only in 5 of 7.

One-click upgrade for existing installations

  • An engine update keeps the installed model and never starts a download. The launcher shows a card with the expected speedup, the download and the space freed; nothing changes until the user clicks.
  • The download is planned and approved exactly as shown, and runs beside the model in use with only the setup lease, so generation can continue.
  • Between videos the launcher stops the ComfyUI it started, because its resident worker keeps SageAttention's DLL loaded and Windows refuses to replace it. Setup checks the int8 kernels before it marks the installation unfinished, so a GPU that fails them keeps its configuration.
  • The retired variant is then removed file by file: only catalogued files at their catalogued sizes, their download sidecars, interrupted partials, staging folders and engine-generated AdaLN tables bound to their identity. Nothing is removed through links or junctions, nothing shared with a LoRA variant and never the model in use; the setup, engine and host-setup leases are held, every fingerprint is rechecked and a receipt is written.
  • A move interrupted before the removal, or an int8 download the GPU could not run, is offered again on the card and in Settings → Storage and removed only after confirmation.

Verified on real machines

  • RTX 3090, Windows 10: fresh v0.2.4 install, engine update, card shown without generating anything, 21.4 GiB download, switch with the int8 probe passing on sm_86, old model removed: 557 files, 21.355 GiB, 0 skipped. The inventory of 129,246 files before and after matched the receipt; the only other removals were the replaced engine package metadata. The leftover path removed a restored copy again after confirmation.
  • RTX 4070: a command-line migration removed 559 files with none skipped, and removing the int8 variant again (302 files) matched the inventory diff exactly.

Setup runs from an open launcher (Install / repair, the switch) no longer fail on Windows after a video, since the launcher stops its idle ComfyUI first. Release notes, README News and the planning docs are updated.

…e-click upgrade (0.3.0)

GeForce RTX 40/50 and RTX 30-series cards run the prepared ConvRot int8
model: activations get the cache's group-256 rotation and one scale per
row, and a Triton GEMM multiplies them with the int8 weight rows.
Workstation and datacenter cards keep FP8. Setup checks the int8 kernels
on the GPU. On an RTX 3090 a 5 s 1344x768 Light request went from 434 s
to 188 s.

Existing installations keep their model until the user upgrades from the
launcher: the new model downloads beside the current one, the switch runs
between videos after the launcher stops its own ComfyUI and setup checks
the int8 kernels, and then the retired variant is removed file by file
with a receipt. Leftovers are offered again and removed only after
confirmation.

Release 0.3.0: release notes, README News and the planning docs.
@KBRASK
KBRASK merged commit 355c7d8 into main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant