How to use from the
Use from the
Diffusers library
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("QuantFunc/LTX-2.5-QuantFunc-4bit", dtype=torch.bfloat16, device_map="cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]

LTX-2.5-QuantFunc-4bit

4-bit quantized Lightricks LTX-2.5 22B distilled โ€” pure 4-bit inference holds up to BF16 quality while cutting weight/VRAM footprint to roughly 1/3. Quantized and loaded by the QuantFunc inference engine.

Why pure 4-bit holds up

Compressing a 22B joint video/audio DiT down to 4-bit activations ร— 4-bit weights (pure a4w4) runs into one main enemy: activation outliers, which blow past 4-bit's dynamic range and show up as blur and lost detail. QuantFunc adds two layers on top of the SVDQuant approach:

  1. A much stronger low-rank branch. Algorithmic improvements let the low-rank branch absorb significantly more activation-outlier energy than standard SVDQuant โ€” outliers get siphoned off before the main path is quantized to 4-bit, so the main path sees a much cleaner distribution.
  2. A smoother residual. A combination of in-house techniques further smooths the outlier distribution in what's left after low-rank absorption, so the 4-bit grid covers the remaining distribution with far less error.

Net effect: at pure a4w4 inference, output video (including audio) clarity, detail retention and motion coherence are on par with the BF16 baseline โ€” at roughly 1/3 the memory.

Files

Quantizes the base checkpoint โ€” ltx-2.5-22b-distilled-transformer-bf16.safetensors from Lightricks LTX-2.5 (22B, distilled variant):

File Format Size Minimum GPU
ltx-2.5-22b-distilled-quantfunc-int4-r128.safetensors INT4 weights / INT4 activations (SVDQuant, rank-128) ~14.06 GiB NVIDIA SM120+ (Blackwell, e.g. RTX 50-series / RTX 6000D)

Produced & loadable ONLY by QuantFunc

This is not a standard diffusers/safetensors checkpoint. The file is produced by, and only loadable by, the QuantFunc inference engine (native C++/CUDA engine) or the ComfyUI-QuantFunc custom-node plugin for ComfyUI. It will not load in vanilla diffusers, ComfyUI's stock LTX loader, or any other inference stack. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors file below โ€” see HF's download-stats docs โ€” not because the file loads via the diffusers Python package.)

Inference settings (required for correct output)

This is the distilled checkpoint, joint video+audio. It must be run with the distilled model's own schedule โ€” a generic/default scheduler will produce degraded output:

  • Sigma schedule (8 steps, fixed): 1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875, 0.0
  • CFG: 1 (no classifier-free guidance โ€” this is a guidance-distilled model)
  • STG: disabled (no Spatio-Temporal Guidance)

Original / base model

This repository redistributes derived (quantized) weights only, produced from the base checkpoint above via the QuantFunc SVDQuant pipeline. No architecture or training changes were made โ€” only post-training quantization (weight + activation, 4-bit, rank-128 SVD residual).

License

Follows the original model's LTX-2 Community License Agreement (license text). This repository distributes quantized weights only; copyright and licensing of the original model belong to Lightricks.

Community

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for QuantFunc/LTX-2.5-QuantFunc-4bit

Quantized
(30)
this model