FastH3 text encoder in NVFP4

The Qwen3-VL text encoder of FastH3, https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree, stored in NVFP4 so that a single NVIDIA DGX Spark with its 128 GB of unified memory keeps about 24 GiB of encoder resident instead of 48 GiB. The 4-bit encoder reads prompts almost exactly like the bf16 one: its layer-50 output vectors differ by about 2 percent.

Drop it into any FastH3 model directory as the text_encoder component, or point FastVideo at it directly:

from fastvideo import VideoGenerator
from fastvideo.api import ComponentConfig, GeneratorConfig, PipelineSelection

config = GeneratorConfig(
    model_path="FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree",
    pipeline=PipelineSelection(components=ComponentConfig(text_encoder_weights="/path/to/FastH3-text-encoder-nvfp4")),
)
generator = VideoGenerator.from_config(config)
hf download KyleNeverGivesUp/FastH3-text-encoder-nvfp4 --local-dir /path/to/FastH3-text-encoder-nvfp4

FastH3 8-Step V2, https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2, ships a text encoder, tokenizer and processor byte-identical to the four-step v1 release, so this checkpoint drops into a V2 model directory the same way. The measurements below were taken with v1.

Measured on one DGX Spark

FastH3 with the FP8 DiT and the rank-16 AdaLN from FastVideo PR 1699, 832x480, 124 frames, one GB10, one process, lazy_module_load on, flashinfer-python 0.6.13rc2.

text encoder encoder resident peak memory end to end layer-50 cosine vs bf16 layer-50 relative error
bf16 48.0 GiB 49.7 GiB 225.7 s 1.000 0
this checkpoint about 24 GiB 41.6 GiB 188 to 195 s 0.995 1.4 to 3.0 percent

Once the encoder is quantized the peak is set by the DiT denoise phase, so the encode phase has headroom to spare. That headroom pays for keeping mlp.down_proj in bf16, which cuts the encoder error by a factor of 8 against a fully quantized variant at no cost in peak memory. Video SSIM against the bf16 run at the same seed is 0.76, above the 0.53 that two bf16 runs at different seeds score and below the 1.00 of a repeated run: a 2 percent change in conditioning moves a 4-step sampler onto a neighboring trajectory, so SSIM against one sample is not a quality ranking. The fully quantized variant also lost the voice track, and this checkpoint keeps it.

Measured on one RTX 5090

KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8, https://huggingface.co/KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8, pairs this encoder with an int8 transformer so FastH3 runs on one 32 GB GeForce RTX 5090. At 864x480, 124 frames and four transformer forwards it takes 27.3 s per clip with a 25.0 GB peak. The transformer streams in block by block on that card, so the peak is set by the encode phase rather than the denoise phase.

What is inside

50 language layers, the ones FastH3 reads, and no lm_head. In every layer the q_proj, k_proj, v_proj, o_proj, gate_proj and up_proj linears are stored as three tensors each:

tensor dtype shape meaning
weight_packed uint8 [out, in / 2] two E2M1 values per byte
weight_scale uint8 [out, in / 16] one E4M3 scale per 16 values, FlashInfer 128x4 layout
weight_global_scale float32 [1] 448 * 6 / amax(W)

mlp.down_proj stays bf16 in every layer. The token embedding, norms and the vision tower are copied unchanged. config.json carries the quantization_config block that makes FastVideo's loader select its serialized NVFP4 text-encoder path; no inference flag is needed.

26 GB on disk against 63 GB for the bf16 encoder.

How it was made

python scripts/checkpoint_conversion/convert_minimax_h3_text_encoder_nvfp4.py \
    --src FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree/text_encoder \
    --dst FastH3-text-encoder-nvfp4 --keep-bf16 mlp.down_proj

Weights are quantized with flashinfer.nvfp4_quantize in the exact layout the loader allocates, so loading reproduces what FastVideo's runtime NVFP4 conversion would build without ever materializing the bf16 weights. Activations are quantized per call and multiplied with flashinfer.mm_fp4. The loading path is FastVideo PR 1837, https://github.com/hao-ai-lab/FastVideo/pull/1837, and the converter is PR 1838, https://github.com/hao-ai-lab/FastVideo/pull/1838. Both were merged into main on 2026-09-13.

Requirements and limits

  • A Blackwell GPU, sm_100 or sm_120 class; the DGX Spark is sm_121 and the RTX 5090 is sm_120. Needs flashinfer-python.
  • On an RTX 5090 the first run spends several minutes compiling FlashInfer's FP4 GEMM kernels. That build needs cublasLt.h; with pip-installed CUDA, put the environment's nvidia/cu13/include and nvidia/cu13/lib on CPATH and LIBRARY_PATH.
  • Single GPU: the packed columns and swizzled scales cannot be split across tensor-parallel ranks.
  • This is a text encoder only. Pair it with a FastH3 model directory for the DiT, VAEs, tokenizer and scheduler.

Credits

Base checkpoint by FastVideo at Hao AI Lab, UC San Diego, https://github.com/hao-ai-lab/FastVideo, built on MiniMax-H3, https://huggingface.co/MiniMaxAI/MiniMax-H3. Licensed under the MiniMax-H3 community license, see LICENSE.

Downloads last month
68
Safetensors
Model size
17B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KyleNeverGivesUp/FastH3-text-encoder-nvfp4