FastH3 text encoder in NVFP4
The Qwen3-VL text encoder of FastH3, https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree, stored in NVFP4 so that a single NVIDIA DGX Spark with its 128 GB of unified memory keeps about 24 GiB of encoder resident instead of 48 GiB. The 4-bit encoder reads prompts almost exactly like the bf16 one: its layer-50 output vectors differ by about 2 percent.
Drop it into any FastH3 model directory as the text_encoder component, or point FastVideo at it directly:
from fastvideo import VideoGenerator
from fastvideo.api import ComponentConfig, GeneratorConfig, PipelineSelection
config = GeneratorConfig(
model_path="FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree",
pipeline=PipelineSelection(components=ComponentConfig(text_encoder_weights="/path/to/FastH3-text-encoder-nvfp4")),
)
generator = VideoGenerator.from_config(config)
hf download KyleNeverGivesUp/FastH3-text-encoder-nvfp4 --local-dir /path/to/FastH3-text-encoder-nvfp4
FastH3 8-Step V2, https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2, ships a text encoder, tokenizer and processor byte-identical to the four-step v1 release, so this checkpoint drops into a V2 model directory the same way. The measurements below were taken with v1.
Measured on one DGX Spark
FastH3 with the FP8 DiT and the rank-16 AdaLN from FastVideo PR 1699, 832x480, 124 frames, one GB10, one process, lazy_module_load on, flashinfer-python 0.6.13rc2.
| text encoder | encoder resident | peak memory | end to end | layer-50 cosine vs bf16 | layer-50 relative error |
|---|---|---|---|---|---|
| bf16 | 48.0 GiB | 49.7 GiB | 225.7 s | 1.000 | 0 |
| this checkpoint | about 24 GiB | 41.6 GiB | 188 to 195 s | 0.995 | 1.4 to 3.0 percent |
Once the encoder is quantized the peak is set by the DiT denoise phase, so the encode phase has headroom to spare. That headroom pays for keeping mlp.down_proj in bf16, which cuts the encoder error by a factor of 8 against a fully quantized variant at no cost in peak memory. Video SSIM against the bf16 run at the same seed is 0.76, above the 0.53 that two bf16 runs at different seeds score and below the 1.00 of a repeated run: a 2 percent change in conditioning moves a 4-step sampler onto a neighboring trajectory, so SSIM against one sample is not a quality ranking. The fully quantized variant also lost the voice track, and this checkpoint keeps it.
Measured on one RTX 5090
KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8, https://huggingface.co/KyleNeverGivesUp/FastH3-4-step-Preview-v1-r16-int8, pairs this encoder with an int8 transformer so FastH3 runs on one 32 GB GeForce RTX 5090. At 864x480, 124 frames and four transformer forwards it takes 27.3 s per clip with a 25.0 GB peak. The transformer streams in block by block on that card, so the peak is set by the encode phase rather than the denoise phase.
What is inside
50 language layers, the ones FastH3 reads, and no lm_head. In every layer the q_proj, k_proj, v_proj, o_proj, gate_proj and up_proj linears are stored as three tensors each:
| tensor | dtype | shape | meaning |
|---|---|---|---|
weight_packed |
uint8 | [out, in / 2] |
two E2M1 values per byte |
weight_scale |
uint8 | [out, in / 16] |
one E4M3 scale per 16 values, FlashInfer 128x4 layout |
weight_global_scale |
float32 | [1] |
448 * 6 / amax(W) |
mlp.down_proj stays bf16 in every layer. The token embedding, norms and the vision tower are copied unchanged. config.json carries the quantization_config block that makes FastVideo's loader select its serialized NVFP4 text-encoder path; no inference flag is needed.
26 GB on disk against 63 GB for the bf16 encoder.
How it was made
python scripts/checkpoint_conversion/convert_minimax_h3_text_encoder_nvfp4.py \
--src FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree/text_encoder \
--dst FastH3-text-encoder-nvfp4 --keep-bf16 mlp.down_proj
Weights are quantized with flashinfer.nvfp4_quantize in the exact layout the loader allocates, so loading reproduces what FastVideo's runtime NVFP4 conversion would build without ever materializing the bf16 weights. Activations are quantized per call and multiplied with flashinfer.mm_fp4. The loading path is FastVideo PR 1837, https://github.com/hao-ai-lab/FastVideo/pull/1837, and the converter is PR 1838, https://github.com/hao-ai-lab/FastVideo/pull/1838. Both were merged into main on 2026-09-13.
Requirements and limits
- A Blackwell GPU, sm_100 or sm_120 class; the DGX Spark is sm_121 and the RTX 5090 is sm_120. Needs
flashinfer-python. - On an RTX 5090 the first run spends several minutes compiling FlashInfer's FP4 GEMM kernels. That build needs
cublasLt.h; with pip-installed CUDA, put the environment'snvidia/cu13/includeandnvidia/cu13/libonCPATHandLIBRARY_PATH. - Single GPU: the packed columns and swizzled scales cannot be split across tensor-parallel ranks.
- This is a text encoder only. Pair it with a FastH3 model directory for the DiT, VAEs, tokenizer and scheduler.
Credits
Base checkpoint by FastVideo at Hao AI Lab, UC San Diego, https://github.com/hao-ai-lab/FastVideo, built on MiniMax-H3, https://huggingface.co/MiniMaxAI/MiniMax-H3. Licensed under the MiniMax-H3 community license, see LICENSE.
- Downloads last month
- 68
Model tree for KyleNeverGivesUp/FastH3-text-encoder-nvfp4
Base model
MiniMaxAI/MiniMax-H3