Spark-H3

Spark-H3

Adaptive block-sparse attention for MiniMax-H3 video generation

Reblock similar tokens · Reweight tail tokens

Technical Blog · GitHub · ComfyUI · Documentation · Base Model

This is a code-only release of the Spark-H3 attention backend. It does not redistribute MiniMax-H3 model weights. Download the base model from MiniMaxAI/MiniMax-H3.

Overview

Spark-H3 is the MiniMax-H3 implementation of Spark-Attn, a block-sparse attention method for accelerating video generation while preserving output fidelity. It improves both branches of block-sparse attention:

  • Spark-Reblock groups tokens with similar attention preferences so the exact branch spends its budget on more relevant interactions.
  • Spark-Reweight builds weighted key/value summaries and corrects their attention mass, reducing the bias introduced by mean-pooled tail blocks.

Spark-H3 is an inference backend rather than a LoRA or a new checkpoint. It does not modify the MiniMax-H3 backbone weights and can be combined with few-step LoRAs and other compatible MiniMax-H3 variants.

Highlights

  • Up to 2.33× attention speedup and 1.73× DiT speedup in a 10-second, 768p benchmark at 10% attention density. All sparse methods share a 20% dense warmup, with the first Transformer layer kept dense throughout.
  • 23.30 dB PSNR at 10% attention density, compared with 20.36 dB for Sol-H3 under the same evaluation protocol.
  • Supports Diffusers, with a ComfyUI preview integration now available.
  • Production fused kernels are available for SM120 GPUs, including the NVIDIA GeForce RTX 50 series and RTX PRO 5000/6000 Blackwell GPUs.

Unless otherwise noted, all latency measurements on this page were collected on a single NVIDIA RTX PRO 6000 Blackwell GPU.

Samples

VDN + LightX2V

The three samples below use VDN prompts 01, 06, and 10 from VDN blog with the LightX2V MiniMax-H3 Turbo LoRA and Spark-H3 at 10% attention density, with a 20% dense warmup and the first Transformer layer kept dense throughout.

VDN Prompt 01 — Japanese city montage

VDN Prompt 06 — Rainy-night camcorder documentary

VDN Prompt 10 — 2D anime fashion sequence

LBH two-stage sampling

The following samples combine Spark-H3 with SelfLift two-stage sampling based on the LBH learned latent upsampler.

LBH Case 01 — Black cat by the pool

Dense Spark-H3, 10%
DiT latency: 108.5 s · 1.00× DiT latency: 59.8 s · 1.81×

LBH Case 02 — Stormtrooper on the beach

Dense Spark-H3, 10%
DiT latency: 108.9 s · 1.00× DiT latency: 59.5 s · 1.83×

Ref2VA

Spark-H3 is also compatible with MiniMax-H3's Ref2VA pipeline, as shown on the official Ref2VA case.

Official Ref2VA case

Dense Spark-H3, 10% Spark-H3, 10% + Compact Cond
DiT latency: 708.5 s · 1.00× DiT latency: 455.0 s · 1.56× DiT latency: 202.6 s · 3.50×

See the technical blog for more results.

Benchmark Results

Evaluation setup. Dense attention, Sol-H3, Spark-H3, and Spark-H3-Lite are evaluated on VBench prompts with 19 denoising steps, 1344×768 resolution, and 240 frames at 24 fps (10 seconds). All sparse methods—Sol-H3, Spark-H3, and Spark-H3-Lite—use a 20% dense warmup (the first 4 of 19 Transformer evaluations) and keep the first Transformer layer dense throughout. Apart from these shared dense settings, Sol-H3 follows its official configuration. Quality metrics compare each method against dense outputs generated with the same prompts and seeds. Higher PSNR and SSIM are better; lower LPIPS is better.

Method Density ↓ DiT latency (s) ↓ DiT speedup ↑ PSNR (dB) ↑ SSIM ↑ LPIPS ↓ ATTN speedup ↑
Dense 100% 583.3 1.00× ∞ 1.00 0.00 1.00×
Sol-H3 — 355.4 1.64× 20.36 0.71 0.20 2.19×
Spark-H3, 10% 10% 337.4 1.73× 23.30 0.80 0.14 2.33×
Spark-H3, 15% 15% 355.4 1.64× 24.45 0.83 0.11 —
Spark-H3, 20% 20% 371.1 1.57× 25.38 0.85 0.09 1.96×
Spark-H3, 30% 30% 404.1 1.44× 27.07 0.88 0.07 1.76×
Spark-H3-Lite, 10% 10% 330.2 1.77× 23.34 0.80 0.14 —
Spark-H3-Lite, 15% 15% 346.1 1.69× 24.47 0.83 0.11 —
Spark-H3-Lite, 20% 20% 362.6 1.61× 25.12 0.85 0.10 —
Spark-H3-Lite, 30% 30% 395.4 1.48× 26.26 0.87 0.08 —

At the 10% density operating point, Spark-H3 improves PSNR by 2.94 dB over Sol-H3 while also reducing DiT latency from 355.4 s to 337.4 s. Increasing the attention density provides a monotonic fidelity trade-off for standard Spark-H3, reaching 27.07 dB PSNR at 30% density.

Quick Start

Set up the upstream MiniMax-H3 pipeline first, then install Spark-H3 in the same environment with a compatible PyTorch/CUDA stack:

git clone https://github.com/zechengtang/Spark-H3.git
cd Spark-H3
python -m pip install -e '.[cuda]'

Wrap an already loaded MiniMax-H3 pipeline with the Spark installer:

from h3_sparse_attention import install_h3_spark_attn

num_denoise_steps = 19
num_inference_steps = num_denoise_steps + 1

with install_h3_spark_attn(
    pipe.transformer,
    num_denoise_steps=num_denoise_steps,
    warmup_mode="warmup_steps",
    warmup_steps=4,
    dense_layers=1,
    min_tokens=8192,
):
    result = pipe(**inputs, num_inference_steps=num_inference_steps)

MiniMax-H3's scheduler includes the terminal sigma value in num_inference_steps, but that point does not execute the Transformer. Pass the actual Transformer evaluation count to Spark as num_denoise_steps.

For configuration options, supported GPU architectures, and implementation details, see the Spark-H3 documentation.

ComfyUI

Spark-H3 includes a ComfyUI preview integration and example workflows. Follow the ComfyUI installation guide, then start with one of the workflows under workflows/.

License

Spark-H3's original code and documentation are released under the Apache License 2.0. Included third-party components retain their respective licenses and attribution notices. MiniMax-H3 model weights are not included and remain subject to the upstream MiniMax-H3 license.

Acknowledgments

Spark-H3 builds on MiniMax-H3 and Sol-Attn / Sol-Engine. The VDN prompt cases are sourced from the Video DeltaNet project.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support