--- license: other license_name: minimax-h3-community license_link: LICENSE base_model: - MiniMaxAI/MiniMax-H3 - larryvrh/MiniMax-H3-Turbo-Lora pipeline_tag: text-to-video tags: - ref-to-video - video - audio - text-to-audio-video - sparse-attention - block-sparse - minimax-h3 - veda - fp8 ---

Miowtion

# Veda-MiniMax-H3 R2VA (Preview) [Demo](https://veda.hanxiao.run/) · [Project Page](https://veda-sparse.github.io/) · [Code](https://github.com/veda-sparse/Miowtion/tree/main) · [Paper](https://arxiv.org/abs/2605.30325) · [Deployment Guide](AGENTS.md) ## Introduction **Veda** is a learned sparse-attention method for video diffusion models. This checkpoint brings Veda to [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) reference-to-audio-video generation. It selects visual reference and generated video tiles independently; text, VLM conditioning, and audio attention remain dense. The predictor is not a LoRA and leaves the backbone weights unchanged. This preview was fine-tuned for 600 updates from the T2VA predictor on 5.17-second R2VA clips. It uses the MiniMax-H3 Turbo LoRA for eight-step generation, with up to 32 visual-reference tiles and 32 current-video tiles per query tile and head. The predictor is stored in FP8 and dequantized to BF16 when loaded. ### Highlights - **Learned sparse attention.** Visual reference and generated-video attention are selected independently; text, VLM conditioning, and audio stay dense. - **Fixed visual budgets.** The released preview uses 32 reference tiles and 32 current-video tiles per query tile and head. - **Plug-and-play predictor.** The FP8 predictor is separate from the H3 backbone and Turbo LoRA weights. ## Samples The examples below use the same seed, prompt, geometry, and Turbo LoRA for Dense and Veda. Each generated output is 1344×768, 124 frames, and about 5.17 seconds. Dense (left) and Veda (right) are stitched into one synchronized comparison video, with attention and denoising speedups in its top title bar. Each comparison uses the original Dense audio track.
TaskReferencesDense (left) · Veda (right)
Pink suit and lamb
A man in a bright pink suit holds a black lamb in a sunlit pasture, speaks a short line, and keeps the source soundtrack.
Video 1 · appearance and scene

Audio 1 · source soundtrack

Audio 2 · male voice reference

Open comparison · Open Dense video · Open Veda video
Anime train by the coast
An anime traveler watches a bright coastline pass outside a moving train window.
Picture 1 · character and coastal scene
Anime traveler beside a coastal train window
open reference

Open comparison · Open Dense video · Open Veda video
Anime twilight carnival
An anime woman walks through a fairground as the Ferris wheel and lights turn on at blue hour.
Picture 1 · character
Anime character reference
open character reference
Picture 2 · fairground
Anime fairground reference
open fairground reference

Open comparison · Open Dense video · Open Veda video
Aurora over the lake
A green-violet aurora slowly grows above a mountain lake and appears in the calm reflection.
Picture 1 · lake and sky
Aurora over a mountain lake
open lake reference
Video 1 · aurora motion

Open comparison · Open Dense video · Open Veda video
Red fox in autumn woods
A red fox crosses a warm autumn woodland, slows, and looks toward the camera.
Picture 1 · fox appearance
Red fox appearance reference
open fox reference
Video 1 · fox motion

Open comparison · Open Dense video · Open Veda video
Perfume by the pool
A glass perfume bottle catches warm light beside rippling water in a stone courtyard.
Picture 1 · product and pool
Perfume bottle beside a pool
open product reference

Open comparison · Open Dense video · Open Veda video
## Performance Measured with eight-step Turbo LoRA inference on one RTX PRO 6000 Blackwell GPU. Veda uses fixed budgets of 32 reference and 32 current-video tiles. | GPU | Clip | Attention speedup | End-to-end speedup | |---|---|---:|---:| | RTX PRO 6000 Blackwell | 16:9 · 5.17 s | 2.42× | 1.69× | | RTX PRO 6000 Blackwell | 9:16 · 5.17 s | 2.65× | 1.67× | | RTX PRO 6000 Blackwell | 1:1 · 5.17 s | 2.06× | 1.37× | | RTX PRO 6000 Blackwell | 4:3 · 5.17 s | 2.19× | 1.54× | | RTX PRO 6000 Blackwell | 16:9 · 10.1 s | 2.84× | 2.05× | | RTX PRO 6000 Blackwell | 16:9 · 14.4 s | 2.90× | 2.20× | ## Usage See [AGENTS.md](AGENTS.md) for installation and GPU-specific deployment notes. The commands below use the R2VA code on the [`main` branch](https://github.com/veda-sparse/Miowtion/tree/main). ```bash git clone --recurse-submodules --branch main https://github.com/veda-sparse/Miowtion.git cd Miowtion pip install -e '.[gpu,encode]' hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "Ref2VA/*" \ --local-dir weights/MiniMax-H3 hf download larryvrh/MiniMax-H3-Turbo-Lora \ minimax_h3_turbo_v4_step600_ema.safetensors --local-dir weights/turbo_lora hf download Veda-Sparse/Minimax-H3-R2VA-Veda-Preview \ --local-dir weights/veda/h3-r2va-preview ``` Prepare the reference files, then encode and generate: ```bash python scripts/prepare_official_r2va_demo.py \ --out artifacts/examples/minimax_h3_ref2va python scripts/encode_samples.py --root weights/MiniMax-H3 \ --manifest artifacts/examples/minimax_h3_ref2va/manifests/prompts.jsonl \ --out artifacts/samples/demo python scripts/generate.py --config configs/infer_r2va_preview.yaml \ --predictor weights/veda/h3-r2va-preview/minimax_h3_r2va_veda_preview_fp8.safetensors \ --attention dense veda ``` ## License This model inherits the [MiniMax H3 Community License](LICENSE) from its base model. ## Citation ```bibtex @inproceedings{han2026veda, title={Veda: Scalable Video Diffusion via Distilled Sparse Attention}, author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng and Jiang, Yi and Qi, Xiaojuan}, booktitle={International Conference on Machine Learning (ICML)}, year={2026} } ```