TAEH3 (mirror)
An unmodified mirror of taeh3.safetensors from madebyollin/taehv,
the tiny autoencoder for the MiniMax-H3 video latent space. It decodes a whole clip of latents to video far
more cheaply than the full H3 VAE, which is what makes per-step previews affordable.
This mirror exists because OzzyGT/minimax_h3_preview_blocks
fetches it at runtime to preview MiniMax-H3's prediction while it denoises.
Use
Needs taehv.py from the upstream repo:
from huggingface_hub import hf_hub_download
from taehv import TAEHV
tae = TAEHV(hf_hub_download("OzzyGT/taeh3", "taeh3.safetensors"), arch_name="taeh3").cuda().half().eval()
frames = tae.decode_video(latents, parallel=False) # (N, T, 24, h, w) normalized latents -> (N, F, 3, h*16, w*16) in [0, 1]
The decoder follows H3's own chunking: 5 * n + 2 latent frames decode to 17 * n + 5 video frames, the same
count the full VAE gives.
Pass arch_name explicitly: upstream guesses the architecture from the checkpoint's filename, which is a
cache path here. parallel=False decodes frame by frame; parallel=True holds every frame's activations at
once, about 25 GB for 209 frames at 768x1344, and is no faster.
Provenance
- Source:
safetensors/taeh3.safetensorsat commit62f7591 sha256 4fd022bfcab08772fe0536b17ea1a3bbb5625be11e397868d1c5d891863d4c13- 22,709,752 bytes, 128 tensors, fp16, 11.3M parameters
- Byte-for-byte the upstream file. Nothing was converted, quantized or retrained.
License
MIT, (c) 2025 Ollin Boer Bohan. The upstream LICENSE is included in this repo unchanged. All credit for
the model goes to the author; this repo only hosts a copy on the Hub, since upstream distributes these
weights through GitHub.