Speech-to-video avatar model, 2B, real-time variant (anonymous ICLR 2027 submission)
Exponential moving average weights of the distilled stage-2 generator described in the paper.
Contents
model.safetensors: DiT weights, fp32, 2.27B parameters, 8.44 GiB, 1015 tensors. Key names match thestate_dictof the DiT class in the released inference code.vae_decoder_x15.safetensors: Turbo-VAED decoder fine-tuned to upscale by 12 instead of 8, fp32, 40.7M parameters, 158 tensors. It decodes the same latents at 1.5x the output resolution and is loaded with--vae-x15.
Usage
from safetensors.torch import load_file
state_dict = load_file("model.safetensors")
dit.load_state_dict(state_dict, strict=True)
See the code repository linked in the paper for the model definition, the conditioning encoders and the sampling loop.
License
CC BY-NC 4.0, research use. Animating a real person without their consent is not permitted.