Speech-to-video avatar model, 2B, real-time variant (anonymous ICLR 2027 submission)

Exponential moving average weights of the distilled stage-2 generator described in the paper.

Contents

  • model.safetensors: DiT weights, fp32, 2.27B parameters, 8.44 GiB, 1015 tensors. Key names match the state_dict of the DiT class in the released inference code.
  • vae_decoder_x15.safetensors: Turbo-VAED decoder fine-tuned to upscale by 12 instead of 8, fp32, 40.7M parameters, 158 tensors. It decodes the same latents at 1.5x the output resolution and is loaded with --vae-x15.

Usage

from safetensors.torch import load_file
state_dict = load_file("model.safetensors")
dit.load_state_dict(state_dict, strict=True)

See the code repository linked in the paper for the model definition, the conditioning encoders and the sampling loop.

License

CC BY-NC 4.0, research use. Animating a real person without their consent is not permitted.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support