FaceLift
Single Image 3D Face Reconstruction with Gaussian Splatting
Run complete 3.96M and 9.36M text-to-waveform models live.
Multilingual CPU-only ASR with a 1.58-bit BitNet decoder
Live interactive world rollout from an image
Codec-native video & image understanding with Mage-VL 4B
Verbatim + intended transcripts with word-level timing
Zero-shot TTS with explicit word-level prosody control
Low-resource speech synthesis with emotional control!
Efficient native-resolution image generation and editing