Vidu S1: A Real-Time Interactive Video Generation Model Paper • 2607.03118 • Published 27 days ago • 144
ProductWebGen: Benchmarking Multimodal Product Webpage Generation Paper • 2606.01022 • Published May 31 • 5
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model Paper • 2604.02097 • Published Apr 2 • 32
SIFT: Grounding LLM Reasoning in Contexts via Stickers Paper • 2502.14922 • Published Feb 19, 2025 • 32
Show-o Turbo: Towards Accelerated Unified Multimodal Understanding and Generation Paper • 2502.05415 • Published Feb 8, 2025 • 20