Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Paper • 2608.23383 • Published Aug 24 • 19
SANA-WM Collection SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer • 6 items • Updated Jun 11 • 7
Self Gradient Forcing: Native Long Video Extrapolation Paper • 2607.20368 • Published Jul 22 • 37
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models Paper • 2606.25041 • Published Jun 23 • 125
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 218
Echo-Memory: A Controlled Study of Memory in Action World Models Paper • 2606.09803 • Published Jun 8 • 33
Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation Paper • 2606.04527 • Published Jun 3 • 29
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments Paper • 2605.22189 • Published May 21 • 6
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments Paper • 2605.22189 • Published May 21 • 6
view article Article Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents nvidia • Apr 28 • 66