Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated 7 days ago • 23
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 6 days ago • 28
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 87
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Paper • 2506.05414 • Published Jun 4, 2025 • 4