Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Paper • 2607.24904 • Published • 24
None defined yet.
Utonia: Toward One Encoder for All Point Clouds
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations