EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos Paper • 2609.39378 • Published 9 days ago • 70
RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents Paper • 2609.32862 • Published 13 days ago • 27
Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs Paper • 2605.13737 • Published May 13 • 2
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 28 days ago • 266
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published Aug 26 • 193
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • Jul 17 • 197
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 78
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 122
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 75
SenseNova-U1 Collection SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-Unify Architecture • 12 items • Updated Aug 14 • 77
FileGram: Grounding Agent Personalization in File-System Behavioral Traces Paper • 2604.04901 • Published Apr 6 • 36
HippoCamp: Benchmarking Contextual Agents on Personal Computers Paper • 2604.01221 • Published Apr 1 • 33
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification Paper • 2603.15726 • Published Mar 16 • 187