OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction Paper • 2610.01762 • Published 6 days ago • 230
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender Paper • 2609.15478 • Published 23 days ago • 34
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published Aug 5 • 28
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 65
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published Jul 31 • 42
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 34
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published Jun 26 • 52
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling Paper • 2605.13301 • Published May 13 • 166
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents Paper • 2604.26752 • Published Apr 29 • 115
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision Paper • 2604.12002 • Published Apr 13 • 12
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification Paper • 2603.26648 • Published Mar 27 • 45
view article Article Welcome Gemma 4: Frontier multimodal intelligence on device +5 merve, pcuenq, sergiopaniego, burtenshaw, Steveeeeeeen, alvarobartt, SaylorTwift • Apr 2 • 930