HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 2 days ago • 66
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 2 days ago • 119
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published 4 days ago • 13
view article Article LFM2.5-Encoders for Fast Long-Context Inference on CPU LiquidAI • 2 days ago • 51
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 4 days ago • 87
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 4 days ago • 26
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published 3 days ago • 58
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 4 days ago • 81
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents Paper • 2607.22798 • Published 7 days ago • 57
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 5 days ago • 119
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published 9 days ago • 191
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 8 days ago • 38
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 7 days ago • 45
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 9 days ago • 30