Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published 8 days ago • 140
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 8 days ago • 29
SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments Paper • 2607.20207 • Published 8 days ago • 4
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion Paper • 2607.20417 • Published 8 days ago • 9
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 12 days ago • 137
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 9 days ago • 57
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 8 days ago • 303
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Paper • 2607.18171 • Published 10 days ago • 6
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models Paper • 2607.11498 • Published 17 days ago • 7
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 14 days ago • 141
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Paper • 2607.16107 • Published 13 days ago • 11
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published 14 days ago • 70
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 14 days ago • 53
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 16 days ago • 227
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling Paper • 2607.10995 • Published 17 days ago • 9
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published 17 days ago • 44
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 16 days ago • 102