HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published 3 days ago • 141
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 9 days ago • 304
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published 11 days ago • 197
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models Paper • 2607.07673 • Published 23 days ago • 14
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 15 days ago • 203
KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Paper • 2607.12625 • Published 16 days ago • 80
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras Paper • 2607.12993 • Published 17 days ago • 131
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published 21 days ago • 9
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published 24 days ago • 95
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams Paper • 2606.21337 • Published Jun 19 • 75
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources Paper • 2605.29250 • Published May 28 • 81
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents Paper • 2605.30723 • Published May 29 • 17
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 252
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models Paper • 2605.31603 • Published May 29 • 8
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Paper • 2605.18018 • Published May 18 • 33
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers Paper • 2603.24414 • Published Mar 25 • 183
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling Paper • 2603.25746 • Published Mar 26 • 155
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 248