ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 8 days ago • 302
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders Paper • 2607.14088 • Published 14 days ago • 14
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published 9 days ago • 197
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 10 days ago • 165
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Paper • 2607.08317 • Published 20 days ago • 35
KronQ: LLM Quantization via Kronecker-Factored Hessian Paper • 2607.07964 • Published 21 days ago • 32
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 13 days ago • 170
AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models Paper • 2607.02269 • Published 27 days ago • 9
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 177
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Paper • 2606.24428 • Published Jun 23 • 52
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 141
DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects Paper • 2606.15133 • Published Jun 13 • 74