RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published 12 days ago • 198
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 13 days ago • 166
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling Paper • 2607.10995 • Published 19 days ago • 9
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published 22 days ago • 9
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors Paper • 2606.32029 • Published Jun 30 • 14
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 177
RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling Paper • 2606.06309 • Published Jun 4 • 11
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Paper • 2605.25659 • Published May 25 • 17
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models Paper • 2606.02277 • Published Jun 1 • 7