UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation Paper • 2610.09823 • Published 2 days ago • 92
EmbodiedSmith: Scaling Embodied Data through Recursive Self-Improvement Flywheel in Simulation Paper • 2610.07969 • Published 3 days ago • 13
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 22 days ago • 139
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published Aug 5 • 61
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published Jul 16 • 72
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published Jul 13 • 43
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Paper • 2606.06155 • Published Jun 4 • 10
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models Paper • 2605.10903 • Published May 11
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark Paper • 2605.10921 • Published May 11 • 3
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation Paper • 2603.16669 • Published Mar 17 • 70
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning Paper • 2601.06943 • Published Jan 11 • 198