Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published 8 days ago • 151
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Paper • 2607.06553 • Published 21 days ago • 20
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs Paper • 2605.23898 • Published May 22 • 7
Can Vision Language Models Infer Human Gaze Direction? A Controlled Study Paper • 2506.05412 • Published Jun 4, 2025 • 5
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning Paper • 2604.04746 • Published Apr 8 • 73
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model Paper • 2603.22281 • Published Mar 23 • 20
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding Paper • 2603.22285 • Published Mar 23 • 49
view article Article Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty? zhangchenxu • Feb 25 • 14
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 65
Less is Enough: Synthesizing Diverse Data in Feature Space of LLMs Paper • 2602.10388 • Published Feb 11 • 246
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning Paper • 2602.02192 • Published Feb 2 • 13
TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents Paper • 2602.07274 • Published Feb 6 • 212
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards Paper • 2510.08529 • Published Oct 9, 2025 • 19
EgoPrivacy: What Your First-Person Camera Says About You? Paper • 2506.12258 • Published Jun 13, 2025 • 3
Core Knowledge Deficits in Multi-Modal Language Models Paper • 2410.10855 • Published Oct 6, 2024 • 4