Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Paper • 2609.27891 • Published Aug 21 • 32
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published Sep 2 • 62
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training Paper • 2609.00188 • Published Aug 31 • 54
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models Paper • 2608.29098 • Published Aug 29 • 10
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models Paper • 2608.29098 • Published Aug 29 • 10
Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching Paper • 2608.28695 • Published Aug 27 • 7
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Paper • 2608.27456 • Published Aug 27 • 88
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published Aug 17 • 41
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Paper • 2608.18852 • Published Aug 19 • 9
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Paper • 2608.18933 • Published Aug 19 • 14
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation Paper • 2608.16765 • Published Aug 17 • 14
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Paper • 2608.05102 • Published Aug 5 • 69
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published Jul 13 • 16
ViT-Up: Faithful Feature Upsampling for Vision Transformers Paper • 2606.14024 • Published Jun 12 • 10
Revisiting Articulated Parts Perception in Robot Manipulation Paper • 2606.08103 • Published Jun 6 • 3