HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents Paper • 2610.03574 • Published 5 days ago • 53
GTR: Gated Token Recurrence for Efficient Dense Prediction Paper • 2609.26590 • Published 14 days ago • 12
Language Models that Play Chess and Explain Their Moves Paper • 2610.03695 • Published 5 days ago • 31
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 268
Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems Paper • 2610.01257 • Published 6 days ago • 40
APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants Paper • 2609.37559 • Published 8 days ago • 46
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 11 days ago • 324
Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents Paper • 2609.37236 • Published 8 days ago • 40
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation Paper • 2609.36601 • Published 8 days ago • 94