Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published Aug 31 • 30
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training Paper • 2609.40111 • Published 6 days ago • 50
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training Paper • 2609.40111 • Published 6 days ago • 50
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published Aug 25 • 152
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published Aug 26 • 69
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 212
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 212
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation Paper • 2606.17628 • Published Jun 16 • 30
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Paper • 2601.18510 • Published Jan 26 • 1
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies Paper • 2604.00830 • Published Apr 2 • 13
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections Paper • 2605.15030 • Published May 14
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video Paper • 2602.03328 • Published Feb 3
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Paper • 2601.18510 • Published Jan 26 • 1
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Paper • 2606.06036 • Published Jun 4 • 78
Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents Paper • 2606.06036 • Published Jun 4 • 78