TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent Paper • 2609.27277 • Published 8 days ago • 29
LastOPD: Taming Collapse in Latent On-Policy Distillation Paper • 2609.28845 • Published 8 days ago • 27
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 4 days ago • 65
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 4 days ago • 65
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 17 days ago • 84
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 114
view article Article FineBooks: are open OCR models good enough to unlock historical knowledge? finebooks • Aug 10 • 26
PUMA Collection The official repository of the paper: Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models • 3 items • Updated May 25 • 2
Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models Paper • 2605.17672 • Published May 17 • 23
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering Paper • 2605.29648 • Published May 28 • 8
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering Paper • 2605.29648 • Published May 28 • 8
PUMA Collection The official repository of the paper: Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models • 3 items • Updated May 25 • 2