Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 8 days ago • 30
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 16 days ago • 138
NeoHorse-1 Collection NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness • 3 items • Updated 24 days ago • 8
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 26 days ago • 327
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published Aug 31 • 63
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities Paper • 2608.28122 • Published Aug 28 • 66
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Paper • 2608.27454 • Published Aug 27 • 36
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published Aug 27 • 36
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published Aug 24 • 65
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 212
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published Aug 19 • 100
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published Aug 15 • 451