Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 14 days ago • 106
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published Sep 8 • 321