Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 6 days ago • 120
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 5 days ago • 496
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 10 days ago • 101
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 11 days ago • 16
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 11 days ago • 130
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 11 days ago • 219
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 28 days ago • 119
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 14 days ago • 138
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 14 days ago • 79
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 18 days ago • 159
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 16 days ago • 133
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 15 days ago • 191
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 25 days ago • 376