Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 9 days ago • 57
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 8 days ago • 56
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 10 days ago • 71
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 17 days ago • 81
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 9 days ago • 79
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published 17 days ago • 107
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 11 days ago • 84
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 8 days ago • 85
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 21 days ago • 114
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 22 days ago • 101
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 24 days ago • 120
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 8 days ago • 130
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 11 days ago • 151
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems Paper • 2609.02750 • Published 23 days ago • 144
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 8 days ago • 178