Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 11 days ago • 45
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 9 days ago • 108
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 22 days ago • 115
GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills Paper • 2609.21749 • Published 8 days ago • 18
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 12 days ago • 152
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 8 days ago • 131
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 9 days ago • 130
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 9 days ago • 85
HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses Paper • 2609.15938 • Published 12 days ago • 31
Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 10 days ago • 57
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Paper • 2609.11682 • Published 16 days ago • 46
Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models Paper • 2609.05779 • Published 21 days ago • 9
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published 18 days ago • 107
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 17 days ago • 44
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published 19 days ago • 146
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Paper • 2609.05903 • Published 21 days ago • 65
Revisiting Complete Reasoning Traces for Post-Training Paper • 2609.07103 • Published 19 days ago • 23
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 18 days ago • 174
When Models Edit Too Much: On the Fidelity of Minimal Code Edits Paper • 2609.04061 • Published 23 days ago • 9