StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published Aug 18 • 138
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 568
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 16 days ago • 274
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 29 days ago • 85