HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 567
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 158
Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue Paper • 2609.31948 • Published 13 days ago • 89
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 16 days ago • 164
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation Paper • 2609.06931 • Published Sep 7 • 27