ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 12 days ago • 214
Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 26 days ago • 174
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 25 days ago • 220
Aspire: Can Models Self-Evolve from Vague Goals? Paper • 2608.31111 • Published 26 days ago • 169
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 26 days ago • 42
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 25 days ago • 220
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 26 days ago • 42
Aspire: Can Models Self-Evolve from Vague Goals? Paper • 2608.31111 • Published 26 days ago • 169
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published about 1 month ago • 174
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published about 1 month ago • 174
Works Mentioning or Using ClawBench Collection Public papers and reports that cite, discuss, or build on ClawBench and related real-world agent evaluation work. • 15 items • Updated Jul 29