WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents Paper • 2609.27490 • Published 1 day ago • 8
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published 18 days ago • 16
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines Paper • 2604.01029 • Published Apr 1 • 6