The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean Paper • 2609.09218 • Published 26 days ago • 2
The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean Paper • 2609.09218 • Published 26 days ago • 2
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations Paper • 2609.30867 • Published 7 days ago • 5
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations Paper • 2609.30867 • Published 7 days ago • 5
Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations Paper • 2609.30867 • Published 7 days ago • 5
Sleeping RL ComtradeBench: An OpenEnv Benchmark for Reliable LLM Tool-Use Under Adversarial API Conditions 📊 Benchmark LLM agents on robust data‑fetching tool use
Sleeping RL ComtradeBench: An OpenEnv Benchmark for Reliable LLM Tool-Use Under Adversarial API Conditions 📊 Benchmark LLM agents on robust data‑fetching tool use