Papers
arxiv:2609.40097

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

Published on Sep 30
· Submitted by
Chenghao Xiao
on Oct 2
Authors:
,
,
,
,
,
,
,
,
,
,
,

Abstract

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.

Community

Paper submitter

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities.

We introduce AutoDataBench, a controlled testbed for evaluating Data Intelligence in LLM research agents.

  • Data Intelligence: We study three complementary capabilities—data diagnosis & repair, data organization, and data construction—across tool use, retrieval, and knowledge injection.
  • Understanding data interventions: Beyond whether an intervention works, we test whether agents can predict its effect before training and refine their understanding through experimental feedback.
  • Benchmark as Data Engine: We reuse AutoDataBench research trajectories as mid-training data and find that they improve downstream coding performance, showing that benchmarks can also serve as sources of useful training data.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.40097
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 3

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.40097 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.40097 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.