INFERENCENET CHALLENGE

Can AI reproduce an empirical finding?

Real research data. Executable analysis. Results you can inspect.

InferenceNet evaluates how models and agent harnesses turn an econometric specification into a reproducible result.

1,000 fixed evaluation tasks
15 journals in this subset
286 distinct article titles
5 estimation classes

Selected_1000 · Counts from the pinned task list, revision 59f9512.

01 / THE CHALLENGE

From a research question to a reproducible result.

  1. 01

    Read the task

    A research specification and its source data.

  2. 02

    Write & run code

    Fit the requested econometric model.

  3. 03

    Inspect & revise

    A harness can use execution feedback.

  4. 04

    Recover the finding

    Return the coefficient, standard error and p-value.

02 / EXPLORE

One project. Three ways in.

01

Explore the data

Where the tasks come from, what they ask, and how to access them.

02

Compare the results

Single-pass models, DeepAgents and an independent DSH run.

03

Inside the agent loop

Follow the code, tools and final artifact of a real replication task.

03 / RESULTS

A shared view of the evidence.

PUBLISHED RESEARCH SNAPSHOT

— experiment groups / — model aliases

The homepage and leaderboard read the same published result file.

View results →

Research results use local scoring. Read the run protocols, scoring definitions and evidence status with every comparison.