YICHEN013's picture
Add the InferenceNet project portal with Home, Data and Agent pages
03c1286 verified
|
Raw History Blame Contribute Delete
7.01 kB
metadata
title: InferenceNet Challenge
emoji: 📊
colorFrom: green
colorTo: blue
sdk: static
pinned: false
license: apache-2.0
short_description: Explore the benchmark data, agent harnesses and leaderboard

InferenceNet Challenge

The project portal brings the benchmark, data documentation, execution walkthrough and published results into one Hugging Face Space.

The original Space URL remains valid. Data files continue to live in the dataset repository; the portal includes only a small, reproducible distribution summary. Home and Leaderboard both read leaderboard-data.json. dataset-summary.json is derived from the same pinned task-list revision.

Agent & Harness Leaderboard

The September 2026 research leaderboard displays six requested model aliases, thirteen groups, and 13,000 archived task records. Switch between the original baseline / DeepAgents paired comparison, a single-pass Agent ranking, and a harness ranking that also includes one independent DSH run. Every group uses the same 1,000 Selected_1000 task IDs. An API alias does not establish the identity of the underlying model weights.

  • Single-pass Agent: one model call to generate code, followed by execution.
  • DeepAgents: up to six model calls and four trial tool calls, followed by final execution.
  • DeepSeek Harness (DSH): official SDK 0.1.5rc1, requested deepseek-v4-pro API alias; displayed separately as DeepSeek API [v4-pro alias] + DSH. Its underlying model weights are unverified, and it is never used in the baseline / DeepAgents paired lift.

Models: GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro Preview, Qwen3.7-Max, and the historical DeepSeek V4 Pro entries. GPT-5.5 is outside this edition.

The DSH run protocol documents its fixed selection of 1,000 results: 913 original attempts and 87 technical replacement attempts (60 infrastructure failures and 27 input-schema failures). The 27 schema recoveries changed the input-schema prompt. All 1,087 attempts remain in the cost accounting; selection is not best-of. The final selection has 699 scored, 271 failed and 30 uncertain task outcomes. Matching per-attempt numeric caps do not establish equal total run budgets: the DSH result includes 87 additional attempts. Full execution-environment and reasoning-setting parity with the prior DeepAgents run has not been verified.

Reading the results

The default metric is full replication, from local-paper-v1: coefficient and standard-error relative errors ≤ 1%, and p-value absolute error ≤ 0.01. Four separately labeled hf-leaderboard-v1 metrics are also available: execution success, coefficient-only partial replication (≤ 5% relative error), coefficient direction, and significance category.

These are locally scored research results. Official scorer parity has not been verified. All rates use 1,000 as the denominator; metric unknowns and invalid evidence are retained and earn no successes. Task-status unknown counts are reported separately from metric unknown counts.

The comparisons preserve historical recovery amendments. Call budgets, time limits, reasoning settings and source protocols differ. Differences are descriptive observations, not equal-cost causal estimates of a harness effect. See each run's configuration before making cross-model claims.

Kimi retains 5 baseline and 4 harness task-status unknowns. Qwen3.7-Max harness retains 6 task-status unknowns and 1 invalid-evidence record (task 453), which is not sealed. DSH retains 30 task-status unknowns, with no invalid-evidence records in the fixed selection. Historical Sol/Opus interruption counts are shown as originally exported. An archive of 1,000 task IDs does not imply 1,000 definitive outcomes.

Data and reproducibility

  • leaderboard-data.json: displayed data, metric profiles, counts, evidence status and source-file SHA-256 hashes.
  • leaderboard.csv: 65 displayed metric rows; thirteen groups × five metrics. One metric profile is explicit on every row.
  • results/: frozen per-task archives, configurations and summaries.
  • legacy.html: original leaderboard with a historical banner. Its original results.csv is unchanged and is not mixed into the new comparison.

The exact source archive commit is recorded as source_revision in leaderboard-data.json; every displayed evidence link is pinned to that commit. Dataset: CamoAiLab/InferenceNet, revision 59f9512a38e594528807744214a60ee00367434e, task list Selected_1000/1000_new.csv.

The per-run archive files keep their original publication metadata. Their historical included_in_displayed_leaderboard: false fields describe the archive upload, not this later display release. The current display manifest is leaderboard-data.json; no experiment evidence is edited to publish this page.

To rebuild the display exports from those archives:

python build_leaderboard.py --source . --output . \
  --source-revision ARCHIVE_COMMIT_SHA --published-date 2026-09-20

Use the 40-character archive commit SHA from leaderboard-data.json above. The builder verifies all task populations and 61 metric counts against archived flags. The four Sol/Opus local full-replication summaries have no per-task local flags in the archive and are preserved from their accepted summaries. This is a display refresh, not a fresh inference run, native artifact audit, or rescore.

Project

Project website · Project code · Evaluation framework

Editing the portal

index.html, data.html, leaderboard.html and agent.html are the four static pages. site-shell.js shares language and theme preferences; portal.css extends the existing leaderboard styles. portal.js renders the publication counts and data distributions. Published metric values still come only from leaderboard-data.json.

To regenerate the data distribution summary from the exact pinned CSV:

python build_dataset_summary.py /path/to/1000_new.csv

The script verifies the CSV SHA-256 before generating dataset-summary.json. Serve the repository with a local HTTP server to preview module scripts and JSON loading; no frontend build step is required.