Download README.md from CamoAiLab/InferenceNet-Leaderboard: direct link, hf CLI and curl.
- Browser
- Download file 7.01 kB
-
https://huggingface.co/spaces/CamoAiLab/InferenceNet-Leaderboard/resolve/main/README.md
- Command line
-
hf download hf://spaces/CamoAiLab/InferenceNet-Leaderboard/README.md
-
curl -L -o README.md https://huggingface.co/spaces/CamoAiLab/InferenceNet-Leaderboard/resolve/main/README.md
title: InferenceNet Challenge
emoji: 📊
colorFrom: green
colorTo: blue
sdk: static
pinned: false
license: apache-2.0
short_description: Explore the benchmark data, agent harnesses and leaderboard
InferenceNet Challenge
The project portal brings the benchmark, data documentation, execution walkthrough and published results into one Hugging Face Space.
The original Space URL remains valid. Data files continue to live in the dataset repository; the portal includes only a small, reproducible distribution summary. Home and Leaderboard both read leaderboard-data.json. dataset-summary.json is derived from the same pinned task-list revision.
Agent & Harness Leaderboard
The September 2026 research leaderboard displays six requested model aliases, thirteen groups, and 13,000 archived task records. Switch between the original baseline / DeepAgents paired comparison, a single-pass Agent ranking, and a harness ranking that also includes one independent DSH run. Every group uses the same 1,000 Selected_1000 task IDs. An API alias does not establish the identity of the underlying model weights.
- Single-pass Agent: one model call to generate code, followed by execution.
- DeepAgents: up to six model calls and four trial tool calls, followed by final execution.
- DeepSeek Harness (DSH): official SDK 0.1.5rc1, requested
deepseek-v4-proAPI alias; displayed separately as DeepSeek API [v4-pro alias] + DSH. Its underlying model weights are unverified, and it is never used in the baseline / DeepAgents paired lift.
Models: GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro Preview, Qwen3.7-Max, and the historical DeepSeek V4 Pro entries. GPT-5.5 is outside this edition.
The DSH run protocol documents its fixed selection of 1,000 results: 913 original attempts and 87 technical replacement attempts (60 infrastructure failures and 27 input-schema failures). The 27 schema recoveries changed the input-schema prompt. All 1,087 attempts remain in the cost accounting; selection is not best-of. The final selection has 699 scored, 271 failed and 30 uncertain task outcomes. Matching per-attempt numeric caps do not establish equal total run budgets: the DSH result includes 87 additional attempts. Full execution-environment and reasoning-setting parity with the prior DeepAgents run has not been verified.
Reading the results
The default metric is full replication, from local-paper-v1: coefficient
and standard-error relative errors ≤ 1%, and p-value absolute error ≤ 0.01.
Four separately labeled hf-leaderboard-v1 metrics are also available:
execution success, coefficient-only partial replication (≤ 5% relative error),
coefficient direction, and significance category.
These are locally scored research results. Official scorer parity has not been verified. All rates use 1,000 as the denominator; metric unknowns and invalid evidence are retained and earn no successes. Task-status unknown counts are reported separately from metric unknown counts.
The comparisons preserve historical recovery amendments. Call budgets, time limits, reasoning settings and source protocols differ. Differences are descriptive observations, not equal-cost causal estimates of a harness effect. See each run's configuration before making cross-model claims.
Kimi retains 5 baseline and 4 harness task-status unknowns. Qwen3.7-Max harness retains 6 task-status unknowns and 1 invalid-evidence record (task 453), which is not sealed. DSH retains 30 task-status unknowns, with no invalid-evidence records in the fixed selection. Historical Sol/Opus interruption counts are shown as originally exported. An archive of 1,000 task IDs does not imply 1,000 definitive outcomes.
Data and reproducibility
leaderboard-data.json: displayed data, metric profiles, counts, evidence status and source-file SHA-256 hashes.leaderboard.csv: 65 displayed metric rows; thirteen groups × five metrics. One metric profile is explicit on every row.results/: frozen per-task archives, configurations and summaries.legacy.html: original leaderboard with a historical banner. Its originalresults.csvis unchanged and is not mixed into the new comparison.
The exact source archive commit is recorded as source_revision in
leaderboard-data.json; every displayed evidence
link is pinned to that commit.
Dataset: CamoAiLab/InferenceNet,
revision 59f9512a38e594528807744214a60ee00367434e, task list
Selected_1000/1000_new.csv.
The per-run archive files keep their original publication metadata. Their
historical included_in_displayed_leaderboard: false fields describe the archive
upload, not this later display release. The current display manifest is
leaderboard-data.json; no experiment evidence is edited to publish this page.
To rebuild the display exports from those archives:
python build_leaderboard.py --source . --output . \
--source-revision ARCHIVE_COMMIT_SHA --published-date 2026-09-20
Use the 40-character archive commit SHA from leaderboard-data.json above.
The builder verifies all task populations and 61 metric counts against archived
flags. The four Sol/Opus local full-replication summaries have no per-task local
flags in the archive and are preserved from their accepted summaries. This is a
display refresh, not a fresh inference run, native artifact audit, or rescore.
Project
Project website · Project code · Evaluation framework
Editing the portal
index.html, data.html, leaderboard.html and agent.html are the four static pages.
site-shell.js shares language and theme preferences; portal.css extends the
existing leaderboard styles. portal.js renders the publication counts and data
distributions. Published metric values still come only from leaderboard-data.json.
To regenerate the data distribution summary from the exact pinned CSV:
python build_dataset_summary.py /path/to/1000_new.csv
The script verifies the CSV SHA-256 before generating dataset-summary.json.
Serve the repository with a local HTTP server to preview module scripts and JSON
loading; no frontend build step is required.