|
Download README.md from CamoAiLab/InferenceNet-Leaderboard: direct link, hf CLI and curl.
- Browser
- Download file 7.01 kB
-
https://huggingface.co/spaces/CamoAiLab/InferenceNet-Leaderboard/resolve/main/README.md
- Command line
-
hf download hf://spaces/CamoAiLab/InferenceNet-Leaderboard/README.md
-
curl -L -o README.md https://huggingface.co/spaces/CamoAiLab/InferenceNet-Leaderboard/resolve/main/README.md
7.01 kB
| title: InferenceNet Challenge | |
| emoji: 📊 | |
| colorFrom: green | |
| colorTo: blue | |
| sdk: static | |
| pinned: false | |
| license: apache-2.0 | |
| short_description: Explore the benchmark data, agent harnesses and leaderboard | |
| # InferenceNet Challenge | |
| The project portal brings the benchmark, data documentation, execution walkthrough and published results into one Hugging Face Space. | |
| - [Home](https://camoailab-inferencenet-leaderboard.static.hf.space/index.html) | |
| - [Data](https://camoailab-inferencenet-leaderboard.static.hf.space/data.html) | |
| - [Leaderboard](https://camoailab-inferencenet-leaderboard.static.hf.space/leaderboard.html) | |
| - [Agent / Harness](https://camoailab-inferencenet-leaderboard.static.hf.space/agent.html) | |
| - [Dataset repository](https://huggingface.co/datasets/CamoAiLab/InferenceNet) | |
| The original Space URL remains valid. Data files continue to live in the dataset repository; the portal includes only a small, reproducible distribution summary. Home and Leaderboard both read `leaderboard-data.json`. `dataset-summary.json` is derived from the same pinned task-list revision. | |
| ## Agent & Harness Leaderboard | |
| The September 2026 research leaderboard displays **six requested model aliases, | |
| thirteen groups, and 13,000 archived task records**. Switch between the original | |
| baseline / DeepAgents paired comparison, a single-pass Agent ranking, and a | |
| harness ranking that also includes one independent DSH run. Every group uses | |
| the same 1,000 Selected_1000 task IDs. An API alias does not establish the | |
| identity of the underlying model weights. | |
| - **Single-pass Agent:** one model call to generate code, followed by execution. | |
| - **DeepAgents:** up to six model calls and four trial tool | |
| calls, followed by final execution. | |
| - **DeepSeek Harness (DSH):** official SDK 0.1.5rc1, requested | |
| `deepseek-v4-pro` API alias; displayed separately as | |
| **DeepSeek API [v4-pro alias] + DSH**. Its underlying model weights are | |
| **unverified**, and it is never used in the baseline / DeepAgents paired lift. | |
| Models: GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro Preview, | |
| Qwen3.7-Max, and the historical DeepSeek V4 Pro entries. GPT-5.5 is outside this edition. | |
| The DSH [run protocol](./results/deepseek-v4-pro/dsh/README.md) documents its | |
| fixed selection of 1,000 results: 913 original attempts and 87 technical | |
| replacement attempts (60 infrastructure failures and 27 input-schema failures). | |
| The 27 schema recoveries changed the input-schema prompt. All 1,087 attempts | |
| remain in the cost accounting; selection is not best-of. The final selection | |
| has 699 scored, 271 failed and 30 uncertain task outcomes. Matching per-attempt | |
| numeric caps do not establish equal total run budgets: the DSH result includes | |
| 87 additional attempts. Full execution-environment and reasoning-setting | |
| parity with the prior DeepAgents run has not been verified. | |
| ## Reading the results | |
| The default metric is **full replication**, from `local-paper-v1`: coefficient | |
| and standard-error relative errors ≤ 1%, and p-value absolute error ≤ 0.01. | |
| Four separately labeled `hf-leaderboard-v1` metrics are also available: | |
| execution success, coefficient-only partial replication (≤ 5% relative error), | |
| coefficient direction, and significance category. | |
| These are **locally scored research results**. Official scorer parity has not | |
| been verified. All rates use 1,000 as the denominator; metric unknowns and invalid | |
| evidence are retained and earn no successes. Task-status unknown counts are | |
| reported separately from metric unknown counts. | |
| The comparisons preserve historical recovery amendments. Call budgets, time | |
| limits, reasoning settings and source protocols differ. Differences are | |
| descriptive observations, not equal-cost causal estimates of a harness effect. | |
| See each run's configuration before making cross-model claims. | |
| Kimi retains 5 baseline and 4 harness task-status unknowns. Qwen3.7-Max harness | |
| retains 6 task-status unknowns and 1 invalid-evidence record (task 453), which is | |
| not sealed. DSH retains 30 task-status unknowns, with no invalid-evidence | |
| records in the fixed selection. Historical Sol/Opus interruption counts are shown as originally | |
| exported. An archive of 1,000 task IDs does not imply 1,000 definitive outcomes. | |
| ## Data and reproducibility | |
| - [`leaderboard-data.json`](./leaderboard-data.json): displayed data, metric | |
| profiles, counts, evidence status and source-file SHA-256 hashes. | |
| - [`leaderboard.csv`](./leaderboard.csv): 65 displayed metric rows; thirteen | |
| groups × five metrics. One metric profile is explicit on every row. | |
| - [`results/`](./results): frozen per-task archives, configurations and summaries. | |
| - [`legacy.html`](./legacy.html): original leaderboard with a historical banner. | |
| Its original [`results.csv`](./results.csv) is unchanged and is not mixed into | |
| the new comparison. | |
| The exact source archive commit is recorded as `source_revision` in | |
| [`leaderboard-data.json`](./leaderboard-data.json); every displayed evidence | |
| link is pinned to that commit. | |
| Dataset: [CamoAiLab/InferenceNet](https://huggingface.co/datasets/CamoAiLab/InferenceNet), | |
| revision `59f9512a38e594528807744214a60ee00367434e`, task list | |
| `Selected_1000/1000_new.csv`. | |
| The per-run archive files keep their original publication metadata. Their | |
| historical `included_in_displayed_leaderboard: false` fields describe the archive | |
| upload, not this later display release. The current display manifest is | |
| `leaderboard-data.json`; no experiment evidence is edited to publish this page. | |
| To rebuild the display exports from those archives: | |
| ```sh | |
| python build_leaderboard.py --source . --output . \ | |
| --source-revision ARCHIVE_COMMIT_SHA --published-date 2026-09-20 | |
| ``` | |
| Use the 40-character archive commit SHA from `leaderboard-data.json` above. | |
| The builder verifies all task populations and 61 metric counts against archived | |
| flags. The four Sol/Opus local full-replication summaries have no per-task local | |
| flags in the archive and are preserved from their accepted summaries. This is a | |
| display refresh, not a fresh inference run, native artifact audit, or rescore. | |
| ## Project | |
| [Project website](https://easonai-5589.github.io/inferencenet-challenge/) · | |
| [Project code](https://github.com/EasonAI-5589/NTU-GIFTS-InferenceNet) · | |
| [Evaluation framework](https://github.com/YTZSR/chatpilot_evaluator/tree/dev-gyc-eval) | |
| ## Editing the portal | |
| `index.html`, `data.html`, `leaderboard.html` and `agent.html` are the four static pages. | |
| `site-shell.js` shares language and theme preferences; `portal.css` extends the | |
| existing leaderboard styles. `portal.js` renders the publication counts and data | |
| distributions. Published metric values still come only from `leaderboard-data.json`. | |
| To regenerate the data distribution summary from the exact pinned CSV: | |
| ```sh | |
| python build_dataset_summary.py /path/to/1000_new.csv | |
| ``` | |
| The script verifies the CSV SHA-256 before generating `dataset-summary.json`. | |
| Serve the repository with a local HTTP server to preview module scripts and JSON | |
| loading; no frontend build step is required. | |