YICHEN013's picture
Add the InferenceNet project portal with Home, Data and Agent pages
03c1286 verified
|
Raw History Blame Contribute Delete
7.01 kB
---
title: InferenceNet Challenge
emoji: 📊
colorFrom: green
colorTo: blue
sdk: static
pinned: false
license: apache-2.0
short_description: Explore the benchmark data, agent harnesses and leaderboard
---
# InferenceNet Challenge
The project portal brings the benchmark, data documentation, execution walkthrough and published results into one Hugging Face Space.
- [Home](https://camoailab-inferencenet-leaderboard.static.hf.space/index.html)
- [Data](https://camoailab-inferencenet-leaderboard.static.hf.space/data.html)
- [Leaderboard](https://camoailab-inferencenet-leaderboard.static.hf.space/leaderboard.html)
- [Agent / Harness](https://camoailab-inferencenet-leaderboard.static.hf.space/agent.html)
- [Dataset repository](https://huggingface.co/datasets/CamoAiLab/InferenceNet)
The original Space URL remains valid. Data files continue to live in the dataset repository; the portal includes only a small, reproducible distribution summary. Home and Leaderboard both read `leaderboard-data.json`. `dataset-summary.json` is derived from the same pinned task-list revision.
## Agent & Harness Leaderboard
The September 2026 research leaderboard displays **six requested model aliases,
thirteen groups, and 13,000 archived task records**. Switch between the original
baseline / DeepAgents paired comparison, a single-pass Agent ranking, and a
harness ranking that also includes one independent DSH run. Every group uses
the same 1,000 Selected_1000 task IDs. An API alias does not establish the
identity of the underlying model weights.
- **Single-pass Agent:** one model call to generate code, followed by execution.
- **DeepAgents:** up to six model calls and four trial tool
calls, followed by final execution.
- **DeepSeek Harness (DSH):** official SDK 0.1.5rc1, requested
`deepseek-v4-pro` API alias; displayed separately as
**DeepSeek API [v4-pro alias] + DSH**. Its underlying model weights are
**unverified**, and it is never used in the baseline / DeepAgents paired lift.
Models: GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro Preview,
Qwen3.7-Max, and the historical DeepSeek V4 Pro entries. GPT-5.5 is outside this edition.
The DSH [run protocol](./results/deepseek-v4-pro/dsh/README.md) documents its
fixed selection of 1,000 results: 913 original attempts and 87 technical
replacement attempts (60 infrastructure failures and 27 input-schema failures).
The 27 schema recoveries changed the input-schema prompt. All 1,087 attempts
remain in the cost accounting; selection is not best-of. The final selection
has 699 scored, 271 failed and 30 uncertain task outcomes. Matching per-attempt
numeric caps do not establish equal total run budgets: the DSH result includes
87 additional attempts. Full execution-environment and reasoning-setting
parity with the prior DeepAgents run has not been verified.
## Reading the results
The default metric is **full replication**, from `local-paper-v1`: coefficient
and standard-error relative errors ≤ 1%, and p-value absolute error ≤ 0.01.
Four separately labeled `hf-leaderboard-v1` metrics are also available:
execution success, coefficient-only partial replication (≤ 5% relative error),
coefficient direction, and significance category.
These are **locally scored research results**. Official scorer parity has not
been verified. All rates use 1,000 as the denominator; metric unknowns and invalid
evidence are retained and earn no successes. Task-status unknown counts are
reported separately from metric unknown counts.
The comparisons preserve historical recovery amendments. Call budgets, time
limits, reasoning settings and source protocols differ. Differences are
descriptive observations, not equal-cost causal estimates of a harness effect.
See each run's configuration before making cross-model claims.
Kimi retains 5 baseline and 4 harness task-status unknowns. Qwen3.7-Max harness
retains 6 task-status unknowns and 1 invalid-evidence record (task 453), which is
not sealed. DSH retains 30 task-status unknowns, with no invalid-evidence
records in the fixed selection. Historical Sol/Opus interruption counts are shown as originally
exported. An archive of 1,000 task IDs does not imply 1,000 definitive outcomes.
## Data and reproducibility
- [`leaderboard-data.json`](./leaderboard-data.json): displayed data, metric
profiles, counts, evidence status and source-file SHA-256 hashes.
- [`leaderboard.csv`](./leaderboard.csv): 65 displayed metric rows; thirteen
groups × five metrics. One metric profile is explicit on every row.
- [`results/`](./results): frozen per-task archives, configurations and summaries.
- [`legacy.html`](./legacy.html): original leaderboard with a historical banner.
Its original [`results.csv`](./results.csv) is unchanged and is not mixed into
the new comparison.
The exact source archive commit is recorded as `source_revision` in
[`leaderboard-data.json`](./leaderboard-data.json); every displayed evidence
link is pinned to that commit.
Dataset: [CamoAiLab/InferenceNet](https://huggingface.co/datasets/CamoAiLab/InferenceNet),
revision `59f9512a38e594528807744214a60ee00367434e`, task list
`Selected_1000/1000_new.csv`.
The per-run archive files keep their original publication metadata. Their
historical `included_in_displayed_leaderboard: false` fields describe the archive
upload, not this later display release. The current display manifest is
`leaderboard-data.json`; no experiment evidence is edited to publish this page.
To rebuild the display exports from those archives:
```sh
python build_leaderboard.py --source . --output . \
--source-revision ARCHIVE_COMMIT_SHA --published-date 2026-09-20
```
Use the 40-character archive commit SHA from `leaderboard-data.json` above.
The builder verifies all task populations and 61 metric counts against archived
flags. The four Sol/Opus local full-replication summaries have no per-task local
flags in the archive and are preserved from their accepted summaries. This is a
display refresh, not a fresh inference run, native artifact audit, or rescore.
## Project
[Project website](https://easonai-5589.github.io/inferencenet-challenge/) ·
[Project code](https://github.com/EasonAI-5589/NTU-GIFTS-InferenceNet) ·
[Evaluation framework](https://github.com/YTZSR/chatpilot_evaluator/tree/dev-gyc-eval)
## Editing the portal
`index.html`, `data.html`, `leaderboard.html` and `agent.html` are the four static pages.
`site-shell.js` shares language and theme preferences; `portal.css` extends the
existing leaderboard styles. `portal.js` renders the publication counts and data
distributions. Published metric values still come only from `leaderboard-data.json`.
To regenerate the data distribution summary from the exact pinned CSV:
```sh
python build_dataset_summary.py /path/to/1000_new.csv
```
The script verifies the CSV SHA-256 before generating `dataset-summary.json`.
Serve the repository with a local HTTP server to preview module scripts and JSON
loading; no frontend build step is required.