Add DeepSeek API + DSH results and independent leaderboard entry

#6
by lry520111 - opened
README.md CHANGED
@@ -6,22 +6,38 @@ colorTo: blue
6
  sdk: static
7
  pinned: false
8
  license: apache-2.0
9
- short_description: Single-pass agents and DeepAgents on econometric replication
10
  ---
11
 
12
  # InferenceNet Agent & Harness Leaderboard
13
 
14
- The September 2026 research leaderboard displays **six models, twelve groups,
15
- and 12,000 archived task records**. Switch between a paired comparison, a
16
- single-pass Agent ranking, and an Agent + Harness ranking. Every group uses
17
- the same 1,000 Selected_1000 task IDs.
 
 
18
 
19
  - **Single-pass Agent:** one model call to generate code, followed by execution.
20
- - **Agent + Harness:** DeepAgents with up to six model calls and four trial tool
21
  calls, followed by final execution.
 
 
 
 
22
 
23
  Models: GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro Preview,
24
- Qwen3.7-Max, and DeepSeek V4 Pro. GPT-5.5 is outside this edition.
 
 
 
 
 
 
 
 
 
 
25
 
26
  ## Reading the results
27
 
@@ -43,21 +59,24 @@ See each run's configuration before making cross-model claims.
43
 
44
  Kimi retains 5 baseline and 4 harness task-status unknowns. Qwen3.7-Max harness
45
  retains 6 task-status unknowns and 1 invalid-evidence record (task 453), which is
46
- not sealed. Historical Sol/Opus interruption counts are shown as originally
 
47
  exported. An archive of 1,000 task IDs does not imply 1,000 definitive outcomes.
48
 
49
  ## Data and reproducibility
50
 
51
  - [`leaderboard-data.json`](./leaderboard-data.json): displayed data, metric
52
  profiles, counts, evidence status and source-file SHA-256 hashes.
53
- - [`leaderboard.csv`](./leaderboard.csv): 60 displayed metric rows; six models ×
54
- two arms × five metrics. One metric profile is explicit on every row.
55
  - [`results/`](./results): frozen per-task archives, configurations and summaries.
56
  - [`legacy.html`](./legacy.html): original leaderboard with a historical banner.
57
  Its original [`results.csv`](./results.csv) is unchanged and is not mixed into
58
  the new comparison.
59
 
60
- Source archive commit: `731111974cf4d1a3c7a4f848a899f0b0708dd12b`.
 
 
61
  Dataset: [CamoAiLab/InferenceNet](https://huggingface.co/datasets/CamoAiLab/InferenceNet),
62
  revision `59f9512a38e594528807744214a60ee00367434e`, task list
63
  `Selected_1000/1000_new.csv`.
@@ -71,10 +90,11 @@ To rebuild the display exports from those archives:
71
 
72
  ```sh
73
  python build_leaderboard.py --source . --output . \
74
- --source-revision 731111974cf4d1a3c7a4f848a899f0b0708dd12b
75
  ```
76
 
77
- The builder verifies all task populations and 56 metric counts against archived
 
78
  flags. The four Sol/Opus local full-replication summaries have no per-task local
79
  flags in the archive and are preserved from their accepted summaries. This is a
80
  display refresh, not a fresh inference run, native artifact audit, or rescore.
 
6
  sdk: static
7
  pinned: false
8
  license: apache-2.0
9
+ short_description: Single-pass, DeepAgents and DSH replication results
10
  ---
11
 
12
  # InferenceNet Agent & Harness Leaderboard
13
 
14
+ The September 2026 research leaderboard displays **six requested model aliases,
15
+ thirteen groups, and 13,000 archived task records**. Switch between the original
16
+ baseline / DeepAgents paired comparison, a single-pass Agent ranking, and a
17
+ harness ranking that also includes one independent DSH run. Every group uses
18
+ the same 1,000 Selected_1000 task IDs. An API alias does not establish the
19
+ identity of the underlying model weights.
20
 
21
  - **Single-pass Agent:** one model call to generate code, followed by execution.
22
+ - **DeepAgents:** up to six model calls and four trial tool
23
  calls, followed by final execution.
24
+ - **DeepSeek Harness (DSH):** official SDK 0.1.5rc1, requested
25
+ `deepseek-v4-pro` API alias; displayed separately as
26
+ **DeepSeek API [v4-pro alias] + DSH**. Its underlying model weights are
27
+ **unverified**, and it is never used in the baseline / DeepAgents paired lift.
28
 
29
  Models: GPT-5.6 Sol, Claude Opus 4.8, Kimi K3, Gemini 3.1 Pro Preview,
30
+ Qwen3.7-Max, and the historical DeepSeek V4 Pro entries. GPT-5.5 is outside this edition.
31
+
32
+ The DSH [run protocol](./results/deepseek-v4-pro/dsh/README.md) documents its
33
+ fixed selection of 1,000 results: 913 original attempts and 87 technical
34
+ replacement attempts (60 infrastructure failures and 27 input-schema failures).
35
+ The 27 schema recoveries changed the input-schema prompt. All 1,087 attempts
36
+ remain in the cost accounting; selection is not best-of. The final selection
37
+ has 699 scored, 271 failed and 30 uncertain task outcomes. Matching per-attempt
38
+ numeric caps do not establish equal total run budgets: the DSH result includes
39
+ 87 additional attempts. Full execution-environment and reasoning-setting
40
+ parity with the prior DeepAgents run has not been verified.
41
 
42
  ## Reading the results
43
 
 
59
 
60
  Kimi retains 5 baseline and 4 harness task-status unknowns. Qwen3.7-Max harness
61
  retains 6 task-status unknowns and 1 invalid-evidence record (task 453), which is
62
+ not sealed. DSH retains 30 task-status unknowns, with no invalid-evidence
63
+ records in the fixed selection. Historical Sol/Opus interruption counts are shown as originally
64
  exported. An archive of 1,000 task IDs does not imply 1,000 definitive outcomes.
65
 
66
  ## Data and reproducibility
67
 
68
  - [`leaderboard-data.json`](./leaderboard-data.json): displayed data, metric
69
  profiles, counts, evidence status and source-file SHA-256 hashes.
70
+ - [`leaderboard.csv`](./leaderboard.csv): 65 displayed metric rows; thirteen
71
+ groups × five metrics. One metric profile is explicit on every row.
72
  - [`results/`](./results): frozen per-task archives, configurations and summaries.
73
  - [`legacy.html`](./legacy.html): original leaderboard with a historical banner.
74
  Its original [`results.csv`](./results.csv) is unchanged and is not mixed into
75
  the new comparison.
76
 
77
+ The exact source archive commit is recorded as `source_revision` in
78
+ [`leaderboard-data.json`](./leaderboard-data.json); every displayed evidence
79
+ link is pinned to that commit.
80
  Dataset: [CamoAiLab/InferenceNet](https://huggingface.co/datasets/CamoAiLab/InferenceNet),
81
  revision `59f9512a38e594528807744214a60ee00367434e`, task list
82
  `Selected_1000/1000_new.csv`.
 
90
 
91
  ```sh
92
  python build_leaderboard.py --source . --output . \
93
+ --source-revision ARCHIVE_COMMIT_SHA --published-date 2026-09-20
94
  ```
95
 
96
+ Use the 40-character archive commit SHA from `leaderboard-data.json` above.
97
+ The builder verifies all task populations and 61 metric counts against archived
98
  flags. The four Sol/Opus local full-replication summaries have no per-task local
99
  flags in the archive and are preserved from their accepted summaries. This is a
100
  display refresh, not a fresh inference run, native artifact audit, or rescore.
build_leaderboard.py CHANGED
@@ -7,6 +7,7 @@ import argparse
7
  import csv
8
  import hashlib
9
  import json
 
10
  from pathlib import Path
11
 
12
  NAMES = {
@@ -18,30 +19,45 @@ NAMES = {
18
  "deepseek-v4-pro": "DeepSeek V4 Pro",
19
  }
20
  METRICS = ("compilation_success", "partial_replication", "coefficient_direction", "significance_level")
 
 
 
 
 
 
 
21
 
22
 
23
- def build(source, output, revision):
24
  assert len(revision) == 40 and all(c in "0123456789abcdef" for c in revision)
25
- index = json.loads((source / "results/index.json").read_text())
26
  entries = index["entries"]
27
- assert {(e["model"], e["arm"]) for e in entries} == {(m, a) for m in NAMES for a in ("baseline", "deepagents")}
28
- assert len(entries) == 12
 
 
29
  models = {m: {"id": m, "name": name, "arms": {}} for m, name in NAMES.items()}
30
  population = None
31
  dataset_revision = None
32
  source_files = {}
 
33
  for entry in entries:
34
  path = Path(entry["path"])
35
  assert path == Path("results") / entry["model"] / entry["arm"]
36
- summary = json.loads((source / path / "summary.json").read_text())
37
- config = json.loads((source / path / "config.json").read_text())
38
- records = [json.loads(line) for line in (source / path / "results.jsonl").read_text().splitlines()]
39
  ids = {r["task_id"] for r in records}
40
  assert len(records) == len(ids) == entry["rows"] == summary["result"]["expected_count"] == 1000
41
  assert ids == set(config["dataset"]["task_ids"])
42
  assert all(r["model"] == entry["model"] and r["arm"] == entry["arm"] for r in records)
43
  assert (summary["model"], summary["arm"]) == (entry["model"], entry["arm"])
44
  assert summary["official_parity_verified"] is False
 
 
 
 
 
45
  if population is None:
46
  population = ids
47
  dataset_revision = summary["dataset_revision"]
@@ -55,10 +71,15 @@ def build(source, output, revision):
55
  assert abs(m["rate"] - count / n) < 1e-12
56
  field = "local_perfect" if local else metric
57
  has_flags = all(field in r for r in records)
 
 
58
  if not local or has_flags:
59
  assert all(r[field] is None or isinstance(r[field], bool) for r in records)
60
  assert sum(r[field] is True for r in records) == count
61
  assert sum(r[field] is None for r in records) == unknown
 
 
 
62
  # Sol/Opus local flags were not exported: preserve their accepted
63
  # summary counts, without claiming a fresh per-task recomputation.
64
  metrics[metric] = {
@@ -77,29 +98,37 @@ def build(source, output, revision):
77
  "interruption_count": summary["result"].get("interruption_count"),
78
  "evidence_class": summary["evidence_class"],
79
  }
 
 
 
 
80
  models[entry["model"]]["arms"][entry["arm"]] = arm
 
81
  for name in ("summary.json", "results.jsonl", "config.json"):
82
  rel = (path / name).as_posix()
83
  source_files[rel] = hashlib.sha256((source / rel).read_bytes()).hexdigest()
84
  data = {
85
  "schema_version": 1, "kind": "research_leaderboard",
86
- "published_date": "2026-09-17", "source_repository": "CamoAiLab/InferenceNet-Leaderboard",
87
  "source_revision": revision, "dataset_revision": dataset_revision,
88
  "task_list": "Selected_1000/1000_new.csv", "tasks_per_group": 1000,
89
- "groups": 12, "records": 12000, "official_parity_verified": False,
 
 
90
  "models": list(models.values()), "source_sha256": source_files,
91
  }
92
  output.mkdir(parents=True, exist_ok=True)
93
- (output / "leaderboard-data.json").write_text(json.dumps(data, indent=2) + "\n")
94
- with (output / "leaderboard.csv").open("w", newline="") as f:
95
  writer = csv.writer(f)
96
- writer.writerow(["model", "arm", "harness", "metric_profile", "metric", "successes", "denominator", "unknown", "score_percent", "official_parity_verified", "source_revision"])
97
  for model in data["models"]:
98
  for arm_id, arm in model["arms"].items():
99
  for key, metric in arm["metrics"].items():
100
- writer.writerow([model["id"], arm_id, arm["harness"], metric["profile"], key, metric["count"], 1000, metric["unknown"], metric["score"], False, revision])
101
- print("Verified: 6 models, 12 arms, 12,000 unique-per-arm records; 60 displayed metric summaries.")
102
- print("56 metric summaries checked against task flags; 4 historical local summaries preserved as accepted.")
 
103
 
104
 
105
  if __name__ == "__main__":
@@ -107,5 +136,6 @@ if __name__ == "__main__":
107
  parser.add_argument("--source", type=Path, default=Path("."))
108
  parser.add_argument("--output", type=Path, default=Path("."))
109
  parser.add_argument("--source-revision", required=True)
 
110
  args = parser.parse_args()
111
- build(args.source, args.output, args.source_revision)
 
7
  import csv
8
  import hashlib
9
  import json
10
+ from datetime import date
11
  from pathlib import Path
12
 
13
  NAMES = {
 
19
  "deepseek-v4-pro": "DeepSeek V4 Pro",
20
  }
21
  METRICS = ("compilation_success", "partial_replication", "coefficient_direction", "significance_level")
22
+ PAIRED_ARMS = ("baseline", "deepagents")
23
+ DSH_ENTRY = ("deepseek-v4-pro", "dsh")
24
+ DSH_NAME = "DeepSeek API [v4-pro alias] + DSH"
25
+ DSH_NOTICE = (
26
+ "Requested API alias: deepseek-v4-pro. Underlying model weights are unverified. "
27
+ "This technical recovery result is a separate harness entry, not a paired DeepAgents gain."
28
+ )
29
 
30
 
31
+ def build(source, output, revision, published_date=None):
32
  assert len(revision) == 40 and all(c in "0123456789abcdef" for c in revision)
33
+ index = json.loads((source / "results/index.json").read_text(encoding="utf-8"))
34
  entries = index["entries"]
35
+ pairs = {(e["model"], e["arm"]) for e in entries}
36
+ original_pairs = {(m, a) for m in NAMES for a in PAIRED_ARMS}
37
+ assert original_pairs <= pairs <= original_pairs | {DSH_ENTRY}
38
+ assert len(entries) == len(pairs), "Duplicate archive entry"
39
  models = {m: {"id": m, "name": name, "arms": {}} for m, name in NAMES.items()}
40
  population = None
41
  dataset_revision = None
42
  source_files = {}
43
+ checked_metric_count = accepted_summary_count = records_count = 0
44
  for entry in entries:
45
  path = Path(entry["path"])
46
  assert path == Path("results") / entry["model"] / entry["arm"]
47
+ summary = json.loads((source / path / "summary.json").read_text(encoding="utf-8"))
48
+ config = json.loads((source / path / "config.json").read_text(encoding="utf-8"))
49
+ records = [json.loads(line) for line in (source / path / "results.jsonl").read_text(encoding="utf-8").splitlines()]
50
  ids = {r["task_id"] for r in records}
51
  assert len(records) == len(ids) == entry["rows"] == summary["result"]["expected_count"] == 1000
52
  assert ids == set(config["dataset"]["task_ids"])
53
  assert all(r["model"] == entry["model"] and r["arm"] == entry["arm"] for r in records)
54
  assert (summary["model"], summary["arm"]) == (entry["model"], entry["arm"])
55
  assert summary["official_parity_verified"] is False
56
+ independent = (entry["model"], entry["arm"]) == DSH_ENTRY
57
+ if independent:
58
+ assert config["model_alias"] == "deepseek-v4-pro"
59
+ assert config["underlying_model_weights_verified"] is False
60
+ assert summary["evidence_class"] == "technical_recovery_view"
61
  if population is None:
62
  population = ids
63
  dataset_revision = summary["dataset_revision"]
 
71
  assert abs(m["rate"] - count / n) < 1e-12
72
  field = "local_perfect" if local else metric
73
  has_flags = all(field in r for r in records)
74
+ if independent:
75
+ assert has_flags, "Every DSH displayed metric requires an archived task flag"
76
  if not local or has_flags:
77
  assert all(r[field] is None or isinstance(r[field], bool) for r in records)
78
  assert sum(r[field] is True for r in records) == count
79
  assert sum(r[field] is None for r in records) == unknown
80
+ checked_metric_count += 1
81
+ else:
82
+ accepted_summary_count += 1
83
  # Sol/Opus local flags were not exported: preserve their accepted
84
  # summary counts, without claiming a fresh per-task recomputation.
85
  metrics[metric] = {
 
98
  "interruption_count": summary["result"].get("interruption_count"),
99
  "evidence_class": summary["evidence_class"],
100
  }
101
+ if independent:
102
+ arm.update(display_name=DSH_NAME, identity_notice=DSH_NOTICE, paired_comparison=False,
103
+ model_alias=config["model_alias"], underlying_model_weights_verified=False,
104
+ comparison_group="independent", protocol_url=(path / "README.md").as_posix())
105
  models[entry["model"]]["arms"][entry["arm"]] = arm
106
+ records_count += len(records)
107
  for name in ("summary.json", "results.jsonl", "config.json"):
108
  rel = (path / name).as_posix()
109
  source_files[rel] = hashlib.sha256((source / rel).read_bytes()).hexdigest()
110
  data = {
111
  "schema_version": 1, "kind": "research_leaderboard",
112
+ "published_date": published_date or date.today().isoformat(), "source_repository": "CamoAiLab/InferenceNet-Leaderboard",
113
  "source_revision": revision, "dataset_revision": dataset_revision,
114
  "task_list": "Selected_1000/1000_new.csv", "tasks_per_group": 1000,
115
+ "groups": len(entries), "records": records_count, "model_aliases": len(models),
116
+ "paired_groups": len(original_pairs), "independent_groups": len(pairs-original_pairs),
117
+ "paired_arms": list(PAIRED_ARMS), "official_parity_verified": False,
118
  "models": list(models.values()), "source_sha256": source_files,
119
  }
120
  output.mkdir(parents=True, exist_ok=True)
121
+ (output / "leaderboard-data.json").write_text(json.dumps(data, indent=2) + "\n", encoding="utf-8")
122
+ with (output / "leaderboard.csv").open("w", newline="", encoding="utf-8") as f:
123
  writer = csv.writer(f)
124
+ writer.writerow(["model", "arm", "harness", "metric_profile", "metric", "successes", "denominator", "unknown", "score_percent", "official_parity_verified", "source_revision", "display_name", "comparison_group", "underlying_model_weights_verified"])
125
  for model in data["models"]:
126
  for arm_id, arm in model["arms"].items():
127
  for key, metric in arm["metrics"].items():
128
+ writer.writerow([model["id"], arm_id, arm["harness"], metric["profile"], key, metric["count"], 1000, metric["unknown"], metric["score"], False, revision, arm.get("display_name", model["name"]), arm.get("comparison_group", "historical_pair"), arm.get("underlying_model_weights_verified", "")])
129
+ print(f"Verified: {len(models)} requested model aliases, {len(entries)} arms, {records_count:,} records; {len(entries)*5} displayed metric summaries.")
130
+ print(f"{checked_metric_count} metric summaries checked against task flags; {accepted_summary_count} historical local summaries preserved as accepted.")
131
+ return data
132
 
133
 
134
  if __name__ == "__main__":
 
136
  parser.add_argument("--source", type=Path, default=Path("."))
137
  parser.add_argument("--output", type=Path, default=Path("."))
138
  parser.add_argument("--source-revision", required=True)
139
+ parser.add_argument("--published-date", help="Publication date, YYYY-MM-DD; defaults to today")
140
  args = parser.parse_args()
141
+ build(args.source, args.output, args.source_revision, args.published_date)
index.html CHANGED
@@ -3,7 +3,7 @@
3
  <head>
4
  <meta charset="utf-8">
5
  <meta name="viewport" content="width=device-width, initial-scale=1">
6
- <meta name="description" content="InferenceNet research leaderboard: six models compared with single-pass generation and a DeepAgents harness on 1,000 econometric replication tasks.">
7
  <meta name="color-scheme" content="dark">
8
  <title>InferenceNet · Agent & Harness Leaderboard</title>
9
  <link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 32 32'%3E%3Crect width='32' height='32' rx='6' fill='%230c1219'/%3E%3Ctext x='5' y='23' font-family='sans-serif' font-size='21' fill='%2378dbc0'%3EiN%3C/text%3E%3C/svg%3E">
@@ -24,30 +24,33 @@
24
 
25
  <header class="hero">
26
  <p class="eyebrow"><span class="dot" aria-hidden="true"></span><span data-i18n="edition">RESEARCH LEADERBOARD · SEPTEMBER 2026</span></p>
27
- <h1 data-i18n="title">Same models.<br><span>Two ways to solve.</span></h1>
28
- <p class="intro" data-i18n="intro">Compare single-pass agents with a DeepAgents harness on real econometric replication tasks.</p>
29
  <div class="facts" aria-label="Dataset coverage">
30
- <span><strong id="model-count">—</strong><span data-i18n="models">models</span></span>
31
  <span><strong id="group-count">—</strong><span data-i18n="groups">experiment groups</span></span>
32
  <span><strong id="task-count">—</strong><span data-i18n="tasks">tasks per group</span></span>
33
- <span class="update" data-i18n="updated">Published 17 Sep 2026</span>
 
34
  </div>
35
  </header>
36
 
37
  <section class="protocols" aria-label="Experimental settings">
38
  <div><p class="protocol-tag baseline"><span aria-hidden="true">A</span><strong data-i18n="single">Single-pass Agent</strong></p><p data-i18n="singleDesc">One code-generation call, followed by execution. No iterative agent loop.</p></div>
39
- <div><p class="protocol-tag harness"><span aria-hidden="true">B</span><strong data-i18n="harness">Agent + Harness</strong></p><p data-i18n="harnessDesc">DeepAgents plans, uses tools and revises code. Up to 6 model calls, plus final execution.</p></div>
 
40
  </section>
41
 
42
  <section id="rankings" aria-labelledby="rankings-title">
43
  <div class="section-head"><div><p class="eyebrow" data-i18n="resultsLabel">01 / RESULTS</p><h2 id="rankings-title" data-i18n="rankings">Agent & Harness leaderboard</h2></div><span class="badge" data-i18n="research">Research results · provisional</span></div>
44
  <p class="scope-note" data-i18n="scope">Locally scored, archived results. Official scorer parity is not yet verified. Different budgets and run protocols mean the lift is descriptive, not an equal-cost comparison.</p>
 
45
 
46
  <div class="toolbar">
47
  <div class="view-switch" role="group" aria-label="Leaderboard view">
48
- <button type="button" data-view="compare" aria-pressed="true" data-i18n="compare">Side by side</button>
49
  <button type="button" data-view="baseline" aria-pressed="false" data-i18n="single">Single-pass Agent</button>
50
- <button type="button" data-view="deepagents" aria-pressed="false" data-i18n="harness">Agent + Harness</button>
51
  </div>
52
  <a class="download" href="./leaderboard.csv" download data-i18n="download">Download CSV ↓</a>
53
  </div>
@@ -63,7 +66,7 @@
63
  <option value="significance_level" data-i18n="significance_level">Significance level · HF-style</option>
64
  </optgroup>
65
  </select></label>
66
- <label id="sort-control"><span data-i18n="sort">Sort by</span><select id="sort"><option value="deepagents" data-i18n="harnessScore">Harness score</option><option value="baseline" data-i18n="singleScore">Single-pass score</option><option value="gain" data-i18n="lift">Lift</option></select></label>
67
  <label class="search-control"><span data-i18n="search">Find a model</span><input id="search" type="search" placeholder="Search models…" autocomplete="off"></label>
68
  </div>
69
  <p id="metric-description" class="metric-description"></p>
@@ -81,13 +84,13 @@
81
  <div class="method-grid">
82
  <article><h3 data-i18n="fixedTasks">Fixed task set</h3><p data-i18n="fixedDesc">The same Selected_1000 task IDs and dataset revision are used for every group. Each score is successes ÷ 1,000; failed or unknown tasks stay in the denominator.</p></article>
83
  <article><h3 data-i18n="separateMetrics">Two scoring profiles</h3><p data-i18n="profilesDesc">Full replication is the local paper metric. The four HF-style metrics reproduce the published descriptions locally; they are not yet verified against the official scorer.</p></article>
84
- <article><h3 data-i18n="budgetTitle">Read lift with the budgets</h3><p data-i18n="budgetDesc">Baseline: 1 model call. DeepAgents: up to 6 model calls and 4 trial tool calls. Recovery amendments, time limits and run protocols differ; these are not equal-cost causal estimates.</p></article>
85
  </div>
86
  <details class="evidence"><summary data-i18n="evidence">Evidence status & metric definitions</summary>
87
  <p data-i18n="unknownDesc">“Metric unknown” means the selected metric cannot be assessed, including failed execution or unavailable outputs. It is different from a task whose final status is unknown. Neither is dropped or counted as a success.</p>
88
  <ul id="definitions"></ul>
89
- <div class="evidence-table-wrap"><table class="evidence-table"><caption data-i18n="evidenceCaption">Status from each accepted archive</caption><thead><tr><th data-i18n="model">Model</th><th data-i18n="single">Single-pass Agent</th><th data-i18n="harness">Agent + Harness</th></tr></thead><tbody id="evidence-body"></tbody></table></div>
90
- <p data-i18n="verification">HF-style counts are checked against archived task flags. Full-replication counts for Sol and Opus come from accepted summaries; the other four models also have per-task local flags. This refresh does not rerun or rescore any experiment.</p>
91
  <p data-i18n="archiveNote">Per-run archive files remain frozen, including their original publication metadata. This page is a new display derived from those archives. GPT-5.5 and unarchived experiments are not included in this edition.</p>
92
  </details>
93
  </section>
 
3
  <head>
4
  <meta charset="utf-8">
5
  <meta name="viewport" content="width=device-width, initial-scale=1">
6
+ <meta name="description" content="InferenceNet research leaderboard: single-pass generation, DeepAgents and an independent DeepSeek Harness result on 1,000 econometric replication tasks.">
7
  <meta name="color-scheme" content="dark">
8
  <title>InferenceNet · Agent & Harness Leaderboard</title>
9
  <link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 32 32'%3E%3Crect width='32' height='32' rx='6' fill='%230c1219'/%3E%3Ctext x='5' y='23' font-family='sans-serif' font-size='21' fill='%2378dbc0'%3EiN%3C/text%3E%3C/svg%3E">
 
24
 
25
  <header class="hero">
26
  <p class="eyebrow"><span class="dot" aria-hidden="true"></span><span data-i18n="edition">RESEARCH LEADERBOARD · SEPTEMBER 2026</span></p>
27
+ <h1 data-i18n="title">Models and harnesses.<br><span>Real research, reproduced.</span></h1>
28
+ <p class="intro" data-i18n="intro">Explore single-pass agents, DeepAgents and a separate DeepSeek Harness result on real econometric replication tasks.</p>
29
  <div class="facts" aria-label="Dataset coverage">
30
+ <span><strong id="model-count">—</strong><span data-i18n="models">requested model aliases</span></span>
31
  <span><strong id="group-count">—</strong><span data-i18n="groups">experiment groups</span></span>
32
  <span><strong id="task-count">—</strong><span data-i18n="tasks">tasks per group</span></span>
33
+ <span><strong id="record-count">—</strong><span data-i18n="records">task records</span></span>
34
+ <span class="update" id="published-date"></span>
35
  </div>
36
  </header>
37
 
38
  <section class="protocols" aria-label="Experimental settings">
39
  <div><p class="protocol-tag baseline"><span aria-hidden="true">A</span><strong data-i18n="single">Single-pass Agent</strong></p><p data-i18n="singleDesc">One code-generation call, followed by execution. No iterative agent loop.</p></div>
40
+ <div><p class="protocol-tag harness"><span aria-hidden="true">B</span><strong data-i18n="deepagents">DeepAgents</strong></p><p data-i18n="harnessDesc">DeepAgents plans, uses tools and revises code. Up to 6 model calls, plus final execution.</p></div>
41
+ <div><p class="protocol-tag independent"><span aria-hidden="true">C</span><strong data-i18n="dsh">DeepSeek Harness · separate run</strong></p><p data-i18n="dshDesc">Official DSH SDK, requested v4-pro API alias. A fixed technical recovery selection; underlying model weights are unverified.</p></div>
42
  </section>
43
 
44
  <section id="rankings" aria-labelledby="rankings-title">
45
  <div class="section-head"><div><p class="eyebrow" data-i18n="resultsLabel">01 / RESULTS</p><h2 id="rankings-title" data-i18n="rankings">Agent & Harness leaderboard</h2></div><span class="badge" data-i18n="research">Research results · provisional</span></div>
46
  <p class="scope-note" data-i18n="scope">Locally scored, archived results. Official scorer parity is not yet verified. Different budgets and run protocols mean the lift is descriptive, not an equal-cost comparison.</p>
47
+ <aside id="independent-notice" class="independent-notice" hidden><strong data-i18n="dshNoticeTitle">DSH is an independent result.</strong><p data-i18n="dshNotice">The request used deepseek-v4-pro; the underlying model weights were not verified. The fixed 1,000-task selection includes 87 technical replacement attempts (27 with changed input-schema prompts), with 30 task outcomes still unknown. It is excluded from the paired DeepAgents lift.</p><a id="dsh-protocol" data-i18n="dshProtocol" target="_blank" rel="noopener noreferrer">Read run protocol & evidence ↗</a></aside>
48
 
49
  <div class="toolbar">
50
  <div class="view-switch" role="group" aria-label="Leaderboard view">
51
+ <button type="button" data-view="compare" aria-pressed="false" data-i18n="compare">Baseline / DeepAgents pairs</button>
52
  <button type="button" data-view="baseline" aria-pressed="false" data-i18n="single">Single-pass Agent</button>
53
+ <button type="button" data-view="harness" aria-pressed="true" data-i18n="harness">All harness runs</button>
54
  </div>
55
  <a class="download" href="./leaderboard.csv" download data-i18n="download">Download CSV ↓</a>
56
  </div>
 
66
  <option value="significance_level" data-i18n="significance_level">Significance level · HF-style</option>
67
  </optgroup>
68
  </select></label>
69
+ <label id="sort-control"><span data-i18n="sort">Sort by</span><select id="sort"><option value="deepagents" data-i18n="harnessScore">DeepAgents score</option><option value="baseline" data-i18n="singleScore">Single-pass score</option><option value="gain" data-i18n="lift">Lift</option></select></label>
70
  <label class="search-control"><span data-i18n="search">Find a model</span><input id="search" type="search" placeholder="Search models…" autocomplete="off"></label>
71
  </div>
72
  <p id="metric-description" class="metric-description"></p>
 
84
  <div class="method-grid">
85
  <article><h3 data-i18n="fixedTasks">Fixed task set</h3><p data-i18n="fixedDesc">The same Selected_1000 task IDs and dataset revision are used for every group. Each score is successes ÷ 1,000; failed or unknown tasks stay in the denominator.</p></article>
86
  <article><h3 data-i18n="separateMetrics">Two scoring profiles</h3><p data-i18n="profilesDesc">Full replication is the local paper metric. The four HF-style metrics reproduce the published descriptions locally; they are not yet verified against the official scorer.</p></article>
87
+ <article><h3 data-i18n="budgetTitle">Read lift with the budgets</h3><p data-i18n="budgetDesc">Baseline: 1 model call. DeepAgents and DSH: up to 6 model calls and 4 trial tool calls per attempt. DSH adds 87 recovery attempts; full environment and reasoning parity is unverified. DSH has no paired lift; these are not equal-cost causal estimates.</p></article>
88
  </div>
89
  <details class="evidence"><summary data-i18n="evidence">Evidence status & metric definitions</summary>
90
  <p data-i18n="unknownDesc">“Metric unknown” means the selected metric cannot be assessed, including failed execution or unavailable outputs. It is different from a task whose final status is unknown. Neither is dropped or counted as a success.</p>
91
  <ul id="definitions"></ul>
92
+ <div class="evidence-table-wrap"><table class="evidence-table"><caption data-i18n="evidenceCaption">Status from each accepted archive</caption><thead><tr><th data-i18n="model">Model / API alias</th><th data-i18n="protocol">Protocol</th><th data-i18n="evidenceCol">Evidence</th></tr></thead><tbody id="evidence-body"></tbody></table></div>
93
+ <p data-i18n="verification">HF-style counts are checked against archived task flags. Full-replication counts for Sol and Opus come from accepted summaries; all other displayed runs, including DSH, have per-task local flags. This refresh does not rerun or rescore any experiment.</p>
94
  <p data-i18n="archiveNote">Per-run archive files remain frozen, including their original publication metadata. This page is a new display derived from those archives. GPT-5.5 and unarchived experiments are not included in this edition.</p>
95
  </details>
96
  </section>
leaderboard-data.json CHANGED
@@ -1,785 +1,860 @@
1
- {
2
- "schema_version": 1,
3
- "kind": "research_leaderboard",
4
- "published_date": "2026-09-17",
5
- "source_repository": "CamoAiLab/InferenceNet-Leaderboard",
6
- "source_revision": "731111974cf4d1a3c7a4f848a899f0b0708dd12b",
7
- "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
8
- "task_list": "Selected_1000/1000_new.csv",
9
- "tasks_per_group": 1000,
10
- "groups": 12,
11
- "records": 12000,
12
- "official_parity_verified": false,
13
- "models": [
14
- {
15
- "id": "gpt-5.6-sol",
16
- "name": "GPT-5.6 Sol",
17
- "arms": {
18
- "baseline": {
19
- "path": "results/gpt-5.6-sol/baseline",
20
- "metrics": {
21
- "perfect": {
22
- "count": 338,
23
- "denominator": 1000,
24
- "unknown": 445,
25
- "failure": 217,
26
- "score": 33.8,
27
- "profile": "local-paper-v1",
28
- "verification": "accepted_summary"
29
- },
30
- "compilation_success": {
31
- "count": 557,
32
- "denominator": 1000,
33
- "unknown": 3,
34
- "failure": 440,
35
- "score": 55.7,
36
- "profile": "hf-leaderboard-v1",
37
- "verification": "archived_task_flags"
38
- },
39
- "partial_replication": {
40
- "count": 430,
41
- "denominator": 1000,
42
- "unknown": 444,
43
- "failure": 126,
44
- "score": 43.0,
45
- "profile": "hf-leaderboard-v1",
46
- "verification": "archived_task_flags"
47
- },
48
- "coefficient_direction": {
49
- "count": 537,
50
- "denominator": 1000,
51
- "unknown": 444,
52
- "failure": 19,
53
- "score": 53.7,
54
- "profile": "hf-leaderboard-v1",
55
- "verification": "archived_task_flags"
56
- },
57
- "significance_level": {
58
- "count": 485,
59
- "denominator": 1000,
60
- "unknown": 444,
61
- "failure": 71,
62
- "score": 48.5,
63
- "profile": "hf-leaderboard-v1",
64
- "verification": "archived_task_flags"
65
- }
66
- },
67
- "generated_at": "2026-09-13T17:49:16.537733+00:00",
68
- "complete": true,
69
- "harness": "none",
70
- "model_call_cap": 1,
71
- "task_unknown_count": null,
72
- "invalid_evidence_count": 0,
73
- "interruption_count": 3,
74
- "evidence_class": "amended_comparison"
75
- },
76
- "deepagents": {
77
- "path": "results/gpt-5.6-sol/deepagents",
78
- "metrics": {
79
- "perfect": {
80
- "count": 518,
81
- "denominator": 1000,
82
- "unknown": 95,
83
- "failure": 387,
84
- "score": 51.8,
85
- "profile": "local-paper-v1",
86
- "verification": "accepted_summary"
87
- },
88
- "compilation_success": {
89
- "count": 910,
90
- "denominator": 1000,
91
- "unknown": 10,
92
- "failure": 80,
93
- "score": 91.0,
94
- "profile": "hf-leaderboard-v1",
95
- "verification": "archived_task_flags"
96
- },
97
- "partial_replication": {
98
- "count": 699,
99
- "denominator": 1000,
100
- "unknown": 92,
101
- "failure": 209,
102
- "score": 69.9,
103
- "profile": "hf-leaderboard-v1",
104
- "verification": "archived_task_flags"
105
- },
106
- "coefficient_direction": {
107
- "count": 879,
108
- "denominator": 1000,
109
- "unknown": 92,
110
- "failure": 29,
111
- "score": 87.9,
112
- "profile": "hf-leaderboard-v1",
113
- "verification": "archived_task_flags"
114
- },
115
- "significance_level": {
116
- "count": 787,
117
- "denominator": 1000,
118
- "unknown": 93,
119
- "failure": 120,
120
- "score": 78.7,
121
- "profile": "hf-leaderboard-v1",
122
- "verification": "archived_task_flags"
123
- }
124
- },
125
- "generated_at": "2026-09-13T17:49:16.537733+00:00",
126
- "complete": true,
127
- "harness": "DeepAgents",
128
- "model_call_cap": 6,
129
- "task_unknown_count": null,
130
- "invalid_evidence_count": 0,
131
- "interruption_count": 10,
132
- "evidence_class": "amended_comparison"
133
- }
134
- }
135
- },
136
- {
137
- "id": "claude-opus-4-8",
138
- "name": "Claude Opus 4.8",
139
- "arms": {
140
- "baseline": {
141
- "path": "results/claude-opus-4-8/baseline",
142
- "metrics": {
143
- "perfect": {
144
- "count": 184,
145
- "denominator": 1000,
146
- "unknown": 538,
147
- "failure": 278,
148
- "score": 18.4,
149
- "profile": "local-paper-v1",
150
- "verification": "accepted_summary"
151
- },
152
- "compilation_success": {
153
- "count": 467,
154
- "denominator": 1000,
155
- "unknown": 0,
156
- "failure": 533,
157
- "score": 46.7,
158
- "profile": "hf-leaderboard-v1",
159
- "verification": "archived_task_flags"
160
- },
161
- "partial_replication": {
162
- "count": 305,
163
- "denominator": 1000,
164
- "unknown": 537,
165
- "failure": 158,
166
- "score": 30.5,
167
- "profile": "hf-leaderboard-v1",
168
- "verification": "archived_task_flags"
169
- },
170
- "coefficient_direction": {
171
- "count": 437,
172
- "denominator": 1000,
173
- "unknown": 537,
174
- "failure": 26,
175
- "score": 43.7,
176
- "profile": "hf-leaderboard-v1",
177
- "verification": "archived_task_flags"
178
- },
179
- "significance_level": {
180
- "count": 370,
181
- "denominator": 1000,
182
- "unknown": 537,
183
- "failure": 93,
184
- "score": 37.0,
185
- "profile": "hf-leaderboard-v1",
186
- "verification": "archived_task_flags"
187
- }
188
- },
189
- "generated_at": "2026-09-13T17:50:43.870921+00:00",
190
- "complete": true,
191
- "harness": "none",
192
- "model_call_cap": 1,
193
- "task_unknown_count": null,
194
- "invalid_evidence_count": 0,
195
- "interruption_count": 0,
196
- "evidence_class": "amended_comparison"
197
- },
198
- "deepagents": {
199
- "path": "results/claude-opus-4-8/deepagents",
200
- "metrics": {
201
- "perfect": {
202
- "count": 392,
203
- "denominator": 1000,
204
- "unknown": 174,
205
- "failure": 434,
206
- "score": 39.2,
207
- "profile": "local-paper-v1",
208
- "verification": "accepted_summary"
209
- },
210
- "compilation_success": {
211
- "count": 834,
212
- "denominator": 1000,
213
- "unknown": 4,
214
- "failure": 162,
215
- "score": 83.4,
216
- "profile": "hf-leaderboard-v1",
217
- "verification": "archived_task_flags"
218
- },
219
- "partial_replication": {
220
- "count": 621,
221
- "denominator": 1000,
222
- "unknown": 170,
223
- "failure": 209,
224
- "score": 62.1,
225
- "profile": "hf-leaderboard-v1",
226
- "verification": "archived_task_flags"
227
- },
228
- "coefficient_direction": {
229
- "count": 793,
230
- "denominator": 1000,
231
- "unknown": 170,
232
- "failure": 37,
233
- "score": 79.3,
234
- "profile": "hf-leaderboard-v1",
235
- "verification": "archived_task_flags"
236
- },
237
- "significance_level": {
238
- "count": 696,
239
- "denominator": 1000,
240
- "unknown": 172,
241
- "failure": 132,
242
- "score": 69.6,
243
- "profile": "hf-leaderboard-v1",
244
- "verification": "archived_task_flags"
245
- }
246
- },
247
- "generated_at": "2026-09-13T17:50:43.870921+00:00",
248
- "complete": true,
249
- "harness": "DeepAgents",
250
- "model_call_cap": 6,
251
- "task_unknown_count": null,
252
- "invalid_evidence_count": 0,
253
- "interruption_count": 4,
254
- "evidence_class": "amended_comparison"
255
- }
256
- }
257
- },
258
- {
259
- "id": "kimi-k3",
260
- "name": "Kimi K3",
261
- "arms": {
262
- "baseline": {
263
- "path": "results/kimi-k3/baseline",
264
- "metrics": {
265
- "perfect": {
266
- "count": 230,
267
- "denominator": 1000,
268
- "unknown": 481,
269
- "failure": 289,
270
- "score": 23.0,
271
- "profile": "local-paper-v1",
272
- "verification": "archived_task_flags"
273
- },
274
- "compilation_success": {
275
- "count": 526,
276
- "denominator": 1000,
277
- "unknown": 5,
278
- "failure": 469,
279
- "score": 52.6,
280
- "profile": "hf-leaderboard-v1",
281
- "verification": "archived_task_flags"
282
- },
283
- "partial_replication": {
284
- "count": 366,
285
- "denominator": 1000,
286
- "unknown": 481,
287
- "failure": 153,
288
- "score": 36.6,
289
- "profile": "hf-leaderboard-v1",
290
- "verification": "archived_task_flags"
291
- },
292
- "coefficient_direction": {
293
- "count": 494,
294
- "denominator": 1000,
295
- "unknown": 481,
296
- "failure": 25,
297
- "score": 49.4,
298
- "profile": "hf-leaderboard-v1",
299
- "verification": "archived_task_flags"
300
- },
301
- "significance_level": {
302
- "count": 431,
303
- "denominator": 1000,
304
- "unknown": 481,
305
- "failure": 88,
306
- "score": 43.1,
307
- "profile": "hf-leaderboard-v1",
308
- "verification": "archived_task_flags"
309
- }
310
- },
311
- "generated_at": "2026-09-16T15:28:00.953066+00:00",
312
- "complete": false,
313
- "harness": "none",
314
- "model_call_cap": 1,
315
- "task_unknown_count": 5,
316
- "invalid_evidence_count": 0,
317
- "interruption_count": null,
318
- "evidence_class": "amended_comparison"
319
- },
320
- "deepagents": {
321
- "path": "results/kimi-k3/deepagents",
322
- "metrics": {
323
- "perfect": {
324
- "count": 390,
325
- "denominator": 1000,
326
- "unknown": 159,
327
- "failure": 451,
328
- "score": 39.0,
329
- "profile": "local-paper-v1",
330
- "verification": "archived_task_flags"
331
- },
332
- "compilation_success": {
333
- "count": 867,
334
- "denominator": 1000,
335
- "unknown": 4,
336
- "failure": 129,
337
- "score": 86.7,
338
- "profile": "hf-leaderboard-v1",
339
- "verification": "archived_task_flags"
340
- },
341
- "partial_replication": {
342
- "count": 635,
343
- "denominator": 1000,
344
- "unknown": 156,
345
- "failure": 209,
346
- "score": 63.5,
347
- "profile": "hf-leaderboard-v1",
348
- "verification": "archived_task_flags"
349
- },
350
- "coefficient_direction": {
351
- "count": 804,
352
- "denominator": 1000,
353
- "unknown": 156,
354
- "failure": 40,
355
- "score": 80.4,
356
- "profile": "hf-leaderboard-v1",
357
- "verification": "archived_task_flags"
358
- },
359
- "significance_level": {
360
- "count": 709,
361
- "denominator": 1000,
362
- "unknown": 157,
363
- "failure": 134,
364
- "score": 70.9,
365
- "profile": "hf-leaderboard-v1",
366
- "verification": "archived_task_flags"
367
- }
368
- },
369
- "generated_at": "2026-09-16T15:28:00.953066+00:00",
370
- "complete": false,
371
- "harness": "DeepAgents 0.7.13",
372
- "model_call_cap": 6,
373
- "task_unknown_count": 4,
374
- "invalid_evidence_count": 0,
375
- "interruption_count": null,
376
- "evidence_class": "amended_comparison"
377
- }
378
- }
379
- },
380
- {
381
- "id": "gemini-3.1-pro-preview",
382
- "name": "Gemini 3.1 Pro Preview",
383
- "arms": {
384
- "baseline": {
385
- "path": "results/gemini-3.1-pro-preview/baseline",
386
- "metrics": {
387
- "perfect": {
388
- "count": 246,
389
- "denominator": 1000,
390
- "unknown": 462,
391
- "failure": 292,
392
- "score": 24.6,
393
- "profile": "local-paper-v1",
394
- "verification": "archived_task_flags"
395
- },
396
- "compilation_success": {
397
- "count": 543,
398
- "denominator": 1000,
399
- "unknown": 0,
400
- "failure": 457,
401
- "score": 54.3,
402
- "profile": "hf-leaderboard-v1",
403
- "verification": "archived_task_flags"
404
- },
405
- "partial_replication": {
406
- "count": 366,
407
- "denominator": 1000,
408
- "unknown": 460,
409
- "failure": 174,
410
- "score": 36.6,
411
- "profile": "hf-leaderboard-v1",
412
- "verification": "archived_task_flags"
413
- },
414
- "coefficient_direction": {
415
- "count": 506,
416
- "denominator": 1000,
417
- "unknown": 460,
418
- "failure": 34,
419
- "score": 50.6,
420
- "profile": "hf-leaderboard-v1",
421
- "verification": "archived_task_flags"
422
- },
423
- "significance_level": {
424
- "count": 437,
425
- "denominator": 1000,
426
- "unknown": 460,
427
- "failure": 103,
428
- "score": 43.7,
429
- "profile": "hf-leaderboard-v1",
430
- "verification": "archived_task_flags"
431
- }
432
- },
433
- "generated_at": "2026-09-16T15:28:00.953066+00:00",
434
- "complete": true,
435
- "harness": "none",
436
- "model_call_cap": 1,
437
- "task_unknown_count": 0,
438
- "invalid_evidence_count": 0,
439
- "interruption_count": null,
440
- "evidence_class": "amended_comparison"
441
- },
442
- "deepagents": {
443
- "path": "results/gemini-3.1-pro-preview/deepagents",
444
- "metrics": {
445
- "perfect": {
446
- "count": 391,
447
- "denominator": 1000,
448
- "unknown": 155,
449
- "failure": 454,
450
- "score": 39.1,
451
- "profile": "local-paper-v1",
452
- "verification": "archived_task_flags"
453
- },
454
- "compilation_success": {
455
- "count": 850,
456
- "denominator": 1000,
457
- "unknown": 0,
458
- "failure": 150,
459
- "score": 85.0,
460
- "profile": "hf-leaderboard-v1",
461
- "verification": "archived_task_flags"
462
- },
463
- "partial_replication": {
464
- "count": 632,
465
- "denominator": 1000,
466
- "unknown": 153,
467
- "failure": 215,
468
- "score": 63.2,
469
- "profile": "hf-leaderboard-v1",
470
- "verification": "archived_task_flags"
471
- },
472
- "coefficient_direction": {
473
- "count": 811,
474
- "denominator": 1000,
475
- "unknown": 153,
476
- "failure": 36,
477
- "score": 81.1,
478
- "profile": "hf-leaderboard-v1",
479
- "verification": "archived_task_flags"
480
- },
481
- "significance_level": {
482
- "count": 718,
483
- "denominator": 1000,
484
- "unknown": 154,
485
- "failure": 128,
486
- "score": 71.8,
487
- "profile": "hf-leaderboard-v1",
488
- "verification": "archived_task_flags"
489
- }
490
- },
491
- "generated_at": "2026-09-16T15:28:00.953066+00:00",
492
- "complete": true,
493
- "harness": "DeepAgents 0.7.13",
494
- "model_call_cap": 6,
495
- "task_unknown_count": 0,
496
- "invalid_evidence_count": 0,
497
- "interruption_count": null,
498
- "evidence_class": "amended_comparison"
499
- }
500
- }
501
- },
502
- {
503
- "id": "qwen3.7-max",
504
- "name": "Qwen3.7-Max",
505
- "arms": {
506
- "baseline": {
507
- "path": "results/qwen3.7-max/baseline",
508
- "metrics": {
509
- "perfect": {
510
- "count": 198,
511
- "denominator": 1000,
512
- "unknown": 519,
513
- "failure": 283,
514
- "score": 19.8,
515
- "profile": "local-paper-v1",
516
- "verification": "archived_task_flags"
517
- },
518
- "compilation_success": {
519
- "count": 486,
520
- "denominator": 1000,
521
- "unknown": 0,
522
- "failure": 514,
523
- "score": 48.6,
524
- "profile": "hf-leaderboard-v1",
525
- "verification": "archived_task_flags"
526
- },
527
- "partial_replication": {
528
- "count": 315,
529
- "denominator": 1000,
530
- "unknown": 518,
531
- "failure": 167,
532
- "score": 31.5,
533
- "profile": "hf-leaderboard-v1",
534
- "verification": "archived_task_flags"
535
- },
536
- "coefficient_direction": {
537
- "count": 454,
538
- "denominator": 1000,
539
- "unknown": 518,
540
- "failure": 28,
541
- "score": 45.4,
542
- "profile": "hf-leaderboard-v1",
543
- "verification": "archived_task_flags"
544
- },
545
- "significance_level": {
546
- "count": 396,
547
- "denominator": 1000,
548
- "unknown": 518,
549
- "failure": 86,
550
- "score": 39.6,
551
- "profile": "hf-leaderboard-v1",
552
- "verification": "archived_task_flags"
553
- }
554
- },
555
- "generated_at": "2026-09-16T17:18:47.373930+00:00",
556
- "complete": true,
557
- "harness": "none",
558
- "model_call_cap": 1,
559
- "task_unknown_count": 0,
560
- "invalid_evidence_count": 0,
561
- "interruption_count": null,
562
- "evidence_class": "amended_comparison"
563
- },
564
- "deepagents": {
565
- "path": "results/qwen3.7-max/deepagents",
566
- "metrics": {
567
- "perfect": {
568
- "count": 248,
569
- "denominator": 1000,
570
- "unknown": 481,
571
- "failure": 271,
572
- "score": 24.8,
573
- "profile": "local-paper-v1",
574
- "verification": "archived_task_flags"
575
- },
576
- "compilation_success": {
577
- "count": 523,
578
- "denominator": 1000,
579
- "unknown": 7,
580
- "failure": 470,
581
- "score": 52.3,
582
- "profile": "hf-leaderboard-v1",
583
- "verification": "archived_task_flags"
584
- },
585
- "partial_replication": {
586
- "count": 379,
587
- "denominator": 1000,
588
- "unknown": 481,
589
- "failure": 140,
590
- "score": 37.9,
591
- "profile": "hf-leaderboard-v1",
592
- "verification": "archived_task_flags"
593
- },
594
- "coefficient_direction": {
595
- "count": 488,
596
- "denominator": 1000,
597
- "unknown": 481,
598
- "failure": 31,
599
- "score": 48.8,
600
- "profile": "hf-leaderboard-v1",
601
- "verification": "archived_task_flags"
602
- },
603
- "significance_level": {
604
- "count": 433,
605
- "denominator": 1000,
606
- "unknown": 481,
607
- "failure": 86,
608
- "score": 43.3,
609
- "profile": "hf-leaderboard-v1",
610
- "verification": "archived_task_flags"
611
- }
612
- },
613
- "generated_at": "2026-09-16T17:18:47.373930+00:00",
614
- "complete": false,
615
- "harness": "DeepAgents 0.7.13",
616
- "model_call_cap": 6,
617
- "task_unknown_count": 6,
618
- "invalid_evidence_count": 1,
619
- "interruption_count": null,
620
- "evidence_class": "amended_comparison"
621
- }
622
- }
623
- },
624
- {
625
- "id": "deepseek-v4-pro",
626
- "name": "DeepSeek V4 Pro",
627
- "arms": {
628
- "baseline": {
629
- "path": "results/deepseek-v4-pro/baseline",
630
- "metrics": {
631
- "perfect": {
632
- "count": 129,
633
- "denominator": 1000,
634
- "unknown": 723,
635
- "failure": 148,
636
- "score": 12.9,
637
- "profile": "local-paper-v1",
638
- "verification": "archived_task_flags"
639
- },
640
- "compilation_success": {
641
- "count": 281,
642
- "denominator": 1000,
643
- "unknown": 0,
644
- "failure": 719,
645
- "score": 28.1,
646
- "profile": "hf-leaderboard-v1",
647
- "verification": "archived_task_flags"
648
- },
649
- "partial_replication": {
650
- "count": 195,
651
- "denominator": 1000,
652
- "unknown": 723,
653
- "failure": 82,
654
- "score": 19.5,
655
- "profile": "hf-leaderboard-v1",
656
- "verification": "archived_task_flags"
657
- },
658
- "coefficient_direction": {
659
- "count": 260,
660
- "denominator": 1000,
661
- "unknown": 723,
662
- "failure": 17,
663
- "score": 26.0,
664
- "profile": "hf-leaderboard-v1",
665
- "verification": "archived_task_flags"
666
- },
667
- "significance_level": {
668
- "count": 236,
669
- "denominator": 1000,
670
- "unknown": 723,
671
- "failure": 41,
672
- "score": 23.6,
673
- "profile": "hf-leaderboard-v1",
674
- "verification": "archived_task_flags"
675
- }
676
- },
677
- "generated_at": "2026-09-16T17:18:47.373930+00:00",
678
- "complete": true,
679
- "harness": "none",
680
- "model_call_cap": 1,
681
- "task_unknown_count": 0,
682
- "invalid_evidence_count": 0,
683
- "interruption_count": null,
684
- "evidence_class": "amended_comparison"
685
- },
686
- "deepagents": {
687
- "path": "results/deepseek-v4-pro/deepagents",
688
- "metrics": {
689
- "perfect": {
690
- "count": 183,
691
- "denominator": 1000,
692
- "unknown": 648,
693
- "failure": 169,
694
- "score": 18.3,
695
- "profile": "local-paper-v1",
696
- "verification": "archived_task_flags"
697
- },
698
- "compilation_success": {
699
- "count": 352,
700
- "denominator": 1000,
701
- "unknown": 0,
702
- "failure": 648,
703
- "score": 35.2,
704
- "profile": "hf-leaderboard-v1",
705
- "verification": "archived_task_flags"
706
- },
707
- "partial_replication": {
708
- "count": 270,
709
- "denominator": 1000,
710
- "unknown": 648,
711
- "failure": 82,
712
- "score": 27.0,
713
- "profile": "hf-leaderboard-v1",
714
- "verification": "archived_task_flags"
715
- },
716
- "coefficient_direction": {
717
- "count": 338,
718
- "denominator": 1000,
719
- "unknown": 648,
720
- "failure": 14,
721
- "score": 33.8,
722
- "profile": "hf-leaderboard-v1",
723
- "verification": "archived_task_flags"
724
- },
725
- "significance_level": {
726
- "count": 308,
727
- "denominator": 1000,
728
- "unknown": 648,
729
- "failure": 44,
730
- "score": 30.8,
731
- "profile": "hf-leaderboard-v1",
732
- "verification": "archived_task_flags"
733
- }
734
- },
735
- "generated_at": "2026-09-16T17:18:47.373930+00:00",
736
- "complete": true,
737
- "harness": "DeepAgents 0.7.13",
738
- "model_call_cap": 6,
739
- "task_unknown_count": 0,
740
- "invalid_evidence_count": 0,
741
- "interruption_count": null,
742
- "evidence_class": "amended_comparison"
743
- }
744
- }
745
- }
746
- ],
747
- "source_sha256": {
748
- "results/gpt-5.6-sol/baseline/summary.json": "6d6ab1aaffe754e08f9656dcb5f51fcd5c772e60660dc34d86cfc21f3c5c03d7",
749
- "results/gpt-5.6-sol/baseline/results.jsonl": "3b12160382043e27531ff5cb3e23bccc9a8f7807e4121a82ffa13df28f65a859",
750
- "results/gpt-5.6-sol/baseline/config.json": "90ec4e0781b57498fef52de31eacf5c8ec9424549643865a7140a1ccdb04c1d4",
751
- "results/gpt-5.6-sol/deepagents/summary.json": "821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5",
752
- "results/gpt-5.6-sol/deepagents/results.jsonl": "378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be",
753
- "results/gpt-5.6-sol/deepagents/config.json": "a3e24175f2395cffbbcd6d9d3c89689bd3a3b22dd4ce61bb78d5da33c6b6cc1b",
754
- "results/claude-opus-4-8/baseline/summary.json": "7b5b06bd9b3d7b47952b1dc424a01fcdf04ddccca0b3f64d04132358d6380cdc",
755
- "results/claude-opus-4-8/baseline/results.jsonl": "b9d6bd51c9eb81a634e78e422c86da16f655d97da87d3f5afa4c9aa0f96809f4",
756
- "results/claude-opus-4-8/baseline/config.json": "8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56",
757
- "results/claude-opus-4-8/deepagents/summary.json": "f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614",
758
- "results/claude-opus-4-8/deepagents/results.jsonl": "3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a",
759
- "results/claude-opus-4-8/deepagents/config.json": "12215dfcba5d2d692268a305364514a03ff7579094b2d4af12f493383b8e0f21",
760
- "results/kimi-k3/baseline/summary.json": "2edd2e96f3b448ceceb6189edb7f23c737b245c00c9820728c69f5237feb34b7",
761
- "results/kimi-k3/baseline/results.jsonl": "0502ae8892dd6c66ee92890a83c678527cff6c3396aa4c13faaadcf41e1857c1",
762
- "results/kimi-k3/baseline/config.json": "7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081",
763
- "results/kimi-k3/deepagents/summary.json": "81c8cabaa58acc65ef6afba792241b70951b666ab90c3f7c25194b61b0d9a228",
764
- "results/kimi-k3/deepagents/results.jsonl": "bdab99f8f1d81885f74326cf9c8c8e57b9bb8a8d616bfbb1b5a2f6bf32b737ca",
765
- "results/kimi-k3/deepagents/config.json": "312c9260c01d5988878e4f2b88da55a16362ee994e8217c63b15ba772da3857c",
766
- "results/gemini-3.1-pro-preview/baseline/summary.json": "a8ba5100a28f1f0aeae6949eb3669c7717c844c4e2eee6fa42ded351aef6b5c7",
767
- "results/gemini-3.1-pro-preview/baseline/results.jsonl": "2f1e2019157d105efee7ba311a49eba99abe4a7b84f0a4b4e0ac174e2d56cdba",
768
- "results/gemini-3.1-pro-preview/baseline/config.json": "ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998",
769
- "results/gemini-3.1-pro-preview/deepagents/summary.json": "bf9ce721c94103f91b0ca576f717f4f5dca7cc9965601658a65fc6d2a15b7045",
770
- "results/gemini-3.1-pro-preview/deepagents/results.jsonl": "c2824f8277b283a6dd60118b6e6a0cb0419e4c8b9193de15a7bcf7aaaad24181",
771
- "results/gemini-3.1-pro-preview/deepagents/config.json": "f09d9d99d26328ae2f54e9173b22f69d68f7d7bb50c3cc5e352c996b98dd57f9",
772
- "results/qwen3.7-max/baseline/summary.json": "e506814d33789673fcd6b8e6869cdcbd52371cd3d0455ce53ae27e1098a52214",
773
- "results/qwen3.7-max/baseline/results.jsonl": "4924e929d67956f9deb6c319605042d09f40b12f7bd2f18baa0464b8f7c9851a",
774
- "results/qwen3.7-max/baseline/config.json": "9f553440e741863e82f32fec11470affe6f0dbffa03c668c647bd63d8f73227f",
775
- "results/qwen3.7-max/deepagents/summary.json": "a5dcb2297ccc31a921e3f4eeb025ed81a764e16556a12068733cfe3bfe0bdbe3",
776
- "results/qwen3.7-max/deepagents/results.jsonl": "e009f8c5c88b05896836330822afbda33af463c1f264f9adb9bbf5b094e61dfa",
777
- "results/qwen3.7-max/deepagents/config.json": "3862d438d381ca3e57b1bf68921ef01a8c38727902152ada2c6e1fcf7b70cc77",
778
- "results/deepseek-v4-pro/baseline/summary.json": "c51e4f1de78eba8f30e4a59276fb7821132cd103f0ae3f9d04c75c98344f2e34",
779
- "results/deepseek-v4-pro/baseline/results.jsonl": "7827218f6f18f39988e8c439f0c93cab611c04d158d8704aa50c9274cd438837",
780
- "results/deepseek-v4-pro/baseline/config.json": "4cef524f37a325029ea4ad325b163f72f9cf725f507f20cd76424d06e71e24ad",
781
- "results/deepseek-v4-pro/deepagents/summary.json": "231e4264b0403e12788136d01af465f86fe9ba3c3ab992510925a7ef4c2a40e9",
782
- "results/deepseek-v4-pro/deepagents/results.jsonl": "a006c8b0b80064295faeb3aa7395a2408f0208677853448816b361bf67e39bb6",
783
- "results/deepseek-v4-pro/deepagents/config.json": "400c28e4008d5241bf50bc71c2674861846d6b6d2e7da21f2a7602a60ecc0ae9"
784
- }
785
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 1,
3
+ "kind": "research_leaderboard",
4
+ "published_date": "2026-09-20",
5
+ "source_repository": "CamoAiLab/InferenceNet-Leaderboard",
6
+ "source_revision": "6ea7c5684664bfe5c6e3d282a5e28f68244acd9e",
7
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
8
+ "task_list": "Selected_1000/1000_new.csv",
9
+ "tasks_per_group": 1000,
10
+ "groups": 13,
11
+ "records": 13000,
12
+ "model_aliases": 6,
13
+ "paired_groups": 12,
14
+ "independent_groups": 1,
15
+ "paired_arms": [
16
+ "baseline",
17
+ "deepagents"
18
+ ],
19
+ "official_parity_verified": false,
20
+ "models": [
21
+ {
22
+ "id": "gpt-5.6-sol",
23
+ "name": "GPT-5.6 Sol",
24
+ "arms": {
25
+ "baseline": {
26
+ "path": "results/gpt-5.6-sol/baseline",
27
+ "metrics": {
28
+ "perfect": {
29
+ "count": 338,
30
+ "denominator": 1000,
31
+ "unknown": 445,
32
+ "failure": 217,
33
+ "score": 33.8,
34
+ "profile": "local-paper-v1",
35
+ "verification": "accepted_summary"
36
+ },
37
+ "compilation_success": {
38
+ "count": 557,
39
+ "denominator": 1000,
40
+ "unknown": 3,
41
+ "failure": 440,
42
+ "score": 55.7,
43
+ "profile": "hf-leaderboard-v1",
44
+ "verification": "archived_task_flags"
45
+ },
46
+ "partial_replication": {
47
+ "count": 430,
48
+ "denominator": 1000,
49
+ "unknown": 444,
50
+ "failure": 126,
51
+ "score": 43.0,
52
+ "profile": "hf-leaderboard-v1",
53
+ "verification": "archived_task_flags"
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 537,
57
+ "denominator": 1000,
58
+ "unknown": 444,
59
+ "failure": 19,
60
+ "score": 53.7,
61
+ "profile": "hf-leaderboard-v1",
62
+ "verification": "archived_task_flags"
63
+ },
64
+ "significance_level": {
65
+ "count": 485,
66
+ "denominator": 1000,
67
+ "unknown": 444,
68
+ "failure": 71,
69
+ "score": 48.5,
70
+ "profile": "hf-leaderboard-v1",
71
+ "verification": "archived_task_flags"
72
+ }
73
+ },
74
+ "generated_at": "2026-09-13T17:49:16.537733+00:00",
75
+ "complete": true,
76
+ "harness": "none",
77
+ "model_call_cap": 1,
78
+ "task_unknown_count": null,
79
+ "invalid_evidence_count": 0,
80
+ "interruption_count": 3,
81
+ "evidence_class": "amended_comparison"
82
+ },
83
+ "deepagents": {
84
+ "path": "results/gpt-5.6-sol/deepagents",
85
+ "metrics": {
86
+ "perfect": {
87
+ "count": 518,
88
+ "denominator": 1000,
89
+ "unknown": 95,
90
+ "failure": 387,
91
+ "score": 51.8,
92
+ "profile": "local-paper-v1",
93
+ "verification": "accepted_summary"
94
+ },
95
+ "compilation_success": {
96
+ "count": 910,
97
+ "denominator": 1000,
98
+ "unknown": 10,
99
+ "failure": 80,
100
+ "score": 91.0,
101
+ "profile": "hf-leaderboard-v1",
102
+ "verification": "archived_task_flags"
103
+ },
104
+ "partial_replication": {
105
+ "count": 699,
106
+ "denominator": 1000,
107
+ "unknown": 92,
108
+ "failure": 209,
109
+ "score": 69.9,
110
+ "profile": "hf-leaderboard-v1",
111
+ "verification": "archived_task_flags"
112
+ },
113
+ "coefficient_direction": {
114
+ "count": 879,
115
+ "denominator": 1000,
116
+ "unknown": 92,
117
+ "failure": 29,
118
+ "score": 87.9,
119
+ "profile": "hf-leaderboard-v1",
120
+ "verification": "archived_task_flags"
121
+ },
122
+ "significance_level": {
123
+ "count": 787,
124
+ "denominator": 1000,
125
+ "unknown": 93,
126
+ "failure": 120,
127
+ "score": 78.7,
128
+ "profile": "hf-leaderboard-v1",
129
+ "verification": "archived_task_flags"
130
+ }
131
+ },
132
+ "generated_at": "2026-09-13T17:49:16.537733+00:00",
133
+ "complete": true,
134
+ "harness": "DeepAgents",
135
+ "model_call_cap": 6,
136
+ "task_unknown_count": null,
137
+ "invalid_evidence_count": 0,
138
+ "interruption_count": 10,
139
+ "evidence_class": "amended_comparison"
140
+ }
141
+ }
142
+ },
143
+ {
144
+ "id": "claude-opus-4-8",
145
+ "name": "Claude Opus 4.8",
146
+ "arms": {
147
+ "baseline": {
148
+ "path": "results/claude-opus-4-8/baseline",
149
+ "metrics": {
150
+ "perfect": {
151
+ "count": 184,
152
+ "denominator": 1000,
153
+ "unknown": 538,
154
+ "failure": 278,
155
+ "score": 18.4,
156
+ "profile": "local-paper-v1",
157
+ "verification": "accepted_summary"
158
+ },
159
+ "compilation_success": {
160
+ "count": 467,
161
+ "denominator": 1000,
162
+ "unknown": 0,
163
+ "failure": 533,
164
+ "score": 46.7,
165
+ "profile": "hf-leaderboard-v1",
166
+ "verification": "archived_task_flags"
167
+ },
168
+ "partial_replication": {
169
+ "count": 305,
170
+ "denominator": 1000,
171
+ "unknown": 537,
172
+ "failure": 158,
173
+ "score": 30.5,
174
+ "profile": "hf-leaderboard-v1",
175
+ "verification": "archived_task_flags"
176
+ },
177
+ "coefficient_direction": {
178
+ "count": 437,
179
+ "denominator": 1000,
180
+ "unknown": 537,
181
+ "failure": 26,
182
+ "score": 43.7,
183
+ "profile": "hf-leaderboard-v1",
184
+ "verification": "archived_task_flags"
185
+ },
186
+ "significance_level": {
187
+ "count": 370,
188
+ "denominator": 1000,
189
+ "unknown": 537,
190
+ "failure": 93,
191
+ "score": 37.0,
192
+ "profile": "hf-leaderboard-v1",
193
+ "verification": "archived_task_flags"
194
+ }
195
+ },
196
+ "generated_at": "2026-09-13T17:50:43.870921+00:00",
197
+ "complete": true,
198
+ "harness": "none",
199
+ "model_call_cap": 1,
200
+ "task_unknown_count": null,
201
+ "invalid_evidence_count": 0,
202
+ "interruption_count": 0,
203
+ "evidence_class": "amended_comparison"
204
+ },
205
+ "deepagents": {
206
+ "path": "results/claude-opus-4-8/deepagents",
207
+ "metrics": {
208
+ "perfect": {
209
+ "count": 392,
210
+ "denominator": 1000,
211
+ "unknown": 174,
212
+ "failure": 434,
213
+ "score": 39.2,
214
+ "profile": "local-paper-v1",
215
+ "verification": "accepted_summary"
216
+ },
217
+ "compilation_success": {
218
+ "count": 834,
219
+ "denominator": 1000,
220
+ "unknown": 4,
221
+ "failure": 162,
222
+ "score": 83.4,
223
+ "profile": "hf-leaderboard-v1",
224
+ "verification": "archived_task_flags"
225
+ },
226
+ "partial_replication": {
227
+ "count": 621,
228
+ "denominator": 1000,
229
+ "unknown": 170,
230
+ "failure": 209,
231
+ "score": 62.1,
232
+ "profile": "hf-leaderboard-v1",
233
+ "verification": "archived_task_flags"
234
+ },
235
+ "coefficient_direction": {
236
+ "count": 793,
237
+ "denominator": 1000,
238
+ "unknown": 170,
239
+ "failure": 37,
240
+ "score": 79.3,
241
+ "profile": "hf-leaderboard-v1",
242
+ "verification": "archived_task_flags"
243
+ },
244
+ "significance_level": {
245
+ "count": 696,
246
+ "denominator": 1000,
247
+ "unknown": 172,
248
+ "failure": 132,
249
+ "score": 69.6,
250
+ "profile": "hf-leaderboard-v1",
251
+ "verification": "archived_task_flags"
252
+ }
253
+ },
254
+ "generated_at": "2026-09-13T17:50:43.870921+00:00",
255
+ "complete": true,
256
+ "harness": "DeepAgents",
257
+ "model_call_cap": 6,
258
+ "task_unknown_count": null,
259
+ "invalid_evidence_count": 0,
260
+ "interruption_count": 4,
261
+ "evidence_class": "amended_comparison"
262
+ }
263
+ }
264
+ },
265
+ {
266
+ "id": "kimi-k3",
267
+ "name": "Kimi K3",
268
+ "arms": {
269
+ "baseline": {
270
+ "path": "results/kimi-k3/baseline",
271
+ "metrics": {
272
+ "perfect": {
273
+ "count": 230,
274
+ "denominator": 1000,
275
+ "unknown": 481,
276
+ "failure": 289,
277
+ "score": 23.0,
278
+ "profile": "local-paper-v1",
279
+ "verification": "archived_task_flags"
280
+ },
281
+ "compilation_success": {
282
+ "count": 526,
283
+ "denominator": 1000,
284
+ "unknown": 5,
285
+ "failure": 469,
286
+ "score": 52.6,
287
+ "profile": "hf-leaderboard-v1",
288
+ "verification": "archived_task_flags"
289
+ },
290
+ "partial_replication": {
291
+ "count": 366,
292
+ "denominator": 1000,
293
+ "unknown": 481,
294
+ "failure": 153,
295
+ "score": 36.6,
296
+ "profile": "hf-leaderboard-v1",
297
+ "verification": "archived_task_flags"
298
+ },
299
+ "coefficient_direction": {
300
+ "count": 494,
301
+ "denominator": 1000,
302
+ "unknown": 481,
303
+ "failure": 25,
304
+ "score": 49.4,
305
+ "profile": "hf-leaderboard-v1",
306
+ "verification": "archived_task_flags"
307
+ },
308
+ "significance_level": {
309
+ "count": 431,
310
+ "denominator": 1000,
311
+ "unknown": 481,
312
+ "failure": 88,
313
+ "score": 43.1,
314
+ "profile": "hf-leaderboard-v1",
315
+ "verification": "archived_task_flags"
316
+ }
317
+ },
318
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
319
+ "complete": false,
320
+ "harness": "none",
321
+ "model_call_cap": 1,
322
+ "task_unknown_count": 5,
323
+ "invalid_evidence_count": 0,
324
+ "interruption_count": null,
325
+ "evidence_class": "amended_comparison"
326
+ },
327
+ "deepagents": {
328
+ "path": "results/kimi-k3/deepagents",
329
+ "metrics": {
330
+ "perfect": {
331
+ "count": 390,
332
+ "denominator": 1000,
333
+ "unknown": 159,
334
+ "failure": 451,
335
+ "score": 39.0,
336
+ "profile": "local-paper-v1",
337
+ "verification": "archived_task_flags"
338
+ },
339
+ "compilation_success": {
340
+ "count": 867,
341
+ "denominator": 1000,
342
+ "unknown": 4,
343
+ "failure": 129,
344
+ "score": 86.7,
345
+ "profile": "hf-leaderboard-v1",
346
+ "verification": "archived_task_flags"
347
+ },
348
+ "partial_replication": {
349
+ "count": 635,
350
+ "denominator": 1000,
351
+ "unknown": 156,
352
+ "failure": 209,
353
+ "score": 63.5,
354
+ "profile": "hf-leaderboard-v1",
355
+ "verification": "archived_task_flags"
356
+ },
357
+ "coefficient_direction": {
358
+ "count": 804,
359
+ "denominator": 1000,
360
+ "unknown": 156,
361
+ "failure": 40,
362
+ "score": 80.4,
363
+ "profile": "hf-leaderboard-v1",
364
+ "verification": "archived_task_flags"
365
+ },
366
+ "significance_level": {
367
+ "count": 709,
368
+ "denominator": 1000,
369
+ "unknown": 157,
370
+ "failure": 134,
371
+ "score": 70.9,
372
+ "profile": "hf-leaderboard-v1",
373
+ "verification": "archived_task_flags"
374
+ }
375
+ },
376
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
377
+ "complete": false,
378
+ "harness": "DeepAgents 0.7.13",
379
+ "model_call_cap": 6,
380
+ "task_unknown_count": 4,
381
+ "invalid_evidence_count": 0,
382
+ "interruption_count": null,
383
+ "evidence_class": "amended_comparison"
384
+ }
385
+ }
386
+ },
387
+ {
388
+ "id": "gemini-3.1-pro-preview",
389
+ "name": "Gemini 3.1 Pro Preview",
390
+ "arms": {
391
+ "baseline": {
392
+ "path": "results/gemini-3.1-pro-preview/baseline",
393
+ "metrics": {
394
+ "perfect": {
395
+ "count": 246,
396
+ "denominator": 1000,
397
+ "unknown": 462,
398
+ "failure": 292,
399
+ "score": 24.6,
400
+ "profile": "local-paper-v1",
401
+ "verification": "archived_task_flags"
402
+ },
403
+ "compilation_success": {
404
+ "count": 543,
405
+ "denominator": 1000,
406
+ "unknown": 0,
407
+ "failure": 457,
408
+ "score": 54.3,
409
+ "profile": "hf-leaderboard-v1",
410
+ "verification": "archived_task_flags"
411
+ },
412
+ "partial_replication": {
413
+ "count": 366,
414
+ "denominator": 1000,
415
+ "unknown": 460,
416
+ "failure": 174,
417
+ "score": 36.6,
418
+ "profile": "hf-leaderboard-v1",
419
+ "verification": "archived_task_flags"
420
+ },
421
+ "coefficient_direction": {
422
+ "count": 506,
423
+ "denominator": 1000,
424
+ "unknown": 460,
425
+ "failure": 34,
426
+ "score": 50.6,
427
+ "profile": "hf-leaderboard-v1",
428
+ "verification": "archived_task_flags"
429
+ },
430
+ "significance_level": {
431
+ "count": 437,
432
+ "denominator": 1000,
433
+ "unknown": 460,
434
+ "failure": 103,
435
+ "score": 43.7,
436
+ "profile": "hf-leaderboard-v1",
437
+ "verification": "archived_task_flags"
438
+ }
439
+ },
440
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
441
+ "complete": true,
442
+ "harness": "none",
443
+ "model_call_cap": 1,
444
+ "task_unknown_count": 0,
445
+ "invalid_evidence_count": 0,
446
+ "interruption_count": null,
447
+ "evidence_class": "amended_comparison"
448
+ },
449
+ "deepagents": {
450
+ "path": "results/gemini-3.1-pro-preview/deepagents",
451
+ "metrics": {
452
+ "perfect": {
453
+ "count": 391,
454
+ "denominator": 1000,
455
+ "unknown": 155,
456
+ "failure": 454,
457
+ "score": 39.1,
458
+ "profile": "local-paper-v1",
459
+ "verification": "archived_task_flags"
460
+ },
461
+ "compilation_success": {
462
+ "count": 850,
463
+ "denominator": 1000,
464
+ "unknown": 0,
465
+ "failure": 150,
466
+ "score": 85.0,
467
+ "profile": "hf-leaderboard-v1",
468
+ "verification": "archived_task_flags"
469
+ },
470
+ "partial_replication": {
471
+ "count": 632,
472
+ "denominator": 1000,
473
+ "unknown": 153,
474
+ "failure": 215,
475
+ "score": 63.2,
476
+ "profile": "hf-leaderboard-v1",
477
+ "verification": "archived_task_flags"
478
+ },
479
+ "coefficient_direction": {
480
+ "count": 811,
481
+ "denominator": 1000,
482
+ "unknown": 153,
483
+ "failure": 36,
484
+ "score": 81.1,
485
+ "profile": "hf-leaderboard-v1",
486
+ "verification": "archived_task_flags"
487
+ },
488
+ "significance_level": {
489
+ "count": 718,
490
+ "denominator": 1000,
491
+ "unknown": 154,
492
+ "failure": 128,
493
+ "score": 71.8,
494
+ "profile": "hf-leaderboard-v1",
495
+ "verification": "archived_task_flags"
496
+ }
497
+ },
498
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
499
+ "complete": true,
500
+ "harness": "DeepAgents 0.7.13",
501
+ "model_call_cap": 6,
502
+ "task_unknown_count": 0,
503
+ "invalid_evidence_count": 0,
504
+ "interruption_count": null,
505
+ "evidence_class": "amended_comparison"
506
+ }
507
+ }
508
+ },
509
+ {
510
+ "id": "qwen3.7-max",
511
+ "name": "Qwen3.7-Max",
512
+ "arms": {
513
+ "baseline": {
514
+ "path": "results/qwen3.7-max/baseline",
515
+ "metrics": {
516
+ "perfect": {
517
+ "count": 198,
518
+ "denominator": 1000,
519
+ "unknown": 519,
520
+ "failure": 283,
521
+ "score": 19.8,
522
+ "profile": "local-paper-v1",
523
+ "verification": "archived_task_flags"
524
+ },
525
+ "compilation_success": {
526
+ "count": 486,
527
+ "denominator": 1000,
528
+ "unknown": 0,
529
+ "failure": 514,
530
+ "score": 48.6,
531
+ "profile": "hf-leaderboard-v1",
532
+ "verification": "archived_task_flags"
533
+ },
534
+ "partial_replication": {
535
+ "count": 315,
536
+ "denominator": 1000,
537
+ "unknown": 518,
538
+ "failure": 167,
539
+ "score": 31.5,
540
+ "profile": "hf-leaderboard-v1",
541
+ "verification": "archived_task_flags"
542
+ },
543
+ "coefficient_direction": {
544
+ "count": 454,
545
+ "denominator": 1000,
546
+ "unknown": 518,
547
+ "failure": 28,
548
+ "score": 45.4,
549
+ "profile": "hf-leaderboard-v1",
550
+ "verification": "archived_task_flags"
551
+ },
552
+ "significance_level": {
553
+ "count": 396,
554
+ "denominator": 1000,
555
+ "unknown": 518,
556
+ "failure": 86,
557
+ "score": 39.6,
558
+ "profile": "hf-leaderboard-v1",
559
+ "verification": "archived_task_flags"
560
+ }
561
+ },
562
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
563
+ "complete": true,
564
+ "harness": "none",
565
+ "model_call_cap": 1,
566
+ "task_unknown_count": 0,
567
+ "invalid_evidence_count": 0,
568
+ "interruption_count": null,
569
+ "evidence_class": "amended_comparison"
570
+ },
571
+ "deepagents": {
572
+ "path": "results/qwen3.7-max/deepagents",
573
+ "metrics": {
574
+ "perfect": {
575
+ "count": 248,
576
+ "denominator": 1000,
577
+ "unknown": 481,
578
+ "failure": 271,
579
+ "score": 24.8,
580
+ "profile": "local-paper-v1",
581
+ "verification": "archived_task_flags"
582
+ },
583
+ "compilation_success": {
584
+ "count": 523,
585
+ "denominator": 1000,
586
+ "unknown": 7,
587
+ "failure": 470,
588
+ "score": 52.3,
589
+ "profile": "hf-leaderboard-v1",
590
+ "verification": "archived_task_flags"
591
+ },
592
+ "partial_replication": {
593
+ "count": 379,
594
+ "denominator": 1000,
595
+ "unknown": 481,
596
+ "failure": 140,
597
+ "score": 37.9,
598
+ "profile": "hf-leaderboard-v1",
599
+ "verification": "archived_task_flags"
600
+ },
601
+ "coefficient_direction": {
602
+ "count": 488,
603
+ "denominator": 1000,
604
+ "unknown": 481,
605
+ "failure": 31,
606
+ "score": 48.8,
607
+ "profile": "hf-leaderboard-v1",
608
+ "verification": "archived_task_flags"
609
+ },
610
+ "significance_level": {
611
+ "count": 433,
612
+ "denominator": 1000,
613
+ "unknown": 481,
614
+ "failure": 86,
615
+ "score": 43.3,
616
+ "profile": "hf-leaderboard-v1",
617
+ "verification": "archived_task_flags"
618
+ }
619
+ },
620
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
621
+ "complete": false,
622
+ "harness": "DeepAgents 0.7.13",
623
+ "model_call_cap": 6,
624
+ "task_unknown_count": 6,
625
+ "invalid_evidence_count": 1,
626
+ "interruption_count": null,
627
+ "evidence_class": "amended_comparison"
628
+ }
629
+ }
630
+ },
631
+ {
632
+ "id": "deepseek-v4-pro",
633
+ "name": "DeepSeek V4 Pro",
634
+ "arms": {
635
+ "baseline": {
636
+ "path": "results/deepseek-v4-pro/baseline",
637
+ "metrics": {
638
+ "perfect": {
639
+ "count": 129,
640
+ "denominator": 1000,
641
+ "unknown": 723,
642
+ "failure": 148,
643
+ "score": 12.9,
644
+ "profile": "local-paper-v1",
645
+ "verification": "archived_task_flags"
646
+ },
647
+ "compilation_success": {
648
+ "count": 281,
649
+ "denominator": 1000,
650
+ "unknown": 0,
651
+ "failure": 719,
652
+ "score": 28.1,
653
+ "profile": "hf-leaderboard-v1",
654
+ "verification": "archived_task_flags"
655
+ },
656
+ "partial_replication": {
657
+ "count": 195,
658
+ "denominator": 1000,
659
+ "unknown": 723,
660
+ "failure": 82,
661
+ "score": 19.5,
662
+ "profile": "hf-leaderboard-v1",
663
+ "verification": "archived_task_flags"
664
+ },
665
+ "coefficient_direction": {
666
+ "count": 260,
667
+ "denominator": 1000,
668
+ "unknown": 723,
669
+ "failure": 17,
670
+ "score": 26.0,
671
+ "profile": "hf-leaderboard-v1",
672
+ "verification": "archived_task_flags"
673
+ },
674
+ "significance_level": {
675
+ "count": 236,
676
+ "denominator": 1000,
677
+ "unknown": 723,
678
+ "failure": 41,
679
+ "score": 23.6,
680
+ "profile": "hf-leaderboard-v1",
681
+ "verification": "archived_task_flags"
682
+ }
683
+ },
684
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
685
+ "complete": true,
686
+ "harness": "none",
687
+ "model_call_cap": 1,
688
+ "task_unknown_count": 0,
689
+ "invalid_evidence_count": 0,
690
+ "interruption_count": null,
691
+ "evidence_class": "amended_comparison"
692
+ },
693
+ "deepagents": {
694
+ "path": "results/deepseek-v4-pro/deepagents",
695
+ "metrics": {
696
+ "perfect": {
697
+ "count": 183,
698
+ "denominator": 1000,
699
+ "unknown": 648,
700
+ "failure": 169,
701
+ "score": 18.3,
702
+ "profile": "local-paper-v1",
703
+ "verification": "archived_task_flags"
704
+ },
705
+ "compilation_success": {
706
+ "count": 352,
707
+ "denominator": 1000,
708
+ "unknown": 0,
709
+ "failure": 648,
710
+ "score": 35.2,
711
+ "profile": "hf-leaderboard-v1",
712
+ "verification": "archived_task_flags"
713
+ },
714
+ "partial_replication": {
715
+ "count": 270,
716
+ "denominator": 1000,
717
+ "unknown": 648,
718
+ "failure": 82,
719
+ "score": 27.0,
720
+ "profile": "hf-leaderboard-v1",
721
+ "verification": "archived_task_flags"
722
+ },
723
+ "coefficient_direction": {
724
+ "count": 338,
725
+ "denominator": 1000,
726
+ "unknown": 648,
727
+ "failure": 14,
728
+ "score": 33.8,
729
+ "profile": "hf-leaderboard-v1",
730
+ "verification": "archived_task_flags"
731
+ },
732
+ "significance_level": {
733
+ "count": 308,
734
+ "denominator": 1000,
735
+ "unknown": 648,
736
+ "failure": 44,
737
+ "score": 30.8,
738
+ "profile": "hf-leaderboard-v1",
739
+ "verification": "archived_task_flags"
740
+ }
741
+ },
742
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
743
+ "complete": true,
744
+ "harness": "DeepAgents 0.7.13",
745
+ "model_call_cap": 6,
746
+ "task_unknown_count": 0,
747
+ "invalid_evidence_count": 0,
748
+ "interruption_count": null,
749
+ "evidence_class": "amended_comparison"
750
+ },
751
+ "dsh": {
752
+ "path": "results/deepseek-v4-pro/dsh",
753
+ "metrics": {
754
+ "perfect": {
755
+ "count": 372,
756
+ "denominator": 1000,
757
+ "unknown": 305,
758
+ "failure": 323,
759
+ "score": 37.2,
760
+ "profile": "local-paper-v1",
761
+ "verification": "archived_task_flags"
762
+ },
763
+ "compilation_success": {
764
+ "count": 699,
765
+ "denominator": 1000,
766
+ "unknown": 30,
767
+ "failure": 271,
768
+ "score": 69.9,
769
+ "profile": "hf-leaderboard-v1",
770
+ "verification": "archived_task_flags"
771
+ },
772
+ "partial_replication": {
773
+ "count": 556,
774
+ "denominator": 1000,
775
+ "unknown": 302,
776
+ "failure": 142,
777
+ "score": 55.6,
778
+ "profile": "hf-leaderboard-v1",
779
+ "verification": "archived_task_flags"
780
+ },
781
+ "coefficient_direction": {
782
+ "count": 683,
783
+ "denominator": 1000,
784
+ "unknown": 302,
785
+ "failure": 15,
786
+ "score": 68.3,
787
+ "profile": "hf-leaderboard-v1",
788
+ "verification": "archived_task_flags"
789
+ },
790
+ "significance_level": {
791
+ "count": 598,
792
+ "denominator": 1000,
793
+ "unknown": 304,
794
+ "failure": 98,
795
+ "score": 59.8,
796
+ "profile": "hf-leaderboard-v1",
797
+ "verification": "archived_task_flags"
798
+ }
799
+ },
800
+ "generated_at": "2026-09-19T18:18:48.070003+00:00",
801
+ "complete": false,
802
+ "harness": "DeepSeek Harness SDK 0.1.5rc1",
803
+ "model_call_cap": 6,
804
+ "task_unknown_count": 30,
805
+ "invalid_evidence_count": 0,
806
+ "interruption_count": null,
807
+ "evidence_class": "technical_recovery_view",
808
+ "display_name": "DeepSeek API [v4-pro alias] + DSH",
809
+ "identity_notice": "Requested API alias: deepseek-v4-pro. Underlying model weights are unverified. This technical recovery result is a separate harness entry, not a paired DeepAgents gain.",
810
+ "paired_comparison": false,
811
+ "model_alias": "deepseek-v4-pro",
812
+ "underlying_model_weights_verified": false,
813
+ "comparison_group": "independent",
814
+ "protocol_url": "results/deepseek-v4-pro/dsh/README.md"
815
+ }
816
+ }
817
+ }
818
+ ],
819
+ "source_sha256": {
820
+ "results/gpt-5.6-sol/baseline/summary.json": "6d6ab1aaffe754e08f9656dcb5f51fcd5c772e60660dc34d86cfc21f3c5c03d7",
821
+ "results/gpt-5.6-sol/baseline/results.jsonl": "3b12160382043e27531ff5cb3e23bccc9a8f7807e4121a82ffa13df28f65a859",
822
+ "results/gpt-5.6-sol/baseline/config.json": "90ec4e0781b57498fef52de31eacf5c8ec9424549643865a7140a1ccdb04c1d4",
823
+ "results/gpt-5.6-sol/deepagents/summary.json": "821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5",
824
+ "results/gpt-5.6-sol/deepagents/results.jsonl": "378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be",
825
+ "results/gpt-5.6-sol/deepagents/config.json": "a3e24175f2395cffbbcd6d9d3c89689bd3a3b22dd4ce61bb78d5da33c6b6cc1b",
826
+ "results/claude-opus-4-8/baseline/summary.json": "7b5b06bd9b3d7b47952b1dc424a01fcdf04ddccca0b3f64d04132358d6380cdc",
827
+ "results/claude-opus-4-8/baseline/results.jsonl": "b9d6bd51c9eb81a634e78e422c86da16f655d97da87d3f5afa4c9aa0f96809f4",
828
+ "results/claude-opus-4-8/baseline/config.json": "8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56",
829
+ "results/claude-opus-4-8/deepagents/summary.json": "f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614",
830
+ "results/claude-opus-4-8/deepagents/results.jsonl": "3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a",
831
+ "results/claude-opus-4-8/deepagents/config.json": "12215dfcba5d2d692268a305364514a03ff7579094b2d4af12f493383b8e0f21",
832
+ "results/kimi-k3/baseline/summary.json": "2edd2e96f3b448ceceb6189edb7f23c737b245c00c9820728c69f5237feb34b7",
833
+ "results/kimi-k3/baseline/results.jsonl": "0502ae8892dd6c66ee92890a83c678527cff6c3396aa4c13faaadcf41e1857c1",
834
+ "results/kimi-k3/baseline/config.json": "7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081",
835
+ "results/kimi-k3/deepagents/summary.json": "81c8cabaa58acc65ef6afba792241b70951b666ab90c3f7c25194b61b0d9a228",
836
+ "results/kimi-k3/deepagents/results.jsonl": "bdab99f8f1d81885f74326cf9c8c8e57b9bb8a8d616bfbb1b5a2f6bf32b737ca",
837
+ "results/kimi-k3/deepagents/config.json": "312c9260c01d5988878e4f2b88da55a16362ee994e8217c63b15ba772da3857c",
838
+ "results/gemini-3.1-pro-preview/baseline/summary.json": "a8ba5100a28f1f0aeae6949eb3669c7717c844c4e2eee6fa42ded351aef6b5c7",
839
+ "results/gemini-3.1-pro-preview/baseline/results.jsonl": "2f1e2019157d105efee7ba311a49eba99abe4a7b84f0a4b4e0ac174e2d56cdba",
840
+ "results/gemini-3.1-pro-preview/baseline/config.json": "ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998",
841
+ "results/gemini-3.1-pro-preview/deepagents/summary.json": "bf9ce721c94103f91b0ca576f717f4f5dca7cc9965601658a65fc6d2a15b7045",
842
+ "results/gemini-3.1-pro-preview/deepagents/results.jsonl": "c2824f8277b283a6dd60118b6e6a0cb0419e4c8b9193de15a7bcf7aaaad24181",
843
+ "results/gemini-3.1-pro-preview/deepagents/config.json": "f09d9d99d26328ae2f54e9173b22f69d68f7d7bb50c3cc5e352c996b98dd57f9",
844
+ "results/qwen3.7-max/baseline/summary.json": "e506814d33789673fcd6b8e6869cdcbd52371cd3d0455ce53ae27e1098a52214",
845
+ "results/qwen3.7-max/baseline/results.jsonl": "4924e929d67956f9deb6c319605042d09f40b12f7bd2f18baa0464b8f7c9851a",
846
+ "results/qwen3.7-max/baseline/config.json": "9f553440e741863e82f32fec11470affe6f0dbffa03c668c647bd63d8f73227f",
847
+ "results/qwen3.7-max/deepagents/summary.json": "a5dcb2297ccc31a921e3f4eeb025ed81a764e16556a12068733cfe3bfe0bdbe3",
848
+ "results/qwen3.7-max/deepagents/results.jsonl": "e009f8c5c88b05896836330822afbda33af463c1f264f9adb9bbf5b094e61dfa",
849
+ "results/qwen3.7-max/deepagents/config.json": "3862d438d381ca3e57b1bf68921ef01a8c38727902152ada2c6e1fcf7b70cc77",
850
+ "results/deepseek-v4-pro/baseline/summary.json": "c51e4f1de78eba8f30e4a59276fb7821132cd103f0ae3f9d04c75c98344f2e34",
851
+ "results/deepseek-v4-pro/baseline/results.jsonl": "7827218f6f18f39988e8c439f0c93cab611c04d158d8704aa50c9274cd438837",
852
+ "results/deepseek-v4-pro/baseline/config.json": "4cef524f37a325029ea4ad325b163f72f9cf725f507f20cd76424d06e71e24ad",
853
+ "results/deepseek-v4-pro/deepagents/summary.json": "231e4264b0403e12788136d01af465f86fe9ba3c3ab992510925a7ef4c2a40e9",
854
+ "results/deepseek-v4-pro/deepagents/results.jsonl": "a006c8b0b80064295faeb3aa7395a2408f0208677853448816b361bf67e39bb6",
855
+ "results/deepseek-v4-pro/deepagents/config.json": "400c28e4008d5241bf50bc71c2674861846d6b6d2e7da21f2a7602a60ecc0ae9",
856
+ "results/deepseek-v4-pro/dsh/summary.json": "62936f76b1ca0de5e761d1ea1f894343695503ce934bf0213dc9fde1ea72a782",
857
+ "results/deepseek-v4-pro/dsh/results.jsonl": "b12427d5f4476351e571e4b29f61e0463540a1701bec51e642753d717dbc1df2",
858
+ "results/deepseek-v4-pro/dsh/config.json": "df322f13761e06dc78cb553c5dac347ce4361599937af4017a902c56a7df9d77"
859
+ }
860
+ }
leaderboard.css CHANGED
@@ -3,3 +3,5 @@
3
  @media(max-width:760px){.page{padding:0 20px}.topbar{min-height:76px}.nav-links{gap:12px}.nav-links a{font-size:11px}.brand{font-size:18px}.brand-mark{display:none}.hero{padding-top:34px}h1{letter-spacing:-1.3px}.intro{font-size:15px}.facts{gap:18px}.facts .update{width:100%;margin-left:0}.protocols{gap:18px}.protocols>div+div{padding-left:18px}.protocols p+p{font-size:12px}.protocol-tag{font-size:12px}h2{font-size:22px}.section-head{align-items:flex-start;flex-direction:column;gap:12px}.toolbar{align-items:flex-start;flex-direction:column}.view-switch{width:100%}.view-switch button{flex:1;font-size:11px;padding:8px}.filters{flex-wrap:wrap;gap:12px}.metric-control{width:100%;min-width:0!important}.filters label:not(.metric-control){flex:1}.search-control{margin-left:0;width:auto}.metric-description{padding-bottom:16px}.method-grid{grid-template-columns:1fr;gap:20px}.method{padding-top:36px}.table-foot{flex-direction:column;gap:5px}.table-foot>span:last-child{text-align:left}.table-wrap td{padding:17px 14px}.method .section-head{gap:14px}footer{flex-direction:column;gap:12px}}
4
  tbody th[scope=row]{background:transparent;color:var(--text);border-bottom-color:#23313b}tbody tr:last-child th{border-bottom:none}
5
  @media(prefers-reduced-motion:no-preference){button,a{transition:background .12s,color .12s}}
 
 
 
3
  @media(max-width:760px){.page{padding:0 20px}.topbar{min-height:76px}.nav-links{gap:12px}.nav-links a{font-size:11px}.brand{font-size:18px}.brand-mark{display:none}.hero{padding-top:34px}h1{letter-spacing:-1.3px}.intro{font-size:15px}.facts{gap:18px}.facts .update{width:100%;margin-left:0}.protocols{gap:18px}.protocols>div+div{padding-left:18px}.protocols p+p{font-size:12px}.protocol-tag{font-size:12px}h2{font-size:22px}.section-head{align-items:flex-start;flex-direction:column;gap:12px}.toolbar{align-items:flex-start;flex-direction:column}.view-switch{width:100%}.view-switch button{flex:1;font-size:11px;padding:8px}.filters{flex-wrap:wrap;gap:12px}.metric-control{width:100%;min-width:0!important}.filters label:not(.metric-control){flex:1}.search-control{margin-left:0;width:auto}.metric-description{padding-bottom:16px}.method-grid{grid-template-columns:1fr;gap:20px}.method{padding-top:36px}.table-foot{flex-direction:column;gap:5px}.table-foot>span:last-child{text-align:left}.table-wrap td{padding:17px 14px}.method .section-head{gap:14px}footer{flex-direction:column;gap:12px}}
4
  tbody th[scope=row]{background:transparent;color:var(--text);border-bottom-color:#23313b}tbody tr:last-child th{border-bottom:none}
5
  @media(prefers-reduced-motion:no-preference){button,a{transition:background .12s,color .12s}}
6
+ .protocols{grid-template-columns:repeat(3,1fr);gap:24px}.protocols>div+div{padding-left:24px}.independent{color:var(--amber)}.independent .fill{background:var(--amber)}.independent-notice{margin-top:18px;padding:16px 20px;border:1px solid #4b422e;border-left:3px solid var(--amber);border-radius:6px;background:#1c1e1b;font-size:12px;color:var(--muted)}.independent-notice strong{color:var(--amber)}.independent-notice p{margin:5px 0 7px;max-width:960px}.independent-notice a{color:var(--amber)}.run-protocol{display:block;font-size:10px;margin-top:5px}.model-name{white-space:normal;max-width:355px}.model-meta{flex-wrap:wrap}.facts{gap:24px}.facts .update{font-size:11px}
7
+ @media(max-width:760px){.protocols{grid-template-columns:1fr;gap:18px}.protocols>div+div{border-left:0;border-top:1px solid var(--line);padding:18px 0 0}.facts{gap:14px 24px}.independent-notice{padding:14px}.view-switch button{flex-basis:100%}.model-cell{min-width:225px}.model-name{max-width:275px}}
leaderboard.csv CHANGED
@@ -1,61 +1,66 @@
1
- model,arm,harness,metric_profile,metric,successes,denominator,unknown,score_percent,official_parity_verified,source_revision
2
- gpt-5.6-sol,baseline,none,local-paper-v1,perfect,338,1000,445,33.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
3
- gpt-5.6-sol,baseline,none,hf-leaderboard-v1,compilation_success,557,1000,3,55.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
4
- gpt-5.6-sol,baseline,none,hf-leaderboard-v1,partial_replication,430,1000,444,43.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
5
- gpt-5.6-sol,baseline,none,hf-leaderboard-v1,coefficient_direction,537,1000,444,53.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
6
- gpt-5.6-sol,baseline,none,hf-leaderboard-v1,significance_level,485,1000,444,48.5,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
7
- gpt-5.6-sol,deepagents,DeepAgents,local-paper-v1,perfect,518,1000,95,51.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
8
- gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,compilation_success,910,1000,10,91.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
9
- gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,partial_replication,699,1000,92,69.9,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
10
- gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,coefficient_direction,879,1000,92,87.9,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
11
- gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,significance_level,787,1000,93,78.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
12
- claude-opus-4-8,baseline,none,local-paper-v1,perfect,184,1000,538,18.4,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
13
- claude-opus-4-8,baseline,none,hf-leaderboard-v1,compilation_success,467,1000,0,46.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
14
- claude-opus-4-8,baseline,none,hf-leaderboard-v1,partial_replication,305,1000,537,30.5,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
15
- claude-opus-4-8,baseline,none,hf-leaderboard-v1,coefficient_direction,437,1000,537,43.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
16
- claude-opus-4-8,baseline,none,hf-leaderboard-v1,significance_level,370,1000,537,37.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
17
- claude-opus-4-8,deepagents,DeepAgents,local-paper-v1,perfect,392,1000,174,39.2,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
18
- claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,compilation_success,834,1000,4,83.4,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
19
- claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,partial_replication,621,1000,170,62.1,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
20
- claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,coefficient_direction,793,1000,170,79.3,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
21
- claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,significance_level,696,1000,172,69.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
22
- kimi-k3,baseline,none,local-paper-v1,perfect,230,1000,481,23.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
23
- kimi-k3,baseline,none,hf-leaderboard-v1,compilation_success,526,1000,5,52.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
24
- kimi-k3,baseline,none,hf-leaderboard-v1,partial_replication,366,1000,481,36.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
25
- kimi-k3,baseline,none,hf-leaderboard-v1,coefficient_direction,494,1000,481,49.4,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
26
- kimi-k3,baseline,none,hf-leaderboard-v1,significance_level,431,1000,481,43.1,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
27
- kimi-k3,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,390,1000,159,39.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
28
- kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,867,1000,4,86.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
29
- kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,635,1000,156,63.5,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
30
- kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,804,1000,156,80.4,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
31
- kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,709,1000,157,70.9,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
32
- gemini-3.1-pro-preview,baseline,none,local-paper-v1,perfect,246,1000,462,24.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
33
- gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,compilation_success,543,1000,0,54.3,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
34
- gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,partial_replication,366,1000,460,36.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
35
- gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,coefficient_direction,506,1000,460,50.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
36
- gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,significance_level,437,1000,460,43.7,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
37
- gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,391,1000,155,39.1,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
38
- gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,850,1000,0,85.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
39
- gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,632,1000,153,63.2,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
40
- gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,811,1000,153,81.1,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
41
- gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,718,1000,154,71.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
42
- qwen3.7-max,baseline,none,local-paper-v1,perfect,198,1000,519,19.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
43
- qwen3.7-max,baseline,none,hf-leaderboard-v1,compilation_success,486,1000,0,48.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
44
- qwen3.7-max,baseline,none,hf-leaderboard-v1,partial_replication,315,1000,518,31.5,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
45
- qwen3.7-max,baseline,none,hf-leaderboard-v1,coefficient_direction,454,1000,518,45.4,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
46
- qwen3.7-max,baseline,none,hf-leaderboard-v1,significance_level,396,1000,518,39.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
47
- qwen3.7-max,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,248,1000,481,24.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
48
- qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,523,1000,7,52.3,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
49
- qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,379,1000,481,37.9,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
50
- qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,488,1000,481,48.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
51
- qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,433,1000,481,43.3,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
52
- deepseek-v4-pro,baseline,none,local-paper-v1,perfect,129,1000,723,12.9,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
53
- deepseek-v4-pro,baseline,none,hf-leaderboard-v1,compilation_success,281,1000,0,28.1,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
54
- deepseek-v4-pro,baseline,none,hf-leaderboard-v1,partial_replication,195,1000,723,19.5,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
55
- deepseek-v4-pro,baseline,none,hf-leaderboard-v1,coefficient_direction,260,1000,723,26.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
56
- deepseek-v4-pro,baseline,none,hf-leaderboard-v1,significance_level,236,1000,723,23.6,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
57
- deepseek-v4-pro,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,183,1000,648,18.3,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
58
- deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,352,1000,0,35.2,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
59
- deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,270,1000,648,27.0,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
60
- deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,338,1000,648,33.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
61
- deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,308,1000,648,30.8,False,731111974cf4d1a3c7a4f848a899f0b0708dd12b
 
 
 
 
 
 
1
+ model,arm,harness,metric_profile,metric,successes,denominator,unknown,score_percent,official_parity_verified,source_revision,display_name,comparison_group,underlying_model_weights_verified
2
+ gpt-5.6-sol,baseline,none,local-paper-v1,perfect,338,1000,445,33.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
3
+ gpt-5.6-sol,baseline,none,hf-leaderboard-v1,compilation_success,557,1000,3,55.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
4
+ gpt-5.6-sol,baseline,none,hf-leaderboard-v1,partial_replication,430,1000,444,43.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
5
+ gpt-5.6-sol,baseline,none,hf-leaderboard-v1,coefficient_direction,537,1000,444,53.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
6
+ gpt-5.6-sol,baseline,none,hf-leaderboard-v1,significance_level,485,1000,444,48.5,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
7
+ gpt-5.6-sol,deepagents,DeepAgents,local-paper-v1,perfect,518,1000,95,51.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
8
+ gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,compilation_success,910,1000,10,91.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
9
+ gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,partial_replication,699,1000,92,69.9,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
10
+ gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,coefficient_direction,879,1000,92,87.9,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
11
+ gpt-5.6-sol,deepagents,DeepAgents,hf-leaderboard-v1,significance_level,787,1000,93,78.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,GPT-5.6 Sol,historical_pair,
12
+ claude-opus-4-8,baseline,none,local-paper-v1,perfect,184,1000,538,18.4,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
13
+ claude-opus-4-8,baseline,none,hf-leaderboard-v1,compilation_success,467,1000,0,46.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
14
+ claude-opus-4-8,baseline,none,hf-leaderboard-v1,partial_replication,305,1000,537,30.5,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
15
+ claude-opus-4-8,baseline,none,hf-leaderboard-v1,coefficient_direction,437,1000,537,43.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
16
+ claude-opus-4-8,baseline,none,hf-leaderboard-v1,significance_level,370,1000,537,37.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
17
+ claude-opus-4-8,deepagents,DeepAgents,local-paper-v1,perfect,392,1000,174,39.2,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
18
+ claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,compilation_success,834,1000,4,83.4,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
19
+ claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,partial_replication,621,1000,170,62.1,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
20
+ claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,coefficient_direction,793,1000,170,79.3,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
21
+ claude-opus-4-8,deepagents,DeepAgents,hf-leaderboard-v1,significance_level,696,1000,172,69.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Claude Opus 4.8,historical_pair,
22
+ kimi-k3,baseline,none,local-paper-v1,perfect,230,1000,481,23.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
23
+ kimi-k3,baseline,none,hf-leaderboard-v1,compilation_success,526,1000,5,52.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
24
+ kimi-k3,baseline,none,hf-leaderboard-v1,partial_replication,366,1000,481,36.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
25
+ kimi-k3,baseline,none,hf-leaderboard-v1,coefficient_direction,494,1000,481,49.4,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
26
+ kimi-k3,baseline,none,hf-leaderboard-v1,significance_level,431,1000,481,43.1,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
27
+ kimi-k3,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,390,1000,159,39.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
28
+ kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,867,1000,4,86.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
29
+ kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,635,1000,156,63.5,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
30
+ kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,804,1000,156,80.4,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
31
+ kimi-k3,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,709,1000,157,70.9,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Kimi K3,historical_pair,
32
+ gemini-3.1-pro-preview,baseline,none,local-paper-v1,perfect,246,1000,462,24.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
33
+ gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,compilation_success,543,1000,0,54.3,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
34
+ gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,partial_replication,366,1000,460,36.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
35
+ gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,coefficient_direction,506,1000,460,50.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
36
+ gemini-3.1-pro-preview,baseline,none,hf-leaderboard-v1,significance_level,437,1000,460,43.7,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
37
+ gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,391,1000,155,39.1,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
38
+ gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,850,1000,0,85.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
39
+ gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,632,1000,153,63.2,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
40
+ gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,811,1000,153,81.1,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
41
+ gemini-3.1-pro-preview,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,718,1000,154,71.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Gemini 3.1 Pro Preview,historical_pair,
42
+ qwen3.7-max,baseline,none,local-paper-v1,perfect,198,1000,519,19.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
43
+ qwen3.7-max,baseline,none,hf-leaderboard-v1,compilation_success,486,1000,0,48.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
44
+ qwen3.7-max,baseline,none,hf-leaderboard-v1,partial_replication,315,1000,518,31.5,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
45
+ qwen3.7-max,baseline,none,hf-leaderboard-v1,coefficient_direction,454,1000,518,45.4,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
46
+ qwen3.7-max,baseline,none,hf-leaderboard-v1,significance_level,396,1000,518,39.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
47
+ qwen3.7-max,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,248,1000,481,24.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
48
+ qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,523,1000,7,52.3,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
49
+ qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,379,1000,481,37.9,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
50
+ qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,488,1000,481,48.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
51
+ qwen3.7-max,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,433,1000,481,43.3,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,Qwen3.7-Max,historical_pair,
52
+ deepseek-v4-pro,baseline,none,local-paper-v1,perfect,129,1000,723,12.9,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
53
+ deepseek-v4-pro,baseline,none,hf-leaderboard-v1,compilation_success,281,1000,0,28.1,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
54
+ deepseek-v4-pro,baseline,none,hf-leaderboard-v1,partial_replication,195,1000,723,19.5,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
55
+ deepseek-v4-pro,baseline,none,hf-leaderboard-v1,coefficient_direction,260,1000,723,26.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
56
+ deepseek-v4-pro,baseline,none,hf-leaderboard-v1,significance_level,236,1000,723,23.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
57
+ deepseek-v4-pro,deepagents,DeepAgents 0.7.13,local-paper-v1,perfect,183,1000,648,18.3,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
58
+ deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,compilation_success,352,1000,0,35.2,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
59
+ deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,partial_replication,270,1000,648,27.0,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
60
+ deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,coefficient_direction,338,1000,648,33.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
61
+ deepseek-v4-pro,deepagents,DeepAgents 0.7.13,hf-leaderboard-v1,significance_level,308,1000,648,30.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek V4 Pro,historical_pair,
62
+ deepseek-v4-pro,dsh,DeepSeek Harness SDK 0.1.5rc1,local-paper-v1,perfect,372,1000,305,37.2,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek API [v4-pro alias] + DSH,independent,False
63
+ deepseek-v4-pro,dsh,DeepSeek Harness SDK 0.1.5rc1,hf-leaderboard-v1,compilation_success,699,1000,30,69.9,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek API [v4-pro alias] + DSH,independent,False
64
+ deepseek-v4-pro,dsh,DeepSeek Harness SDK 0.1.5rc1,hf-leaderboard-v1,partial_replication,556,1000,302,55.6,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek API [v4-pro alias] + DSH,independent,False
65
+ deepseek-v4-pro,dsh,DeepSeek Harness SDK 0.1.5rc1,hf-leaderboard-v1,coefficient_direction,683,1000,302,68.3,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek API [v4-pro alias] + DSH,independent,False
66
+ deepseek-v4-pro,dsh,DeepSeek Harness SDK 0.1.5rc1,hf-leaderboard-v1,significance_level,598,1000,304,59.8,False,6ea7c5684664bfe5c6e3d282a5e28f68244acd9e,DeepSeek API [v4-pro alias] + DSH,independent,False
leaderboard.js CHANGED
@@ -3,7 +3,7 @@ const $ = id => document.getElementById(id);
3
  const metricKeys = ['perfect', 'partial_replication', 'compilation_success', 'coefficient_direction', 'significance_level'];
4
  const en = Object.fromEntries([...document.querySelectorAll('[data-i18n]')].map(el => [el.dataset.i18n, el.innerHTML]));
5
  Object.assign(en, {
6
- model: 'Model', rank: '#', score: 'Score', evidenceCol: 'Evidence', pp: 'pp',
7
  passed: 'successes', metricUnknown: 'metric unknown', review: 'Evidence gaps',
8
  config: 'Config', recordsLink: 'Results', allSealed: 'Archived',
9
  taskUnknown: 'task status unknown', invalid: 'invalid evidence', interruptions: 'interruptions',
@@ -11,6 +11,8 @@ Object.assign(en, {
11
  groupLocal: 'Local paper profile', groupHF: 'HF-style profile · parity unverified',
12
  loading: 'Loading archived results…', searchPlaceholder: 'Search models…',
13
  matched: '{n} of {total} models', localProfile: 'local-paper-v1', hfProfile: 'hf-leaderboard-v1 · parity unverified',
 
 
14
  perfectDesc: 'Coefficient and standard-error relative errors ≤ 1%, and p-value absolute error ≤ 0.01, all at once.',
15
  partial_replicationDesc: 'Coefficient relative error ≤ 5%. No standard-error or p-value condition.',
16
  compilation_successDesc: 'Verified code execution completed successfully; a valid prediction JSON is not required.',
@@ -19,35 +21,41 @@ Object.assign(en, {
19
  });
20
  const zh = {
21
  dataset: '数据集 ↗', archive: '结果存档 ↗', edition: '研究榜单 · 2026 年 9 月',
22
- title: '同样的模型,<br><span>两种执行方式。</span>',
23
- intro: '在真实计量经济学复现任务上,对比单次 Agent 与加入 DeepAgents Harness 后的表现。',
24
- models: '个模型', groups: '组实验', tasks: '题 / 组', updated: '发布于 2026 年 9 月 17 日',
25
- single: '单次 Agent', harness: 'Agent + Harness',
 
 
 
 
 
26
  singleDesc: '一次模型调用生成代码,然后执行;不进行多轮 Agent 循环。',
27
  harnessDesc: 'DeepAgents 进行规划、工具调用与代码修订;最多 6 次模型调用,最后执行。',
28
  resultsLabel: '01 / 评测结果', rankings: 'Agent 与 Harness 对照榜单', research: '研究结果 · 阶段版',
29
  scope: '基于已存档结果的本地评分,尚未验证与官方评分器完全一致。两组预算及不同批次协议存在差异,提升值用于描述观察结果,不代表等成本对比。',
30
- compare: '双组对照', download: '下载 CSV ↓', metric: '评分指标',
31
  perfect: '完整复现率 · 本地口径', partial_replication: '部分复现率 · HF 口径',
32
  compilation_success: '执行成功率 · HF 口径', coefficient_direction: '系数方向正确率 · HF 口径',
33
  significance_level: '显著性水平正确率 · HF 口径',
34
- sort: '排序方式', harnessScore: 'Harness 组成绩', singleScore: '单次组成绩', lift: '提升', search: '查找模型',
35
  error: '结果加载失败,请重试或打开结果存档。', retry: '重试', empty: '没有匹配的模型。',
36
  denominator: '所有指标固定以 1,000 题为分母,未知和无效记录不计为成功。',
37
  methodLabel: '02 / 如何解读', methodTitle: '每个成绩都有据可查。', historical: '查看历史榜单 ↗',
38
  fixedTasks: '固定的题目集合', fixedDesc: '所有组使用相同的 Selected_1000 题号和数据集版本。成绩 = 成功数 ÷ 1,000,失败和未知记录始终保留在分母中。',
39
  separateMetrics: '两套评分口径', profilesDesc: '完整复现率采用本地论文口径。其余四项 HF 指标由本地实现复现网页定义,尚未验证与官方评分器一致。',
40
- budgetTitle: '结合预算看提升', budgetDesc: '单次组使用 1 次模型调用;DeepAgents 最多使用 6 次模型调用和 4 次试运行工具调用。历史恢复、时限和运行协议有差异,不能作为等成本的因果估计。',
41
  evidence: '证据状态与指标定义',
42
  unknownDesc: '“指标未知”表示该指标无法判断,可能源于执行失败或输出缺失;它与“任务最终状态未知”不同。两者均不剔除,也不计为成功。',
43
- evidenceCaption: '各组已接收存档中的状态', model: '模型',
44
- verification: '四项 HF 指标已与逐题存档标记核对。Sol 和 Opus 的完整复现率来自已接收汇总;其余四个模型还可逐题核对本地评分标记。本次网页更新没有重新运行或重新评分实验。',
45
  archiveNote: '各组原始存档及当时的发布元数据保持不变,本页是基于存档的新展示。本版不包含 GPT-5.5 和尚未存档的实验。',
46
  footer: '衡量计量经济学复现能力。', source: '来源快照 ↗', json: '数据 JSON ↗', project: '项目页 ↗',
47
  rank: '#', score: '成绩', evidenceCol: '证据', pp: '百分点', passed: '题成功', metricUnknown: '指标未知', review: '存在证据缺口',
48
  config: '配置', recordsLink: '逐题结果', allSealed: '已存档', taskUnknown: '任务状态未知', invalid: '无效证据', interruptions: '中断记录',
49
  historicalStatus: '历史存档,未导出任务未知总数', groupLocal: '本地论文口径', groupHF: 'HF 网页口径 · 一致性待验证',
50
  loading: '正在加载存档结果…', searchPlaceholder: '搜索模型…', matched: '显示 {n} / {total} 个模型',
 
51
  localProfile: 'local-paper-v1 · 本地口径', hfProfile: 'hf-leaderboard-v1 · 一致性待验证',
52
  perfectDesc: '系数与标准误的相对误差均 ≤ 1%,同时 p 值绝对误差 ≤ 0.01。',
53
  partial_replicationDesc: '系数相对误差 ≤ 5%,不要求标准误或 p 值达标。',
@@ -58,7 +66,7 @@ const zh = {
58
  let language = 'en';
59
  try { language = localStorage.getItem('inferencenet-language') === 'zh' ? 'zh' : 'en'; } catch (_) {}
60
  let data = null;
61
- let view = 'compare';
62
  let sortBy = 'deepagents';
63
  const t = key => (language === 'zh' ? zh[key] : en[key]) ?? en[key] ?? key;
64
  const esc = value => String(value).replace(/[&<>"']/g, char => ({'&':'&amp;', '<':'&lt;', '>':'&gt;', '"':'&quot;', "'":'&#39;'}[char]));
@@ -79,23 +87,40 @@ function applyLanguage() {
79
  }
80
 
81
  function validate(candidate) {
82
- if (candidate.schema_version !== 1 || candidate.models?.length !== 6 || candidate.groups !== 12 || candidate.records !== 12000 || candidate.tasks_per_group !== 1000 || candidate.official_parity_verified !== false || !/^[a-f0-9]{40}$/.test(candidate.source_revision)) throw new Error('Invalid leaderboard snapshot');
83
  const ids = new Set();
 
84
  for (const model of candidate.models) {
85
  if (!/^[a-z0-9.-]+$/.test(model.id) || ids.has(model.id) || typeof model.name !== 'string') throw new Error('Invalid model');
86
  ids.add(model.id);
87
- for (const armId of ['baseline', 'deepagents']) {
 
 
88
  const arm = model.arms?.[armId];
89
  if (arm?.path !== `results/${model.id}/${armId}` || typeof arm.complete !== 'boolean') throw new Error('Invalid arm');
 
 
90
  for (const key of metricKeys) {
91
  const metric = arm.metrics?.[key];
92
  if (!metric || metric.denominator !== 1000 || !Number.isInteger(metric.count) || !Number.isInteger(metric.unknown) || metric.count < 0 || metric.unknown < 0 || metric.count + metric.unknown > 1000 || !Number.isFinite(metric.score) || Math.abs(metric.score - metric.count / 10) > 1e-8 || metric.profile !== (key === 'perfect' ? 'local-paper-v1' : 'hf-leaderboard-v1')) throw new Error('Invalid metric');
93
  }
94
  }
95
  }
 
96
  return candidate;
97
  }
98
 
 
 
 
 
 
 
 
 
 
 
 
99
  function scoreCell(arm, key, style) {
100
  const m = arm.metrics[key];
101
  return `<td class="score-cell ${style}"><span class="score">${m.score.toFixed(1)}<small>%</small></span><div class="track" aria-hidden="true"><div class="fill" style="width:${m.score}%"></div></div><span class="counts">${m.count} / 1,000 ${esc(t('passed'))}</span><span class="unknown">${esc(t('metricUnknown'))}: ${m.unknown}</span></td>`;
@@ -112,38 +137,43 @@ function render() {
112
  const key = $('metric').value;
113
  $('metric-description').innerHTML = `<span class="profile">${esc(t(key === 'perfect' ? 'localProfile' : 'hfProfile'))}</span>${esc(t(key + 'Desc'))}`;
114
  if (!data) return;
115
- const search = $('search').value.toLowerCase().trim();
116
  const count = (m, arm) => m.arms[arm].metrics[key].count;
117
- const value = m => view !== 'compare' ? count(m, view) : sortBy === 'gain' ? count(m, 'deepagents') - count(m, 'baseline') : count(m, sortBy);
118
- const models = data.models.filter(m => `${m.name} ${m.id}`.toLowerCase().includes(search)).sort((a, b) => value(b) - value(a) || a.name.localeCompare(b.name));
119
  $('sort-control').hidden = view !== 'compare';
120
- $('table-wrap').hidden = models.length === 0;
121
- $('empty').hidden = models.length !== 0;
122
- $('result-count').textContent = t('matched').replace('{n}', models.length).replace('{total}', data.models.length);
123
  $('caption').textContent = `${t(view === 'compare' ? 'compare' : view === 'baseline' ? 'single' : 'harness')} — ${t(key)}`;
124
  const th = (text, cls = '') => `<th scope="col" class="${cls}">${esc(text)}</th>`;
125
- $('table-head').innerHTML = '<tr>' + th(t('rank'), 'rank') + th(t('model')) + (view === 'compare' ? th(t('single'), 'baseline') + th(t('harness'), 'harness') + th(t('lift')) : th(t('score'), view === 'baseline' ? 'baseline' : 'harness')) + th(t('evidenceCol')) + '</tr>';
126
  let lastValue = null, rank = 0;
127
- $('table-body').innerHTML = models.map((model, index) => {
128
- const currentValue = value(model);
129
  if (currentValue !== lastValue) rank = index + 1;
130
  lastValue = currentValue;
131
  const arms = model.arms;
132
- const gap = view === 'compare' ? !arms.baseline.complete || !arms.deepagents.complete : !arms[view].complete;
133
  const badge = gap ? `<span class="tag">${esc(t('review'))}</span>` : '';
134
- const identity = `<td class="rank">${String(rank).padStart(2, '0')}</td><th scope="row" class="model-cell"><div class="model-name">${esc(model.name)}</div><div class="model-meta">1,000 ${language === 'zh' ? '题 / 组' : 'tasks / group'} ${badge}</div></th>`;
 
 
135
  let scores = '', sources = '';
136
  if (view === 'compare') {
137
  const gain = (count(model, 'deepagents') - count(model, 'baseline')) / 10;
138
  scores = scoreCell(arms.baseline, key, 'baseline') + scoreCell(arms.deepagents, key, 'harness') + `<td class="gain ${gain < 0 ? 'negative' : ''}">${gain > 0 ? '+' : ''}${gain.toFixed(1)}<small>${esc(t('pp'))}</small></td>`;
139
  sources = external(sourceURL(arms.baseline.path, true), 'A') + external(sourceURL(arms.deepagents.path, true), 'B');
140
  } else {
141
- scores = scoreCell(arms[view], key, view === 'baseline' ? 'baseline' : 'harness');
142
- sources = external(sourceURL(arms[view].path + '/results.jsonl'), t('recordsLink')) + external(sourceURL(arms[view].path + '/config.json'), t('config'));
 
143
  }
144
- return `<tr data-model="${esc(model.id)}">${identity}${scores}<td class="source-cell">${sources}</td></tr>`;
145
  }).join('');
146
- $('evidence-body').innerHTML = data.models.map(m => `<tr><th scope="row">${esc(m.name)}</th>${evidenceState(m.arms.baseline)}${evidenceState(m.arms.deepagents)}</tr>`).join('');
 
 
 
147
  }
148
 
149
  async function loadData() {
@@ -155,7 +185,8 @@ async function loadData() {
155
  $('empty').hidden = true;
156
  $('result-count').textContent = '';
157
  $('evidence-body').replaceChildren();
158
- for (const id of ['model-count', 'group-count', 'task-count']) $(id).textContent = '—';
 
159
  try {
160
  const response = await fetch('./leaderboard-data.json', {cache: 'no-cache'});
161
  if (!response.ok) throw new Error(`HTTP ${response.status}`);
@@ -163,6 +194,7 @@ async function loadData() {
163
  $('model-count').textContent = data.models.length;
164
  $('group-count').textContent = data.groups;
165
  $('task-count').textContent = data.tasks_per_group.toLocaleString('en-US');
 
166
  $('revision-link').href = sourceURL('results', true);
167
  render();
168
  } catch (error) {
 
3
  const metricKeys = ['perfect', 'partial_replication', 'compilation_success', 'coefficient_direction', 'significance_level'];
4
  const en = Object.fromEntries([...document.querySelectorAll('[data-i18n]')].map(el => [el.dataset.i18n, el.innerHTML]));
5
  Object.assign(en, {
6
+ model: 'Model / API alias', rank: '#', score: 'Score', evidenceCol: 'Evidence', pp: 'pp',
7
  passed: 'successes', metricUnknown: 'metric unknown', review: 'Evidence gaps',
8
  config: 'Config', recordsLink: 'Results', allSealed: 'Archived',
9
  taskUnknown: 'task status unknown', invalid: 'invalid evidence', interruptions: 'interruptions',
 
11
  groupLocal: 'Local paper profile', groupHF: 'HF-style profile · parity unverified',
12
  loading: 'Loading archived results…', searchPlaceholder: 'Search models…',
13
  matched: '{n} of {total} models', localProfile: 'local-paper-v1', hfProfile: 'hf-leaderboard-v1 · parity unverified',
14
+ matchedRuns: '{n} of {total} runs', protocol: 'Protocol', independentRun: 'Independent · alias unverified',
15
+ published: 'Published {date}',
16
  perfectDesc: 'Coefficient and standard-error relative errors ≤ 1%, and p-value absolute error ≤ 0.01, all at once.',
17
  partial_replicationDesc: 'Coefficient relative error ≤ 5%. No standard-error or p-value condition.',
18
  compilation_successDesc: 'Verified code execution completed successfully; a valid prediction JSON is not required.',
 
21
  });
22
  const zh = {
23
  dataset: '数据集 ↗', archive: '结果存档 ↗', edition: '研究榜单 · 2026 年 9 月',
24
+ title: '模型与 Harness,<br><span>复现真实研究。</span>',
25
+ intro: '查看单次 Agent、DeepAgents 和独立 DeepSeek Harness 实验在真实计量经济学复现任务上的表现。',
26
+ models: '种请求模型别名', groups: '组实验', tasks: '题 / 组', records: '条任务记录', published: '发布于 {date}',
27
+ single: '单次 Agent', harness: '全部 Harness 实验', deepagents: 'DeepAgents',
28
+ dsh: 'DeepSeek Harness · 独立实验',
29
+ dshDesc: '使用官方 DSH SDK,请求 v4-pro API 别名。按固定规则纳入技术补测,底层模型权重未经核实。',
30
+ dshNoticeTitle: 'DSH 作为独立实验展示。',
31
+ dshNotice: '请求使用 deepseek-v4-pro 别名,底层模型权重未经核实。固定的 1,000 题结果包含 87 次技术补测替换(其中 27 题调整了输入 schema 提示),仍有 30 题状态未知;不纳入 DeepAgents 配对提升计算。',
32
+ dshProtocol: '查看运行协议与证据 ↗',
33
  singleDesc: '一次模型调用生成代码,然后执行;不进行多轮 Agent 循环。',
34
  harnessDesc: 'DeepAgents 进行规划、工具调用与代码修订;最多 6 次模型调用,最后执行。',
35
  resultsLabel: '01 / 评测结果', rankings: 'Agent 与 Harness 对照榜单', research: '研究结果 · 阶段版',
36
  scope: '基于已存档结果的本地评分,尚未验证与官方评分器完全一致。两组预算及不同批次协议存在差异,提升值用于描述观察结果,不代表等成本对比。',
37
+ compare: 'Baseline / DeepAgents 配对', download: '下载 CSV ↓', metric: '评分指标',
38
  perfect: '完整复现率 · 本地口径', partial_replication: '部分复现率 · HF 口径',
39
  compilation_success: '执行成功率 · HF 口径', coefficient_direction: '系数方向正确率 · HF 口径',
40
  significance_level: '显著性水平正确率 · HF 口径',
41
+ sort: '排序方式', harnessScore: 'DeepAgents 组成绩', singleScore: '单次组成绩', lift: '提升', search: '查找模型',
42
  error: '结果加载失败,请重试或打开结果存档。', retry: '重试', empty: '没有匹配的模型。',
43
  denominator: '所有指标固定以 1,000 题为分母,未知和无效记录不计为成功。',
44
  methodLabel: '02 / 如何解读', methodTitle: '每个成绩都有据可查。', historical: '查看历史榜单 ↗',
45
  fixedTasks: '固定的题目集合', fixedDesc: '所有组使用相同的 Selected_1000 题号和数据集版本。成绩 = 成功数 ÷ 1,000,失败和未知记录始终保留在分母中。',
46
  separateMetrics: '两套评分口径', profilesDesc: '完整复现率采用本地论文口径。其余四项 HF 指标由本地实现复现网页定义,尚未验证与官方评分器一致。',
47
+ budgetTitle: '结合预算看提升', budgetDesc: '单次组使用 1 次模型调用;DeepAgents 和 DSH 每次尝试最多 6 次模型调用、4 次试运行工具调用。DSH 额外进行了 87 次技术补测,完整环境和推理设定是否一致尚未核实;不计算 DSH 配对提升,也不能据此判断等成本的因果效果。',
48
  evidence: '证据状态与指标定义',
49
  unknownDesc: '“指标未知”表示该指标无法判断,可能源于执行失败或输出缺失;它与“任务最终状态未知”不同。两者均不剔除,也不计为成功。',
50
+ evidenceCaption: '各组已接收存档中的状态', model: '模型 / API 别名', protocol: '执行方式',
51
+ verification: '四项 HF 指标已与逐题存档标记核对。Sol 和 Opus 的完整复现率来自已接收汇总;其余实验(含 DSH)均可逐题核对本地评分标记。本次网页更新没有重新运行或重新评分实验。',
52
  archiveNote: '各组原始存档及当时的发布元数据保持不变,本页是基于存档的新展示。本版不包含 GPT-5.5 和尚未存档的实验。',
53
  footer: '衡量计量经济学复现能力。', source: '来源快照 ↗', json: '数据 JSON ↗', project: '项目页 ↗',
54
  rank: '#', score: '成绩', evidenceCol: '证据', pp: '百分点', passed: '题成功', metricUnknown: '指标未知', review: '存在证据缺口',
55
  config: '配置', recordsLink: '逐题结果', allSealed: '已存档', taskUnknown: '任务状态未知', invalid: '无效证据', interruptions: '中断记录',
56
  historicalStatus: '历史存档,未导出任务未知总数', groupLocal: '本地论文口径', groupHF: 'HF 网页口径 · 一致性待验证',
57
  loading: '正在加载存档结果…', searchPlaceholder: '搜索模型…', matched: '显示 {n} / {total} 个模型',
58
+ matchedRuns: '显示 {n} / {total} 组实验', independentRun: '独立实验 · 别名权重未核实',
59
  localProfile: 'local-paper-v1 · 本地口径', hfProfile: 'hf-leaderboard-v1 · 一致性待验证',
60
  perfectDesc: '系数与标准误的相对误差均 ≤ 1%,同时 p 值绝对误差 ≤ 0.01。',
61
  partial_replicationDesc: '系数相对误差 ≤ 5%,不要求标准误或 p 值达标。',
 
66
  let language = 'en';
67
  try { language = localStorage.getItem('inferencenet-language') === 'zh' ? 'zh' : 'en'; } catch (_) {}
68
  let data = null;
69
+ let view = 'harness';
70
  let sortBy = 'deepagents';
71
  const t = key => (language === 'zh' ? zh[key] : en[key]) ?? en[key] ?? key;
72
  const esc = value => String(value).replace(/[&<>"']/g, char => ({'&':'&amp;', '<':'&lt;', '>':'&gt;', '"':'&quot;', "'":'&#39;'}[char]));
 
87
  }
88
 
89
  function validate(candidate) {
90
+ if (candidate.schema_version !== 1 || candidate.models?.length !== 6 || ![12, 13].includes(candidate.groups) || candidate.records !== candidate.groups * 1000 || candidate.tasks_per_group !== 1000 || candidate.official_parity_verified !== false || !/^[a-f0-9]{40}$/.test(candidate.source_revision) || !/^\d{4}-\d{2}-\d{2}$/.test(candidate.published_date)) throw new Error('Invalid leaderboard snapshot');
91
  const ids = new Set();
92
+ let groups = 0;
93
  for (const model of candidate.models) {
94
  if (!/^[a-z0-9.-]+$/.test(model.id) || ids.has(model.id) || typeof model.name !== 'string') throw new Error('Invalid model');
95
  ids.add(model.id);
96
+ if (!model.arms?.baseline || !model.arms?.deepagents) throw new Error('Missing paired arm');
97
+ for (const armId of Object.keys(model.arms)) {
98
+ if (!['baseline', 'deepagents'].includes(armId) && !(model.id === 'deepseek-v4-pro' && armId === 'dsh')) throw new Error('Unexpected arm');
99
  const arm = model.arms?.[armId];
100
  if (arm?.path !== `results/${model.id}/${armId}` || typeof arm.complete !== 'boolean') throw new Error('Invalid arm');
101
+ if (armId === 'dsh' && (arm.paired_comparison !== false || arm.comparison_group !== 'independent' || arm.underlying_model_weights_verified !== false || arm.model_alias !== 'deepseek-v4-pro' || typeof arm.display_name !== 'string' || arm.evidence_class !== 'technical_recovery_view' || arm.protocol_url !== arm.path + '/README.md')) throw new Error('Invalid independent result');
102
+ groups += 1;
103
  for (const key of metricKeys) {
104
  const metric = arm.metrics?.[key];
105
  if (!metric || metric.denominator !== 1000 || !Number.isInteger(metric.count) || !Number.isInteger(metric.unknown) || metric.count < 0 || metric.unknown < 0 || metric.count + metric.unknown > 1000 || !Number.isFinite(metric.score) || Math.abs(metric.score - metric.count / 10) > 1e-8 || metric.profile !== (key === 'perfect' ? 'local-paper-v1' : 'hf-leaderboard-v1')) throw new Error('Invalid metric');
106
  }
107
  }
108
  }
109
+ if (groups !== candidate.groups) throw new Error('Inconsistent group count');
110
  return candidate;
111
  }
112
 
113
+ // Pairing is intentionally limited to the two historical arms. Independent
114
+ // harness entries have their own rank and can never enter a paired lift.
115
+ function rankedEntries(snapshot, selectedView, key, selectedSort, query) {
116
+ const entries = selectedView === 'harness'
117
+ ? snapshot.models.flatMap(model => Object.entries(model.arms).filter(([id]) => id !== 'baseline').map(([armId, arm]) => ({model, armId, arm, name: arm.display_name || model.name, value: arm.metrics[key].count})))
118
+ : snapshot.models.map(model => ({model, armId: selectedView === 'compare' ? null : 'baseline', arm: model.arms.baseline, name: model.name,
119
+ value: selectedView === 'baseline' ? model.arms.baseline.metrics[key].count : selectedSort === 'gain' ? model.arms.deepagents.metrics[key].count - model.arms.baseline.metrics[key].count : model.arms[selectedSort].metrics[key].count}));
120
+ const search = query.toLowerCase().trim();
121
+ return {total: entries.length, entries: entries.filter(entry => `${entry.name} ${entry.model.id} ${entry.armId || ''}`.toLowerCase().includes(search)).sort((a, b) => b.value - a.value || a.name.localeCompare(b.name))};
122
+ }
123
+
124
  function scoreCell(arm, key, style) {
125
  const m = arm.metrics[key];
126
  return `<td class="score-cell ${style}"><span class="score">${m.score.toFixed(1)}<small>%</small></span><div class="track" aria-hidden="true"><div class="fill" style="width:${m.score}%"></div></div><span class="counts">${m.count} / 1,000 ${esc(t('passed'))}</span><span class="unknown">${esc(t('metricUnknown'))}: ${m.unknown}</span></td>`;
 
137
  const key = $('metric').value;
138
  $('metric-description').innerHTML = `<span class="profile">${esc(t(key === 'perfect' ? 'localProfile' : 'hfProfile'))}</span>${esc(t(key + 'Desc'))}`;
139
  if (!data) return;
 
140
  const count = (m, arm) => m.arms[arm].metrics[key].count;
141
+ const {total, entries} = rankedEntries(data, view, key, sortBy, $('search').value);
142
+ $('published-date').textContent = t('published').replace('{date}', data.published_date);
143
  $('sort-control').hidden = view !== 'compare';
144
+ $('table-wrap').hidden = entries.length === 0;
145
+ $('empty').hidden = entries.length !== 0;
146
+ $('result-count').textContent = t(view === 'compare' ? 'matched' : 'matchedRuns').replace('{n}', entries.length).replace('{total}', total);
147
  $('caption').textContent = `${t(view === 'compare' ? 'compare' : view === 'baseline' ? 'single' : 'harness')} — ${t(key)}`;
148
  const th = (text, cls = '') => `<th scope="col" class="${cls}">${esc(text)}</th>`;
149
+ $('table-head').innerHTML = '<tr>' + th(t('rank'), 'rank') + th(t('model')) + (view === 'compare' ? th(t('single'), 'baseline') + th(t('deepagents'), 'harness') + th(t('lift')) : th(t('score'), view === 'baseline' ? 'baseline' : 'harness')) + th(t('evidenceCol')) + '</tr>';
150
  let lastValue = null, rank = 0;
151
+ $('table-body').innerHTML = entries.map((entry, index) => {
152
+ const {model, arm, armId, name, value: currentValue} = entry;
153
  if (currentValue !== lastValue) rank = index + 1;
154
  lastValue = currentValue;
155
  const arms = model.arms;
156
+ const gap = view === 'compare' ? !arms.baseline.complete || !arms.deepagents.complete : !arm.complete;
157
  const badge = gap ? `<span class="tag">${esc(t('review'))}</span>` : '';
158
+ const independent = armId === 'dsh';
159
+ const protocol = view === 'harness' ? `<span class="run-protocol ${independent ? 'independent' : 'harness'}">${esc(independent ? t('independentRun') : 'DeepAgents')}</span>` : '';
160
+ const identity = `<td class="rank">${String(rank).padStart(2, '0')}</td><th scope="row" class="model-cell"><div class="model-name">${esc(name)}</div>${protocol}<div class="model-meta">1,000 ${language === 'zh' ? '题 / 组' : 'tasks / group'} ${badge}</div></th>`;
161
  let scores = '', sources = '';
162
  if (view === 'compare') {
163
  const gain = (count(model, 'deepagents') - count(model, 'baseline')) / 10;
164
  scores = scoreCell(arms.baseline, key, 'baseline') + scoreCell(arms.deepagents, key, 'harness') + `<td class="gain ${gain < 0 ? 'negative' : ''}">${gain > 0 ? '+' : ''}${gain.toFixed(1)}<small>${esc(t('pp'))}</small></td>`;
165
  sources = external(sourceURL(arms.baseline.path, true), 'A') + external(sourceURL(arms.deepagents.path, true), 'B');
166
  } else {
167
+ scores = scoreCell(arm, key, view === 'baseline' ? 'baseline' : independent ? 'independent' : 'harness');
168
+ sources = external(sourceURL(arm.path + '/results.jsonl'), t('recordsLink')) + external(sourceURL(arm.path + '/config.json'), t('config'));
169
+ if (independent) sources += external(sourceURL(arm.protocol_url), t('protocol'));
170
  }
171
+ return `<tr data-model="${esc(model.id)}" data-arm="${esc(armId || 'paired')}">${identity}${scores}<td class="source-cell">${sources}</td></tr>`;
172
  }).join('');
173
+ $('evidence-body').innerHTML = data.models.flatMap(m => Object.entries(m.arms).map(([id, arm]) => `<tr><th scope="row">${esc(arm.display_name || m.name)}</th><td>${esc(id === 'baseline' ? t('single') : id === 'dsh' ? t('dsh') : 'DeepAgents')}</td>${evidenceState(arm)}</tr>`)).join('');
174
+ const dsh = data.models.find(m => m.id === 'deepseek-v4-pro')?.arms.dsh;
175
+ $('independent-notice').hidden = !dsh;
176
+ if (dsh) $('dsh-protocol').href = sourceURL(dsh.protocol_url);
177
  }
178
 
179
  async function loadData() {
 
185
  $('empty').hidden = true;
186
  $('result-count').textContent = '';
187
  $('evidence-body').replaceChildren();
188
+ $('independent-notice').hidden = true;
189
+ for (const id of ['model-count', 'group-count', 'task-count', 'record-count']) $(id).textContent = '—';
190
  try {
191
  const response = await fetch('./leaderboard-data.json', {cache: 'no-cache'});
192
  if (!response.ok) throw new Error(`HTTP ${response.status}`);
 
194
  $('model-count').textContent = data.models.length;
195
  $('group-count').textContent = data.groups;
196
  $('task-count').textContent = data.tasks_per_group.toLocaleString('en-US');
197
+ $('record-count').textContent = data.records.toLocaleString('en-US');
198
  $('revision-link').href = sourceURL('results', true);
199
  render();
200
  } catch (error) {
results/README.md CHANGED
@@ -1,3 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # InferenceNet 实验结果归档
2
 
3
  目录按 `results/<model>/<baseline|deepagents>/` 排列,不增加日期层。实验批次和时间记录在各目录的 config.json;更新历史由 Space 的 Git commit 保留。
 
1
+ # Current archive addition: DeepSeek API + DeepSeek Harness
2
+
3
+ The archive now includes a separate [DSH group](deepseek-v4-pro/dsh/README.md)
4
+ alongside all twelve existing baseline / DeepAgents groups: 13 groups and 13,000
5
+ task records. The six existing baseline / DeepAgents comparisons are unchanged.
6
+
7
+ This new group requested `deepseek-v4-pro` through the DeepSeek API; the actual
8
+ served weights are unconfirmed because the provider's routing announcement and
9
+ model table conflict. It is not a verified same-model, equal-cost comparison
10
+ against the historical DeepSeek DeepAgents group.
11
+
12
+ The DSH archive fixes 1,000 selected task results after 87 preselected technical
13
+ supplements, retains 30 uncertain tasks, and accounts for all 1,087 attempts.
14
+ Full replication is 37.2%; the four separate HF-style scores are 69.9%, 55.6%,
15
+ 68.3% and 59.8%. See its README and configuration for recovery, input-profile
16
+ amendments, unknown usage and source identities. No new model calls or rescoring
17
+ were performed for publication.
18
+
19
+ The current display is defined by the root `leaderboard-data.json` and root
20
+ README. Publication flags and page references in the historical archive text
21
+ below describe those earlier uploads, not necessarily the current display.
22
+
23
+ ## Preserved historical archive notes
24
+
25
  # InferenceNet 实验结果归档
26
 
27
  目录按 `results/<model>/<baseline|deepagents>/` 排列,不增加日期层。实验批次和时间记录在各目录的 config.json;更新历史由 Space 的 Git commit 保留。
results/deepseek-v4-pro/dsh/README.md ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DeepSeek API (deepseek-v4-pro alias) + DSH
2
+
3
+ This is a **technical-recovery research view**, not an equal-cost controlled estimate of a harness effect. The API alias is `deepseek-v4-pro`; **underlying model weights are not verified**. Conflicting official routing/model descriptions are linked in `config.json`.
4
+
5
+ | Metric | Successes / 1000 | Score | Unknown |
6
+ |---|---:|---:|---:|
7
+ | Execution success | 699 | 69.9% | 30 |
8
+ | Coefficient-only partial replication | 556 | 55.6% | 302 |
9
+ | Coefficient direction | 683 | 68.3% | 302 |
10
+ | Significance category | 598 | 59.8% | 304 |
11
+ | Full replication — local-paper-v1 | 372 | 37.2% | 305 |
12
+
13
+ The selected 1000 tasks contain **699 scored, 271 failed, and 30 uncertain** outcomes. Every rate retains all 1000 tasks. Four HF-style metrics and five local-paper-v1 metrics have separate profiles; unknowns are never dropped. Local full replication requires coefficient and standard-error relative errors <= 1%, and p-value absolute error <= 0.01. Official scorer parity remains unverified.
14
+
15
+ ## Attempt selection and amendments
16
+
17
+ The original 1000 attempts remain unchanged. Eligibility for **one extra attempt for 87 tasks** was frozen before their supplementary outcomes: 60 infrastructure-interrupted tasks and 27 pre-model input-metadata failures. The final view selects the second attempt for all 87, regardless of outcome, and the original attempt for the other 913. No best-of selection or fallback to a better result is used.
18
+
19
+ The 27 schema tasks explicitly change `raw-role-schema` to `raw-role`, allowing the existing Python tool to inspect input headers. No tools or per-attempt budgets are added, but exact prompt parity is not claimed. Controller transport, cancellation and writer-drain repairs are retained in source hashes; the original historical attempts are not relabelled as if those repairs had already applied.
20
+
21
+ Each attempt permits 6 model calls, 4 Python tool calls, 32768 cumulative output tokens and 1800 seconds; execution timeout is 90 seconds. Execution uses bwrap and a frozen rootfs, not Docker. The historical baseline/DeepAgents configurations and complete environment/prompt parity are not fully established. This DSH entry does not replace either historical group.
22
+
23
+ ## All-attempt accounting
24
+
25
+ All **1087 attempts** are charged: **5944 model calls, 3999 tool calls**. There are **38 attempts with unknown usage**, so full input/output/total token sums and monetary cost remain unknown. The **60580246-token subtotal** covers only the 1049 attempts with complete known usage and is not a complete invoice. Task wall-time sums are cumulative worker time, not elapsed campaign duration.
26
+
27
+ Task 0907 retains its original sealed record plus a reviewed late-timeout receipt; effective usage is unknown. `all-attempts.jsonl` preserves original recorded budgets alongside effective accounting rather than rewriting history.
28
+
29
+ ## Files and evidence boundary
30
+
31
+ - `results.jsonl` / `results.csv`: the fixed final selection, five local flags plus four HF-style flags, result/score SHA256, source commit, status and usage.
32
+ - `summary.json` / `config.json`: metrics, fixed task IDs, runtime/package pins, per-attempt protocol and disclosures.
33
+ - `first-pass-results.jsonl` / `first-pass-summary.json`: unchanged first-pass four-metric view, independently labelled.
34
+ - `all-attempts.jsonl` / `accounting.json`: all original and supplementary costs, including unknown usage.
35
+ - `provenance.json` / `export-verification.json` / `manifest.sha256`: source bindings and export checks.
36
+ - `diagnostics.json`: failure categories and unknown IDs, without raw error messages.
37
+
38
+ `resolved=true` means sealed, not necessarily successful or fully known. A later-stage pending/failed status may follow an earlier failure. The selected-results export includes 699 verified prediction triplets and explicitly marks 301 records without a selected final prediction. Prediction fields in the separate first-pass export are omitted (`prediction_export_status=public_export_omission`); this does **not** imply that scored first-pass attempts lacked predictions. Hidden references, inputs, raw prompts/responses, credentials and machine paths are not included. The export copies verified flags and does not call a model, run a program or rescore an attempt.
results/deepseek-v4-pro/dsh/accounting.json ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "accounting_note": "Known subtotals are not complete token totals or invoices.",
3
+ "all_attempts": {
4
+ "attempt_count": 1087,
5
+ "known_usage_subtotal": {
6
+ "input_tokens": 42956609,
7
+ "output_tokens": 17623637,
8
+ "total_tokens": 60580246
9
+ },
10
+ "known_usage_task_count": 1049,
11
+ "model_calls": 5944,
12
+ "task_wall_seconds_sum": 273576.778,
13
+ "tool_calls": 3999,
14
+ "unknown_usage_task_count": 38,
15
+ "usage": {
16
+ "input_tokens": null,
17
+ "output_tokens": null,
18
+ "total_tokens": null
19
+ }
20
+ },
21
+ "infrastructure": {
22
+ "attempt_count": 60,
23
+ "known_usage_subtotal": {
24
+ "input_tokens": 2161547,
25
+ "output_tokens": 804095,
26
+ "total_tokens": 2965642
27
+ },
28
+ "known_usage_task_count": 49,
29
+ "model_calls": 310,
30
+ "task_wall_seconds_sum": 28521.539,
31
+ "tool_calls": 207,
32
+ "unknown_usage_task_count": 11,
33
+ "usage": {
34
+ "input_tokens": null,
35
+ "output_tokens": null,
36
+ "total_tokens": null
37
+ }
38
+ },
39
+ "monetary_cost": null,
40
+ "primary": {
41
+ "attempt_count": 1000,
42
+ "known_usage_subtotal": {
43
+ "input_tokens": 39691223,
44
+ "output_tokens": 16146935,
45
+ "total_tokens": 55838158
46
+ },
47
+ "known_usage_task_count": 974,
48
+ "model_calls": 5476,
49
+ "task_wall_seconds_sum": 236238.034,
50
+ "tool_calls": 3687,
51
+ "unknown_usage_task_count": 26,
52
+ "usage": {
53
+ "input_tokens": null,
54
+ "output_tokens": null,
55
+ "total_tokens": null
56
+ }
57
+ },
58
+ "schema": {
59
+ "attempt_count": 27,
60
+ "known_usage_subtotal": {
61
+ "input_tokens": 1103839,
62
+ "output_tokens": 672607,
63
+ "total_tokens": 1776446
64
+ },
65
+ "known_usage_task_count": 26,
66
+ "model_calls": 158,
67
+ "task_wall_seconds_sum": 8817.205,
68
+ "tool_calls": 105,
69
+ "unknown_usage_task_count": 1,
70
+ "usage": {
71
+ "input_tokens": null,
72
+ "output_tokens": null,
73
+ "total_tokens": null
74
+ }
75
+ }
76
+ }
results/deepseek-v4-pro/dsh/all-attempts.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/dsh/config.json ADDED
@@ -0,0 +1,1205 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "arm": "dsh",
3
+ "campaign": {
4
+ "execution_workers": 6,
5
+ "max_attempts": 1,
6
+ "workers": 12
7
+ },
8
+ "comparison": {
9
+ "historical_parity_verified": false,
10
+ "unknown_historical_settings": [
11
+ "complete_system_prompt_and_tool_schema",
12
+ "provider_endpoint_and_deployment_version",
13
+ "effective_thinking_policy",
14
+ "full_rootfs_and_dependency_lock",
15
+ "controller_source_and_retry_semantics"
16
+ ]
17
+ },
18
+ "dataset": {
19
+ "expected_tasks": 1000,
20
+ "prior_results_imported": false,
21
+ "repo_id": "CamoAiLab/InferenceNet",
22
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
23
+ "task_directory": "Selected_1000",
24
+ "task_ids": [
25
+ 1,
26
+ 2,
27
+ 3,
28
+ 4,
29
+ 5,
30
+ 6,
31
+ 7,
32
+ 8,
33
+ 9,
34
+ 10,
35
+ 11,
36
+ 12,
37
+ 13,
38
+ 14,
39
+ 15,
40
+ 16,
41
+ 17,
42
+ 18,
43
+ 19,
44
+ 20,
45
+ 21,
46
+ 22,
47
+ 23,
48
+ 24,
49
+ 25,
50
+ 26,
51
+ 27,
52
+ 28,
53
+ 29,
54
+ 30,
55
+ 31,
56
+ 32,
57
+ 33,
58
+ 34,
59
+ 35,
60
+ 36,
61
+ 37,
62
+ 38,
63
+ 39,
64
+ 40,
65
+ 41,
66
+ 42,
67
+ 43,
68
+ 44,
69
+ 45,
70
+ 46,
71
+ 47,
72
+ 48,
73
+ 49,
74
+ 50,
75
+ 51,
76
+ 52,
77
+ 53,
78
+ 54,
79
+ 55,
80
+ 56,
81
+ 57,
82
+ 58,
83
+ 59,
84
+ 60,
85
+ 61,
86
+ 62,
87
+ 63,
88
+ 64,
89
+ 65,
90
+ 66,
91
+ 67,
92
+ 68,
93
+ 69,
94
+ 70,
95
+ 71,
96
+ 72,
97
+ 73,
98
+ 74,
99
+ 75,
100
+ 76,
101
+ 77,
102
+ 78,
103
+ 79,
104
+ 80,
105
+ 81,
106
+ 82,
107
+ 83,
108
+ 84,
109
+ 85,
110
+ 86,
111
+ 87,
112
+ 88,
113
+ 89,
114
+ 90,
115
+ 91,
116
+ 92,
117
+ 93,
118
+ 94,
119
+ 95,
120
+ 96,
121
+ 97,
122
+ 98,
123
+ 100,
124
+ 101,
125
+ 102,
126
+ 103,
127
+ 104,
128
+ 105,
129
+ 106,
130
+ 107,
131
+ 108,
132
+ 109,
133
+ 110,
134
+ 111,
135
+ 112,
136
+ 113,
137
+ 114,
138
+ 115,
139
+ 116,
140
+ 117,
141
+ 118,
142
+ 119,
143
+ 120,
144
+ 121,
145
+ 122,
146
+ 123,
147
+ 124,
148
+ 125,
149
+ 126,
150
+ 127,
151
+ 128,
152
+ 129,
153
+ 130,
154
+ 131,
155
+ 132,
156
+ 133,
157
+ 134,
158
+ 135,
159
+ 136,
160
+ 137,
161
+ 138,
162
+ 139,
163
+ 140,
164
+ 141,
165
+ 142,
166
+ 143,
167
+ 144,
168
+ 145,
169
+ 146,
170
+ 147,
171
+ 148,
172
+ 149,
173
+ 150,
174
+ 151,
175
+ 152,
176
+ 153,
177
+ 154,
178
+ 155,
179
+ 156,
180
+ 157,
181
+ 158,
182
+ 159,
183
+ 160,
184
+ 161,
185
+ 162,
186
+ 163,
187
+ 164,
188
+ 165,
189
+ 166,
190
+ 167,
191
+ 168,
192
+ 169,
193
+ 170,
194
+ 171,
195
+ 172,
196
+ 173,
197
+ 174,
198
+ 175,
199
+ 176,
200
+ 177,
201
+ 178,
202
+ 179,
203
+ 180,
204
+ 181,
205
+ 182,
206
+ 183,
207
+ 184,
208
+ 185,
209
+ 186,
210
+ 187,
211
+ 188,
212
+ 189,
213
+ 190,
214
+ 191,
215
+ 192,
216
+ 193,
217
+ 194,
218
+ 195,
219
+ 196,
220
+ 197,
221
+ 198,
222
+ 199,
223
+ 200,
224
+ 201,
225
+ 202,
226
+ 203,
227
+ 204,
228
+ 205,
229
+ 206,
230
+ 207,
231
+ 208,
232
+ 209,
233
+ 210,
234
+ 211,
235
+ 212,
236
+ 213,
237
+ 214,
238
+ 215,
239
+ 216,
240
+ 217,
241
+ 218,
242
+ 219,
243
+ 220,
244
+ 221,
245
+ 222,
246
+ 223,
247
+ 224,
248
+ 225,
249
+ 226,
250
+ 227,
251
+ 228,
252
+ 229,
253
+ 230,
254
+ 231,
255
+ 232,
256
+ 233,
257
+ 234,
258
+ 235,
259
+ 236,
260
+ 237,
261
+ 238,
262
+ 239,
263
+ 240,
264
+ 241,
265
+ 242,
266
+ 248,
267
+ 249,
268
+ 250,
269
+ 251,
270
+ 252,
271
+ 253,
272
+ 254,
273
+ 255,
274
+ 256,
275
+ 257,
276
+ 258,
277
+ 259,
278
+ 261,
279
+ 262,
280
+ 263,
281
+ 264,
282
+ 265,
283
+ 266,
284
+ 290,
285
+ 291,
286
+ 292,
287
+ 293,
288
+ 294,
289
+ 295,
290
+ 302,
291
+ 303,
292
+ 304,
293
+ 305,
294
+ 306,
295
+ 307,
296
+ 308,
297
+ 309,
298
+ 310,
299
+ 311,
300
+ 312,
301
+ 313,
302
+ 314,
303
+ 315,
304
+ 316,
305
+ 317,
306
+ 318,
307
+ 319,
308
+ 320,
309
+ 321,
310
+ 322,
311
+ 323,
312
+ 324,
313
+ 325,
314
+ 326,
315
+ 327,
316
+ 328,
317
+ 329,
318
+ 330,
319
+ 331,
320
+ 332,
321
+ 333,
322
+ 334,
323
+ 335,
324
+ 336,
325
+ 337,
326
+ 338,
327
+ 339,
328
+ 340,
329
+ 341,
330
+ 342,
331
+ 343,
332
+ 344,
333
+ 345,
334
+ 346,
335
+ 347,
336
+ 348,
337
+ 349,
338
+ 350,
339
+ 351,
340
+ 352,
341
+ 353,
342
+ 354,
343
+ 355,
344
+ 356,
345
+ 357,
346
+ 358,
347
+ 359,
348
+ 360,
349
+ 361,
350
+ 362,
351
+ 363,
352
+ 364,
353
+ 365,
354
+ 366,
355
+ 367,
356
+ 368,
357
+ 369,
358
+ 370,
359
+ 371,
360
+ 372,
361
+ 373,
362
+ 374,
363
+ 375,
364
+ 376,
365
+ 377,
366
+ 378,
367
+ 379,
368
+ 380,
369
+ 381,
370
+ 382,
371
+ 383,
372
+ 384,
373
+ 385,
374
+ 386,
375
+ 387,
376
+ 388,
377
+ 389,
378
+ 390,
379
+ 391,
380
+ 392,
381
+ 393,
382
+ 394,
383
+ 395,
384
+ 396,
385
+ 397,
386
+ 398,
387
+ 399,
388
+ 400,
389
+ 401,
390
+ 402,
391
+ 403,
392
+ 404,
393
+ 405,
394
+ 406,
395
+ 407,
396
+ 408,
397
+ 409,
398
+ 410,
399
+ 411,
400
+ 412,
401
+ 413,
402
+ 414,
403
+ 415,
404
+ 416,
405
+ 417,
406
+ 418,
407
+ 419,
408
+ 420,
409
+ 421,
410
+ 422,
411
+ 423,
412
+ 424,
413
+ 425,
414
+ 426,
415
+ 427,
416
+ 428,
417
+ 429,
418
+ 430,
419
+ 431,
420
+ 432,
421
+ 433,
422
+ 434,
423
+ 435,
424
+ 436,
425
+ 437,
426
+ 438,
427
+ 439,
428
+ 440,
429
+ 441,
430
+ 442,
431
+ 443,
432
+ 444,
433
+ 445,
434
+ 446,
435
+ 447,
436
+ 448,
437
+ 449,
438
+ 450,
439
+ 451,
440
+ 452,
441
+ 453,
442
+ 454,
443
+ 455,
444
+ 456,
445
+ 457,
446
+ 458,
447
+ 459,
448
+ 460,
449
+ 461,
450
+ 462,
451
+ 463,
452
+ 464,
453
+ 465,
454
+ 466,
455
+ 467,
456
+ 468,
457
+ 469,
458
+ 470,
459
+ 471,
460
+ 472,
461
+ 473,
462
+ 474,
463
+ 475,
464
+ 476,
465
+ 477,
466
+ 478,
467
+ 479,
468
+ 480,
469
+ 481,
470
+ 482,
471
+ 483,
472
+ 484,
473
+ 485,
474
+ 486,
475
+ 487,
476
+ 488,
477
+ 489,
478
+ 490,
479
+ 491,
480
+ 492,
481
+ 493,
482
+ 494,
483
+ 495,
484
+ 496,
485
+ 497,
486
+ 498,
487
+ 499,
488
+ 500,
489
+ 501,
490
+ 502,
491
+ 503,
492
+ 504,
493
+ 505,
494
+ 506,
495
+ 507,
496
+ 508,
497
+ 509,
498
+ 510,
499
+ 511,
500
+ 512,
501
+ 513,
502
+ 514,
503
+ 515,
504
+ 516,
505
+ 517,
506
+ 518,
507
+ 519,
508
+ 520,
509
+ 521,
510
+ 522,
511
+ 523,
512
+ 524,
513
+ 525,
514
+ 526,
515
+ 527,
516
+ 528,
517
+ 529,
518
+ 530,
519
+ 531,
520
+ 532,
521
+ 533,
522
+ 534,
523
+ 535,
524
+ 536,
525
+ 537,
526
+ 538,
527
+ 539,
528
+ 540,
529
+ 541,
530
+ 542,
531
+ 543,
532
+ 544,
533
+ 545,
534
+ 546,
535
+ 547,
536
+ 548,
537
+ 549,
538
+ 550,
539
+ 551,
540
+ 552,
541
+ 553,
542
+ 554,
543
+ 555,
544
+ 556,
545
+ 557,
546
+ 558,
547
+ 559,
548
+ 560,
549
+ 561,
550
+ 562,
551
+ 563,
552
+ 565,
553
+ 566,
554
+ 567,
555
+ 568,
556
+ 570,
557
+ 571,
558
+ 572,
559
+ 573,
560
+ 574,
561
+ 575,
562
+ 576,
563
+ 577,
564
+ 578,
565
+ 579,
566
+ 580,
567
+ 581,
568
+ 582,
569
+ 583,
570
+ 584,
571
+ 585,
572
+ 586,
573
+ 587,
574
+ 588,
575
+ 589,
576
+ 590,
577
+ 591,
578
+ 592,
579
+ 593,
580
+ 594,
581
+ 595,
582
+ 596,
583
+ 597,
584
+ 598,
585
+ 601,
586
+ 602,
587
+ 603,
588
+ 604,
589
+ 605,
590
+ 606,
591
+ 607,
592
+ 608,
593
+ 609,
594
+ 610,
595
+ 611,
596
+ 612,
597
+ 613,
598
+ 614,
599
+ 615,
600
+ 616,
601
+ 617,
602
+ 618,
603
+ 619,
604
+ 620,
605
+ 621,
606
+ 622,
607
+ 623,
608
+ 624,
609
+ 664,
610
+ 665,
611
+ 666,
612
+ 667,
613
+ 668,
614
+ 669,
615
+ 670,
616
+ 671,
617
+ 672,
618
+ 673,
619
+ 674,
620
+ 675,
621
+ 676,
622
+ 677,
623
+ 678,
624
+ 679,
625
+ 680,
626
+ 681,
627
+ 682,
628
+ 683,
629
+ 684,
630
+ 685,
631
+ 686,
632
+ 687,
633
+ 688,
634
+ 689,
635
+ 690,
636
+ 691,
637
+ 692,
638
+ 693,
639
+ 694,
640
+ 695,
641
+ 696,
642
+ 697,
643
+ 698,
644
+ 699,
645
+ 700,
646
+ 701,
647
+ 702,
648
+ 703,
649
+ 704,
650
+ 705,
651
+ 706,
652
+ 707,
653
+ 708,
654
+ 709,
655
+ 710,
656
+ 711,
657
+ 712,
658
+ 713,
659
+ 714,
660
+ 715,
661
+ 716,
662
+ 717,
663
+ 718,
664
+ 719,
665
+ 720,
666
+ 721,
667
+ 722,
668
+ 723,
669
+ 724,
670
+ 725,
671
+ 726,
672
+ 727,
673
+ 728,
674
+ 729,
675
+ 730,
676
+ 731,
677
+ 732,
678
+ 733,
679
+ 734,
680
+ 735,
681
+ 736,
682
+ 737,
683
+ 739,
684
+ 740,
685
+ 741,
686
+ 742,
687
+ 743,
688
+ 744,
689
+ 745,
690
+ 746,
691
+ 747,
692
+ 748,
693
+ 749,
694
+ 750,
695
+ 751,
696
+ 752,
697
+ 753,
698
+ 754,
699
+ 755,
700
+ 756,
701
+ 757,
702
+ 758,
703
+ 759,
704
+ 760,
705
+ 761,
706
+ 762,
707
+ 763,
708
+ 764,
709
+ 765,
710
+ 766,
711
+ 767,
712
+ 768,
713
+ 772,
714
+ 773,
715
+ 774,
716
+ 775,
717
+ 776,
718
+ 777,
719
+ 778,
720
+ 779,
721
+ 780,
722
+ 781,
723
+ 782,
724
+ 783,
725
+ 784,
726
+ 785,
727
+ 786,
728
+ 787,
729
+ 788,
730
+ 789,
731
+ 790,
732
+ 791,
733
+ 793,
734
+ 794,
735
+ 795,
736
+ 796,
737
+ 797,
738
+ 798,
739
+ 799,
740
+ 800,
741
+ 801,
742
+ 802,
743
+ 803,
744
+ 804,
745
+ 805,
746
+ 806,
747
+ 807,
748
+ 808,
749
+ 825,
750
+ 826,
751
+ 827,
752
+ 828,
753
+ 829,
754
+ 830,
755
+ 831,
756
+ 832,
757
+ 833,
758
+ 834,
759
+ 835,
760
+ 836,
761
+ 837,
762
+ 838,
763
+ 839,
764
+ 840,
765
+ 841,
766
+ 842,
767
+ 843,
768
+ 844,
769
+ 845,
770
+ 846,
771
+ 847,
772
+ 848,
773
+ 849,
774
+ 850,
775
+ 851,
776
+ 852,
777
+ 853,
778
+ 854,
779
+ 855,
780
+ 856,
781
+ 857,
782
+ 858,
783
+ 859,
784
+ 860,
785
+ 861,
786
+ 862,
787
+ 863,
788
+ 864,
789
+ 865,
790
+ 866,
791
+ 867,
792
+ 868,
793
+ 869,
794
+ 870,
795
+ 871,
796
+ 872,
797
+ 873,
798
+ 874,
799
+ 875,
800
+ 876,
801
+ 877,
802
+ 878,
803
+ 879,
804
+ 880,
805
+ 881,
806
+ 882,
807
+ 883,
808
+ 884,
809
+ 885,
810
+ 886,
811
+ 887,
812
+ 888,
813
+ 889,
814
+ 890,
815
+ 891,
816
+ 892,
817
+ 893,
818
+ 894,
819
+ 895,
820
+ 896,
821
+ 897,
822
+ 898,
823
+ 899,
824
+ 900,
825
+ 901,
826
+ 902,
827
+ 903,
828
+ 906,
829
+ 907,
830
+ 913,
831
+ 914,
832
+ 915,
833
+ 917,
834
+ 920,
835
+ 921,
836
+ 922,
837
+ 923,
838
+ 924,
839
+ 925,
840
+ 926,
841
+ 927,
842
+ 928,
843
+ 929,
844
+ 930,
845
+ 931,
846
+ 932,
847
+ 933,
848
+ 934,
849
+ 935,
850
+ 936,
851
+ 937,
852
+ 938,
853
+ 939,
854
+ 940,
855
+ 941,
856
+ 942,
857
+ 943,
858
+ 944,
859
+ 945,
860
+ 946,
861
+ 947,
862
+ 948,
863
+ 949,
864
+ 950,
865
+ 951,
866
+ 952,
867
+ 953,
868
+ 954,
869
+ 955,
870
+ 956,
871
+ 957,
872
+ 958,
873
+ 959,
874
+ 960,
875
+ 961,
876
+ 962,
877
+ 963,
878
+ 964,
879
+ 965,
880
+ 966,
881
+ 967,
882
+ 968,
883
+ 969,
884
+ 970,
885
+ 971,
886
+ 972,
887
+ 973,
888
+ 974,
889
+ 975,
890
+ 976,
891
+ 977,
892
+ 978,
893
+ 979,
894
+ 980,
895
+ 981,
896
+ 982,
897
+ 983,
898
+ 984,
899
+ 985,
900
+ 1001,
901
+ 1002,
902
+ 1003,
903
+ 1004,
904
+ 1005,
905
+ 1006,
906
+ 1007,
907
+ 1008,
908
+ 1009,
909
+ 1010,
910
+ 1011,
911
+ 1012,
912
+ 1013,
913
+ 1014,
914
+ 1015,
915
+ 1016,
916
+ 1017,
917
+ 1018,
918
+ 1019,
919
+ 1020,
920
+ 1021,
921
+ 1022,
922
+ 1023,
923
+ 1024,
924
+ 1025,
925
+ 1026,
926
+ 1027,
927
+ 1028,
928
+ 1029,
929
+ 1030,
930
+ 1031,
931
+ 1032,
932
+ 1033,
933
+ 1034,
934
+ 1035,
935
+ 1036,
936
+ 1037,
937
+ 1038,
938
+ 1039,
939
+ 1040,
940
+ 1041,
941
+ 1042,
942
+ 1043,
943
+ 1044,
944
+ 1045,
945
+ 1046,
946
+ 1047,
947
+ 1048,
948
+ 1049,
949
+ 1050,
950
+ 1051,
951
+ 1052,
952
+ 1053,
953
+ 1054,
954
+ 1055,
955
+ 1056,
956
+ 1057,
957
+ 1058,
958
+ 1059,
959
+ 1060,
960
+ 1061,
961
+ 1062,
962
+ 1063,
963
+ 1064,
964
+ 1065,
965
+ 1066,
966
+ 1067,
967
+ 1068,
968
+ 1069,
969
+ 1070,
970
+ 1071,
971
+ 1072,
972
+ 1073,
973
+ 1074,
974
+ 1075,
975
+ 1076,
976
+ 1077,
977
+ 1078,
978
+ 1079,
979
+ 1080,
980
+ 1081,
981
+ 1082,
982
+ 1083,
983
+ 1084,
984
+ 1085,
985
+ 1086,
986
+ 1087,
987
+ 1088,
988
+ 1089,
989
+ 1090,
990
+ 1091,
991
+ 1092,
992
+ 1093,
993
+ 1094,
994
+ 1095,
995
+ 1096,
996
+ 1097,
997
+ 1098,
998
+ 1099,
999
+ 1100,
1000
+ 1101,
1001
+ 1102,
1002
+ 1103,
1003
+ 1104,
1004
+ 1105,
1005
+ 1106,
1006
+ 1107,
1007
+ 1108,
1008
+ 1109,
1009
+ 1110,
1010
+ 1111,
1011
+ 1112,
1012
+ 1113,
1013
+ 1114,
1014
+ 1115,
1015
+ 1116,
1016
+ 1117,
1017
+ 1118,
1018
+ 1119,
1019
+ 1120,
1020
+ 1121,
1021
+ 1122,
1022
+ 1123,
1023
+ 1124,
1024
+ 1125
1025
+ ],
1026
+ "task_list": "Selected_1000/1000_new.csv",
1027
+ "task_set_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901"
1028
+ },
1029
+ "display_name": "DeepSeek API (deepseek-v4-pro alias) + DSH",
1030
+ "effective_model_call_cap": 6,
1031
+ "evidence_class": "technical_recovery_view",
1032
+ "execution": {
1033
+ "backend": "bwrap",
1034
+ "bwrap_use_rootfs_loader": true,
1035
+ "cpus": 2.0,
1036
+ "memory": "6g",
1037
+ "proc_mode": "empty",
1038
+ "rootfs_inventory_sha256": "5cdd4e51297f5a092a5b53cbe18e5d0ffd4cd9a9f219451bb2561289487848dd",
1039
+ "timeout": 90,
1040
+ "tmpfs": "256m"
1041
+ },
1042
+ "generation": {
1043
+ "max_output_tokens": 32768,
1044
+ "model": "deepseek-v4-pro",
1045
+ "prompt_profile": "raw-role-schema",
1046
+ "provider": "chat-completions",
1047
+ "reasoning_effort": "high",
1048
+ "schema_context_chars": 24000,
1049
+ "stream": true,
1050
+ "timeout": 1800
1051
+ },
1052
+ "harness": "DeepSeek Harness SDK 0.1.5rc1",
1053
+ "limits": {
1054
+ "model_calls": 6,
1055
+ "output_tokens": 32768,
1056
+ "tool_calls": 4,
1057
+ "tool_timeout": 90,
1058
+ "wall_seconds": 1800
1059
+ },
1060
+ "model": "deepseek-v4-pro",
1061
+ "model_alias": "deepseek-v4-pro",
1062
+ "model_identity": {
1063
+ "api_request_alias": "deepseek-v4-pro",
1064
+ "conflicting_official_descriptions": [
1065
+ {
1066
+ "claim": "Announcement says alias routes to V4.1 Flash from 2026-09-14 12:00 Beijing time until V4.1 Pro launches.",
1067
+ "url": "https://api-docs.deepseek.com/zh-cn/news/news260910/"
1068
+ },
1069
+ {
1070
+ "claim": "Model table lists alias as DeepSeek-V4-Pro-0813.",
1071
+ "url": "https://api-docs.deepseek.com/quick_start/pricing/"
1072
+ }
1073
+ ],
1074
+ "interpretation": "Conflicting official descriptions; neither underlying weight version is independently established.",
1075
+ "successful_response_model_label": "deepseek-v4-pro",
1076
+ "underlying_model_weights_verified": false
1077
+ },
1078
+ "notes": [
1079
+ "Cumulative per-attempt output cap includes provider-reported reasoning tokens; no over-limit response is admitted for scoring.",
1080
+ "Execution uses bwrap with a frozen rootfs, not Docker; historical rootfs equivalence is unverified.",
1081
+ "27 supplementary tasks use raw-role instead of raw-role-schema. These are not exact-prompt-parity results.",
1082
+ "API alias is established; underlying weights are unverified. Do not label as confirmed original V4-Pro-0813.",
1083
+ "Selected prediction triplets and sealed local flags are allowlisted from original evidence; hidden references and raw model responses are excluded.",
1084
+ "No independent same-protocol baseline is supplied; preserve the existing historical baseline and DeepAgents groups."
1085
+ ],
1086
+ "official_parity_verified": false,
1087
+ "recovery": {
1088
+ "best_of_attempts": false,
1089
+ "infrastructure_tasks": 60,
1090
+ "maximum_supplementary_attempts_per_task": 1,
1091
+ "original_attempts": 1000,
1092
+ "original_results_modified": false,
1093
+ "policy": "Freeze eligibility before supplementary outcomes; select attempt 2 regardless of its score or status. Never best-of or fall back to a better attempt.",
1094
+ "schema_amendment": {
1095
+ "exact_prompt_parity": false,
1096
+ "generation.prompt_profile": {
1097
+ "from": "raw-role-schema",
1098
+ "to": "raw-role"
1099
+ },
1100
+ "tools_and_per_attempt_budgets_changed": false
1101
+ },
1102
+ "schema_tasks": 27,
1103
+ "selected_primary_attempts": 913,
1104
+ "selected_supplementary_attempts": 87,
1105
+ "supplementary_attempts": 87,
1106
+ "total_attempts_charged": 1087
1107
+ },
1108
+ "request_policy": {
1109
+ "reasoning_effort": "high",
1110
+ "thinking": "enabled"
1111
+ },
1112
+ "runtime": {
1113
+ "binary_sha256": "6f68ce88d98307533ee8fa58a8125de4dc019ab16fac8b512cec141a2d1961f8",
1114
+ "packages": {
1115
+ "annotated-types": {
1116
+ "installed_wheel_sha256": "f072f4d804ea359e4eaf198b1af7a8b0943881a87f31bb764f8bf219bb9419e0",
1117
+ "record_verified": true,
1118
+ "version": "0.8.0"
1119
+ },
1120
+ "deepseek-harness-runtime-bin": {
1121
+ "installed_wheel_sha256": "ad877237382baff4628969a9ebdbfe38486d971b7987c032d0e29775759b5229",
1122
+ "record_verified": true,
1123
+ "version": "0.1.5rc1"
1124
+ },
1125
+ "deepseek-harness-sdk": {
1126
+ "installed_wheel_sha256": "8d9897e97e8a95fd738cb2c3e2c563951d824e43847ec6f3b0fb172c0729b4d8",
1127
+ "record_verified": true,
1128
+ "version": "0.1.5rc1"
1129
+ },
1130
+ "pydantic": {
1131
+ "installed_wheel_sha256": "346a034f080da3755d8e9cb5e00e8b07de1d39e4f6e2c87d8ab7cafa0b269a73",
1132
+ "record_verified": true,
1133
+ "version": "2.13.5"
1134
+ },
1135
+ "pydantic-core": {
1136
+ "installed_wheel_sha256": "0fc5be0abd4a407e200d844b404e33639a554e7bd0d448e7b9ae181be4789ac2",
1137
+ "record_verified": true,
1138
+ "version": "2.46.5"
1139
+ },
1140
+ "typing-extensions": {
1141
+ "installed_wheel_sha256": "481caa481374e813c1b176ada14e97f1f67a4539ce9cfeb3f350d78d6370c2e8",
1142
+ "record_verified": true,
1143
+ "version": "4.16.0"
1144
+ },
1145
+ "typing-inspection": {
1146
+ "installed_wheel_sha256": "65b8397ba37ccbce054456aaccddfc91e6e3083c92824df348d96ca832f3f147",
1147
+ "record_verified": true,
1148
+ "version": "0.4.4"
1149
+ }
1150
+ },
1151
+ "profile_files": {
1152
+ "__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
1153
+ "profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
1154
+ "requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
1155
+ "requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
1156
+ "run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
1157
+ "runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
1158
+ "runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7"
1159
+ },
1160
+ "python_version": "3.12.12",
1161
+ "research_source_commit": "0d1f50007f9bca3f52b06e1c3074fa14d5fb0720",
1162
+ "research_source_matches_installed_build_verified": false
1163
+ },
1164
+ "schedule_seed": 20260918,
1165
+ "schema_version": 3,
1166
+ "scoring": {
1167
+ "profile": "local-paper-v1"
1168
+ },
1169
+ "source_identity": {
1170
+ "controller_files": {
1171
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
1172
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
1173
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
1174
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
1175
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
1176
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
1177
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
1178
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
1179
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
1180
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
1181
+ "ntu_gifts/dsh/bridge.py": "f56ef6509cd831ec7de2c61b5331a7496771cf0754fcfaf0b29aeb88fa9be0f9",
1182
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
1183
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
1184
+ "ntu_gifts/dsh/evidence.py": "bb0f5be46a80c4fcaf993e8071b22160fc04d5d871a9e446340a4041a9ef7591",
1185
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
1186
+ "ntu_gifts/dsh/failure_policy.py": "50f39ce6228e1eaca24278dd46774d13a39ecd2e7773a5c36e952ee70451083f",
1187
+ "ntu_gifts/dsh/late_evidence.py": "5f6231a93f7668107f6d15501177013a64e27af42a904b8a8acba502edcab197",
1188
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
1189
+ "ntu_gifts/dsh/reporting.py": "823067b7a99f03cc682fb448aa26992b06b75ee15096567d93d85b31e715e9f9",
1190
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
1191
+ "ntu_gifts/dsh/study.py": "57c3495a97f12baae8aa2bb74c2289bfc5a7cbad0cc70205f34e66c734c28008",
1192
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
1193
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
1194
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
1195
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
1196
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
1197
+ },
1198
+ "evaluator_revision": "92e56c82d1c91b1297d08dd331e449362eca30f8",
1199
+ "latest_controller_commit": "3ced22b417aea99b0956b504794561964dbf7b33"
1200
+ },
1201
+ "tools": [
1202
+ "run_python"
1203
+ ],
1204
+ "underlying_model_weights_verified": false
1205
+ }
results/deepseek-v4-pro/dsh/diagnostics.json ADDED
@@ -0,0 +1,1303 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "categories": {
3
+ "failed": {
4
+ "compilation_evidence": {
5
+ "confirmed_final_execution_error": 43,
6
+ "confirmed_final_execution_timeout": 3,
7
+ "confirmed_model_failure_without_final_execution": 225
8
+ },
9
+ "count": 271,
10
+ "final_exception_types": {
11
+ "AttributeError": 1,
12
+ "IndexError": 3,
13
+ "KeyError": 4,
14
+ "RuntimeError": 1,
15
+ "TypeError": 5,
16
+ "ValueError": 26,
17
+ "none": 3,
18
+ "numpy._core._exceptions._ArrayMemoryError": 1,
19
+ "numpy.linalg.LinAlgError": 1,
20
+ "statsmodels.tools.sm_exceptions.MissingDataError": 1
21
+ },
22
+ "incomplete_response_finish_reasons": {
23
+ "['length']": 182
24
+ },
25
+ "reasons": {
26
+ "Expected one fenced Python program": 1,
27
+ "final_execution_failed": 46,
28
+ "model_calls_exhausted": 42,
29
+ "model_response_not_completed": 182
30
+ }
31
+ },
32
+ "uncertain": {
33
+ "compilation_evidence": {
34
+ "pending_or_uncertain_task": 30
35
+ },
36
+ "count": 30,
37
+ "final_exception_types": {},
38
+ "incomplete_response_finish_reasons": {},
39
+ "reasons": {
40
+ "bridge_work_cleanup_unconfirmed": 1,
41
+ "model_budget_or_deadline_exceeded_after_dispatch": 18,
42
+ "model_outcome_unknown": 6,
43
+ "provider_exceeded_output_limit": 1,
44
+ "study_stopped": 4
45
+ }
46
+ }
47
+ },
48
+ "chosen_task_count": 1000,
49
+ "failed_task_ids": [
50
+ 1,
51
+ 5,
52
+ 8,
53
+ 23,
54
+ 35,
55
+ 40,
56
+ 41,
57
+ 55,
58
+ 61,
59
+ 69,
60
+ 82,
61
+ 86,
62
+ 87,
63
+ 88,
64
+ 89,
65
+ 93,
66
+ 94,
67
+ 95,
68
+ 96,
69
+ 97,
70
+ 100,
71
+ 101,
72
+ 102,
73
+ 107,
74
+ 115,
75
+ 122,
76
+ 124,
77
+ 126,
78
+ 131,
79
+ 132,
80
+ 133,
81
+ 134,
82
+ 135,
83
+ 136,
84
+ 137,
85
+ 138,
86
+ 139,
87
+ 140,
88
+ 145,
89
+ 147,
90
+ 148,
91
+ 149,
92
+ 150,
93
+ 156,
94
+ 157,
95
+ 159,
96
+ 160,
97
+ 161,
98
+ 165,
99
+ 166,
100
+ 167,
101
+ 168,
102
+ 170,
103
+ 171,
104
+ 172,
105
+ 182,
106
+ 183,
107
+ 186,
108
+ 187,
109
+ 189,
110
+ 191,
111
+ 192,
112
+ 198,
113
+ 199,
114
+ 211,
115
+ 216,
116
+ 236,
117
+ 238,
118
+ 249,
119
+ 251,
120
+ 252,
121
+ 253,
122
+ 254,
123
+ 255,
124
+ 256,
125
+ 257,
126
+ 258,
127
+ 308,
128
+ 320,
129
+ 321,
130
+ 323,
131
+ 324,
132
+ 325,
133
+ 327,
134
+ 331,
135
+ 332,
136
+ 334,
137
+ 336,
138
+ 337,
139
+ 338,
140
+ 339,
141
+ 341,
142
+ 345,
143
+ 349,
144
+ 352,
145
+ 362,
146
+ 365,
147
+ 366,
148
+ 373,
149
+ 374,
150
+ 376,
151
+ 379,
152
+ 380,
153
+ 384,
154
+ 387,
155
+ 388,
156
+ 393,
157
+ 413,
158
+ 414,
159
+ 419,
160
+ 422,
161
+ 426,
162
+ 427,
163
+ 428,
164
+ 431,
165
+ 434,
166
+ 435,
167
+ 438,
168
+ 443,
169
+ 445,
170
+ 446,
171
+ 447,
172
+ 453,
173
+ 454,
174
+ 455,
175
+ 456,
176
+ 458,
177
+ 462,
178
+ 466,
179
+ 467,
180
+ 473,
181
+ 474,
182
+ 475,
183
+ 492,
184
+ 495,
185
+ 496,
186
+ 498,
187
+ 507,
188
+ 509,
189
+ 519,
190
+ 527,
191
+ 529,
192
+ 532,
193
+ 542,
194
+ 543,
195
+ 553,
196
+ 556,
197
+ 559,
198
+ 568,
199
+ 570,
200
+ 571,
201
+ 572,
202
+ 581,
203
+ 592,
204
+ 601,
205
+ 602,
206
+ 603,
207
+ 604,
208
+ 607,
209
+ 608,
210
+ 609,
211
+ 616,
212
+ 618,
213
+ 621,
214
+ 622,
215
+ 623,
216
+ 667,
217
+ 668,
218
+ 675,
219
+ 676,
220
+ 678,
221
+ 718,
222
+ 732,
223
+ 734,
224
+ 735,
225
+ 739,
226
+ 740,
227
+ 743,
228
+ 744,
229
+ 745,
230
+ 754,
231
+ 760,
232
+ 763,
233
+ 766,
234
+ 767,
235
+ 778,
236
+ 779,
237
+ 780,
238
+ 781,
239
+ 785,
240
+ 789,
241
+ 791,
242
+ 794,
243
+ 798,
244
+ 800,
245
+ 801,
246
+ 802,
247
+ 803,
248
+ 804,
249
+ 806,
250
+ 807,
251
+ 834,
252
+ 842,
253
+ 848,
254
+ 852,
255
+ 854,
256
+ 855,
257
+ 859,
258
+ 862,
259
+ 876,
260
+ 878,
261
+ 880,
262
+ 881,
263
+ 882,
264
+ 885,
265
+ 886,
266
+ 887,
267
+ 892,
268
+ 898,
269
+ 900,
270
+ 906,
271
+ 922,
272
+ 923,
273
+ 924,
274
+ 925,
275
+ 926,
276
+ 931,
277
+ 932,
278
+ 934,
279
+ 943,
280
+ 945,
281
+ 967,
282
+ 968,
283
+ 970,
284
+ 971,
285
+ 973,
286
+ 974,
287
+ 975,
288
+ 976,
289
+ 977,
290
+ 979,
291
+ 980,
292
+ 981,
293
+ 982,
294
+ 983,
295
+ 984,
296
+ 985,
297
+ 1001,
298
+ 1008,
299
+ 1009,
300
+ 1010,
301
+ 1016,
302
+ 1019,
303
+ 1033,
304
+ 1040,
305
+ 1042,
306
+ 1060,
307
+ 1064,
308
+ 1081,
309
+ 1082,
310
+ 1084,
311
+ 1091,
312
+ 1094,
313
+ 1099,
314
+ 1109,
315
+ 1112,
316
+ 1116,
317
+ 1118,
318
+ 1119,
319
+ 1122,
320
+ 1123
321
+ ],
322
+ "scope": "Audited categories only. Raw error text, prompts, model responses, prediction/reference structures and machine paths are excluded.",
323
+ "uncertain_task_ids": [
324
+ 57,
325
+ 64,
326
+ 103,
327
+ 105,
328
+ 146,
329
+ 163,
330
+ 250,
331
+ 294,
332
+ 353,
333
+ 378,
334
+ 424,
335
+ 448,
336
+ 560,
337
+ 577,
338
+ 579,
339
+ 590,
340
+ 598,
341
+ 606,
342
+ 611,
343
+ 787,
344
+ 835,
345
+ 883,
346
+ 897,
347
+ 907,
348
+ 913,
349
+ 921,
350
+ 928,
351
+ 952,
352
+ 962,
353
+ 1004
354
+ ],
355
+ "unknown_metric_task_ids": {
356
+ "coefficient_direction": [
357
+ 1,
358
+ 5,
359
+ 8,
360
+ 23,
361
+ 35,
362
+ 40,
363
+ 41,
364
+ 55,
365
+ 57,
366
+ 61,
367
+ 64,
368
+ 69,
369
+ 82,
370
+ 86,
371
+ 87,
372
+ 88,
373
+ 89,
374
+ 93,
375
+ 94,
376
+ 95,
377
+ 96,
378
+ 97,
379
+ 100,
380
+ 101,
381
+ 102,
382
+ 103,
383
+ 105,
384
+ 107,
385
+ 115,
386
+ 122,
387
+ 124,
388
+ 126,
389
+ 131,
390
+ 132,
391
+ 133,
392
+ 134,
393
+ 135,
394
+ 136,
395
+ 137,
396
+ 138,
397
+ 139,
398
+ 140,
399
+ 145,
400
+ 146,
401
+ 147,
402
+ 148,
403
+ 149,
404
+ 150,
405
+ 156,
406
+ 157,
407
+ 159,
408
+ 160,
409
+ 161,
410
+ 163,
411
+ 165,
412
+ 166,
413
+ 167,
414
+ 168,
415
+ 170,
416
+ 171,
417
+ 172,
418
+ 182,
419
+ 183,
420
+ 186,
421
+ 187,
422
+ 189,
423
+ 191,
424
+ 192,
425
+ 198,
426
+ 199,
427
+ 211,
428
+ 216,
429
+ 236,
430
+ 238,
431
+ 249,
432
+ 250,
433
+ 251,
434
+ 252,
435
+ 253,
436
+ 254,
437
+ 255,
438
+ 256,
439
+ 257,
440
+ 258,
441
+ 294,
442
+ 308,
443
+ 320,
444
+ 321,
445
+ 323,
446
+ 324,
447
+ 325,
448
+ 327,
449
+ 331,
450
+ 332,
451
+ 334,
452
+ 336,
453
+ 337,
454
+ 338,
455
+ 339,
456
+ 341,
457
+ 345,
458
+ 349,
459
+ 352,
460
+ 353,
461
+ 362,
462
+ 365,
463
+ 366,
464
+ 373,
465
+ 374,
466
+ 376,
467
+ 378,
468
+ 379,
469
+ 380,
470
+ 384,
471
+ 387,
472
+ 388,
473
+ 393,
474
+ 413,
475
+ 414,
476
+ 419,
477
+ 422,
478
+ 424,
479
+ 426,
480
+ 427,
481
+ 428,
482
+ 431,
483
+ 434,
484
+ 435,
485
+ 438,
486
+ 443,
487
+ 445,
488
+ 446,
489
+ 447,
490
+ 448,
491
+ 453,
492
+ 454,
493
+ 455,
494
+ 456,
495
+ 458,
496
+ 462,
497
+ 466,
498
+ 467,
499
+ 473,
500
+ 474,
501
+ 475,
502
+ 492,
503
+ 495,
504
+ 496,
505
+ 498,
506
+ 507,
507
+ 509,
508
+ 519,
509
+ 527,
510
+ 529,
511
+ 532,
512
+ 542,
513
+ 543,
514
+ 553,
515
+ 556,
516
+ 559,
517
+ 560,
518
+ 568,
519
+ 570,
520
+ 571,
521
+ 572,
522
+ 577,
523
+ 579,
524
+ 581,
525
+ 590,
526
+ 592,
527
+ 598,
528
+ 601,
529
+ 602,
530
+ 603,
531
+ 604,
532
+ 606,
533
+ 607,
534
+ 608,
535
+ 609,
536
+ 611,
537
+ 616,
538
+ 618,
539
+ 621,
540
+ 622,
541
+ 623,
542
+ 667,
543
+ 668,
544
+ 675,
545
+ 676,
546
+ 678,
547
+ 718,
548
+ 732,
549
+ 734,
550
+ 735,
551
+ 739,
552
+ 740,
553
+ 743,
554
+ 744,
555
+ 745,
556
+ 754,
557
+ 760,
558
+ 763,
559
+ 766,
560
+ 767,
561
+ 778,
562
+ 779,
563
+ 780,
564
+ 781,
565
+ 785,
566
+ 787,
567
+ 789,
568
+ 791,
569
+ 794,
570
+ 798,
571
+ 800,
572
+ 801,
573
+ 802,
574
+ 803,
575
+ 804,
576
+ 806,
577
+ 807,
578
+ 834,
579
+ 835,
580
+ 842,
581
+ 848,
582
+ 852,
583
+ 854,
584
+ 855,
585
+ 859,
586
+ 862,
587
+ 876,
588
+ 878,
589
+ 880,
590
+ 881,
591
+ 882,
592
+ 883,
593
+ 885,
594
+ 886,
595
+ 887,
596
+ 892,
597
+ 897,
598
+ 898,
599
+ 900,
600
+ 906,
601
+ 907,
602
+ 913,
603
+ 921,
604
+ 922,
605
+ 923,
606
+ 924,
607
+ 925,
608
+ 926,
609
+ 928,
610
+ 931,
611
+ 932,
612
+ 934,
613
+ 943,
614
+ 945,
615
+ 952,
616
+ 962,
617
+ 967,
618
+ 968,
619
+ 970,
620
+ 971,
621
+ 973,
622
+ 974,
623
+ 975,
624
+ 976,
625
+ 977,
626
+ 979,
627
+ 980,
628
+ 981,
629
+ 982,
630
+ 983,
631
+ 984,
632
+ 985,
633
+ 1001,
634
+ 1004,
635
+ 1008,
636
+ 1009,
637
+ 1010,
638
+ 1016,
639
+ 1019,
640
+ 1033,
641
+ 1040,
642
+ 1042,
643
+ 1050,
644
+ 1060,
645
+ 1064,
646
+ 1081,
647
+ 1082,
648
+ 1084,
649
+ 1091,
650
+ 1094,
651
+ 1099,
652
+ 1109,
653
+ 1112,
654
+ 1116,
655
+ 1118,
656
+ 1119,
657
+ 1122,
658
+ 1123
659
+ ],
660
+ "compilation_success": [
661
+ 57,
662
+ 64,
663
+ 103,
664
+ 105,
665
+ 146,
666
+ 163,
667
+ 250,
668
+ 294,
669
+ 353,
670
+ 378,
671
+ 424,
672
+ 448,
673
+ 560,
674
+ 577,
675
+ 579,
676
+ 590,
677
+ 598,
678
+ 606,
679
+ 611,
680
+ 787,
681
+ 835,
682
+ 883,
683
+ 897,
684
+ 907,
685
+ 913,
686
+ 921,
687
+ 928,
688
+ 952,
689
+ 962,
690
+ 1004
691
+ ],
692
+ "partial_replication": [
693
+ 1,
694
+ 5,
695
+ 8,
696
+ 23,
697
+ 35,
698
+ 40,
699
+ 41,
700
+ 55,
701
+ 57,
702
+ 61,
703
+ 64,
704
+ 69,
705
+ 82,
706
+ 86,
707
+ 87,
708
+ 88,
709
+ 89,
710
+ 93,
711
+ 94,
712
+ 95,
713
+ 96,
714
+ 97,
715
+ 100,
716
+ 101,
717
+ 102,
718
+ 103,
719
+ 105,
720
+ 107,
721
+ 115,
722
+ 122,
723
+ 124,
724
+ 126,
725
+ 131,
726
+ 132,
727
+ 133,
728
+ 134,
729
+ 135,
730
+ 136,
731
+ 137,
732
+ 138,
733
+ 139,
734
+ 140,
735
+ 145,
736
+ 146,
737
+ 147,
738
+ 148,
739
+ 149,
740
+ 150,
741
+ 156,
742
+ 157,
743
+ 159,
744
+ 160,
745
+ 161,
746
+ 163,
747
+ 165,
748
+ 166,
749
+ 167,
750
+ 168,
751
+ 170,
752
+ 171,
753
+ 172,
754
+ 182,
755
+ 183,
756
+ 186,
757
+ 187,
758
+ 189,
759
+ 191,
760
+ 192,
761
+ 198,
762
+ 199,
763
+ 211,
764
+ 216,
765
+ 236,
766
+ 238,
767
+ 249,
768
+ 250,
769
+ 251,
770
+ 252,
771
+ 253,
772
+ 254,
773
+ 255,
774
+ 256,
775
+ 257,
776
+ 258,
777
+ 294,
778
+ 308,
779
+ 320,
780
+ 321,
781
+ 323,
782
+ 324,
783
+ 325,
784
+ 327,
785
+ 331,
786
+ 332,
787
+ 334,
788
+ 336,
789
+ 337,
790
+ 338,
791
+ 339,
792
+ 341,
793
+ 345,
794
+ 349,
795
+ 352,
796
+ 353,
797
+ 362,
798
+ 365,
799
+ 366,
800
+ 373,
801
+ 374,
802
+ 376,
803
+ 378,
804
+ 379,
805
+ 380,
806
+ 384,
807
+ 387,
808
+ 388,
809
+ 393,
810
+ 413,
811
+ 414,
812
+ 419,
813
+ 422,
814
+ 424,
815
+ 426,
816
+ 427,
817
+ 428,
818
+ 431,
819
+ 434,
820
+ 435,
821
+ 438,
822
+ 443,
823
+ 445,
824
+ 446,
825
+ 447,
826
+ 448,
827
+ 453,
828
+ 454,
829
+ 455,
830
+ 456,
831
+ 458,
832
+ 462,
833
+ 466,
834
+ 467,
835
+ 473,
836
+ 474,
837
+ 475,
838
+ 492,
839
+ 495,
840
+ 496,
841
+ 498,
842
+ 507,
843
+ 509,
844
+ 519,
845
+ 527,
846
+ 529,
847
+ 532,
848
+ 542,
849
+ 543,
850
+ 553,
851
+ 556,
852
+ 559,
853
+ 560,
854
+ 568,
855
+ 570,
856
+ 571,
857
+ 572,
858
+ 577,
859
+ 579,
860
+ 581,
861
+ 590,
862
+ 592,
863
+ 598,
864
+ 601,
865
+ 602,
866
+ 603,
867
+ 604,
868
+ 606,
869
+ 607,
870
+ 608,
871
+ 609,
872
+ 611,
873
+ 616,
874
+ 618,
875
+ 621,
876
+ 622,
877
+ 623,
878
+ 667,
879
+ 668,
880
+ 675,
881
+ 676,
882
+ 678,
883
+ 718,
884
+ 732,
885
+ 734,
886
+ 735,
887
+ 739,
888
+ 740,
889
+ 743,
890
+ 744,
891
+ 745,
892
+ 754,
893
+ 760,
894
+ 763,
895
+ 766,
896
+ 767,
897
+ 778,
898
+ 779,
899
+ 780,
900
+ 781,
901
+ 785,
902
+ 787,
903
+ 789,
904
+ 791,
905
+ 794,
906
+ 798,
907
+ 800,
908
+ 801,
909
+ 802,
910
+ 803,
911
+ 804,
912
+ 806,
913
+ 807,
914
+ 834,
915
+ 835,
916
+ 842,
917
+ 848,
918
+ 852,
919
+ 854,
920
+ 855,
921
+ 859,
922
+ 862,
923
+ 876,
924
+ 878,
925
+ 880,
926
+ 881,
927
+ 882,
928
+ 883,
929
+ 885,
930
+ 886,
931
+ 887,
932
+ 892,
933
+ 897,
934
+ 898,
935
+ 900,
936
+ 906,
937
+ 907,
938
+ 913,
939
+ 921,
940
+ 922,
941
+ 923,
942
+ 924,
943
+ 925,
944
+ 926,
945
+ 928,
946
+ 931,
947
+ 932,
948
+ 934,
949
+ 943,
950
+ 945,
951
+ 952,
952
+ 962,
953
+ 967,
954
+ 968,
955
+ 970,
956
+ 971,
957
+ 973,
958
+ 974,
959
+ 975,
960
+ 976,
961
+ 977,
962
+ 979,
963
+ 980,
964
+ 981,
965
+ 982,
966
+ 983,
967
+ 984,
968
+ 985,
969
+ 1001,
970
+ 1004,
971
+ 1008,
972
+ 1009,
973
+ 1010,
974
+ 1016,
975
+ 1019,
976
+ 1033,
977
+ 1040,
978
+ 1042,
979
+ 1050,
980
+ 1060,
981
+ 1064,
982
+ 1081,
983
+ 1082,
984
+ 1084,
985
+ 1091,
986
+ 1094,
987
+ 1099,
988
+ 1109,
989
+ 1112,
990
+ 1116,
991
+ 1118,
992
+ 1119,
993
+ 1122,
994
+ 1123
995
+ ],
996
+ "significance_level": [
997
+ 1,
998
+ 5,
999
+ 8,
1000
+ 23,
1001
+ 35,
1002
+ 40,
1003
+ 41,
1004
+ 55,
1005
+ 57,
1006
+ 61,
1007
+ 64,
1008
+ 69,
1009
+ 82,
1010
+ 86,
1011
+ 87,
1012
+ 88,
1013
+ 89,
1014
+ 93,
1015
+ 94,
1016
+ 95,
1017
+ 96,
1018
+ 97,
1019
+ 100,
1020
+ 101,
1021
+ 102,
1022
+ 103,
1023
+ 105,
1024
+ 107,
1025
+ 115,
1026
+ 122,
1027
+ 124,
1028
+ 126,
1029
+ 131,
1030
+ 132,
1031
+ 133,
1032
+ 134,
1033
+ 135,
1034
+ 136,
1035
+ 137,
1036
+ 138,
1037
+ 139,
1038
+ 140,
1039
+ 145,
1040
+ 146,
1041
+ 147,
1042
+ 148,
1043
+ 149,
1044
+ 150,
1045
+ 156,
1046
+ 157,
1047
+ 159,
1048
+ 160,
1049
+ 161,
1050
+ 163,
1051
+ 165,
1052
+ 166,
1053
+ 167,
1054
+ 168,
1055
+ 170,
1056
+ 171,
1057
+ 172,
1058
+ 182,
1059
+ 183,
1060
+ 186,
1061
+ 187,
1062
+ 189,
1063
+ 191,
1064
+ 192,
1065
+ 198,
1066
+ 199,
1067
+ 211,
1068
+ 216,
1069
+ 236,
1070
+ 238,
1071
+ 249,
1072
+ 250,
1073
+ 251,
1074
+ 252,
1075
+ 253,
1076
+ 254,
1077
+ 255,
1078
+ 256,
1079
+ 257,
1080
+ 258,
1081
+ 294,
1082
+ 308,
1083
+ 320,
1084
+ 321,
1085
+ 323,
1086
+ 324,
1087
+ 325,
1088
+ 327,
1089
+ 331,
1090
+ 332,
1091
+ 334,
1092
+ 336,
1093
+ 337,
1094
+ 338,
1095
+ 339,
1096
+ 341,
1097
+ 345,
1098
+ 349,
1099
+ 352,
1100
+ 353,
1101
+ 362,
1102
+ 365,
1103
+ 366,
1104
+ 373,
1105
+ 374,
1106
+ 376,
1107
+ 378,
1108
+ 379,
1109
+ 380,
1110
+ 384,
1111
+ 387,
1112
+ 388,
1113
+ 393,
1114
+ 413,
1115
+ 414,
1116
+ 419,
1117
+ 422,
1118
+ 424,
1119
+ 426,
1120
+ 427,
1121
+ 428,
1122
+ 431,
1123
+ 434,
1124
+ 435,
1125
+ 438,
1126
+ 443,
1127
+ 445,
1128
+ 446,
1129
+ 447,
1130
+ 448,
1131
+ 453,
1132
+ 454,
1133
+ 455,
1134
+ 456,
1135
+ 458,
1136
+ 462,
1137
+ 466,
1138
+ 467,
1139
+ 473,
1140
+ 474,
1141
+ 475,
1142
+ 492,
1143
+ 495,
1144
+ 496,
1145
+ 498,
1146
+ 507,
1147
+ 509,
1148
+ 519,
1149
+ 527,
1150
+ 529,
1151
+ 532,
1152
+ 542,
1153
+ 543,
1154
+ 553,
1155
+ 556,
1156
+ 559,
1157
+ 560,
1158
+ 568,
1159
+ 570,
1160
+ 571,
1161
+ 572,
1162
+ 577,
1163
+ 579,
1164
+ 581,
1165
+ 590,
1166
+ 592,
1167
+ 598,
1168
+ 601,
1169
+ 602,
1170
+ 603,
1171
+ 604,
1172
+ 606,
1173
+ 607,
1174
+ 608,
1175
+ 609,
1176
+ 611,
1177
+ 616,
1178
+ 618,
1179
+ 621,
1180
+ 622,
1181
+ 623,
1182
+ 667,
1183
+ 668,
1184
+ 675,
1185
+ 676,
1186
+ 678,
1187
+ 718,
1188
+ 732,
1189
+ 734,
1190
+ 735,
1191
+ 739,
1192
+ 740,
1193
+ 743,
1194
+ 744,
1195
+ 745,
1196
+ 754,
1197
+ 760,
1198
+ 763,
1199
+ 766,
1200
+ 767,
1201
+ 778,
1202
+ 779,
1203
+ 780,
1204
+ 781,
1205
+ 785,
1206
+ 787,
1207
+ 789,
1208
+ 791,
1209
+ 794,
1210
+ 798,
1211
+ 800,
1212
+ 801,
1213
+ 802,
1214
+ 803,
1215
+ 804,
1216
+ 806,
1217
+ 807,
1218
+ 834,
1219
+ 835,
1220
+ 842,
1221
+ 848,
1222
+ 852,
1223
+ 854,
1224
+ 855,
1225
+ 859,
1226
+ 862,
1227
+ 876,
1228
+ 878,
1229
+ 880,
1230
+ 881,
1231
+ 882,
1232
+ 883,
1233
+ 885,
1234
+ 886,
1235
+ 887,
1236
+ 892,
1237
+ 897,
1238
+ 898,
1239
+ 900,
1240
+ 906,
1241
+ 907,
1242
+ 913,
1243
+ 915,
1244
+ 917,
1245
+ 921,
1246
+ 922,
1247
+ 923,
1248
+ 924,
1249
+ 925,
1250
+ 926,
1251
+ 928,
1252
+ 931,
1253
+ 932,
1254
+ 934,
1255
+ 943,
1256
+ 945,
1257
+ 952,
1258
+ 962,
1259
+ 967,
1260
+ 968,
1261
+ 970,
1262
+ 971,
1263
+ 973,
1264
+ 974,
1265
+ 975,
1266
+ 976,
1267
+ 977,
1268
+ 979,
1269
+ 980,
1270
+ 981,
1271
+ 982,
1272
+ 983,
1273
+ 984,
1274
+ 985,
1275
+ 1001,
1276
+ 1004,
1277
+ 1008,
1278
+ 1009,
1279
+ 1010,
1280
+ 1016,
1281
+ 1019,
1282
+ 1033,
1283
+ 1040,
1284
+ 1042,
1285
+ 1050,
1286
+ 1060,
1287
+ 1064,
1288
+ 1081,
1289
+ 1082,
1290
+ 1084,
1291
+ 1091,
1292
+ 1094,
1293
+ 1099,
1294
+ 1109,
1295
+ 1112,
1296
+ 1116,
1297
+ 1118,
1298
+ 1119,
1299
+ 1122,
1300
+ 1123
1301
+ ]
1302
+ }
1303
+ }
results/deepseek-v4-pro/dsh/export-verification.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "all_attempts": 1087,
3
+ "kind": "dsh-public-export-verification-v1",
4
+ "known_and_unknown_costs_preserved": true,
5
+ "local_flags_read_from_original_sealed_scoring": true,
6
+ "metric_counts_verified_against_exported_flags": true,
7
+ "model_calls_performed": 0,
8
+ "official_parity_verified": false,
9
+ "original_attempts": 1000,
10
+ "output_sha256": {
11
+ "README.md": "ccf9f54ce467c8a977074498eeb86a2ce1b7b69a73b370b025663f9582a8d991",
12
+ "accounting.json": "6f92c5376ab6b96b9ab5ddab6faff88cd792d0ec999f1cc96cdd48e32302c9e1",
13
+ "all-attempts.jsonl": "33faf46bce3d4a152bb04cd3eb7cded2d948b2bf5ec2beac49941b145e92c522",
14
+ "config.json": "df322f13761e06dc78cb553c5dac347ce4361599937af4017a902c56a7df9d77",
15
+ "diagnostics.json": "d6500c012259ed5e3b514d6f62597fdcadfb7548e41308eed442a1b130f0e336",
16
+ "first-pass-results.jsonl": "153d15ff14c3dba6ece8fd9c81ae2d9fa1f0b6a240e90efd6b1d6e46f40893e2",
17
+ "first-pass-summary.json": "b264e4796ea32220a7c67fe84d990ac6e040784f7331c662e66ca5d91c71e9c1",
18
+ "provenance.json": "8e03b3fd76310ab80643f51ce791e85545062d2d85ac3c6121cc1143ebfb5a28",
19
+ "results.csv": "954a6adff76d6bae31c77cd969b17fa5203ae7ba38abf76130c6f066e39a9c58",
20
+ "results.jsonl": "b12427d5f4476351e571e4b29f61e0463540a1701bec51e642753d717dbc1df2",
21
+ "summary.json": "62936f76b1ca0de5e761d1ea1f894343695503ce934bf0213dc9fde1ea72a782"
22
+ },
23
+ "private_fields_allowlisted_out": true,
24
+ "rescoring_performed": false,
25
+ "selected_tasks": 1000,
26
+ "selection_verified_against_frozen_plan": true,
27
+ "source_manifest_sha256": "89b41604d0f1fd5c8b35b5a41f2a83b24c27f2ea093444c64b57ab34a6216b0c",
28
+ "supplementary_attempts": 87,
29
+ "underlying_model_weights_verified": false,
30
+ "unique_selected_tasks": 1000,
31
+ "unknown_usage_attempts": 38
32
+ }
results/deepseek-v4-pro/dsh/first-pass-results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/dsh/first-pass-summary.json ADDED
@@ -0,0 +1,113 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "all_slots_sealed": true,
3
+ "arm": "dsh",
4
+ "assumptions": {
5
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy.",
6
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
7
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
8
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
9
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
10
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
11
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
12
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
13
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect."
14
+ },
15
+ "complete": false,
16
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
17
+ "definitions": {
18
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
19
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
20
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
21
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
22
+ },
23
+ "display_name": "DeepSeek API (deepseek-v4-pro alias) + DSH",
24
+ "display_role": "supporting_original_report",
25
+ "evidence_class": "original_first_pass",
26
+ "generated_at": "2026-09-19T18:18:48.070003+00:00",
27
+ "kind": "inferencenet-offline-leaderboard-export",
28
+ "legacy_local_paper_export_status": "not_exported_for_original_first_pass",
29
+ "metric_profile": "hf-leaderboard-v1",
30
+ "model": "deepseek-v4-pro",
31
+ "model_alias": "deepseek-v4-pro",
32
+ "official_parity_verified": false,
33
+ "original_protocol_complete": false,
34
+ "publication_kind": "research_results_with_technical_recovery",
35
+ "result": {
36
+ "columns": [
37
+ "Model ID",
38
+ "Compilation Success",
39
+ "Partial Replication",
40
+ "Correct Coefficient Direction",
41
+ "Significant Level Correctness"
42
+ ],
43
+ "definite_outcome_count": 928,
44
+ "expected_count": 1000,
45
+ "invalid_evidence_count": 0,
46
+ "metrics": {
47
+ "coefficient_direction": {
48
+ "assessable": 661,
49
+ "count": 648,
50
+ "coverage_percent": 66.1,
51
+ "denominator": 1000,
52
+ "failure_count": 13,
53
+ "rate": 0.648,
54
+ "score": 64.8,
55
+ "unknown_count": 339
56
+ },
57
+ "compilation_success": {
58
+ "assessable": 900,
59
+ "count": 662,
60
+ "coverage_percent": 90.0,
61
+ "denominator": 1000,
62
+ "failure_count": 238,
63
+ "rate": 0.662,
64
+ "score": 66.2,
65
+ "unknown_count": 100
66
+ },
67
+ "partial_replication": {
68
+ "assessable": 661,
69
+ "count": 527,
70
+ "coverage_percent": 66.1,
71
+ "denominator": 1000,
72
+ "failure_count": 134,
73
+ "rate": 0.527,
74
+ "score": 52.7,
75
+ "unknown_count": 339
76
+ },
77
+ "significance_level": {
78
+ "assessable": 659,
79
+ "count": 566,
80
+ "coverage_percent": 65.9,
81
+ "denominator": 1000,
82
+ "failure_count": 93,
83
+ "rate": 0.566,
84
+ "score": 56.6,
85
+ "unknown_count": 341
86
+ }
87
+ },
88
+ "resolved_count": 1000,
89
+ "row": {
90
+ "Compilation Success": 66.2,
91
+ "Correct Coefficient Direction": 64.8,
92
+ "Model ID": "DeepSeek API (deepseek-v4-pro alias) + DSH",
93
+ "Partial Replication": 52.7,
94
+ "Significant Level Correctness": 56.6
95
+ },
96
+ "sealed_count": 1000,
97
+ "task_counts": {
98
+ "completed": 1000,
99
+ "failed": 266,
100
+ "invalid_evidence": 0,
101
+ "not_started": 0,
102
+ "planned": 1000,
103
+ "succeeded": 662,
104
+ "unknown": 72,
105
+ "unsealed": 0
106
+ },
107
+ "task_unknown_count": 72,
108
+ "valid_prediction_count": 662
109
+ },
110
+ "schema_version": 3,
111
+ "underlying_model_weights_verified": false,
112
+ "verification_scope": "Preserved verified flags and result/score hashes; no new inference, execution or rescoring."
113
+ }
results/deepseek-v4-pro/dsh/manifest.sha256 ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ccf9f54ce467c8a977074498eeb86a2ce1b7b69a73b370b025663f9582a8d991 README.md
2
+ 6f92c5376ab6b96b9ab5ddab6faff88cd792d0ec999f1cc96cdd48e32302c9e1 accounting.json
3
+ 33faf46bce3d4a152bb04cd3eb7cded2d948b2bf5ec2beac49941b145e92c522 all-attempts.jsonl
4
+ df322f13761e06dc78cb553c5dac347ce4361599937af4017a902c56a7df9d77 config.json
5
+ d6500c012259ed5e3b514d6f62597fdcadfb7548e41308eed442a1b130f0e336 diagnostics.json
6
+ 8d7cae2ea2f33130636367f5a5e8f92b5df0471aa9889aa54094e2d1664dd48b export-verification.json
7
+ 153d15ff14c3dba6ece8fd9c81ae2d9fa1f0b6a240e90efd6b1d6e46f40893e2 first-pass-results.jsonl
8
+ b264e4796ea32220a7c67fe84d990ac6e040784f7331c662e66ca5d91c71e9c1 first-pass-summary.json
9
+ 8e03b3fd76310ab80643f51ce791e85545062d2d85ac3c6121cc1143ebfb5a28 provenance.json
10
+ 954a6adff76d6bae31c77cd969b17fa5203ae7ba38abf76130c6f066e39a9c58 results.csv
11
+ b12427d5f4476351e571e4b29f61e0463540a1701bec51e642753d717dbc1df2 results.jsonl
12
+ 62936f76b1ca0de5e761d1ea1f894343695503ce934bf0213dc9fde1ea72a782 summary.json
results/deepseek-v4-pro/dsh/provenance.json ADDED
@@ -0,0 +1,454 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "arm": "dsh",
3
+ "kind": "dsh-public-recovery-provenance-v1",
4
+ "model": "deepseek-v4-pro",
5
+ "model_identity": {
6
+ "api_request_alias": "deepseek-v4-pro",
7
+ "conflicting_official_descriptions": [
8
+ {
9
+ "claim": "Announcement says alias routes to V4.1 Flash from 2026-09-14 12:00 Beijing time until V4.1 Pro launches.",
10
+ "url": "https://api-docs.deepseek.com/zh-cn/news/news260910/"
11
+ },
12
+ {
13
+ "claim": "Model table lists alias as DeepSeek-V4-Pro-0813.",
14
+ "url": "https://api-docs.deepseek.com/quick_start/pricing/"
15
+ }
16
+ ],
17
+ "interpretation": "Conflicting official descriptions; neither underlying weight version is independently established.",
18
+ "successful_response_model_label": "deepseek-v4-pro",
19
+ "underlying_model_weights_verified": false
20
+ },
21
+ "original_artifacts_modified": false,
22
+ "paid_calls_performed_for_export": 0,
23
+ "plan_id": "15486e52db982d80aaa0f6582f3d67469757919e5b8605e28ce7e98698d15ac3",
24
+ "plan_sha256": "4f903cda0cc6cd0f01f194d743d687eb9cc29fa84dcb7f1d68d35f76d9b21fe9",
25
+ "rescoring_performed": false,
26
+ "selected_internal_results_sha256": "12861e6fdfe58c5432b08c09fb9f14306c3b29c55e49cdc32dfb0fb1526b1343",
27
+ "selection": {
28
+ "best_of_attempts": false,
29
+ "infrastructure_tasks": 60,
30
+ "maximum_supplementary_attempts_per_task": 1,
31
+ "original_attempts": 1000,
32
+ "original_results_modified": false,
33
+ "policy": "Freeze eligibility before supplementary outcomes; select attempt 2 regardless of its score or status. Never best-of or fall back to a better attempt.",
34
+ "schema_amendment": {
35
+ "exact_prompt_parity": false,
36
+ "generation.prompt_profile": {
37
+ "from": "raw-role-schema",
38
+ "to": "raw-role"
39
+ },
40
+ "tools_and_per_attempt_budgets_changed": false
41
+ },
42
+ "schema_tasks": 27,
43
+ "selected_primary_attempts": 913,
44
+ "selected_supplementary_attempts": 87,
45
+ "supplementary_attempts": 87,
46
+ "total_attempts_charged": 1087
47
+ },
48
+ "source_hashes": {
49
+ "failure_audit": "1a75f87f904026051d0e3eca1cc3d8cfe68bbc6828e7de071b36e72be7d2ee47",
50
+ "legacy_flags_export": "3f22a49f736a0679aef8dc758521c10f9ec8e4505fbab689a4a2b5f28021831b",
51
+ "legacy_metrics_audit": "abb056be811678b198f8f95c8961c4766ee3543dec9493441946b8113a276baf",
52
+ "public_exporter": "ea1f1f27ae1f06ff0419bf60157222c9d6ffa6b6d58ded2b2e5f980ddfab151d",
53
+ "recovered_manifest": "89b41604d0f1fd5c8b35b5a41f2a83b24c27f2ea093444c64b57ab34a6216b0c"
54
+ },
55
+ "sources": {
56
+ "infrastructure": {
57
+ "bindings": [
58
+ {
59
+ "continuation_receipt_sha256": null,
60
+ "frozen_identity": {
61
+ "inputs.manifest.json": "58e63dddeed8f6bb39de2abd4c37dc1789fd101a876628229988cd24bdf5495d",
62
+ "protocol.json": "08e2cb843107daa4f85be7d2f9d7b21899b983b4e29a5b8e9193bce6e9bec058",
63
+ "schedule.json": "670c71226010832fe118390cb0cdec9475f4262bd2e3a235f44681b908bcdb75",
64
+ "study.json": "63694ebaf42f4387f264281d57214ecc1ba40742bc5f91a72d452cb8d0e400f7"
65
+ },
66
+ "run_id": "infrastructure",
67
+ "source_commit": "3ced22b417aea99b0956b504794561964dbf7b33",
68
+ "source_identity": {
69
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
70
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
71
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
72
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
73
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
74
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
75
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
76
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
77
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
78
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
79
+ "ntu_gifts/dsh/bridge.py": "f56ef6509cd831ec7de2c61b5331a7496771cf0754fcfaf0b29aeb88fa9be0f9",
80
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
81
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
82
+ "ntu_gifts/dsh/evidence.py": "bb0f5be46a80c4fcaf993e8071b22160fc04d5d871a9e446340a4041a9ef7591",
83
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
84
+ "ntu_gifts/dsh/failure_policy.py": "50f39ce6228e1eaca24278dd46774d13a39ecd2e7773a5c36e952ee70451083f",
85
+ "ntu_gifts/dsh/late_evidence.py": "5f6231a93f7668107f6d15501177013a64e27af42a904b8a8acba502edcab197",
86
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
87
+ "ntu_gifts/dsh/reporting.py": "823067b7a99f03cc682fb448aa26992b06b75ee15096567d93d85b31e715e9f9",
88
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
89
+ "ntu_gifts/dsh/study.py": "57c3495a97f12baae8aa2bb74c2289bfc5a7cbad0cc70205f34e66c734c28008",
90
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
91
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
92
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
93
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
94
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
95
+ }
96
+ },
97
+ {
98
+ "continuation_receipt_sha256": "775f69ede2c79404271e44fe2832251cad21ba233c5add0a6fb6fd31d5c6f5cf",
99
+ "frozen_identity": {
100
+ "inputs.manifest.json": "f325f884f3c1e8fe67934f73c1274c9e10713440e15fc397d93099b35a513833",
101
+ "protocol.json": "e85ff3bc2c4c8eef0deb48567a2e13642cd546bd64296621fe3d10640a629ffb",
102
+ "schedule.json": "d2cc1c2d9d23e6732e78879872970d927351c6da0d73ba9fe64a614c1a756370",
103
+ "study.json": "78afa10631db305ba2edc76dd5a4ee664fa2b71abc592b4745cfeae7d908b749"
104
+ },
105
+ "run_id": "infrastructure-auto01",
106
+ "source_commit": "3ced22b417aea99b0956b504794561964dbf7b33",
107
+ "source_identity": {
108
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
109
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
110
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
111
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
112
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
113
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
114
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
115
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
116
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
117
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
118
+ "ntu_gifts/dsh/bridge.py": "f56ef6509cd831ec7de2c61b5331a7496771cf0754fcfaf0b29aeb88fa9be0f9",
119
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
120
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
121
+ "ntu_gifts/dsh/evidence.py": "bb0f5be46a80c4fcaf993e8071b22160fc04d5d871a9e446340a4041a9ef7591",
122
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
123
+ "ntu_gifts/dsh/failure_policy.py": "50f39ce6228e1eaca24278dd46774d13a39ecd2e7773a5c36e952ee70451083f",
124
+ "ntu_gifts/dsh/late_evidence.py": "5f6231a93f7668107f6d15501177013a64e27af42a904b8a8acba502edcab197",
125
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
126
+ "ntu_gifts/dsh/reporting.py": "823067b7a99f03cc682fb448aa26992b06b75ee15096567d93d85b31e715e9f9",
127
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
128
+ "ntu_gifts/dsh/study.py": "57c3495a97f12baae8aa2bb74c2289bfc5a7cbad0cc70205f34e66c734c28008",
129
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
130
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
131
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
132
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
133
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
134
+ }
135
+ }
136
+ ],
137
+ "recovery_receipt_sha256": "d14de0ebd364f8ad6b2f1d639027bb7d5f80784a70ec76f626044d7651d9940c",
138
+ "report_sha256": "4746df24341152fc33b21fbb25eb43e5055e16835b53548aa5e3c471f235bd4f"
139
+ },
140
+ "primary": {
141
+ "bindings": [
142
+ {
143
+ "continuation_receipt_sha256": null,
144
+ "frozen_identity": {
145
+ "inputs.manifest.json": "f03642df73a8ceeca15f61d2ad329291b896f19148e134aa0cd9edfb14b04ccf",
146
+ "protocol.json": "04528e17dc02f945fb8d673925d66ddc0d51bd0f653751001b43f20dc80b263e",
147
+ "schedule.json": "7144a62e17ebfd6920a1de5f1bc49f4b610a81c95cdf4528dabf451df187da26",
148
+ "study.json": "8c1b83b2b8af544e78423e0ff3f76a1e0725065b465a424416bd14e2788494bc"
149
+ },
150
+ "run_id": "dsh-v4-pro-1000-20260918",
151
+ "source_commit": "5079157d791e2213e6f45a04f82182cc8c700f89",
152
+ "source_identity": {
153
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
154
+ "dsh/profile.patch.yml": "418253f372e18e8e64d57aab19c0d28bdb01ca0eb0a9092cbf1953602066edb6",
155
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
156
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
157
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
158
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
159
+ "dsh/runtime.py": "994b59ab0620f8f0210e89572cd053b7d26d347358c73ad105ddce2ccb8b3119",
160
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
161
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
162
+ "ntu_gifts/dsh/adapter.py": "cfd0cbea8fe52ff9a41077c8a8ffeead186ff42322b76aa34d77b2e32ca24815",
163
+ "ntu_gifts/dsh/bridge.py": "2b8a440717456012f51201a557f6ca4f86947a75ce5498cb0f473700d8ee612a",
164
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
165
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
166
+ "ntu_gifts/dsh/evidence.py": "a3ad48ec6163ddd2844265272f487d733a9f38f7ee886bf361ec240f5a69d224",
167
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
168
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
169
+ "ntu_gifts/dsh/reporting.py": "6fa53a97eec2d89c9eb2793e603564a7c00f65b8f6b5c15c4b043de28341ef0a",
170
+ "ntu_gifts/dsh/runner.py": "3ae5c3495e63fbf7266bfa071cc79a969fc8f3a8494e2fc86f98ec0f3d0ac534",
171
+ "ntu_gifts/dsh/study.py": "e0ba44f4fa4be343b1eae4c56f4c5db7e0424da350e371c5e93840883295683f",
172
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
173
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
174
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
175
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
176
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
177
+ }
178
+ },
179
+ {
180
+ "continuation_receipt_sha256": "36289c1785f2877a672f3128f6bc716910ca8316539d30a83c7df34bc160c69b",
181
+ "frozen_identity": {
182
+ "inputs.manifest.json": "522833a923004b38dfaa821fc0ea5f9c9f8370aee2e7c98a4741775e3c71e876",
183
+ "protocol.json": "bcb08cf8e47a5484629cecb7f897894f15ffffad8ff08695246c330711da476c",
184
+ "schedule.json": "aeff62de51a9e04a8c721af99fc31238dbf6436bf248953e871288b335e3f44f",
185
+ "study.json": "158716e4d5d5e4b6ac5015f35114d0bb338ed9db46934b6491d9993348b3a15f"
186
+ },
187
+ "run_id": "dsh-v4-pro-1000-cont01-20260918",
188
+ "source_commit": "2e26cfc6835a19792fcd09173fde1d5e2524c259",
189
+ "source_identity": {
190
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
191
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
192
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
193
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
194
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
195
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
196
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
197
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
198
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
199
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
200
+ "ntu_gifts/dsh/bridge.py": "4a1f2d2f9ab71a9f6575b1c3da4ca181f09c8d73e6eaac0fec8e00943e4cab7c",
201
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
202
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
203
+ "ntu_gifts/dsh/evidence.py": "a3ad48ec6163ddd2844265272f487d733a9f38f7ee886bf361ec240f5a69d224",
204
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
205
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
206
+ "ntu_gifts/dsh/reporting.py": "6fa53a97eec2d89c9eb2793e603564a7c00f65b8f6b5c15c4b043de28341ef0a",
207
+ "ntu_gifts/dsh/runner.py": "3ae5c3495e63fbf7266bfa071cc79a969fc8f3a8494e2fc86f98ec0f3d0ac534",
208
+ "ntu_gifts/dsh/study.py": "e0ba44f4fa4be343b1eae4c56f4c5db7e0424da350e371c5e93840883295683f",
209
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
210
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
211
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
212
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
213
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
214
+ }
215
+ },
216
+ {
217
+ "continuation_receipt_sha256": "454d4ee4568805c4149f9c03de2e5d7501e8a31e85d2a75d781b59b2b6fb4394",
218
+ "frozen_identity": {
219
+ "inputs.manifest.json": "2f14277588845c559e695c60499e3e37ad79274262bbc80d68660f5473609780",
220
+ "protocol.json": "fad42532a34c4f12fe7a2de80c1b69ad93b024e789d07c0861e828e52e301e56",
221
+ "schedule.json": "a7f1e2d5616efdc5c975ed941b0ff91240f39585130ad38057b115109802b53b",
222
+ "study.json": "63b6ea5de0bc970b913942f88d7e5568f45c336075619177471eeb02287b6192"
223
+ },
224
+ "run_id": "dsh-v4-pro-1000-cont02-20260919",
225
+ "source_commit": "88dfd135a6258d235a3d496c9e8044c4fec6ef2e",
226
+ "source_identity": {
227
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
228
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
229
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
230
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
231
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
232
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
233
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
234
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
235
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
236
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
237
+ "ntu_gifts/dsh/bridge.py": "a2aae1fce301be15cf57faea045800ce8f311fbfc65968807c3da9218221cbb2",
238
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
239
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
240
+ "ntu_gifts/dsh/evidence.py": "a3ad48ec6163ddd2844265272f487d733a9f38f7ee886bf361ec240f5a69d224",
241
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
242
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
243
+ "ntu_gifts/dsh/reporting.py": "6fa53a97eec2d89c9eb2793e603564a7c00f65b8f6b5c15c4b043de28341ef0a",
244
+ "ntu_gifts/dsh/runner.py": "3ae5c3495e63fbf7266bfa071cc79a969fc8f3a8494e2fc86f98ec0f3d0ac534",
245
+ "ntu_gifts/dsh/study.py": "e0ba44f4fa4be343b1eae4c56f4c5db7e0424da350e371c5e93840883295683f",
246
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
247
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
248
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
249
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
250
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
251
+ }
252
+ },
253
+ {
254
+ "continuation_receipt_sha256": "c43ec808dfa7a0f2dc7eb6ef46749648f8642e1086e6724341b36ddca39eb637",
255
+ "frozen_identity": {
256
+ "inputs.manifest.json": "98566f98dee3eda6f9ed372cb16f2f7370829b0fedc83c9f95c80d2be9a25ad0",
257
+ "protocol.json": "e80906d5329ed81db3d7e5afeae6480a512a40961a8768fa3d5974e2026e4413",
258
+ "schedule.json": "4c5f6a7539534a64b27438d49cfa50b259d89be3dea91e83e4b61fdc2aac57e7",
259
+ "study.json": "9c6d9f22e04b1a8b5535e35d60119213476cc5d7b296d0247fab18f5e3a974a2"
260
+ },
261
+ "run_id": "dsh-v4-pro-1000-cont02-20260919-auto01",
262
+ "source_commit": "88dfd135a6258d235a3d496c9e8044c4fec6ef2e",
263
+ "source_identity": {
264
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
265
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
266
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
267
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
268
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
269
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
270
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
271
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
272
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
273
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
274
+ "ntu_gifts/dsh/bridge.py": "a2aae1fce301be15cf57faea045800ce8f311fbfc65968807c3da9218221cbb2",
275
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
276
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
277
+ "ntu_gifts/dsh/evidence.py": "a3ad48ec6163ddd2844265272f487d733a9f38f7ee886bf361ec240f5a69d224",
278
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
279
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
280
+ "ntu_gifts/dsh/reporting.py": "6fa53a97eec2d89c9eb2793e603564a7c00f65b8f6b5c15c4b043de28341ef0a",
281
+ "ntu_gifts/dsh/runner.py": "3ae5c3495e63fbf7266bfa071cc79a969fc8f3a8494e2fc86f98ec0f3d0ac534",
282
+ "ntu_gifts/dsh/study.py": "e0ba44f4fa4be343b1eae4c56f4c5db7e0424da350e371c5e93840883295683f",
283
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
284
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
285
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
286
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
287
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
288
+ }
289
+ },
290
+ {
291
+ "continuation_receipt_sha256": "e59e8c8f8fe0cc912afbda777fe529919480e360641dcbfbe2e9e49cb259b910",
292
+ "frozen_identity": {
293
+ "inputs.manifest.json": "8014f57fc0ed0a4ffd656803b8224d221a0b4e47972248fc58d97dd894ff9af4",
294
+ "protocol.json": "04da83504c6f92d1caccaec0642f964cd0db2ae13816c69606bc749875e8c822",
295
+ "schedule.json": "a9387ccfc5201ef3b0b39525241dd655710734239275b4ae0d33528265389098",
296
+ "study.json": "a854e9dbad49b0350009b757bd0d843550ac543159e9bf87b65fab00689de254"
297
+ },
298
+ "run_id": "dsh-v4-pro-1000-cont03-20260919",
299
+ "source_commit": "e6c90cb95a9c11592d2dd6107304cf8a66c072c4",
300
+ "source_identity": {
301
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
302
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
303
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
304
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
305
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
306
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
307
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
308
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
309
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
310
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
311
+ "ntu_gifts/dsh/bridge.py": "a2aae1fce301be15cf57faea045800ce8f311fbfc65968807c3da9218221cbb2",
312
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
313
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
314
+ "ntu_gifts/dsh/evidence.py": "a3ad48ec6163ddd2844265272f487d733a9f38f7ee886bf361ec240f5a69d224",
315
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
316
+ "ntu_gifts/dsh/failure_policy.py": "ad3f5c6f5dc4b6a4dce7d2373b0f155293826bbd61caa4a0cbb06cd475dd8764",
317
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
318
+ "ntu_gifts/dsh/reporting.py": "6fa53a97eec2d89c9eb2793e603564a7c00f65b8f6b5c15c4b043de28341ef0a",
319
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
320
+ "ntu_gifts/dsh/study.py": "e0ba44f4fa4be343b1eae4c56f4c5db7e0424da350e371c5e93840883295683f",
321
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
322
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
323
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
324
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
325
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
326
+ }
327
+ },
328
+ {
329
+ "continuation_receipt_sha256": "19b3accee1764accedf3bc27e764d655bd69ab4c46f3fa8583c5853d77065039",
330
+ "frozen_identity": {
331
+ "inputs.manifest.json": "f62e5d69edd087491fa54177903ed8fd0839f435bc290b0bc85461e95bb8c409",
332
+ "protocol.json": "acdd1263b564964b2aa91c8801db32351492d694be5bcc65a82e0a31a879ece2",
333
+ "schedule.json": "ffbfe1f4b6e21cf869947e9d251afdb5dec86c10bcfbfd3d8fda53464a93c987",
334
+ "study.json": "a9258462247e545bf96b4d187b4f89108660cd41654947ce630b5198f7955069"
335
+ },
336
+ "run_id": "dsh-v4-pro-1000-cont04-20260919",
337
+ "source_commit": "a1d3d8cbd349104bf292e03226b336faadcd461d",
338
+ "source_identity": {
339
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
340
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
341
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
342
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
343
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
344
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
345
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
346
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
347
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
348
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
349
+ "ntu_gifts/dsh/bridge.py": "92a3d3685609ef7823294fdb43eae24d92e3a75373508a28ad21d1de561d88e7",
350
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
351
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
352
+ "ntu_gifts/dsh/evidence.py": "a3ad48ec6163ddd2844265272f487d733a9f38f7ee886bf361ec240f5a69d224",
353
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
354
+ "ntu_gifts/dsh/failure_policy.py": "50f39ce6228e1eaca24278dd46774d13a39ecd2e7773a5c36e952ee70451083f",
355
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
356
+ "ntu_gifts/dsh/reporting.py": "6fa53a97eec2d89c9eb2793e603564a7c00f65b8f6b5c15c4b043de28341ef0a",
357
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
358
+ "ntu_gifts/dsh/study.py": "e0ba44f4fa4be343b1eae4c56f4c5db7e0424da350e371c5e93840883295683f",
359
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
360
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
361
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
362
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
363
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
364
+ }
365
+ },
366
+ {
367
+ "continuation_receipt_sha256": "cbe103993cf9c9e66f8971cda0e19adce752964577d468f192747d0c34c1c233",
368
+ "frozen_identity": {
369
+ "inputs.manifest.json": "373f9f9b230367f43ad862c448743d2f78ba6211686fb065f8078ac14985f568",
370
+ "protocol.json": "3b0ef361824b426e08c6a6aec0c959314c0bfc2c600b8eec6db7ab36e6060017",
371
+ "schedule.json": "cb0f1f05568a79b1ffdf94ef9866422942879bba91e099ea1064021800f7e0aa",
372
+ "study.json": "4cdd10f36498d393fa216c4dcfad667983b3a438386a60c25c92bd35c10a3df6"
373
+ },
374
+ "run_id": "dsh-v4-pro-1000-cont05-20260919",
375
+ "source_commit": "3ced22b417aea99b0956b504794561964dbf7b33",
376
+ "source_identity": {
377
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
378
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
379
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
380
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
381
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
382
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
383
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
384
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
385
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
386
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
387
+ "ntu_gifts/dsh/bridge.py": "f56ef6509cd831ec7de2c61b5331a7496771cf0754fcfaf0b29aeb88fa9be0f9",
388
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
389
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
390
+ "ntu_gifts/dsh/evidence.py": "bb0f5be46a80c4fcaf993e8071b22160fc04d5d871a9e446340a4041a9ef7591",
391
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
392
+ "ntu_gifts/dsh/failure_policy.py": "50f39ce6228e1eaca24278dd46774d13a39ecd2e7773a5c36e952ee70451083f",
393
+ "ntu_gifts/dsh/late_evidence.py": "5f6231a93f7668107f6d15501177013a64e27af42a904b8a8acba502edcab197",
394
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
395
+ "ntu_gifts/dsh/reporting.py": "823067b7a99f03cc682fb448aa26992b06b75ee15096567d93d85b31e715e9f9",
396
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
397
+ "ntu_gifts/dsh/study.py": "57c3495a97f12baae8aa2bb74c2289bfc5a7cbad0cc70205f34e66c734c28008",
398
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
399
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
400
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
401
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
402
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
403
+ }
404
+ }
405
+ ],
406
+ "report_sha256": "f8f0c4f5b40578f15b40be5b1965ff1a7c4007ca4154b8e95da82e788a3e1c3a"
407
+ },
408
+ "schema": {
409
+ "bindings": [
410
+ {
411
+ "continuation_receipt_sha256": null,
412
+ "frozen_identity": {
413
+ "inputs.manifest.json": "38ce82733f9e2d7a9551ef3ca46a8799579b3d45337c9cf7e4d4d3354322eb63",
414
+ "protocol.json": "63d6e27ba6ac078fcfa809380f926b30b774ed4c70bd64420c406ea190f413a5",
415
+ "schedule.json": "cde6543225e535909044c596d0e0a969d72ff6742eca60ea38dd5b02f83a7daa",
416
+ "study.json": "ef2cdb869ccf2a4618ef829570a3776292be47b6e9e30c115cb5c3b7edce8dbb"
417
+ },
418
+ "run_id": "schema",
419
+ "source_commit": "3ced22b417aea99b0956b504794561964dbf7b33",
420
+ "source_identity": {
421
+ "dsh/__init__.py": "8e24d82c627405e0a8f9de89ebdbf1ce23db63337ef44aacd4a789cea9697908",
422
+ "dsh/profile.patch.yml": "81a92895c5f3c4c85bb602e3349a509262eb144c63a566331a43fa73081e2731",
423
+ "dsh/requirements-controller.txt": "aee6d7d23cbeecc3a101ec56687af9cc7aecbd6e96443295042a5c36307587d4",
424
+ "dsh/requirements.lock.txt": "ed57fde608bf1cdabaa1fc5d23e5b254983606652be2f732ce4ed2f2c69ea7a1",
425
+ "dsh/run-python.mjs": "3681bcbdfc6fa10cbb0742fadcc8f3f94a745b428662c94214b370b01bbea67e",
426
+ "dsh/runtime.lock.json": "e241e17af63e1f2b5ebaab3845a40c8f35c8643c3b5622b778541663428aba44",
427
+ "dsh/runtime.py": "2223a5573a3bb6c7fec46e264920f5cacf3dedaffb48f6a732e7f4371380c5e7",
428
+ "ntu_gifts/dsh/__init__.py": "11855c8b3a9276249180fa9ddf0f4d967d21140d57fff4bc4a638642864483ff",
429
+ "ntu_gifts/dsh/__main__.py": "b54ebc80f16c01be4b84a44474b3567e61acdf1e694813764869060ed002a762",
430
+ "ntu_gifts/dsh/adapter.py": "c0184cbecaba44c58d2779ac4972e33912b4c5235b00d52b20ae0eacb1aa93c9",
431
+ "ntu_gifts/dsh/bridge.py": "f56ef6509cd831ec7de2c61b5331a7496771cf0754fcfaf0b29aeb88fa9be0f9",
432
+ "ntu_gifts/dsh/bwrap_payload.py": "02b1eb6067e4ab4d283d98148bd6b260bcc9a0ac465dfd490acb93f7ab79542e",
433
+ "ntu_gifts/dsh/dependencies.py": "e0618ffbb448082e1b5ee5eb6a3ca384375f0ddbf042be72726ee26053227254",
434
+ "ntu_gifts/dsh/evidence.py": "bb0f5be46a80c4fcaf993e8071b22160fc04d5d871a9e446340a4041a9ef7591",
435
+ "ntu_gifts/dsh/execution_bwrap.py": "f37c3818c3b38bad1de8885c02f1c644a2e9abf7abe9b31a2823ee4ae941b632",
436
+ "ntu_gifts/dsh/failure_policy.py": "50f39ce6228e1eaca24278dd46774d13a39ecd2e7773a5c36e952ee70451083f",
437
+ "ntu_gifts/dsh/late_evidence.py": "5f6231a93f7668107f6d15501177013a64e27af42a904b8a8acba502edcab197",
438
+ "ntu_gifts/dsh/protocol.py": "9467ea64de8a60f60988c2dd2ddd804efea53a9bedab8e6c1dfa28ef8b963a5d",
439
+ "ntu_gifts/dsh/reporting.py": "823067b7a99f03cc682fb448aa26992b06b75ee15096567d93d85b31e715e9f9",
440
+ "ntu_gifts/dsh/runner.py": "082f3878ee2c24ecf15b0b5229e5feb4eab801dbca78fa1fc1c14aa09c8ed42e",
441
+ "ntu_gifts/dsh/study.py": "57c3495a97f12baae8aa2bb74c2289bfc5a7cbad0cc70205f34e66c734c28008",
442
+ "ntu_gifts/dsh/worker.py": "4d5ed66f60161f19e197bf702d4acee3aee37a9bfb99ef0a52a448b0de76810f",
443
+ "ntu_gifts/harness/budget.py": "84e27094f08676328c0899c2abd14097238efde43a5f38c110f28d32bcb2c385",
444
+ "ntu_gifts/harness/study.py": "7264540b065df45f76d4d6705bf8485f4a9d778b900035efbcfe39a7520abc6c",
445
+ "ntu_gifts/workspace.py": "2b5295ade967f7b586ec8b44d4f55a31efb50a27cf75290ada0ee00e9c756346",
446
+ "scripts/setup_dsh.py": "fe835598584191d1a10c680afd769201efc6a7cf689967f681d8bc64c0b914e4"
447
+ }
448
+ }
449
+ ],
450
+ "recovery_receipt_sha256": "a3b879571ae3b3c95aeeb3fe7a3bf59ed775dd5bc19272b6f20e2b899d42b1f3",
451
+ "report_sha256": "53f433a20463235254142f509580b7afb8640caeb282a63b629f2551738c2fca"
452
+ }
453
+ }
454
+ }
results/deepseek-v4-pro/dsh/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/dsh/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/dsh/summary.json ADDED
@@ -0,0 +1,175 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "all_slots_sealed": true,
3
+ "arm": "dsh",
4
+ "assumptions": {
5
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy.",
6
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
7
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
8
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
9
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
10
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
11
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
12
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
13
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect."
14
+ },
15
+ "complete": false,
16
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
17
+ "definitions": {
18
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
19
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
20
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
21
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
22
+ },
23
+ "display_name": "DeepSeek API (deepseek-v4-pro alias) + DSH",
24
+ "display_role": "standalone_dsh_harness_entry",
25
+ "evidence_class": "technical_recovery_view",
26
+ "generated_at": "2026-09-19T18:18:48.070003+00:00",
27
+ "kind": "inferencenet-offline-leaderboard-export",
28
+ "legacy_local_paper": {
29
+ "definitions": {
30
+ "coefficient_only": "coefficient relative error <= .05",
31
+ "direction": "matching strictly positive or strictly negative coefficient signs",
32
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
33
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
34
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
35
+ },
36
+ "metrics": {
37
+ "coefficient_only": {
38
+ "assessable": 695,
39
+ "count": 555,
40
+ "coverage_percent": 69.5,
41
+ "denominator": 1000,
42
+ "failure_count": 140,
43
+ "rate": 0.555,
44
+ "score": 55.5,
45
+ "unknown_count": 305
46
+ },
47
+ "direction": {
48
+ "assessable": 695,
49
+ "count": 680,
50
+ "coverage_percent": 69.5,
51
+ "denominator": 1000,
52
+ "failure_count": 15,
53
+ "rate": 0.68,
54
+ "score": 68.0,
55
+ "unknown_count": 305
56
+ },
57
+ "partial": {
58
+ "assessable": 695,
59
+ "count": 460,
60
+ "coverage_percent": 69.5,
61
+ "denominator": 1000,
62
+ "failure_count": 235,
63
+ "rate": 0.46,
64
+ "score": 46.0,
65
+ "unknown_count": 305
66
+ },
67
+ "perfect": {
68
+ "assessable": 695,
69
+ "count": 372,
70
+ "coverage_percent": 69.5,
71
+ "denominator": 1000,
72
+ "failure_count": 323,
73
+ "rate": 0.372,
74
+ "score": 37.2,
75
+ "unknown_count": 305
76
+ },
77
+ "significance": {
78
+ "assessable": 695,
79
+ "count": 591,
80
+ "coverage_percent": 69.5,
81
+ "denominator": 1000,
82
+ "failure_count": 104,
83
+ "rate": 0.591,
84
+ "score": 59.1,
85
+ "unknown_count": 305
86
+ }
87
+ },
88
+ "profile": "local-paper-v1",
89
+ "verification": "Original sealed scoring.json flags; no rescoring."
90
+ },
91
+ "metric_profile": "hf-leaderboard-v1",
92
+ "model": "deepseek-v4-pro",
93
+ "model_alias": "deepseek-v4-pro",
94
+ "official_parity_verified": false,
95
+ "original_protocol_complete": false,
96
+ "publication_kind": "research_results_with_technical_recovery",
97
+ "result": {
98
+ "columns": [
99
+ "Model ID",
100
+ "Compilation Success",
101
+ "Partial Replication",
102
+ "Correct Coefficient Direction",
103
+ "Significant Level Correctness"
104
+ ],
105
+ "definite_outcome_count": 970,
106
+ "expected_count": 1000,
107
+ "invalid_evidence_count": 0,
108
+ "metrics": {
109
+ "coefficient_direction": {
110
+ "assessable": 698,
111
+ "count": 683,
112
+ "coverage_percent": 69.8,
113
+ "denominator": 1000,
114
+ "failure_count": 15,
115
+ "rate": 0.683,
116
+ "score": 68.3,
117
+ "unknown_count": 302
118
+ },
119
+ "compilation_success": {
120
+ "assessable": 970,
121
+ "count": 699,
122
+ "coverage_percent": 97.0,
123
+ "denominator": 1000,
124
+ "failure_count": 271,
125
+ "rate": 0.699,
126
+ "score": 69.9,
127
+ "unknown_count": 30
128
+ },
129
+ "partial_replication": {
130
+ "assessable": 698,
131
+ "count": 556,
132
+ "coverage_percent": 69.8,
133
+ "denominator": 1000,
134
+ "failure_count": 142,
135
+ "rate": 0.556,
136
+ "score": 55.6,
137
+ "unknown_count": 302
138
+ },
139
+ "significance_level": {
140
+ "assessable": 696,
141
+ "count": 598,
142
+ "coverage_percent": 69.6,
143
+ "denominator": 1000,
144
+ "failure_count": 98,
145
+ "rate": 0.598,
146
+ "score": 59.8,
147
+ "unknown_count": 304
148
+ }
149
+ },
150
+ "resolved_count": 1000,
151
+ "row": {
152
+ "Compilation Success": 69.9,
153
+ "Correct Coefficient Direction": 68.3,
154
+ "Model ID": "DeepSeek API (deepseek-v4-pro alias) + DSH",
155
+ "Partial Replication": 55.6,
156
+ "Significant Level Correctness": 59.8
157
+ },
158
+ "sealed_count": 1000,
159
+ "task_counts": {
160
+ "completed": 1000,
161
+ "failed": 271,
162
+ "invalid_evidence": 0,
163
+ "not_started": 0,
164
+ "planned": 1000,
165
+ "succeeded": 699,
166
+ "unknown": 30,
167
+ "unsealed": 0
168
+ },
169
+ "task_unknown_count": 30,
170
+ "valid_prediction_count": 699
171
+ },
172
+ "schema_version": 3,
173
+ "underlying_model_weights_verified": false,
174
+ "verification_scope": "Preserved verified flags and result/score hashes; no new inference, execution or rescoring."
175
+ }
results/index.json CHANGED
@@ -1,81 +1,87 @@
1
- {
2
- "schema_version": 1,
3
- "kind": "results_archive",
4
- "archived_at_utc": "2026-09-15T07:29:39.569697+00:00",
5
- "displayed_leaderboard_updated": false,
6
- "entries": [
7
- {
8
- "model": "gpt-5.6-sol",
9
- "arm": "baseline",
10
- "rows": 1000,
11
- "path": "results/gpt-5.6-sol/baseline"
12
- },
13
- {
14
- "model": "gpt-5.6-sol",
15
- "arm": "deepagents",
16
- "rows": 1000,
17
- "path": "results/gpt-5.6-sol/deepagents"
18
- },
19
- {
20
- "model": "claude-opus-4-8",
21
- "arm": "baseline",
22
- "rows": 1000,
23
- "path": "results/claude-opus-4-8/baseline"
24
- },
25
- {
26
- "model": "claude-opus-4-8",
27
- "arm": "deepagents",
28
- "rows": 1000,
29
- "path": "results/claude-opus-4-8/deepagents"
30
- },
31
- {
32
- "model": "kimi-k3",
33
- "arm": "baseline",
34
- "rows": 1000,
35
- "path": "results/kimi-k3/baseline"
36
- },
37
- {
38
- "model": "kimi-k3",
39
- "arm": "deepagents",
40
- "rows": 1000,
41
- "path": "results/kimi-k3/deepagents"
42
- },
43
- {
44
- "model": "gemini-3.1-pro-preview",
45
- "arm": "baseline",
46
- "rows": 1000,
47
- "path": "results/gemini-3.1-pro-preview/baseline"
48
- },
49
- {
50
- "model": "gemini-3.1-pro-preview",
51
- "arm": "deepagents",
52
- "rows": 1000,
53
- "path": "results/gemini-3.1-pro-preview/deepagents"
54
- },
55
- {
56
- "model": "qwen3.7-max",
57
- "arm": "baseline",
58
- "rows": 1000,
59
- "path": "results/qwen3.7-max/baseline"
60
- },
61
- {
62
- "model": "qwen3.7-max",
63
- "arm": "deepagents",
64
- "rows": 1000,
65
- "path": "results/qwen3.7-max/deepagents"
66
- },
67
- {
68
- "model": "deepseek-v4-pro",
69
- "arm": "baseline",
70
- "rows": 1000,
71
- "path": "results/deepseek-v4-pro/baseline"
72
- },
73
- {
74
- "model": "deepseek-v4-pro",
75
- "arm": "deepagents",
76
- "rows": 1000,
77
- "path": "results/deepseek-v4-pro/deepagents"
78
- }
79
- ],
80
- "updated_at_utc": "2026-09-16T17:18:47.373930+00:00"
81
- }
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 1,
3
+ "kind": "results_archive",
4
+ "archived_at_utc": "2026-09-15T07:29:39.569697+00:00",
5
+ "displayed_leaderboard_updated": false,
6
+ "entries": [
7
+ {
8
+ "model": "gpt-5.6-sol",
9
+ "arm": "baseline",
10
+ "rows": 1000,
11
+ "path": "results/gpt-5.6-sol/baseline"
12
+ },
13
+ {
14
+ "model": "gpt-5.6-sol",
15
+ "arm": "deepagents",
16
+ "rows": 1000,
17
+ "path": "results/gpt-5.6-sol/deepagents"
18
+ },
19
+ {
20
+ "model": "claude-opus-4-8",
21
+ "arm": "baseline",
22
+ "rows": 1000,
23
+ "path": "results/claude-opus-4-8/baseline"
24
+ },
25
+ {
26
+ "model": "claude-opus-4-8",
27
+ "arm": "deepagents",
28
+ "rows": 1000,
29
+ "path": "results/claude-opus-4-8/deepagents"
30
+ },
31
+ {
32
+ "model": "kimi-k3",
33
+ "arm": "baseline",
34
+ "rows": 1000,
35
+ "path": "results/kimi-k3/baseline"
36
+ },
37
+ {
38
+ "model": "kimi-k3",
39
+ "arm": "deepagents",
40
+ "rows": 1000,
41
+ "path": "results/kimi-k3/deepagents"
42
+ },
43
+ {
44
+ "model": "gemini-3.1-pro-preview",
45
+ "arm": "baseline",
46
+ "rows": 1000,
47
+ "path": "results/gemini-3.1-pro-preview/baseline"
48
+ },
49
+ {
50
+ "model": "gemini-3.1-pro-preview",
51
+ "arm": "deepagents",
52
+ "rows": 1000,
53
+ "path": "results/gemini-3.1-pro-preview/deepagents"
54
+ },
55
+ {
56
+ "model": "qwen3.7-max",
57
+ "arm": "baseline",
58
+ "rows": 1000,
59
+ "path": "results/qwen3.7-max/baseline"
60
+ },
61
+ {
62
+ "model": "qwen3.7-max",
63
+ "arm": "deepagents",
64
+ "rows": 1000,
65
+ "path": "results/qwen3.7-max/deepagents"
66
+ },
67
+ {
68
+ "model": "deepseek-v4-pro",
69
+ "arm": "baseline",
70
+ "rows": 1000,
71
+ "path": "results/deepseek-v4-pro/baseline"
72
+ },
73
+ {
74
+ "model": "deepseek-v4-pro",
75
+ "arm": "deepagents",
76
+ "rows": 1000,
77
+ "path": "results/deepseek-v4-pro/deepagents"
78
+ },
79
+ {
80
+ "model": "deepseek-v4-pro",
81
+ "arm": "dsh",
82
+ "rows": 1000,
83
+ "path": "results/deepseek-v4-pro/dsh"
84
+ }
85
+ ],
86
+ "updated_at_utc": "2026-09-19T18:47:27.943320+00:00"
87
+ }
results/manifest.sha256 CHANGED
@@ -1,80 +1,93 @@
1
- 4ca82d5601f12b48b24e22beb1ef624a17f1369074d296a6151bfd6f4a5d0668 README.md
2
- 2ad22f4778525b1460f1857e1095f2eea5620063fba1ffa6314ca7fd50f14703 claude-opus-4-8/baseline/README.md
3
- 8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56 claude-opus-4-8/baseline/config.json
4
- 4d5f859e472b79195c9be6b1f0e9c5feecf59a81dfe32ac08eb4c1d4967ed98e claude-opus-4-8/baseline/results.csv
5
- b9d6bd51c9eb81a634e78e422c86da16f655d97da87d3f5afa4c9aa0f96809f4 claude-opus-4-8/baseline/results.jsonl
6
- 7b5b06bd9b3d7b47952b1dc424a01fcdf04ddccca0b3f64d04132358d6380cdc claude-opus-4-8/baseline/summary.json
7
- dfb00967c2a2ff320d5a6d8c142cbc6ef37d4fb144abcc7a451673515c2b4c50 claude-opus-4-8/deepagents/README.md
8
- 12215dfcba5d2d692268a305364514a03ff7579094b2d4af12f493383b8e0f21 claude-opus-4-8/deepagents/config.json
9
- bafc0a72cd4d8f53c1da33030250171f36ce441488bc8456cba8c3f52b80e6bd claude-opus-4-8/deepagents/results.csv
10
- 3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a claude-opus-4-8/deepagents/results.jsonl
11
- f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614 claude-opus-4-8/deepagents/summary.json
12
- c738e75419e319515b15c823300a16e9bdf1cbe790612bbbc24ff93c71438b3f deepseek-v4-pro/baseline/README.md
13
- 4cef524f37a325029ea4ad325b163f72f9cf725f507f20cd76424d06e71e24ad deepseek-v4-pro/baseline/config.json
14
- 98f4f77003f4788509360a2cbe41554279df408fd20d65058b85836a25858ecb deepseek-v4-pro/baseline/results.csv
15
- 7827218f6f18f39988e8c439f0c93cab611c04d158d8704aa50c9274cd438837 deepseek-v4-pro/baseline/results.jsonl
16
- c51e4f1de78eba8f30e4a59276fb7821132cd103f0ae3f9d04c75c98344f2e34 deepseek-v4-pro/baseline/summary.json
17
- f1297897da23ab01751289cd93da69a944834af59204cd73f7fca64a112bd6a9 deepseek-v4-pro/deepagents/README.md
18
- 400c28e4008d5241bf50bc71c2674861846d6b6d2e7da21f2a7602a60ecc0ae9 deepseek-v4-pro/deepagents/config.json
19
- a842bac2b7b608ab124d95883194e01700c6fba648593a2128db6f9d40511cf8 deepseek-v4-pro/deepagents/results.csv
20
- a006c8b0b80064295faeb3aa7395a2408f0208677853448816b361bf67e39bb6 deepseek-v4-pro/deepagents/results.jsonl
21
- 231e4264b0403e12788136d01af465f86fe9ba3c3ab992510925a7ef4c2a40e9 deepseek-v4-pro/deepagents/summary.json
22
- eb38fada3bb400d40e37fd97aeb41c8eea6a1475406395e12ad9ac1c14f71e34 deepseek-v4-pro/diagnostics.json
23
- d811ae94edc6589e2b7ee8c39ac6c4da1c66b5c3dad8acbc589b17767f9c5690 deepseek-v4-pro/export-verification.json
24
- 891a8e0c9e1be24652b56d477bbdc0c98cf961c18010177535d33395dc43a132 deepseek-v4-pro/paired_comparison.csv
25
- 8be2d803fc96ad6a7840c0c3eb52b52e0e1d8e8523f6a4349b4adfbdc4c320f3 deepseek-v4-pro/provenance.json
26
- ac7253c934e43a88d0a2f6b105d1bb13424ef004e6838d2a632d5e9c366c72f2 gemini-3.1-pro-preview/baseline/README.md
27
- ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998 gemini-3.1-pro-preview/baseline/config.json
28
- fc708f327a76b22ebe6f453206b0c42e4cf46176afb8cb12eda2c1b71fe90bde gemini-3.1-pro-preview/baseline/results.csv
29
- 2f1e2019157d105efee7ba311a49eba99abe4a7b84f0a4b4e0ac174e2d56cdba gemini-3.1-pro-preview/baseline/results.jsonl
30
- a8ba5100a28f1f0aeae6949eb3669c7717c844c4e2eee6fa42ded351aef6b5c7 gemini-3.1-pro-preview/baseline/summary.json
31
- 49ff37264a63f3b5615632213bb9248d2b6b8feb66b4ec989ed655b45e5faed2 gemini-3.1-pro-preview/deepagents/README.md
32
- f09d9d99d26328ae2f54e9173b22f69d68f7d7bb50c3cc5e352c996b98dd57f9 gemini-3.1-pro-preview/deepagents/config.json
33
- af16c74fdc1163682a9a98f597ac0d101bc0de64d60c6a8582f69312948c8069 gemini-3.1-pro-preview/deepagents/results.csv
34
- c2824f8277b283a6dd60118b6e6a0cb0419e4c8b9193de15a7bcf7aaaad24181 gemini-3.1-pro-preview/deepagents/results.jsonl
35
- bf9ce721c94103f91b0ca576f717f4f5dca7cc9965601658a65fc6d2a15b7045 gemini-3.1-pro-preview/deepagents/summary.json
36
- 67b24d52eeabbcd84baf00c7395e72d4efe7cfcc82d7de1b0f8ee30f8018a1f7 gemini-3.1-pro-preview/diagnostics.json
37
- 885b4de7e83334c416f33ea1535a72f1f70dfe61834d1b622e01ddb3cccfaed4 gemini-3.1-pro-preview/export-verification.json
38
- 930a9ae67dc0dd4210e7e0c96098cab24899c765e4b3d041f75455f58fc072fe gemini-3.1-pro-preview/paired_comparison.csv
39
- 84378a81ca5533439b21dc1cf683925f5accc23532d392d140a4477e59c4edfb gemini-3.1-pro-preview/provenance.json
40
- ed22b41edf2afe7a500f4caf1b61f0b4c6e6c73759522c543fb86ce478eb4d04 gpt-5.6-sol/baseline/README.md
41
- 90ec4e0781b57498fef52de31eacf5c8ec9424549643865a7140a1ccdb04c1d4 gpt-5.6-sol/baseline/config.json
42
- a54171a3d3e7ecda3e61311414a84131e74a680bfa1342e8627cfaa32fc7c873 gpt-5.6-sol/baseline/results.csv
43
- 3b12160382043e27531ff5cb3e23bccc9a8f7807e4121a82ffa13df28f65a859 gpt-5.6-sol/baseline/results.jsonl
44
- 6d6ab1aaffe754e08f9656dcb5f51fcd5c772e60660dc34d86cfc21f3c5c03d7 gpt-5.6-sol/baseline/summary.json
45
- 99a2c4edc5f8a2984e09211c7608f08cc29131e37241cc79b6c77b774d3b21fe gpt-5.6-sol/deepagents/README.md
46
- a3e24175f2395cffbbcd6d9d3c89689bd3a3b22dd4ce61bb78d5da33c6b6cc1b gpt-5.6-sol/deepagents/config.json
47
- 27abe57119d4be49c56b29c318a1bc2757a97f44427764d7fc8d188be4d10671 gpt-5.6-sol/deepagents/results.csv
48
- 378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be gpt-5.6-sol/deepagents/results.jsonl
49
- 821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5 gpt-5.6-sol/deepagents/summary.json
50
- b6376d43e3107a6ac45647cdd1b0fa382437f92e33655715146d4cf6e89467d5 index.json
51
- 29825997a250cc8fe292450006b82a392fe90c8e0493a1335f74eea72c1de104 kimi-k3/baseline/README.md
52
- 7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081 kimi-k3/baseline/config.json
53
- f87d2dd3b390695f6b420a78e126b504abc3d292180a3443deaead174ec1b464 kimi-k3/baseline/results.csv
54
- 0502ae8892dd6c66ee92890a83c678527cff6c3396aa4c13faaadcf41e1857c1 kimi-k3/baseline/results.jsonl
55
- 2edd2e96f3b448ceceb6189edb7f23c737b245c00c9820728c69f5237feb34b7 kimi-k3/baseline/summary.json
56
- d6fe2ba624876d49373d4a0c180e0c27067405501b247e59c2d8ca5c74427ff7 kimi-k3/deepagents/README.md
57
- 312c9260c01d5988878e4f2b88da55a16362ee994e8217c63b15ba772da3857c kimi-k3/deepagents/config.json
58
- 572c3270a3ec255544b587aada9d8bd56cc2555d5246ce26feea2a3d35eb8f94 kimi-k3/deepagents/results.csv
59
- bdab99f8f1d81885f74326cf9c8c8e57b9bb8a8d616bfbb1b5a2f6bf32b737ca kimi-k3/deepagents/results.jsonl
60
- 81c8cabaa58acc65ef6afba792241b70951b666ab90c3f7c25194b61b0d9a228 kimi-k3/deepagents/summary.json
61
- bfaac06563fc53d5f564197fdfb69e24ddbdff622b85d67bbf2a13fb7fb5d9ba kimi-k3/diagnostics.json
62
- b96e0cc3352efac6140c33cef30873e522db547891101c3c49c6aca42085becb kimi-k3/export-verification.json
63
- 3fe1378a23cfe59a37558521ec5bf47ea299a6ad4903a14684147a1b8066ee03 kimi-k3/paired_comparison.csv
64
- 4a0f6dde38c5410a6f379192f30eac6d330403c80fcdbc2b357dbe15427a331b kimi-k3/provenance.json
65
- 1606c3170bec39e7c074659fd8c8d011e02e98b1e69161338554a9e1c416a363 qwen3.7-max/baseline/README.md
66
- 9f553440e741863e82f32fec11470affe6f0dbffa03c668c647bd63d8f73227f qwen3.7-max/baseline/config.json
67
- f88769bb10cbebd47ed1101a7e41c4cc5ba680bf49a768ef010aa3681bfc3845 qwen3.7-max/baseline/results.csv
68
- 4924e929d67956f9deb6c319605042d09f40b12f7bd2f18baa0464b8f7c9851a qwen3.7-max/baseline/results.jsonl
69
- e506814d33789673fcd6b8e6869cdcbd52371cd3d0455ce53ae27e1098a52214 qwen3.7-max/baseline/summary.json
70
- c032c3f9953bbd6e7685028f1b08be3539aee4fc9d186c941b5a248895c2289c qwen3.7-max/deepagents/README.md
71
- 3862d438d381ca3e57b1bf68921ef01a8c38727902152ada2c6e1fcf7b70cc77 qwen3.7-max/deepagents/config.json
72
- 09d09b2411c67b5c026e5d82662b016f36bd7373563db541e40b2fa0429998dc qwen3.7-max/deepagents/results.csv
73
- e009f8c5c88b05896836330822afbda33af463c1f264f9adb9bbf5b094e61dfa qwen3.7-max/deepagents/results.jsonl
74
- a5dcb2297ccc31a921e3f4eeb025ed81a764e16556a12068733cfe3bfe0bdbe3 qwen3.7-max/deepagents/summary.json
75
- af5df4bd67b09d31325e71769cdedf0d2d8c4ed41ee7ce3a6eb10a2919737589 qwen3.7-max/diagnostics.json
76
- 7b2d6aa947e21941e412e38f91aa01b75ff91b0668fea9071b5a9787f3a77347 qwen3.7-max/export-verification.json
77
- 2600a5ed895e9cf217f96ba7cf2bf61762012886cb8f0918693888ab4bb084d3 qwen3.7-max/paired_comparison.csv
78
- 0b4d4df9d186cb3ab5abd0d0a52bfcf89bad5efb4a135039277b282797d9580e qwen3.7-max/provenance.json
79
- 06ed669caaba3ca5e252d12af3de32223019d50aefd9c46751ca3cb35110fce2 verify_kimi_gemini.py
80
- 9c6d1d6b44998a459141969065e4e23fb0126b3556074ce35803f970ce8c0869 verify_qwen_deepseek.py
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 0aed0cba6dfcb66b4f22440aeb7e5721e015d26c3e195a99e944cb85c72108ef README.md
2
+ 2ad22f4778525b1460f1857e1095f2eea5620063fba1ffa6314ca7fd50f14703 claude-opus-4-8/baseline/README.md
3
+ 8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56 claude-opus-4-8/baseline/config.json
4
+ 4d5f859e472b79195c9be6b1f0e9c5feecf59a81dfe32ac08eb4c1d4967ed98e claude-opus-4-8/baseline/results.csv
5
+ b9d6bd51c9eb81a634e78e422c86da16f655d97da87d3f5afa4c9aa0f96809f4 claude-opus-4-8/baseline/results.jsonl
6
+ 7b5b06bd9b3d7b47952b1dc424a01fcdf04ddccca0b3f64d04132358d6380cdc claude-opus-4-8/baseline/summary.json
7
+ dfb00967c2a2ff320d5a6d8c142cbc6ef37d4fb144abcc7a451673515c2b4c50 claude-opus-4-8/deepagents/README.md
8
+ 12215dfcba5d2d692268a305364514a03ff7579094b2d4af12f493383b8e0f21 claude-opus-4-8/deepagents/config.json
9
+ bafc0a72cd4d8f53c1da33030250171f36ce441488bc8456cba8c3f52b80e6bd claude-opus-4-8/deepagents/results.csv
10
+ 3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a claude-opus-4-8/deepagents/results.jsonl
11
+ f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614 claude-opus-4-8/deepagents/summary.json
12
+ c738e75419e319515b15c823300a16e9bdf1cbe790612bbbc24ff93c71438b3f deepseek-v4-pro/baseline/README.md
13
+ 4cef524f37a325029ea4ad325b163f72f9cf725f507f20cd76424d06e71e24ad deepseek-v4-pro/baseline/config.json
14
+ 98f4f77003f4788509360a2cbe41554279df408fd20d65058b85836a25858ecb deepseek-v4-pro/baseline/results.csv
15
+ 7827218f6f18f39988e8c439f0c93cab611c04d158d8704aa50c9274cd438837 deepseek-v4-pro/baseline/results.jsonl
16
+ c51e4f1de78eba8f30e4a59276fb7821132cd103f0ae3f9d04c75c98344f2e34 deepseek-v4-pro/baseline/summary.json
17
+ f1297897da23ab01751289cd93da69a944834af59204cd73f7fca64a112bd6a9 deepseek-v4-pro/deepagents/README.md
18
+ 400c28e4008d5241bf50bc71c2674861846d6b6d2e7da21f2a7602a60ecc0ae9 deepseek-v4-pro/deepagents/config.json
19
+ a842bac2b7b608ab124d95883194e01700c6fba648593a2128db6f9d40511cf8 deepseek-v4-pro/deepagents/results.csv
20
+ a006c8b0b80064295faeb3aa7395a2408f0208677853448816b361bf67e39bb6 deepseek-v4-pro/deepagents/results.jsonl
21
+ 231e4264b0403e12788136d01af465f86fe9ba3c3ab992510925a7ef4c2a40e9 deepseek-v4-pro/deepagents/summary.json
22
+ eb38fada3bb400d40e37fd97aeb41c8eea6a1475406395e12ad9ac1c14f71e34 deepseek-v4-pro/diagnostics.json
23
+ ccf9f54ce467c8a977074498eeb86a2ce1b7b69a73b370b025663f9582a8d991 deepseek-v4-pro/dsh/README.md
24
+ 6f92c5376ab6b96b9ab5ddab6faff88cd792d0ec999f1cc96cdd48e32302c9e1 deepseek-v4-pro/dsh/accounting.json
25
+ 33faf46bce3d4a152bb04cd3eb7cded2d948b2bf5ec2beac49941b145e92c522 deepseek-v4-pro/dsh/all-attempts.jsonl
26
+ df322f13761e06dc78cb553c5dac347ce4361599937af4017a902c56a7df9d77 deepseek-v4-pro/dsh/config.json
27
+ d6500c012259ed5e3b514d6f62597fdcadfb7548e41308eed442a1b130f0e336 deepseek-v4-pro/dsh/diagnostics.json
28
+ 8d7cae2ea2f33130636367f5a5e8f92b5df0471aa9889aa54094e2d1664dd48b deepseek-v4-pro/dsh/export-verification.json
29
+ 153d15ff14c3dba6ece8fd9c81ae2d9fa1f0b6a240e90efd6b1d6e46f40893e2 deepseek-v4-pro/dsh/first-pass-results.jsonl
30
+ b264e4796ea32220a7c67fe84d990ac6e040784f7331c662e66ca5d91c71e9c1 deepseek-v4-pro/dsh/first-pass-summary.json
31
+ fa67ea7aba6f55ddceb6e10695acc0c4f5f4957ff5495f0ba83ef564478e7381 deepseek-v4-pro/dsh/manifest.sha256
32
+ 8e03b3fd76310ab80643f51ce791e85545062d2d85ac3c6121cc1143ebfb5a28 deepseek-v4-pro/dsh/provenance.json
33
+ 954a6adff76d6bae31c77cd969b17fa5203ae7ba38abf76130c6f066e39a9c58 deepseek-v4-pro/dsh/results.csv
34
+ b12427d5f4476351e571e4b29f61e0463540a1701bec51e642753d717dbc1df2 deepseek-v4-pro/dsh/results.jsonl
35
+ 62936f76b1ca0de5e761d1ea1f894343695503ce934bf0213dc9fde1ea72a782 deepseek-v4-pro/dsh/summary.json
36
+ d811ae94edc6589e2b7ee8c39ac6c4da1c66b5c3dad8acbc589b17767f9c5690 deepseek-v4-pro/export-verification.json
37
+ 891a8e0c9e1be24652b56d477bbdc0c98cf961c18010177535d33395dc43a132 deepseek-v4-pro/paired_comparison.csv
38
+ 8be2d803fc96ad6a7840c0c3eb52b52e0e1d8e8523f6a4349b4adfbdc4c320f3 deepseek-v4-pro/provenance.json
39
+ ac7253c934e43a88d0a2f6b105d1bb13424ef004e6838d2a632d5e9c366c72f2 gemini-3.1-pro-preview/baseline/README.md
40
+ ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998 gemini-3.1-pro-preview/baseline/config.json
41
+ fc708f327a76b22ebe6f453206b0c42e4cf46176afb8cb12eda2c1b71fe90bde gemini-3.1-pro-preview/baseline/results.csv
42
+ 2f1e2019157d105efee7ba311a49eba99abe4a7b84f0a4b4e0ac174e2d56cdba gemini-3.1-pro-preview/baseline/results.jsonl
43
+ a8ba5100a28f1f0aeae6949eb3669c7717c844c4e2eee6fa42ded351aef6b5c7 gemini-3.1-pro-preview/baseline/summary.json
44
+ 49ff37264a63f3b5615632213bb9248d2b6b8feb66b4ec989ed655b45e5faed2 gemini-3.1-pro-preview/deepagents/README.md
45
+ f09d9d99d26328ae2f54e9173b22f69d68f7d7bb50c3cc5e352c996b98dd57f9 gemini-3.1-pro-preview/deepagents/config.json
46
+ af16c74fdc1163682a9a98f597ac0d101bc0de64d60c6a8582f69312948c8069 gemini-3.1-pro-preview/deepagents/results.csv
47
+ c2824f8277b283a6dd60118b6e6a0cb0419e4c8b9193de15a7bcf7aaaad24181 gemini-3.1-pro-preview/deepagents/results.jsonl
48
+ bf9ce721c94103f91b0ca576f717f4f5dca7cc9965601658a65fc6d2a15b7045 gemini-3.1-pro-preview/deepagents/summary.json
49
+ 67b24d52eeabbcd84baf00c7395e72d4efe7cfcc82d7de1b0f8ee30f8018a1f7 gemini-3.1-pro-preview/diagnostics.json
50
+ 885b4de7e83334c416f33ea1535a72f1f70dfe61834d1b622e01ddb3cccfaed4 gemini-3.1-pro-preview/export-verification.json
51
+ 930a9ae67dc0dd4210e7e0c96098cab24899c765e4b3d041f75455f58fc072fe gemini-3.1-pro-preview/paired_comparison.csv
52
+ 84378a81ca5533439b21dc1cf683925f5accc23532d392d140a4477e59c4edfb gemini-3.1-pro-preview/provenance.json
53
+ ed22b41edf2afe7a500f4caf1b61f0b4c6e6c73759522c543fb86ce478eb4d04 gpt-5.6-sol/baseline/README.md
54
+ 90ec4e0781b57498fef52de31eacf5c8ec9424549643865a7140a1ccdb04c1d4 gpt-5.6-sol/baseline/config.json
55
+ a54171a3d3e7ecda3e61311414a84131e74a680bfa1342e8627cfaa32fc7c873 gpt-5.6-sol/baseline/results.csv
56
+ 3b12160382043e27531ff5cb3e23bccc9a8f7807e4121a82ffa13df28f65a859 gpt-5.6-sol/baseline/results.jsonl
57
+ 6d6ab1aaffe754e08f9656dcb5f51fcd5c772e60660dc34d86cfc21f3c5c03d7 gpt-5.6-sol/baseline/summary.json
58
+ 99a2c4edc5f8a2984e09211c7608f08cc29131e37241cc79b6c77b774d3b21fe gpt-5.6-sol/deepagents/README.md
59
+ a3e24175f2395cffbbcd6d9d3c89689bd3a3b22dd4ce61bb78d5da33c6b6cc1b gpt-5.6-sol/deepagents/config.json
60
+ 27abe57119d4be49c56b29c318a1bc2757a97f44427764d7fc8d188be4d10671 gpt-5.6-sol/deepagents/results.csv
61
+ 378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be gpt-5.6-sol/deepagents/results.jsonl
62
+ 821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5 gpt-5.6-sol/deepagents/summary.json
63
+ 8ce61775c0f51f4d53181887131a2ac83d7615dfbda59a8fb8706b4d2abcd5bd index.json
64
+ 29825997a250cc8fe292450006b82a392fe90c8e0493a1335f74eea72c1de104 kimi-k3/baseline/README.md
65
+ 7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081 kimi-k3/baseline/config.json
66
+ f87d2dd3b390695f6b420a78e126b504abc3d292180a3443deaead174ec1b464 kimi-k3/baseline/results.csv
67
+ 0502ae8892dd6c66ee92890a83c678527cff6c3396aa4c13faaadcf41e1857c1 kimi-k3/baseline/results.jsonl
68
+ 2edd2e96f3b448ceceb6189edb7f23c737b245c00c9820728c69f5237feb34b7 kimi-k3/baseline/summary.json
69
+ d6fe2ba624876d49373d4a0c180e0c27067405501b247e59c2d8ca5c74427ff7 kimi-k3/deepagents/README.md
70
+ 312c9260c01d5988878e4f2b88da55a16362ee994e8217c63b15ba772da3857c kimi-k3/deepagents/config.json
71
+ 572c3270a3ec255544b587aada9d8bd56cc2555d5246ce26feea2a3d35eb8f94 kimi-k3/deepagents/results.csv
72
+ bdab99f8f1d81885f74326cf9c8c8e57b9bb8a8d616bfbb1b5a2f6bf32b737ca kimi-k3/deepagents/results.jsonl
73
+ 81c8cabaa58acc65ef6afba792241b70951b666ab90c3f7c25194b61b0d9a228 kimi-k3/deepagents/summary.json
74
+ bfaac06563fc53d5f564197fdfb69e24ddbdff622b85d67bbf2a13fb7fb5d9ba kimi-k3/diagnostics.json
75
+ b96e0cc3352efac6140c33cef30873e522db547891101c3c49c6aca42085becb kimi-k3/export-verification.json
76
+ 3fe1378a23cfe59a37558521ec5bf47ea299a6ad4903a14684147a1b8066ee03 kimi-k3/paired_comparison.csv
77
+ 4a0f6dde38c5410a6f379192f30eac6d330403c80fcdbc2b357dbe15427a331b kimi-k3/provenance.json
78
+ 1606c3170bec39e7c074659fd8c8d011e02e98b1e69161338554a9e1c416a363 qwen3.7-max/baseline/README.md
79
+ 9f553440e741863e82f32fec11470affe6f0dbffa03c668c647bd63d8f73227f qwen3.7-max/baseline/config.json
80
+ f88769bb10cbebd47ed1101a7e41c4cc5ba680bf49a768ef010aa3681bfc3845 qwen3.7-max/baseline/results.csv
81
+ 4924e929d67956f9deb6c319605042d09f40b12f7bd2f18baa0464b8f7c9851a qwen3.7-max/baseline/results.jsonl
82
+ e506814d33789673fcd6b8e6869cdcbd52371cd3d0455ce53ae27e1098a52214 qwen3.7-max/baseline/summary.json
83
+ c032c3f9953bbd6e7685028f1b08be3539aee4fc9d186c941b5a248895c2289c qwen3.7-max/deepagents/README.md
84
+ 3862d438d381ca3e57b1bf68921ef01a8c38727902152ada2c6e1fcf7b70cc77 qwen3.7-max/deepagents/config.json
85
+ 09d09b2411c67b5c026e5d82662b016f36bd7373563db541e40b2fa0429998dc qwen3.7-max/deepagents/results.csv
86
+ e009f8c5c88b05896836330822afbda33af463c1f264f9adb9bbf5b094e61dfa qwen3.7-max/deepagents/results.jsonl
87
+ a5dcb2297ccc31a921e3f4eeb025ed81a764e16556a12068733cfe3bfe0bdbe3 qwen3.7-max/deepagents/summary.json
88
+ af5df4bd67b09d31325e71769cdedf0d2d8c4ed41ee7ce3a6eb10a2919737589 qwen3.7-max/diagnostics.json
89
+ 7b2d6aa947e21941e412e38f91aa01b75ff91b0668fea9071b5a9787f3a77347 qwen3.7-max/export-verification.json
90
+ 2600a5ed895e9cf217f96ba7cf2bf61762012886cb8f0918693888ab4bb084d3 qwen3.7-max/paired_comparison.csv
91
+ 0b4d4df9d186cb3ab5abd0d0a52bfcf89bad5efb4a135039277b282797d9580e qwen3.7-max/provenance.json
92
+ 06ed669caaba3ca5e252d12af3de32223019d50aefd9c46751ca3cb35110fce2 verify_kimi_gemini.py
93
+ 9c6d1d6b44998a459141969065e4e23fb0126b3556074ce35803f970ce8c0869 verify_qwen_deepseek.py