Archive Qwen3.7-Max and DeepSeek V4 Pro paired results

#3
by YICHEN013 - opened
Files changed (32) hide show
  1. results/README.md +10 -0
  2. results/deepseek-v4-pro/baseline/README.md +11 -0
  3. results/deepseek-v4-pro/baseline/config.json +1199 -0
  4. results/deepseek-v4-pro/baseline/results.csv +0 -0
  5. results/deepseek-v4-pro/baseline/results.jsonl +0 -0
  6. results/deepseek-v4-pro/baseline/summary.json +191 -0
  7. results/deepseek-v4-pro/deepagents/README.md +11 -0
  8. results/deepseek-v4-pro/deepagents/config.json +1199 -0
  9. results/deepseek-v4-pro/deepagents/results.csv +0 -0
  10. results/deepseek-v4-pro/deepagents/results.jsonl +0 -0
  11. results/deepseek-v4-pro/deepagents/summary.json +191 -0
  12. results/deepseek-v4-pro/diagnostics.json +47 -0
  13. results/deepseek-v4-pro/export-verification.json +44 -0
  14. results/deepseek-v4-pro/paired_comparison.csv +0 -0
  15. results/deepseek-v4-pro/provenance.json +42 -0
  16. results/index.json +25 -1
  17. results/manifest.sha256 +31 -2
  18. results/qwen3.7-max/baseline/README.md +11 -0
  19. results/qwen3.7-max/baseline/config.json +1199 -0
  20. results/qwen3.7-max/baseline/results.csv +0 -0
  21. results/qwen3.7-max/baseline/results.jsonl +0 -0
  22. results/qwen3.7-max/baseline/summary.json +191 -0
  23. results/qwen3.7-max/deepagents/README.md +11 -0
  24. results/qwen3.7-max/deepagents/config.json +1199 -0
  25. results/qwen3.7-max/deepagents/results.csv +0 -0
  26. results/qwen3.7-max/deepagents/results.jsonl +0 -0
  27. results/qwen3.7-max/deepagents/summary.json +191 -0
  28. results/qwen3.7-max/diagnostics.json +74 -0
  29. results/qwen3.7-max/export-verification.json +44 -0
  30. results/qwen3.7-max/paired_comparison.csv +0 -0
  31. results/qwen3.7-max/provenance.json +24 -0
  32. results/verify_qwen_deepseek.py +71 -0
results/README.md CHANGED
@@ -8,6 +8,8 @@
8
  | Claude Opus 4.8 | [1,000 题](claude-opus-4-8/baseline/) | [1,000 题](claude-opus-4-8/deepagents/) |
9
  | Kimi K3 | [1,000 题](kimi-k3/baseline/) | [1,000 题](kimi-k3/deepagents/) |
10
  | Gemini 3.1 Pro Preview | [1,000 题](gemini-3.1-pro-preview/baseline/) | [1,000 题](gemini-3.1-pro-preview/deepagents/) |
 
 
11
 
12
  这里先保存完成的实验产物。页面仍只读取根目录 `/results.csv`,本次没有更新它或修改页面代码。等各模型结果齐备,再核对数据、评分和协议后统一整理正式榜单。
13
 
@@ -22,3 +24,11 @@
22
  K3 保留两组 5 / 4 条任务 unknown,Gemini 无任务 unknown;每组固定 1,000 题分母。指标级 null 与任务级 unknown 分别记账。新增记录的原始预测不在源评分快照中,prediction 字段因此留空并带有导出状态标记;不把导出缺失解释为模型未生成预测。详细协议、源版本、阶段诊断和核验范围见各模型目录。
23
 
24
  运行 `python3 -B results/verify_kimi_gemini.py` 可检查归档清单、4,000 条新增逐题记录及 36 项指标汇总;这是离线导出核验,不是重新运行实验或官方 scorer 一致性认证。
 
 
 
 
 
 
 
 
 
8
  | Claude Opus 4.8 | [1,000 题](claude-opus-4-8/baseline/) | [1,000 题](claude-opus-4-8/deepagents/) |
9
  | Kimi K3 | [1,000 题](kimi-k3/baseline/) | [1,000 题](kimi-k3/deepagents/) |
10
  | Gemini 3.1 Pro Preview | [1,000 题](gemini-3.1-pro-preview/baseline/) | [1,000 题](gemini-3.1-pro-preview/deepagents/) |
11
+ | Qwen3.7-Max | [1,000 题](qwen3.7-max/baseline/) | [1,000 题](qwen3.7-max/deepagents/) |
12
+ | DeepSeek V4 Pro | [1,000 题](deepseek-v4-pro/baseline/) | [1,000 题](deepseek-v4-pro/deepagents/) |
13
 
14
  这里先保存完成的实验产物。页面仍只读取根目录 `/results.csv`,本次没有更新它或修改页面代码。等各模型结果齐备,再核对数据、评分和协议后统一整理正式榜单。
15
 
 
24
  K3 保留两组 5 / 4 条任务 unknown,Gemini 无任务 unknown;每组固定 1,000 题分母。指标级 null 与任务级 unknown 分别记账。新增记录的原始预测不在源评分快照中,prediction 字段因此留空并带有导出状态标记;不把导出缺失解释为模型未生成预测。详细协议、源版本、阶段诊断和核验范围见各模型目录。
25
 
26
  运行 `python3 -B results/verify_kimi_gemini.py` 可检查归档清单、4,000 条新增逐题记录及 36 项指标汇总;这是离线导出核验,不是重新运行实验或官方 scorer 一致性认证。
27
+
28
+ ## Qwen3.7-Max 与 DeepSeek V4 Pro
29
+
30
+ 新增四组、4,000 条记录后,存档覆盖六个模型、12 组、12,000 条记录。GPT-5.5 不在本次范围。
31
+
32
+ Qwen3.7-Max 原生完整复现率为 19.8% / 24.8%,DeepSeek V4 Pro 为 12.9% / 18.3%;本地 HF 四维与原生五维分开保存。Qwen DeepAgents 保留 6 条任务 unknown 和 task 453 的 1 条 invalid evidence:999 条有效封存、1 条无效证据,共 1,000 个固定题号。无效证据不标为失败或有效封存,指标均为空;未重跑或删除。DeepSeek 双组没有任务 unknown 或 invalid。
33
+
34
+ 运行 `python3 -B results/verify_qwen_deepseek.py` 可核对新增四组逐题记录、配对表、36 项汇总和归档清单。这是已接受产物的离线导出验证,不代表重新进行原生审计或官方评分。网页根目录 results.csv 与页面代码由后续网页更新单独处理。
results/deepseek-v4-pro/baseline/README.md ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DeepSeek V4 Pro / baseline
2
+
3
+ 固定 1,000 题:成功 277、确定失败 723、任务 unknown 0、invalid evidence 0。已封存 1000 条;无待开始或在途记录。原生完整复现率 12.9%。
4
+
5
+ `results.jsonl` 与 `results.csv` 保存逐题状态、指标、用量和结果哈希;`summary.json` 分别保存本地 HF 四维和 local-paper-v1 五维;`config.json` 保存任务清单、协议、源代码版本和接续来源。上一级含配对表、来源与导出核验。
6
+
7
+ `resolved` 仅表示有效封存,`definite_outcome` 表示成功或确定失败。任务 unknown、无效证据与指标 null 分开保留,均不从 1,000 的分母移除。无效记录仅有 `unverified_result_sha256`,不可当作有效执行结果。后续阶段 pending 可能表示前序失败后没有执行到,不表示任务仍待派发。
8
+
9
+ 原始预测三元组不在源评分快照中,prediction 字段为空,并有明确导出状态;不能据此认为所有模型都没有生成预测。配对比较含历史协议修订,非严格等预算因果实验。本地评分尚未核实官方一致性。
10
+
11
+ 本次仅归档已接受的评分和状态,没有重新执行模型或程序。不包含 benchmark 输入、隐藏答案、原始提示词/API 轨迹、凭据或机器路径。
results/deepseek-v4-pro/baseline/config.json ADDED
@@ -0,0 +1,1199 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "deepseek-v4-pro",
4
+ "arm": "baseline",
5
+ "harness": "none",
6
+ "archived_at_utc": "2026-09-16T17:18:47.373930+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "engineering_pilot": false,
14
+ "expected_tasks": 1000,
15
+ "task_ids": [
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4,
20
+ 5,
21
+ 6,
22
+ 7,
23
+ 8,
24
+ 9,
25
+ 10,
26
+ 11,
27
+ 12,
28
+ 13,
29
+ 14,
30
+ 15,
31
+ 16,
32
+ 17,
33
+ 18,
34
+ 19,
35
+ 20,
36
+ 21,
37
+ 22,
38
+ 23,
39
+ 24,
40
+ 25,
41
+ 26,
42
+ 27,
43
+ 28,
44
+ 29,
45
+ 30,
46
+ 31,
47
+ 32,
48
+ 33,
49
+ 34,
50
+ 35,
51
+ 36,
52
+ 37,
53
+ 38,
54
+ 39,
55
+ 40,
56
+ 41,
57
+ 42,
58
+ 43,
59
+ 44,
60
+ 45,
61
+ 46,
62
+ 47,
63
+ 48,
64
+ 49,
65
+ 50,
66
+ 51,
67
+ 52,
68
+ 53,
69
+ 54,
70
+ 55,
71
+ 56,
72
+ 57,
73
+ 58,
74
+ 59,
75
+ 60,
76
+ 61,
77
+ 62,
78
+ 63,
79
+ 64,
80
+ 65,
81
+ 66,
82
+ 67,
83
+ 68,
84
+ 69,
85
+ 70,
86
+ 71,
87
+ 72,
88
+ 73,
89
+ 74,
90
+ 75,
91
+ 76,
92
+ 77,
93
+ 78,
94
+ 79,
95
+ 80,
96
+ 81,
97
+ 82,
98
+ 83,
99
+ 84,
100
+ 85,
101
+ 86,
102
+ 87,
103
+ 88,
104
+ 89,
105
+ 90,
106
+ 91,
107
+ 92,
108
+ 93,
109
+ 94,
110
+ 95,
111
+ 96,
112
+ 97,
113
+ 98,
114
+ 100,
115
+ 101,
116
+ 102,
117
+ 103,
118
+ 104,
119
+ 105,
120
+ 106,
121
+ 107,
122
+ 108,
123
+ 109,
124
+ 110,
125
+ 111,
126
+ 112,
127
+ 113,
128
+ 114,
129
+ 115,
130
+ 116,
131
+ 117,
132
+ 118,
133
+ 119,
134
+ 120,
135
+ 121,
136
+ 122,
137
+ 123,
138
+ 124,
139
+ 125,
140
+ 126,
141
+ 127,
142
+ 128,
143
+ 129,
144
+ 130,
145
+ 131,
146
+ 132,
147
+ 133,
148
+ 134,
149
+ 135,
150
+ 136,
151
+ 137,
152
+ 138,
153
+ 139,
154
+ 140,
155
+ 141,
156
+ 142,
157
+ 143,
158
+ 144,
159
+ 145,
160
+ 146,
161
+ 147,
162
+ 148,
163
+ 149,
164
+ 150,
165
+ 151,
166
+ 152,
167
+ 153,
168
+ 154,
169
+ 155,
170
+ 156,
171
+ 157,
172
+ 158,
173
+ 159,
174
+ 160,
175
+ 161,
176
+ 162,
177
+ 163,
178
+ 164,
179
+ 165,
180
+ 166,
181
+ 167,
182
+ 168,
183
+ 169,
184
+ 170,
185
+ 171,
186
+ 172,
187
+ 173,
188
+ 174,
189
+ 175,
190
+ 176,
191
+ 177,
192
+ 178,
193
+ 179,
194
+ 180,
195
+ 181,
196
+ 182,
197
+ 183,
198
+ 184,
199
+ 185,
200
+ 186,
201
+ 187,
202
+ 188,
203
+ 189,
204
+ 190,
205
+ 191,
206
+ 192,
207
+ 193,
208
+ 194,
209
+ 195,
210
+ 196,
211
+ 197,
212
+ 198,
213
+ 199,
214
+ 200,
215
+ 201,
216
+ 202,
217
+ 203,
218
+ 204,
219
+ 205,
220
+ 206,
221
+ 207,
222
+ 208,
223
+ 209,
224
+ 210,
225
+ 211,
226
+ 212,
227
+ 213,
228
+ 214,
229
+ 215,
230
+ 216,
231
+ 217,
232
+ 218,
233
+ 219,
234
+ 220,
235
+ 221,
236
+ 222,
237
+ 223,
238
+ 224,
239
+ 225,
240
+ 226,
241
+ 227,
242
+ 228,
243
+ 229,
244
+ 230,
245
+ 231,
246
+ 232,
247
+ 233,
248
+ 234,
249
+ 235,
250
+ 236,
251
+ 237,
252
+ 238,
253
+ 239,
254
+ 240,
255
+ 241,
256
+ 242,
257
+ 248,
258
+ 249,
259
+ 250,
260
+ 251,
261
+ 252,
262
+ 253,
263
+ 254,
264
+ 255,
265
+ 256,
266
+ 257,
267
+ 258,
268
+ 259,
269
+ 261,
270
+ 262,
271
+ 263,
272
+ 264,
273
+ 265,
274
+ 266,
275
+ 290,
276
+ 291,
277
+ 292,
278
+ 293,
279
+ 294,
280
+ 295,
281
+ 302,
282
+ 303,
283
+ 304,
284
+ 305,
285
+ 306,
286
+ 307,
287
+ 308,
288
+ 309,
289
+ 310,
290
+ 311,
291
+ 312,
292
+ 313,
293
+ 314,
294
+ 315,
295
+ 316,
296
+ 317,
297
+ 318,
298
+ 319,
299
+ 320,
300
+ 321,
301
+ 322,
302
+ 323,
303
+ 324,
304
+ 325,
305
+ 326,
306
+ 327,
307
+ 328,
308
+ 329,
309
+ 330,
310
+ 331,
311
+ 332,
312
+ 333,
313
+ 334,
314
+ 335,
315
+ 336,
316
+ 337,
317
+ 338,
318
+ 339,
319
+ 340,
320
+ 341,
321
+ 342,
322
+ 343,
323
+ 344,
324
+ 345,
325
+ 346,
326
+ 347,
327
+ 348,
328
+ 349,
329
+ 350,
330
+ 351,
331
+ 352,
332
+ 353,
333
+ 354,
334
+ 355,
335
+ 356,
336
+ 357,
337
+ 358,
338
+ 359,
339
+ 360,
340
+ 361,
341
+ 362,
342
+ 363,
343
+ 364,
344
+ 365,
345
+ 366,
346
+ 367,
347
+ 368,
348
+ 369,
349
+ 370,
350
+ 371,
351
+ 372,
352
+ 373,
353
+ 374,
354
+ 375,
355
+ 376,
356
+ 377,
357
+ 378,
358
+ 379,
359
+ 380,
360
+ 381,
361
+ 382,
362
+ 383,
363
+ 384,
364
+ 385,
365
+ 386,
366
+ 387,
367
+ 388,
368
+ 389,
369
+ 390,
370
+ 391,
371
+ 392,
372
+ 393,
373
+ 394,
374
+ 395,
375
+ 396,
376
+ 397,
377
+ 398,
378
+ 399,
379
+ 400,
380
+ 401,
381
+ 402,
382
+ 403,
383
+ 404,
384
+ 405,
385
+ 406,
386
+ 407,
387
+ 408,
388
+ 409,
389
+ 410,
390
+ 411,
391
+ 412,
392
+ 413,
393
+ 414,
394
+ 415,
395
+ 416,
396
+ 417,
397
+ 418,
398
+ 419,
399
+ 420,
400
+ 421,
401
+ 422,
402
+ 423,
403
+ 424,
404
+ 425,
405
+ 426,
406
+ 427,
407
+ 428,
408
+ 429,
409
+ 430,
410
+ 431,
411
+ 432,
412
+ 433,
413
+ 434,
414
+ 435,
415
+ 436,
416
+ 437,
417
+ 438,
418
+ 439,
419
+ 440,
420
+ 441,
421
+ 442,
422
+ 443,
423
+ 444,
424
+ 445,
425
+ 446,
426
+ 447,
427
+ 448,
428
+ 449,
429
+ 450,
430
+ 451,
431
+ 452,
432
+ 453,
433
+ 454,
434
+ 455,
435
+ 456,
436
+ 457,
437
+ 458,
438
+ 459,
439
+ 460,
440
+ 461,
441
+ 462,
442
+ 463,
443
+ 464,
444
+ 465,
445
+ 466,
446
+ 467,
447
+ 468,
448
+ 469,
449
+ 470,
450
+ 471,
451
+ 472,
452
+ 473,
453
+ 474,
454
+ 475,
455
+ 476,
456
+ 477,
457
+ 478,
458
+ 479,
459
+ 480,
460
+ 481,
461
+ 482,
462
+ 483,
463
+ 484,
464
+ 485,
465
+ 486,
466
+ 487,
467
+ 488,
468
+ 489,
469
+ 490,
470
+ 491,
471
+ 492,
472
+ 493,
473
+ 494,
474
+ 495,
475
+ 496,
476
+ 497,
477
+ 498,
478
+ 499,
479
+ 500,
480
+ 501,
481
+ 502,
482
+ 503,
483
+ 504,
484
+ 505,
485
+ 506,
486
+ 507,
487
+ 508,
488
+ 509,
489
+ 510,
490
+ 511,
491
+ 512,
492
+ 513,
493
+ 514,
494
+ 515,
495
+ 516,
496
+ 517,
497
+ 518,
498
+ 519,
499
+ 520,
500
+ 521,
501
+ 522,
502
+ 523,
503
+ 524,
504
+ 525,
505
+ 526,
506
+ 527,
507
+ 528,
508
+ 529,
509
+ 530,
510
+ 531,
511
+ 532,
512
+ 533,
513
+ 534,
514
+ 535,
515
+ 536,
516
+ 537,
517
+ 538,
518
+ 539,
519
+ 540,
520
+ 541,
521
+ 542,
522
+ 543,
523
+ 544,
524
+ 545,
525
+ 546,
526
+ 547,
527
+ 548,
528
+ 549,
529
+ 550,
530
+ 551,
531
+ 552,
532
+ 553,
533
+ 554,
534
+ 555,
535
+ 556,
536
+ 557,
537
+ 558,
538
+ 559,
539
+ 560,
540
+ 561,
541
+ 562,
542
+ 563,
543
+ 565,
544
+ 566,
545
+ 567,
546
+ 568,
547
+ 570,
548
+ 571,
549
+ 572,
550
+ 573,
551
+ 574,
552
+ 575,
553
+ 576,
554
+ 577,
555
+ 578,
556
+ 579,
557
+ 580,
558
+ 581,
559
+ 582,
560
+ 583,
561
+ 584,
562
+ 585,
563
+ 586,
564
+ 587,
565
+ 588,
566
+ 589,
567
+ 590,
568
+ 591,
569
+ 592,
570
+ 593,
571
+ 594,
572
+ 595,
573
+ 596,
574
+ 597,
575
+ 598,
576
+ 601,
577
+ 602,
578
+ 603,
579
+ 604,
580
+ 605,
581
+ 606,
582
+ 607,
583
+ 608,
584
+ 609,
585
+ 610,
586
+ 611,
587
+ 612,
588
+ 613,
589
+ 614,
590
+ 615,
591
+ 616,
592
+ 617,
593
+ 618,
594
+ 619,
595
+ 620,
596
+ 621,
597
+ 622,
598
+ 623,
599
+ 624,
600
+ 664,
601
+ 665,
602
+ 666,
603
+ 667,
604
+ 668,
605
+ 669,
606
+ 670,
607
+ 671,
608
+ 672,
609
+ 673,
610
+ 674,
611
+ 675,
612
+ 676,
613
+ 677,
614
+ 678,
615
+ 679,
616
+ 680,
617
+ 681,
618
+ 682,
619
+ 683,
620
+ 684,
621
+ 685,
622
+ 686,
623
+ 687,
624
+ 688,
625
+ 689,
626
+ 690,
627
+ 691,
628
+ 692,
629
+ 693,
630
+ 694,
631
+ 695,
632
+ 696,
633
+ 697,
634
+ 698,
635
+ 699,
636
+ 700,
637
+ 701,
638
+ 702,
639
+ 703,
640
+ 704,
641
+ 705,
642
+ 706,
643
+ 707,
644
+ 708,
645
+ 709,
646
+ 710,
647
+ 711,
648
+ 712,
649
+ 713,
650
+ 714,
651
+ 715,
652
+ 716,
653
+ 717,
654
+ 718,
655
+ 719,
656
+ 720,
657
+ 721,
658
+ 722,
659
+ 723,
660
+ 724,
661
+ 725,
662
+ 726,
663
+ 727,
664
+ 728,
665
+ 729,
666
+ 730,
667
+ 731,
668
+ 732,
669
+ 733,
670
+ 734,
671
+ 735,
672
+ 736,
673
+ 737,
674
+ 739,
675
+ 740,
676
+ 741,
677
+ 742,
678
+ 743,
679
+ 744,
680
+ 745,
681
+ 746,
682
+ 747,
683
+ 748,
684
+ 749,
685
+ 750,
686
+ 751,
687
+ 752,
688
+ 753,
689
+ 754,
690
+ 755,
691
+ 756,
692
+ 757,
693
+ 758,
694
+ 759,
695
+ 760,
696
+ 761,
697
+ 762,
698
+ 763,
699
+ 764,
700
+ 765,
701
+ 766,
702
+ 767,
703
+ 768,
704
+ 772,
705
+ 773,
706
+ 774,
707
+ 775,
708
+ 776,
709
+ 777,
710
+ 778,
711
+ 779,
712
+ 780,
713
+ 781,
714
+ 782,
715
+ 783,
716
+ 784,
717
+ 785,
718
+ 786,
719
+ 787,
720
+ 788,
721
+ 789,
722
+ 790,
723
+ 791,
724
+ 793,
725
+ 794,
726
+ 795,
727
+ 796,
728
+ 797,
729
+ 798,
730
+ 799,
731
+ 800,
732
+ 801,
733
+ 802,
734
+ 803,
735
+ 804,
736
+ 805,
737
+ 806,
738
+ 807,
739
+ 808,
740
+ 825,
741
+ 826,
742
+ 827,
743
+ 828,
744
+ 829,
745
+ 830,
746
+ 831,
747
+ 832,
748
+ 833,
749
+ 834,
750
+ 835,
751
+ 836,
752
+ 837,
753
+ 838,
754
+ 839,
755
+ 840,
756
+ 841,
757
+ 842,
758
+ 843,
759
+ 844,
760
+ 845,
761
+ 846,
762
+ 847,
763
+ 848,
764
+ 849,
765
+ 850,
766
+ 851,
767
+ 852,
768
+ 853,
769
+ 854,
770
+ 855,
771
+ 856,
772
+ 857,
773
+ 858,
774
+ 859,
775
+ 860,
776
+ 861,
777
+ 862,
778
+ 863,
779
+ 864,
780
+ 865,
781
+ 866,
782
+ 867,
783
+ 868,
784
+ 869,
785
+ 870,
786
+ 871,
787
+ 872,
788
+ 873,
789
+ 874,
790
+ 875,
791
+ 876,
792
+ 877,
793
+ 878,
794
+ 879,
795
+ 880,
796
+ 881,
797
+ 882,
798
+ 883,
799
+ 884,
800
+ 885,
801
+ 886,
802
+ 887,
803
+ 888,
804
+ 889,
805
+ 890,
806
+ 891,
807
+ 892,
808
+ 893,
809
+ 894,
810
+ 895,
811
+ 896,
812
+ 897,
813
+ 898,
814
+ 899,
815
+ 900,
816
+ 901,
817
+ 902,
818
+ 903,
819
+ 906,
820
+ 907,
821
+ 913,
822
+ 914,
823
+ 915,
824
+ 917,
825
+ 920,
826
+ 921,
827
+ 922,
828
+ 923,
829
+ 924,
830
+ 925,
831
+ 926,
832
+ 927,
833
+ 928,
834
+ 929,
835
+ 930,
836
+ 931,
837
+ 932,
838
+ 933,
839
+ 934,
840
+ 935,
841
+ 936,
842
+ 937,
843
+ 938,
844
+ 939,
845
+ 940,
846
+ 941,
847
+ 942,
848
+ 943,
849
+ 944,
850
+ 945,
851
+ 946,
852
+ 947,
853
+ 948,
854
+ 949,
855
+ 950,
856
+ 951,
857
+ 952,
858
+ 953,
859
+ 954,
860
+ 955,
861
+ 956,
862
+ 957,
863
+ 958,
864
+ 959,
865
+ 960,
866
+ 961,
867
+ 962,
868
+ 963,
869
+ 964,
870
+ 965,
871
+ 966,
872
+ 967,
873
+ 968,
874
+ 969,
875
+ 970,
876
+ 971,
877
+ 972,
878
+ 973,
879
+ 974,
880
+ 975,
881
+ 976,
882
+ 977,
883
+ 978,
884
+ 979,
885
+ 980,
886
+ 981,
887
+ 982,
888
+ 983,
889
+ 984,
890
+ 985,
891
+ 1001,
892
+ 1002,
893
+ 1003,
894
+ 1004,
895
+ 1005,
896
+ 1006,
897
+ 1007,
898
+ 1008,
899
+ 1009,
900
+ 1010,
901
+ 1011,
902
+ 1012,
903
+ 1013,
904
+ 1014,
905
+ 1015,
906
+ 1016,
907
+ 1017,
908
+ 1018,
909
+ 1019,
910
+ 1020,
911
+ 1021,
912
+ 1022,
913
+ 1023,
914
+ 1024,
915
+ 1025,
916
+ 1026,
917
+ 1027,
918
+ 1028,
919
+ 1029,
920
+ 1030,
921
+ 1031,
922
+ 1032,
923
+ 1033,
924
+ 1034,
925
+ 1035,
926
+ 1036,
927
+ 1037,
928
+ 1038,
929
+ 1039,
930
+ 1040,
931
+ 1041,
932
+ 1042,
933
+ 1043,
934
+ 1044,
935
+ 1045,
936
+ 1046,
937
+ 1047,
938
+ 1048,
939
+ 1049,
940
+ 1050,
941
+ 1051,
942
+ 1052,
943
+ 1053,
944
+ 1054,
945
+ 1055,
946
+ 1056,
947
+ 1057,
948
+ 1058,
949
+ 1059,
950
+ 1060,
951
+ 1061,
952
+ 1062,
953
+ 1063,
954
+ 1064,
955
+ 1065,
956
+ 1066,
957
+ 1067,
958
+ 1068,
959
+ 1069,
960
+ 1070,
961
+ 1071,
962
+ 1072,
963
+ 1073,
964
+ 1074,
965
+ 1075,
966
+ 1076,
967
+ 1077,
968
+ 1078,
969
+ 1079,
970
+ 1080,
971
+ 1081,
972
+ 1082,
973
+ 1083,
974
+ 1084,
975
+ 1085,
976
+ 1086,
977
+ 1087,
978
+ 1088,
979
+ 1089,
980
+ 1090,
981
+ 1091,
982
+ 1092,
983
+ 1093,
984
+ 1094,
985
+ 1095,
986
+ 1096,
987
+ 1097,
988
+ 1098,
989
+ 1099,
990
+ 1100,
991
+ 1101,
992
+ 1102,
993
+ 1103,
994
+ 1104,
995
+ 1105,
996
+ 1106,
997
+ 1107,
998
+ 1108,
999
+ 1109,
1000
+ 1110,
1001
+ 1111,
1002
+ 1112,
1003
+ 1113,
1004
+ 1114,
1005
+ 1115,
1006
+ 1116,
1007
+ 1117,
1008
+ 1118,
1009
+ 1119,
1010
+ 1120,
1011
+ 1121,
1012
+ 1122,
1013
+ 1123,
1014
+ 1124,
1015
+ 1125
1016
+ ]
1017
+ },
1018
+ "generation": {
1019
+ "model": "deepseek-v4-pro",
1020
+ "provider": "chat-completions",
1021
+ "reasoning_effort": null,
1022
+ "stream": true,
1023
+ "timeout": 1800
1024
+ },
1025
+ "effective_model_call_cap": 1,
1026
+ "recorded_protocols": [
1027
+ {
1028
+ "core_sha256": {
1029
+ "protocol.json": "86b59ee1ca819f74a6d0c5211583a34e1b9665460ab0c965b52da81956bbec35",
1030
+ "inputs.manifest.json": "5e74e33491e38eda692fe3209d81ab85150861cb56bccc5b726712c19d5e49a2",
1031
+ "schedule.json": "11f80f6caf79724f907ef2133a24095d93404494c32232dedab39ef5c8871542",
1032
+ "study.json": "f85ecac78dac79575bfc8b8f3cb4e68c36fa3b9f58b45c2b4c196e4ad2f8847e",
1033
+ "environment.json": "04f141c80a13f3620f971c3217dccd28c2ee58f22a3d1fc5f7db2bfd1f4cc2ef"
1034
+ },
1035
+ "population_sha256": "05d427330ec89d1449dfdfd620fd238fa99321dbc954af8fe934b127887d65a1",
1036
+ "task_count": 999,
1037
+ "task_input_blobs_rehashed": false,
1038
+ "source_and_dependencies_match": true,
1039
+ "limits": {
1040
+ "output_tokens": 32768,
1041
+ "wall_seconds": 1800,
1042
+ "model_calls": 6,
1043
+ "tool_calls": 4,
1044
+ "tool_timeout": 90
1045
+ },
1046
+ "generation": {
1047
+ "model": "deepseek-v4-pro",
1048
+ "provider": "chat-completions",
1049
+ "reasoning_effort": null,
1050
+ "stream": true,
1051
+ "timeout": 1800
1052
+ },
1053
+ "dataset_provenance": {
1054
+ "repo_id": "CamoAiLab/InferenceNet",
1055
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1056
+ "task_directory": "Selected_1000",
1057
+ "task_list": "Selected_1000/1000_new.csv",
1058
+ "prior_results_imported": false,
1059
+ "engineering_pilot": false
1060
+ },
1061
+ "execution_identity": {
1062
+ "backend": "dsw-bwrap-v3",
1063
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1064
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1065
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1066
+ "engine_memory_bytes": 38654705664,
1067
+ "engine_cpus": 12,
1068
+ "cpu_slots": [
1069
+ [
1070
+ 4,
1071
+ 5
1072
+ ],
1073
+ [
1074
+ 6,
1075
+ 7
1076
+ ],
1077
+ [
1078
+ 8,
1079
+ 9
1080
+ ],
1081
+ [
1082
+ 10,
1083
+ 11
1084
+ ],
1085
+ [
1086
+ 12,
1087
+ 13
1088
+ ],
1089
+ [
1090
+ 14,
1091
+ 15
1092
+ ]
1093
+ ],
1094
+ "resources": {
1095
+ "cpus": 2.0,
1096
+ "memory": "6g",
1097
+ "tmpfs": "256m"
1098
+ },
1099
+ "memory_enforcement": "RLIMIT_AS",
1100
+ "single_payload_process": true,
1101
+ "network": "none",
1102
+ "uid": 10001,
1103
+ "capacity_profile": "independent-six-slots-20260916",
1104
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1105
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1106
+ },
1107
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1108
+ "scorer_sha256": {
1109
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1110
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1111
+ },
1112
+ "origin": "child"
1113
+ },
1114
+ {
1115
+ "core_sha256": {
1116
+ "protocol.json": "39af792a5f92b7424f22a55602657d6f930ab37bc5607263c91c48f3b46c6c72",
1117
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1118
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1119
+ "study.json": "57177147bfcead6ff01f6c86d5225aa8542d72e019986b83a9fb4a6c437abb61",
1120
+ "environment.json": "f0e702a8a418d683763e0a0d35acd4377d246e006d1c7533c9bf6cac38fd70bf"
1121
+ },
1122
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1123
+ "task_count": 1000,
1124
+ "task_input_blobs_rehashed": false,
1125
+ "source_and_dependencies_match": true,
1126
+ "limits": {
1127
+ "output_tokens": 32768,
1128
+ "wall_seconds": 1800,
1129
+ "model_calls": 6,
1130
+ "tool_calls": 4,
1131
+ "tool_timeout": 90
1132
+ },
1133
+ "generation": {
1134
+ "model": "deepseek-v4-pro",
1135
+ "provider": "chat-completions",
1136
+ "reasoning_effort": null,
1137
+ "stream": true,
1138
+ "timeout": 1800
1139
+ },
1140
+ "dataset_provenance": {
1141
+ "repo_id": "CamoAiLab/InferenceNet",
1142
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1143
+ "task_directory": "Selected_1000",
1144
+ "task_list": "Selected_1000/1000_new.csv",
1145
+ "prior_results_imported": false,
1146
+ "engineering_pilot": false
1147
+ },
1148
+ "execution_identity": {
1149
+ "backend": "dsw-bwrap-v3",
1150
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1151
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1152
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1153
+ "engine_memory_bytes": 12884901888,
1154
+ "engine_cpus": 4,
1155
+ "cpu_slots": [
1156
+ [
1157
+ 16,
1158
+ 17
1159
+ ],
1160
+ [
1161
+ 18,
1162
+ 19
1163
+ ]
1164
+ ],
1165
+ "resources": {
1166
+ "memory": "6g",
1167
+ "cpus": 2.0,
1168
+ "tmpfs": "256m"
1169
+ },
1170
+ "memory_enforcement": "RLIMIT_AS",
1171
+ "single_payload_process": true,
1172
+ "network": "none",
1173
+ "uid": 10001
1174
+ },
1175
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1176
+ "scorer_sha256": {
1177
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1178
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1179
+ },
1180
+ "origin": "parent"
1181
+ }
1182
+ ],
1183
+ "selected_origin_counts": {
1184
+ "child": 999,
1185
+ "parent": 1
1186
+ },
1187
+ "source_commits": [
1188
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1189
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1190
+ ],
1191
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1192
+ "official_parity_verified": false,
1193
+ "notes": [
1194
+ "Historical disjoint parent/child continuation; original source and protocol hashes retained.",
1195
+ "Different baseline/harness call budgets; not an equal-cost causal comparison.",
1196
+ "Unknown and invalid evidence remain in the denominator; no replacement or rescoring.",
1197
+ "Invalid evidence is not sealed; its unverified hash is distinct from a verified result hash."
1198
+ ]
1199
+ }
results/deepseek-v4-pro/baseline/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/baseline/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/baseline/summary.json ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
5
+ "model": "deepseek-v4-pro",
6
+ "arm": "baseline",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": true,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 281,
37
+ "denominator": 1000,
38
+ "rate": 0.281,
39
+ "score": 28.1,
40
+ "unknown_count": 0,
41
+ "failure_count": 719,
42
+ "assessable": 1000,
43
+ "coverage_percent": 100.0
44
+ },
45
+ "partial_replication": {
46
+ "count": 195,
47
+ "denominator": 1000,
48
+ "rate": 0.195,
49
+ "score": 19.5,
50
+ "unknown_count": 723,
51
+ "failure_count": 82,
52
+ "assessable": 277,
53
+ "coverage_percent": 27.7
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 260,
57
+ "denominator": 1000,
58
+ "rate": 0.26,
59
+ "score": 26.0,
60
+ "unknown_count": 723,
61
+ "failure_count": 17,
62
+ "assessable": 277,
63
+ "coverage_percent": 27.7
64
+ },
65
+ "significance_level": {
66
+ "count": 236,
67
+ "denominator": 1000,
68
+ "rate": 0.236,
69
+ "score": 23.6,
70
+ "unknown_count": 723,
71
+ "failure_count": 41,
72
+ "assessable": 277,
73
+ "coverage_percent": 27.7
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "deepseek-v4-pro baseline",
78
+ "Compilation Success": 28.1,
79
+ "Partial Replication": 19.5,
80
+ "Correct Coefficient Direction": 26.0,
81
+ "Significant Level Correctness": 23.6
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 1000,
94
+ "task_unknown_count": 0,
95
+ "invalid_evidence_count": 0,
96
+ "valid_prediction_count": 277,
97
+ "task_counts": {
98
+ "succeeded": 277,
99
+ "failed": 723,
100
+ "unknown": 0,
101
+ "unsealed": 0,
102
+ "not_started": 0,
103
+ "invalid_evidence": 0,
104
+ "completed": 1000,
105
+ "planned": 1000
106
+ }
107
+ },
108
+ "legacy_local_paper": {
109
+ "profile": "local-paper-v1",
110
+ "definitions": {
111
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
112
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
113
+ "coefficient_only": "coefficient relative error <= .05",
114
+ "direction": "matching strictly positive or strictly negative coefficient signs",
115
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
116
+ },
117
+ "metrics": {
118
+ "perfect": {
119
+ "count": 129,
120
+ "denominator": 1000,
121
+ "rate": 0.129,
122
+ "score": 12.9,
123
+ "unknown_count": 723,
124
+ "failure_count": 148,
125
+ "assessable": 277,
126
+ "coverage_percent": 27.7
127
+ },
128
+ "partial": {
129
+ "count": 159,
130
+ "denominator": 1000,
131
+ "rate": 0.159,
132
+ "score": 15.9,
133
+ "unknown_count": 723,
134
+ "failure_count": 118,
135
+ "assessable": 277,
136
+ "coverage_percent": 27.7
137
+ },
138
+ "coefficient_only": {
139
+ "count": 195,
140
+ "denominator": 1000,
141
+ "rate": 0.195,
142
+ "score": 19.5,
143
+ "unknown_count": 723,
144
+ "failure_count": 82,
145
+ "assessable": 277,
146
+ "coverage_percent": 27.7
147
+ },
148
+ "direction": {
149
+ "count": 260,
150
+ "denominator": 1000,
151
+ "rate": 0.26,
152
+ "score": 26.0,
153
+ "unknown_count": 723,
154
+ "failure_count": 17,
155
+ "assessable": 277,
156
+ "coverage_percent": 27.7
157
+ },
158
+ "significance": {
159
+ "count": 227,
160
+ "denominator": 1000,
161
+ "rate": 0.227,
162
+ "score": 22.7,
163
+ "unknown_count": 723,
164
+ "failure_count": 50,
165
+ "assessable": 277,
166
+ "coverage_percent": 27.7
167
+ }
168
+ }
169
+ },
170
+ "verification_scope": "Offline equality to accepted flags, hashes and states; raw predictions unavailable; not a fresh native audit or official scorer parity check.",
171
+ "usage": {
172
+ "sealed_model_calls": 1000,
173
+ "unknown_usage_calls": 0,
174
+ "unsealed_episodes_excluded": 0,
175
+ "invalid_evidence_episodes_excluded": 0,
176
+ "known_token_subtotal": {
177
+ "input_tokens": 421568,
178
+ "output_tokens": 3425749,
179
+ "total_tokens": 3847317
180
+ },
181
+ "unknown_by_field": {
182
+ "input_tokens": 0,
183
+ "output_tokens": 0,
184
+ "total_tokens": 0
185
+ },
186
+ "total_tokens_complete": true,
187
+ "usd": null,
188
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
189
+ "summed_episode_seconds": 109668.68000000007
190
+ }
191
+ }
results/deepseek-v4-pro/deepagents/README.md ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # DeepSeek V4 Pro / deepagents
2
+
3
+ 固定 1,000 题:成功 352、确定失败 648、任务 unknown 0、invalid evidence 0。已封存 1000 条;无待开始或在途记录。原生完整复现率 18.3%。
4
+
5
+ `results.jsonl` 与 `results.csv` 保存逐题状态、指标、用量和结果哈希;`summary.json` 分别保存本地 HF 四维和 local-paper-v1 五维;`config.json` 保存任务清单、协议、源代码版本和接续来源。上一级含配对表、来源与导出核验。
6
+
7
+ `resolved` 仅表示有效封存,`definite_outcome` 表示成功或确定失败。任务 unknown、无效证据与指标 null 分开保留,均不从 1,000 的分母移除。无效记录仅有 `unverified_result_sha256`,不可当作有效执行结果。后续阶段 pending 可能表示前序失败后没有执行到,不表示任务仍待派发。
8
+
9
+ 原始预测三元组不在源评分快照中,prediction 字段为空,并有明确导出状态;不能据此认为所有模型都没有生成预测。配对比较含历史协议修订,非严格等预算因果实验。本地评分尚未核实官方一致性。
10
+
11
+ 本次仅归档已接受的评分和状态,没有重新执行模型或程序。不包含 benchmark 输入、隐藏答案、原始提示词/API 轨迹、凭据或机器路径。
results/deepseek-v4-pro/deepagents/config.json ADDED
@@ -0,0 +1,1199 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "deepseek-v4-pro",
4
+ "arm": "deepagents",
5
+ "harness": "DeepAgents 0.7.13",
6
+ "archived_at_utc": "2026-09-16T17:18:47.373930+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "engineering_pilot": false,
14
+ "expected_tasks": 1000,
15
+ "task_ids": [
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4,
20
+ 5,
21
+ 6,
22
+ 7,
23
+ 8,
24
+ 9,
25
+ 10,
26
+ 11,
27
+ 12,
28
+ 13,
29
+ 14,
30
+ 15,
31
+ 16,
32
+ 17,
33
+ 18,
34
+ 19,
35
+ 20,
36
+ 21,
37
+ 22,
38
+ 23,
39
+ 24,
40
+ 25,
41
+ 26,
42
+ 27,
43
+ 28,
44
+ 29,
45
+ 30,
46
+ 31,
47
+ 32,
48
+ 33,
49
+ 34,
50
+ 35,
51
+ 36,
52
+ 37,
53
+ 38,
54
+ 39,
55
+ 40,
56
+ 41,
57
+ 42,
58
+ 43,
59
+ 44,
60
+ 45,
61
+ 46,
62
+ 47,
63
+ 48,
64
+ 49,
65
+ 50,
66
+ 51,
67
+ 52,
68
+ 53,
69
+ 54,
70
+ 55,
71
+ 56,
72
+ 57,
73
+ 58,
74
+ 59,
75
+ 60,
76
+ 61,
77
+ 62,
78
+ 63,
79
+ 64,
80
+ 65,
81
+ 66,
82
+ 67,
83
+ 68,
84
+ 69,
85
+ 70,
86
+ 71,
87
+ 72,
88
+ 73,
89
+ 74,
90
+ 75,
91
+ 76,
92
+ 77,
93
+ 78,
94
+ 79,
95
+ 80,
96
+ 81,
97
+ 82,
98
+ 83,
99
+ 84,
100
+ 85,
101
+ 86,
102
+ 87,
103
+ 88,
104
+ 89,
105
+ 90,
106
+ 91,
107
+ 92,
108
+ 93,
109
+ 94,
110
+ 95,
111
+ 96,
112
+ 97,
113
+ 98,
114
+ 100,
115
+ 101,
116
+ 102,
117
+ 103,
118
+ 104,
119
+ 105,
120
+ 106,
121
+ 107,
122
+ 108,
123
+ 109,
124
+ 110,
125
+ 111,
126
+ 112,
127
+ 113,
128
+ 114,
129
+ 115,
130
+ 116,
131
+ 117,
132
+ 118,
133
+ 119,
134
+ 120,
135
+ 121,
136
+ 122,
137
+ 123,
138
+ 124,
139
+ 125,
140
+ 126,
141
+ 127,
142
+ 128,
143
+ 129,
144
+ 130,
145
+ 131,
146
+ 132,
147
+ 133,
148
+ 134,
149
+ 135,
150
+ 136,
151
+ 137,
152
+ 138,
153
+ 139,
154
+ 140,
155
+ 141,
156
+ 142,
157
+ 143,
158
+ 144,
159
+ 145,
160
+ 146,
161
+ 147,
162
+ 148,
163
+ 149,
164
+ 150,
165
+ 151,
166
+ 152,
167
+ 153,
168
+ 154,
169
+ 155,
170
+ 156,
171
+ 157,
172
+ 158,
173
+ 159,
174
+ 160,
175
+ 161,
176
+ 162,
177
+ 163,
178
+ 164,
179
+ 165,
180
+ 166,
181
+ 167,
182
+ 168,
183
+ 169,
184
+ 170,
185
+ 171,
186
+ 172,
187
+ 173,
188
+ 174,
189
+ 175,
190
+ 176,
191
+ 177,
192
+ 178,
193
+ 179,
194
+ 180,
195
+ 181,
196
+ 182,
197
+ 183,
198
+ 184,
199
+ 185,
200
+ 186,
201
+ 187,
202
+ 188,
203
+ 189,
204
+ 190,
205
+ 191,
206
+ 192,
207
+ 193,
208
+ 194,
209
+ 195,
210
+ 196,
211
+ 197,
212
+ 198,
213
+ 199,
214
+ 200,
215
+ 201,
216
+ 202,
217
+ 203,
218
+ 204,
219
+ 205,
220
+ 206,
221
+ 207,
222
+ 208,
223
+ 209,
224
+ 210,
225
+ 211,
226
+ 212,
227
+ 213,
228
+ 214,
229
+ 215,
230
+ 216,
231
+ 217,
232
+ 218,
233
+ 219,
234
+ 220,
235
+ 221,
236
+ 222,
237
+ 223,
238
+ 224,
239
+ 225,
240
+ 226,
241
+ 227,
242
+ 228,
243
+ 229,
244
+ 230,
245
+ 231,
246
+ 232,
247
+ 233,
248
+ 234,
249
+ 235,
250
+ 236,
251
+ 237,
252
+ 238,
253
+ 239,
254
+ 240,
255
+ 241,
256
+ 242,
257
+ 248,
258
+ 249,
259
+ 250,
260
+ 251,
261
+ 252,
262
+ 253,
263
+ 254,
264
+ 255,
265
+ 256,
266
+ 257,
267
+ 258,
268
+ 259,
269
+ 261,
270
+ 262,
271
+ 263,
272
+ 264,
273
+ 265,
274
+ 266,
275
+ 290,
276
+ 291,
277
+ 292,
278
+ 293,
279
+ 294,
280
+ 295,
281
+ 302,
282
+ 303,
283
+ 304,
284
+ 305,
285
+ 306,
286
+ 307,
287
+ 308,
288
+ 309,
289
+ 310,
290
+ 311,
291
+ 312,
292
+ 313,
293
+ 314,
294
+ 315,
295
+ 316,
296
+ 317,
297
+ 318,
298
+ 319,
299
+ 320,
300
+ 321,
301
+ 322,
302
+ 323,
303
+ 324,
304
+ 325,
305
+ 326,
306
+ 327,
307
+ 328,
308
+ 329,
309
+ 330,
310
+ 331,
311
+ 332,
312
+ 333,
313
+ 334,
314
+ 335,
315
+ 336,
316
+ 337,
317
+ 338,
318
+ 339,
319
+ 340,
320
+ 341,
321
+ 342,
322
+ 343,
323
+ 344,
324
+ 345,
325
+ 346,
326
+ 347,
327
+ 348,
328
+ 349,
329
+ 350,
330
+ 351,
331
+ 352,
332
+ 353,
333
+ 354,
334
+ 355,
335
+ 356,
336
+ 357,
337
+ 358,
338
+ 359,
339
+ 360,
340
+ 361,
341
+ 362,
342
+ 363,
343
+ 364,
344
+ 365,
345
+ 366,
346
+ 367,
347
+ 368,
348
+ 369,
349
+ 370,
350
+ 371,
351
+ 372,
352
+ 373,
353
+ 374,
354
+ 375,
355
+ 376,
356
+ 377,
357
+ 378,
358
+ 379,
359
+ 380,
360
+ 381,
361
+ 382,
362
+ 383,
363
+ 384,
364
+ 385,
365
+ 386,
366
+ 387,
367
+ 388,
368
+ 389,
369
+ 390,
370
+ 391,
371
+ 392,
372
+ 393,
373
+ 394,
374
+ 395,
375
+ 396,
376
+ 397,
377
+ 398,
378
+ 399,
379
+ 400,
380
+ 401,
381
+ 402,
382
+ 403,
383
+ 404,
384
+ 405,
385
+ 406,
386
+ 407,
387
+ 408,
388
+ 409,
389
+ 410,
390
+ 411,
391
+ 412,
392
+ 413,
393
+ 414,
394
+ 415,
395
+ 416,
396
+ 417,
397
+ 418,
398
+ 419,
399
+ 420,
400
+ 421,
401
+ 422,
402
+ 423,
403
+ 424,
404
+ 425,
405
+ 426,
406
+ 427,
407
+ 428,
408
+ 429,
409
+ 430,
410
+ 431,
411
+ 432,
412
+ 433,
413
+ 434,
414
+ 435,
415
+ 436,
416
+ 437,
417
+ 438,
418
+ 439,
419
+ 440,
420
+ 441,
421
+ 442,
422
+ 443,
423
+ 444,
424
+ 445,
425
+ 446,
426
+ 447,
427
+ 448,
428
+ 449,
429
+ 450,
430
+ 451,
431
+ 452,
432
+ 453,
433
+ 454,
434
+ 455,
435
+ 456,
436
+ 457,
437
+ 458,
438
+ 459,
439
+ 460,
440
+ 461,
441
+ 462,
442
+ 463,
443
+ 464,
444
+ 465,
445
+ 466,
446
+ 467,
447
+ 468,
448
+ 469,
449
+ 470,
450
+ 471,
451
+ 472,
452
+ 473,
453
+ 474,
454
+ 475,
455
+ 476,
456
+ 477,
457
+ 478,
458
+ 479,
459
+ 480,
460
+ 481,
461
+ 482,
462
+ 483,
463
+ 484,
464
+ 485,
465
+ 486,
466
+ 487,
467
+ 488,
468
+ 489,
469
+ 490,
470
+ 491,
471
+ 492,
472
+ 493,
473
+ 494,
474
+ 495,
475
+ 496,
476
+ 497,
477
+ 498,
478
+ 499,
479
+ 500,
480
+ 501,
481
+ 502,
482
+ 503,
483
+ 504,
484
+ 505,
485
+ 506,
486
+ 507,
487
+ 508,
488
+ 509,
489
+ 510,
490
+ 511,
491
+ 512,
492
+ 513,
493
+ 514,
494
+ 515,
495
+ 516,
496
+ 517,
497
+ 518,
498
+ 519,
499
+ 520,
500
+ 521,
501
+ 522,
502
+ 523,
503
+ 524,
504
+ 525,
505
+ 526,
506
+ 527,
507
+ 528,
508
+ 529,
509
+ 530,
510
+ 531,
511
+ 532,
512
+ 533,
513
+ 534,
514
+ 535,
515
+ 536,
516
+ 537,
517
+ 538,
518
+ 539,
519
+ 540,
520
+ 541,
521
+ 542,
522
+ 543,
523
+ 544,
524
+ 545,
525
+ 546,
526
+ 547,
527
+ 548,
528
+ 549,
529
+ 550,
530
+ 551,
531
+ 552,
532
+ 553,
533
+ 554,
534
+ 555,
535
+ 556,
536
+ 557,
537
+ 558,
538
+ 559,
539
+ 560,
540
+ 561,
541
+ 562,
542
+ 563,
543
+ 565,
544
+ 566,
545
+ 567,
546
+ 568,
547
+ 570,
548
+ 571,
549
+ 572,
550
+ 573,
551
+ 574,
552
+ 575,
553
+ 576,
554
+ 577,
555
+ 578,
556
+ 579,
557
+ 580,
558
+ 581,
559
+ 582,
560
+ 583,
561
+ 584,
562
+ 585,
563
+ 586,
564
+ 587,
565
+ 588,
566
+ 589,
567
+ 590,
568
+ 591,
569
+ 592,
570
+ 593,
571
+ 594,
572
+ 595,
573
+ 596,
574
+ 597,
575
+ 598,
576
+ 601,
577
+ 602,
578
+ 603,
579
+ 604,
580
+ 605,
581
+ 606,
582
+ 607,
583
+ 608,
584
+ 609,
585
+ 610,
586
+ 611,
587
+ 612,
588
+ 613,
589
+ 614,
590
+ 615,
591
+ 616,
592
+ 617,
593
+ 618,
594
+ 619,
595
+ 620,
596
+ 621,
597
+ 622,
598
+ 623,
599
+ 624,
600
+ 664,
601
+ 665,
602
+ 666,
603
+ 667,
604
+ 668,
605
+ 669,
606
+ 670,
607
+ 671,
608
+ 672,
609
+ 673,
610
+ 674,
611
+ 675,
612
+ 676,
613
+ 677,
614
+ 678,
615
+ 679,
616
+ 680,
617
+ 681,
618
+ 682,
619
+ 683,
620
+ 684,
621
+ 685,
622
+ 686,
623
+ 687,
624
+ 688,
625
+ 689,
626
+ 690,
627
+ 691,
628
+ 692,
629
+ 693,
630
+ 694,
631
+ 695,
632
+ 696,
633
+ 697,
634
+ 698,
635
+ 699,
636
+ 700,
637
+ 701,
638
+ 702,
639
+ 703,
640
+ 704,
641
+ 705,
642
+ 706,
643
+ 707,
644
+ 708,
645
+ 709,
646
+ 710,
647
+ 711,
648
+ 712,
649
+ 713,
650
+ 714,
651
+ 715,
652
+ 716,
653
+ 717,
654
+ 718,
655
+ 719,
656
+ 720,
657
+ 721,
658
+ 722,
659
+ 723,
660
+ 724,
661
+ 725,
662
+ 726,
663
+ 727,
664
+ 728,
665
+ 729,
666
+ 730,
667
+ 731,
668
+ 732,
669
+ 733,
670
+ 734,
671
+ 735,
672
+ 736,
673
+ 737,
674
+ 739,
675
+ 740,
676
+ 741,
677
+ 742,
678
+ 743,
679
+ 744,
680
+ 745,
681
+ 746,
682
+ 747,
683
+ 748,
684
+ 749,
685
+ 750,
686
+ 751,
687
+ 752,
688
+ 753,
689
+ 754,
690
+ 755,
691
+ 756,
692
+ 757,
693
+ 758,
694
+ 759,
695
+ 760,
696
+ 761,
697
+ 762,
698
+ 763,
699
+ 764,
700
+ 765,
701
+ 766,
702
+ 767,
703
+ 768,
704
+ 772,
705
+ 773,
706
+ 774,
707
+ 775,
708
+ 776,
709
+ 777,
710
+ 778,
711
+ 779,
712
+ 780,
713
+ 781,
714
+ 782,
715
+ 783,
716
+ 784,
717
+ 785,
718
+ 786,
719
+ 787,
720
+ 788,
721
+ 789,
722
+ 790,
723
+ 791,
724
+ 793,
725
+ 794,
726
+ 795,
727
+ 796,
728
+ 797,
729
+ 798,
730
+ 799,
731
+ 800,
732
+ 801,
733
+ 802,
734
+ 803,
735
+ 804,
736
+ 805,
737
+ 806,
738
+ 807,
739
+ 808,
740
+ 825,
741
+ 826,
742
+ 827,
743
+ 828,
744
+ 829,
745
+ 830,
746
+ 831,
747
+ 832,
748
+ 833,
749
+ 834,
750
+ 835,
751
+ 836,
752
+ 837,
753
+ 838,
754
+ 839,
755
+ 840,
756
+ 841,
757
+ 842,
758
+ 843,
759
+ 844,
760
+ 845,
761
+ 846,
762
+ 847,
763
+ 848,
764
+ 849,
765
+ 850,
766
+ 851,
767
+ 852,
768
+ 853,
769
+ 854,
770
+ 855,
771
+ 856,
772
+ 857,
773
+ 858,
774
+ 859,
775
+ 860,
776
+ 861,
777
+ 862,
778
+ 863,
779
+ 864,
780
+ 865,
781
+ 866,
782
+ 867,
783
+ 868,
784
+ 869,
785
+ 870,
786
+ 871,
787
+ 872,
788
+ 873,
789
+ 874,
790
+ 875,
791
+ 876,
792
+ 877,
793
+ 878,
794
+ 879,
795
+ 880,
796
+ 881,
797
+ 882,
798
+ 883,
799
+ 884,
800
+ 885,
801
+ 886,
802
+ 887,
803
+ 888,
804
+ 889,
805
+ 890,
806
+ 891,
807
+ 892,
808
+ 893,
809
+ 894,
810
+ 895,
811
+ 896,
812
+ 897,
813
+ 898,
814
+ 899,
815
+ 900,
816
+ 901,
817
+ 902,
818
+ 903,
819
+ 906,
820
+ 907,
821
+ 913,
822
+ 914,
823
+ 915,
824
+ 917,
825
+ 920,
826
+ 921,
827
+ 922,
828
+ 923,
829
+ 924,
830
+ 925,
831
+ 926,
832
+ 927,
833
+ 928,
834
+ 929,
835
+ 930,
836
+ 931,
837
+ 932,
838
+ 933,
839
+ 934,
840
+ 935,
841
+ 936,
842
+ 937,
843
+ 938,
844
+ 939,
845
+ 940,
846
+ 941,
847
+ 942,
848
+ 943,
849
+ 944,
850
+ 945,
851
+ 946,
852
+ 947,
853
+ 948,
854
+ 949,
855
+ 950,
856
+ 951,
857
+ 952,
858
+ 953,
859
+ 954,
860
+ 955,
861
+ 956,
862
+ 957,
863
+ 958,
864
+ 959,
865
+ 960,
866
+ 961,
867
+ 962,
868
+ 963,
869
+ 964,
870
+ 965,
871
+ 966,
872
+ 967,
873
+ 968,
874
+ 969,
875
+ 970,
876
+ 971,
877
+ 972,
878
+ 973,
879
+ 974,
880
+ 975,
881
+ 976,
882
+ 977,
883
+ 978,
884
+ 979,
885
+ 980,
886
+ 981,
887
+ 982,
888
+ 983,
889
+ 984,
890
+ 985,
891
+ 1001,
892
+ 1002,
893
+ 1003,
894
+ 1004,
895
+ 1005,
896
+ 1006,
897
+ 1007,
898
+ 1008,
899
+ 1009,
900
+ 1010,
901
+ 1011,
902
+ 1012,
903
+ 1013,
904
+ 1014,
905
+ 1015,
906
+ 1016,
907
+ 1017,
908
+ 1018,
909
+ 1019,
910
+ 1020,
911
+ 1021,
912
+ 1022,
913
+ 1023,
914
+ 1024,
915
+ 1025,
916
+ 1026,
917
+ 1027,
918
+ 1028,
919
+ 1029,
920
+ 1030,
921
+ 1031,
922
+ 1032,
923
+ 1033,
924
+ 1034,
925
+ 1035,
926
+ 1036,
927
+ 1037,
928
+ 1038,
929
+ 1039,
930
+ 1040,
931
+ 1041,
932
+ 1042,
933
+ 1043,
934
+ 1044,
935
+ 1045,
936
+ 1046,
937
+ 1047,
938
+ 1048,
939
+ 1049,
940
+ 1050,
941
+ 1051,
942
+ 1052,
943
+ 1053,
944
+ 1054,
945
+ 1055,
946
+ 1056,
947
+ 1057,
948
+ 1058,
949
+ 1059,
950
+ 1060,
951
+ 1061,
952
+ 1062,
953
+ 1063,
954
+ 1064,
955
+ 1065,
956
+ 1066,
957
+ 1067,
958
+ 1068,
959
+ 1069,
960
+ 1070,
961
+ 1071,
962
+ 1072,
963
+ 1073,
964
+ 1074,
965
+ 1075,
966
+ 1076,
967
+ 1077,
968
+ 1078,
969
+ 1079,
970
+ 1080,
971
+ 1081,
972
+ 1082,
973
+ 1083,
974
+ 1084,
975
+ 1085,
976
+ 1086,
977
+ 1087,
978
+ 1088,
979
+ 1089,
980
+ 1090,
981
+ 1091,
982
+ 1092,
983
+ 1093,
984
+ 1094,
985
+ 1095,
986
+ 1096,
987
+ 1097,
988
+ 1098,
989
+ 1099,
990
+ 1100,
991
+ 1101,
992
+ 1102,
993
+ 1103,
994
+ 1104,
995
+ 1105,
996
+ 1106,
997
+ 1107,
998
+ 1108,
999
+ 1109,
1000
+ 1110,
1001
+ 1111,
1002
+ 1112,
1003
+ 1113,
1004
+ 1114,
1005
+ 1115,
1006
+ 1116,
1007
+ 1117,
1008
+ 1118,
1009
+ 1119,
1010
+ 1120,
1011
+ 1121,
1012
+ 1122,
1013
+ 1123,
1014
+ 1124,
1015
+ 1125
1016
+ ]
1017
+ },
1018
+ "generation": {
1019
+ "model": "deepseek-v4-pro",
1020
+ "provider": "chat-completions",
1021
+ "reasoning_effort": null,
1022
+ "stream": true,
1023
+ "timeout": 1800
1024
+ },
1025
+ "effective_model_call_cap": 6,
1026
+ "recorded_protocols": [
1027
+ {
1028
+ "core_sha256": {
1029
+ "protocol.json": "86b59ee1ca819f74a6d0c5211583a34e1b9665460ab0c965b52da81956bbec35",
1030
+ "inputs.manifest.json": "5e74e33491e38eda692fe3209d81ab85150861cb56bccc5b726712c19d5e49a2",
1031
+ "schedule.json": "11f80f6caf79724f907ef2133a24095d93404494c32232dedab39ef5c8871542",
1032
+ "study.json": "f85ecac78dac79575bfc8b8f3cb4e68c36fa3b9f58b45c2b4c196e4ad2f8847e",
1033
+ "environment.json": "04f141c80a13f3620f971c3217dccd28c2ee58f22a3d1fc5f7db2bfd1f4cc2ef"
1034
+ },
1035
+ "population_sha256": "05d427330ec89d1449dfdfd620fd238fa99321dbc954af8fe934b127887d65a1",
1036
+ "task_count": 999,
1037
+ "task_input_blobs_rehashed": false,
1038
+ "source_and_dependencies_match": true,
1039
+ "limits": {
1040
+ "output_tokens": 32768,
1041
+ "wall_seconds": 1800,
1042
+ "model_calls": 6,
1043
+ "tool_calls": 4,
1044
+ "tool_timeout": 90
1045
+ },
1046
+ "generation": {
1047
+ "model": "deepseek-v4-pro",
1048
+ "provider": "chat-completions",
1049
+ "reasoning_effort": null,
1050
+ "stream": true,
1051
+ "timeout": 1800
1052
+ },
1053
+ "dataset_provenance": {
1054
+ "repo_id": "CamoAiLab/InferenceNet",
1055
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1056
+ "task_directory": "Selected_1000",
1057
+ "task_list": "Selected_1000/1000_new.csv",
1058
+ "prior_results_imported": false,
1059
+ "engineering_pilot": false
1060
+ },
1061
+ "execution_identity": {
1062
+ "backend": "dsw-bwrap-v3",
1063
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1064
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1065
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1066
+ "engine_memory_bytes": 38654705664,
1067
+ "engine_cpus": 12,
1068
+ "cpu_slots": [
1069
+ [
1070
+ 4,
1071
+ 5
1072
+ ],
1073
+ [
1074
+ 6,
1075
+ 7
1076
+ ],
1077
+ [
1078
+ 8,
1079
+ 9
1080
+ ],
1081
+ [
1082
+ 10,
1083
+ 11
1084
+ ],
1085
+ [
1086
+ 12,
1087
+ 13
1088
+ ],
1089
+ [
1090
+ 14,
1091
+ 15
1092
+ ]
1093
+ ],
1094
+ "resources": {
1095
+ "cpus": 2.0,
1096
+ "memory": "6g",
1097
+ "tmpfs": "256m"
1098
+ },
1099
+ "memory_enforcement": "RLIMIT_AS",
1100
+ "single_payload_process": true,
1101
+ "network": "none",
1102
+ "uid": 10001,
1103
+ "capacity_profile": "independent-six-slots-20260916",
1104
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1105
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1106
+ },
1107
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1108
+ "scorer_sha256": {
1109
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1110
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1111
+ },
1112
+ "origin": "child"
1113
+ },
1114
+ {
1115
+ "core_sha256": {
1116
+ "protocol.json": "39af792a5f92b7424f22a55602657d6f930ab37bc5607263c91c48f3b46c6c72",
1117
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1118
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1119
+ "study.json": "57177147bfcead6ff01f6c86d5225aa8542d72e019986b83a9fb4a6c437abb61",
1120
+ "environment.json": "f0e702a8a418d683763e0a0d35acd4377d246e006d1c7533c9bf6cac38fd70bf"
1121
+ },
1122
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1123
+ "task_count": 1000,
1124
+ "task_input_blobs_rehashed": false,
1125
+ "source_and_dependencies_match": true,
1126
+ "limits": {
1127
+ "output_tokens": 32768,
1128
+ "wall_seconds": 1800,
1129
+ "model_calls": 6,
1130
+ "tool_calls": 4,
1131
+ "tool_timeout": 90
1132
+ },
1133
+ "generation": {
1134
+ "model": "deepseek-v4-pro",
1135
+ "provider": "chat-completions",
1136
+ "reasoning_effort": null,
1137
+ "stream": true,
1138
+ "timeout": 1800
1139
+ },
1140
+ "dataset_provenance": {
1141
+ "repo_id": "CamoAiLab/InferenceNet",
1142
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1143
+ "task_directory": "Selected_1000",
1144
+ "task_list": "Selected_1000/1000_new.csv",
1145
+ "prior_results_imported": false,
1146
+ "engineering_pilot": false
1147
+ },
1148
+ "execution_identity": {
1149
+ "backend": "dsw-bwrap-v3",
1150
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1151
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1152
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1153
+ "engine_memory_bytes": 12884901888,
1154
+ "engine_cpus": 4,
1155
+ "cpu_slots": [
1156
+ [
1157
+ 16,
1158
+ 17
1159
+ ],
1160
+ [
1161
+ 18,
1162
+ 19
1163
+ ]
1164
+ ],
1165
+ "resources": {
1166
+ "memory": "6g",
1167
+ "cpus": 2.0,
1168
+ "tmpfs": "256m"
1169
+ },
1170
+ "memory_enforcement": "RLIMIT_AS",
1171
+ "single_payload_process": true,
1172
+ "network": "none",
1173
+ "uid": 10001
1174
+ },
1175
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1176
+ "scorer_sha256": {
1177
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1178
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1179
+ },
1180
+ "origin": "parent"
1181
+ }
1182
+ ],
1183
+ "selected_origin_counts": {
1184
+ "child": 999,
1185
+ "parent": 1
1186
+ },
1187
+ "source_commits": [
1188
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1189
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1190
+ ],
1191
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1192
+ "official_parity_verified": false,
1193
+ "notes": [
1194
+ "Historical disjoint parent/child continuation; original source and protocol hashes retained.",
1195
+ "Different baseline/harness call budgets; not an equal-cost causal comparison.",
1196
+ "Unknown and invalid evidence remain in the denominator; no replacement or rescoring.",
1197
+ "Invalid evidence is not sealed; its unverified hash is distinct from a verified result hash."
1198
+ ]
1199
+ }
results/deepseek-v4-pro/deepagents/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/deepagents/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/deepagents/summary.json ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
5
+ "model": "deepseek-v4-pro",
6
+ "arm": "deepagents",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": true,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 352,
37
+ "denominator": 1000,
38
+ "rate": 0.352,
39
+ "score": 35.2,
40
+ "unknown_count": 0,
41
+ "failure_count": 648,
42
+ "assessable": 1000,
43
+ "coverage_percent": 100.0
44
+ },
45
+ "partial_replication": {
46
+ "count": 270,
47
+ "denominator": 1000,
48
+ "rate": 0.27,
49
+ "score": 27.0,
50
+ "unknown_count": 648,
51
+ "failure_count": 82,
52
+ "assessable": 352,
53
+ "coverage_percent": 35.2
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 338,
57
+ "denominator": 1000,
58
+ "rate": 0.338,
59
+ "score": 33.8,
60
+ "unknown_count": 648,
61
+ "failure_count": 14,
62
+ "assessable": 352,
63
+ "coverage_percent": 35.2
64
+ },
65
+ "significance_level": {
66
+ "count": 308,
67
+ "denominator": 1000,
68
+ "rate": 0.308,
69
+ "score": 30.8,
70
+ "unknown_count": 648,
71
+ "failure_count": 44,
72
+ "assessable": 352,
73
+ "coverage_percent": 35.2
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "deepseek-v4-pro deepagents",
78
+ "Compilation Success": 35.2,
79
+ "Partial Replication": 27.0,
80
+ "Correct Coefficient Direction": 33.8,
81
+ "Significant Level Correctness": 30.8
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 1000,
94
+ "task_unknown_count": 0,
95
+ "invalid_evidence_count": 0,
96
+ "valid_prediction_count": 352,
97
+ "task_counts": {
98
+ "succeeded": 352,
99
+ "failed": 648,
100
+ "unknown": 0,
101
+ "unsealed": 0,
102
+ "not_started": 0,
103
+ "invalid_evidence": 0,
104
+ "completed": 1000,
105
+ "planned": 1000
106
+ }
107
+ },
108
+ "legacy_local_paper": {
109
+ "profile": "local-paper-v1",
110
+ "definitions": {
111
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
112
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
113
+ "coefficient_only": "coefficient relative error <= .05",
114
+ "direction": "matching strictly positive or strictly negative coefficient signs",
115
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
116
+ },
117
+ "metrics": {
118
+ "perfect": {
119
+ "count": 183,
120
+ "denominator": 1000,
121
+ "rate": 0.183,
122
+ "score": 18.3,
123
+ "unknown_count": 648,
124
+ "failure_count": 169,
125
+ "assessable": 352,
126
+ "coverage_percent": 35.2
127
+ },
128
+ "partial": {
129
+ "count": 232,
130
+ "denominator": 1000,
131
+ "rate": 0.232,
132
+ "score": 23.2,
133
+ "unknown_count": 648,
134
+ "failure_count": 120,
135
+ "assessable": 352,
136
+ "coverage_percent": 35.2
137
+ },
138
+ "coefficient_only": {
139
+ "count": 270,
140
+ "denominator": 1000,
141
+ "rate": 0.27,
142
+ "score": 27.0,
143
+ "unknown_count": 648,
144
+ "failure_count": 82,
145
+ "assessable": 352,
146
+ "coverage_percent": 35.2
147
+ },
148
+ "direction": {
149
+ "count": 338,
150
+ "denominator": 1000,
151
+ "rate": 0.338,
152
+ "score": 33.8,
153
+ "unknown_count": 648,
154
+ "failure_count": 14,
155
+ "assessable": 352,
156
+ "coverage_percent": 35.2
157
+ },
158
+ "significance": {
159
+ "count": 299,
160
+ "denominator": 1000,
161
+ "rate": 0.299,
162
+ "score": 29.9,
163
+ "unknown_count": 648,
164
+ "failure_count": 53,
165
+ "assessable": 352,
166
+ "coverage_percent": 35.2
167
+ }
168
+ }
169
+ },
170
+ "verification_scope": "Offline equality to accepted flags, hashes and states; raw predictions unavailable; not a fresh native audit or official scorer parity check.",
171
+ "usage": {
172
+ "sealed_model_calls": 5812,
173
+ "unknown_usage_calls": 0,
174
+ "unsealed_episodes_excluded": 0,
175
+ "invalid_evidence_episodes_excluded": 0,
176
+ "known_token_subtotal": {
177
+ "input_tokens": 35729626,
178
+ "output_tokens": 3614764,
179
+ "total_tokens": 39344390
180
+ },
181
+ "unknown_by_field": {
182
+ "input_tokens": 0,
183
+ "output_tokens": 0,
184
+ "total_tokens": 0
185
+ },
186
+ "total_tokens_complete": true,
187
+ "usd": null,
188
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
189
+ "summed_episode_seconds": 218448.4960000001
190
+ }
191
+ }
results/deepseek-v4-pro/diagnostics.json ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage_counts": {
3
+ "baseline": [
4
+ {
5
+ "generation": "succeeded",
6
+ "execution": "failed",
7
+ "scoring": "pending",
8
+ "count": 723
9
+ },
10
+ {
11
+ "generation": "succeeded",
12
+ "execution": "succeeded",
13
+ "scoring": "scored",
14
+ "count": 277
15
+ }
16
+ ],
17
+ "deepagents": [
18
+ {
19
+ "generation": "failed",
20
+ "execution": "pending",
21
+ "scoring": "pending",
22
+ "count": 635
23
+ },
24
+ {
25
+ "generation": "succeeded",
26
+ "execution": "succeeded",
27
+ "scoring": "scored",
28
+ "count": 352
29
+ },
30
+ {
31
+ "generation": "succeeded",
32
+ "execution": "failed",
33
+ "scoring": "pending",
34
+ "count": 13
35
+ }
36
+ ]
37
+ },
38
+ "task_unknown_ids": {
39
+ "baseline": [],
40
+ "deepagents": []
41
+ },
42
+ "invalid_evidence_ids": {
43
+ "baseline": [],
44
+ "deepagents": []
45
+ },
46
+ "raw_error_reasons_verified": false
47
+ }
results/deepseek-v4-pro/export-verification.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "passed_offline",
3
+ "per_arm": {
4
+ "baseline": {
5
+ "records": 1000,
6
+ "unique_task_ids": 1000,
7
+ "counts": {
8
+ "succeeded": 277,
9
+ "failed": 723,
10
+ "unknown": 0,
11
+ "unsealed": 0,
12
+ "not_started": 0,
13
+ "invalid_evidence": 0,
14
+ "completed": 1000,
15
+ "planned": 1000
16
+ },
17
+ "sealed_count": 1000,
18
+ "all_slots_sealed": true,
19
+ "metrics_recomputed": 9,
20
+ "matches_accepted_flags_states_hashes": true
21
+ },
22
+ "deepagents": {
23
+ "records": 1000,
24
+ "unique_task_ids": 1000,
25
+ "counts": {
26
+ "succeeded": 352,
27
+ "failed": 648,
28
+ "unknown": 0,
29
+ "unsealed": 0,
30
+ "not_started": 0,
31
+ "invalid_evidence": 0,
32
+ "completed": 1000,
33
+ "planned": 1000
34
+ },
35
+ "sealed_count": 1000,
36
+ "all_slots_sealed": true,
37
+ "metrics_recomputed": 9,
38
+ "matches_accepted_flags_states_hashes": true
39
+ }
40
+ },
41
+ "paired_rows": 1000,
42
+ "metrics_recomputed": 18,
43
+ "official_parity_verified": false
44
+ }
results/deepseek-v4-pro/paired_comparison.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/deepseek-v4-pro/provenance.json ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "deepseek-v4-pro",
3
+ "archived_at_utc": "2026-09-16T17:18:47.373930+00:00",
4
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
5
+ "accepted_results_sha256": "dd72f352be1895b11914dd2aea1fdceb21c453932877c588ecaf81ec09861105",
6
+ "accepted_task_records_sha256": "32a797b4da8dfe90d44b37d7715fe726a29c5b4f93a1b24d7d5ffbffb81e9749",
7
+ "record_count": 2000,
8
+ "selected_task_count": 1000,
9
+ "source_commits": [
10
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
11
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
12
+ ],
13
+ "historical_native_audit": {
14
+ "mode": "historical_final_cache",
15
+ "snapshot_sha256": "7de17ac129bbb07ac7d1fa08adb92642eeb7cfadab4d4a68946661308d8adde1",
16
+ "journal_sha256": "d35e92010e0f85de995ef587f3d64616260a70600be6749644c33fb287d1a32b",
17
+ "acceptance_sha256": "8ae98f599a7874a5c5f5c30c2cd735ad074a2790efe0c726c0e06c2686b27b94",
18
+ "collector_sha256": "7963b4fd6ef312307d5471d031e6851429f139b03a3e6d7ffc7ca9bd44d98c89",
19
+ "renderer_sha256": "b2cb2f0ae0523431e3be13ac23a6730c8e6a511e7c645921dea79a1cacffe2fd",
20
+ "started_at": "2026-09-16T01:01:30.976377+00:00",
21
+ "finished_at": "2026-09-16T02:24:03.769359+00:00",
22
+ "model_sha256": "1ea1971a9930edbc23942fdcc07bd847b9de00b90d5e9245618b03a02e49c302",
23
+ "source_checkpoint": {
24
+ "record_sha256": "b607ce017fbb746b127d6cbd9a09edbbd43d9e57e6e20d448c32895ff38b2fe7",
25
+ "record_bytes": 4406227
26
+ },
27
+ "current_identity_check": {
28
+ "kind": "current_source_core_result_hashes_parent_absence",
29
+ "started_at": "2026-09-16T07:43:01.435804+00:00",
30
+ "finished_at": "2026-09-16T07:43:33.877870+00:00",
31
+ "result_hashes_checked": 2000,
32
+ "absent_parent_slots_checked": 1998,
33
+ "native_artifacts_reaudited": false,
34
+ "input_blobs_rehashed": false
35
+ },
36
+ "native_audit_scope": "Historical native artifact/execution/scoring verification; input blobs were not rehashed by that collector."
37
+ },
38
+ "official_parity_verified": false,
39
+ "new_model_calls": 0,
40
+ "new_benchmark_executions": 0,
41
+ "verification_scope": "Offline equality and metric reaggregation of accepted score exports. No fresh input/native audit; no raw prediction rescoring."
42
+ }
results/index.json CHANGED
@@ -51,7 +51,31 @@
51
  "arm": "deepagents",
52
  "rows": 1000,
53
  "path": "results/gemini-3.1-pro-preview/deepagents"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
54
  }
55
  ],
56
- "updated_at_utc": "2026-09-16T15:28:00.953066+00:00"
57
  }
 
51
  "arm": "deepagents",
52
  "rows": 1000,
53
  "path": "results/gemini-3.1-pro-preview/deepagents"
54
+ },
55
+ {
56
+ "model": "qwen3.7-max",
57
+ "arm": "baseline",
58
+ "rows": 1000,
59
+ "path": "results/qwen3.7-max/baseline"
60
+ },
61
+ {
62
+ "model": "qwen3.7-max",
63
+ "arm": "deepagents",
64
+ "rows": 1000,
65
+ "path": "results/qwen3.7-max/deepagents"
66
+ },
67
+ {
68
+ "model": "deepseek-v4-pro",
69
+ "arm": "baseline",
70
+ "rows": 1000,
71
+ "path": "results/deepseek-v4-pro/baseline"
72
+ },
73
+ {
74
+ "model": "deepseek-v4-pro",
75
+ "arm": "deepagents",
76
+ "rows": 1000,
77
+ "path": "results/deepseek-v4-pro/deepagents"
78
  }
79
  ],
80
+ "updated_at_utc": "2026-09-16T17:18:47.373930+00:00"
81
  }
results/manifest.sha256 CHANGED
@@ -1,4 +1,4 @@
1
- f09b15f887f461339ba94542280056bbc9c519eeb3128c64cbab5a00390e7e2e README.md
2
  2ad22f4778525b1460f1857e1095f2eea5620063fba1ffa6314ca7fd50f14703 claude-opus-4-8/baseline/README.md
3
  8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56 claude-opus-4-8/baseline/config.json
4
  4d5f859e472b79195c9be6b1f0e9c5feecf59a81dfe32ac08eb4c1d4967ed98e claude-opus-4-8/baseline/results.csv
@@ -9,6 +9,20 @@ dfb00967c2a2ff320d5a6d8c142cbc6ef37d4fb144abcc7a451673515c2b4c50 claude-opus-4-
9
  bafc0a72cd4d8f53c1da33030250171f36ce441488bc8456cba8c3f52b80e6bd claude-opus-4-8/deepagents/results.csv
10
  3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a claude-opus-4-8/deepagents/results.jsonl
11
  f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614 claude-opus-4-8/deepagents/summary.json
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12
  ac7253c934e43a88d0a2f6b105d1bb13424ef004e6838d2a632d5e9c366c72f2 gemini-3.1-pro-preview/baseline/README.md
13
  ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998 gemini-3.1-pro-preview/baseline/config.json
14
  fc708f327a76b22ebe6f453206b0c42e4cf46176afb8cb12eda2c1b71fe90bde gemini-3.1-pro-preview/baseline/results.csv
@@ -33,7 +47,7 @@ a3e24175f2395cffbbcd6d9d3c89689bd3a3b22dd4ce61bb78d5da33c6b6cc1b gpt-5.6-sol/de
33
  27abe57119d4be49c56b29c318a1bc2757a97f44427764d7fc8d188be4d10671 gpt-5.6-sol/deepagents/results.csv
34
  378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be gpt-5.6-sol/deepagents/results.jsonl
35
  821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5 gpt-5.6-sol/deepagents/summary.json
36
- 2b6c21661ef74f692d64095f1e3fd64f9e696cd9876b5186810ed340aff0c24e index.json
37
  29825997a250cc8fe292450006b82a392fe90c8e0493a1335f74eea72c1de104 kimi-k3/baseline/README.md
38
  7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081 kimi-k3/baseline/config.json
39
  f87d2dd3b390695f6b420a78e126b504abc3d292180a3443deaead174ec1b464 kimi-k3/baseline/results.csv
@@ -48,4 +62,19 @@ bfaac06563fc53d5f564197fdfb69e24ddbdff622b85d67bbf2a13fb7fb5d9ba kimi-k3/diagno
48
  b96e0cc3352efac6140c33cef30873e522db547891101c3c49c6aca42085becb kimi-k3/export-verification.json
49
  3fe1378a23cfe59a37558521ec5bf47ea299a6ad4903a14684147a1b8066ee03 kimi-k3/paired_comparison.csv
50
  4a0f6dde38c5410a6f379192f30eac6d330403c80fcdbc2b357dbe15427a331b kimi-k3/provenance.json
 
 
 
 
 
 
 
 
 
 
 
 
 
 
51
  06ed669caaba3ca5e252d12af3de32223019d50aefd9c46751ca3cb35110fce2 verify_kimi_gemini.py
 
 
1
+ 4ca82d5601f12b48b24e22beb1ef624a17f1369074d296a6151bfd6f4a5d0668 README.md
2
  2ad22f4778525b1460f1857e1095f2eea5620063fba1ffa6314ca7fd50f14703 claude-opus-4-8/baseline/README.md
3
  8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56 claude-opus-4-8/baseline/config.json
4
  4d5f859e472b79195c9be6b1f0e9c5feecf59a81dfe32ac08eb4c1d4967ed98e claude-opus-4-8/baseline/results.csv
 
9
  bafc0a72cd4d8f53c1da33030250171f36ce441488bc8456cba8c3f52b80e6bd claude-opus-4-8/deepagents/results.csv
10
  3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a claude-opus-4-8/deepagents/results.jsonl
11
  f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614 claude-opus-4-8/deepagents/summary.json
12
+ c738e75419e319515b15c823300a16e9bdf1cbe790612bbbc24ff93c71438b3f deepseek-v4-pro/baseline/README.md
13
+ 4cef524f37a325029ea4ad325b163f72f9cf725f507f20cd76424d06e71e24ad deepseek-v4-pro/baseline/config.json
14
+ 98f4f77003f4788509360a2cbe41554279df408fd20d65058b85836a25858ecb deepseek-v4-pro/baseline/results.csv
15
+ 7827218f6f18f39988e8c439f0c93cab611c04d158d8704aa50c9274cd438837 deepseek-v4-pro/baseline/results.jsonl
16
+ c51e4f1de78eba8f30e4a59276fb7821132cd103f0ae3f9d04c75c98344f2e34 deepseek-v4-pro/baseline/summary.json
17
+ f1297897da23ab01751289cd93da69a944834af59204cd73f7fca64a112bd6a9 deepseek-v4-pro/deepagents/README.md
18
+ 400c28e4008d5241bf50bc71c2674861846d6b6d2e7da21f2a7602a60ecc0ae9 deepseek-v4-pro/deepagents/config.json
19
+ a842bac2b7b608ab124d95883194e01700c6fba648593a2128db6f9d40511cf8 deepseek-v4-pro/deepagents/results.csv
20
+ a006c8b0b80064295faeb3aa7395a2408f0208677853448816b361bf67e39bb6 deepseek-v4-pro/deepagents/results.jsonl
21
+ 231e4264b0403e12788136d01af465f86fe9ba3c3ab992510925a7ef4c2a40e9 deepseek-v4-pro/deepagents/summary.json
22
+ eb38fada3bb400d40e37fd97aeb41c8eea6a1475406395e12ad9ac1c14f71e34 deepseek-v4-pro/diagnostics.json
23
+ d811ae94edc6589e2b7ee8c39ac6c4da1c66b5c3dad8acbc589b17767f9c5690 deepseek-v4-pro/export-verification.json
24
+ 891a8e0c9e1be24652b56d477bbdc0c98cf961c18010177535d33395dc43a132 deepseek-v4-pro/paired_comparison.csv
25
+ 8be2d803fc96ad6a7840c0c3eb52b52e0e1d8e8523f6a4349b4adfbdc4c320f3 deepseek-v4-pro/provenance.json
26
  ac7253c934e43a88d0a2f6b105d1bb13424ef004e6838d2a632d5e9c366c72f2 gemini-3.1-pro-preview/baseline/README.md
27
  ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998 gemini-3.1-pro-preview/baseline/config.json
28
  fc708f327a76b22ebe6f453206b0c42e4cf46176afb8cb12eda2c1b71fe90bde gemini-3.1-pro-preview/baseline/results.csv
 
47
  27abe57119d4be49c56b29c318a1bc2757a97f44427764d7fc8d188be4d10671 gpt-5.6-sol/deepagents/results.csv
48
  378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be gpt-5.6-sol/deepagents/results.jsonl
49
  821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5 gpt-5.6-sol/deepagents/summary.json
50
+ b6376d43e3107a6ac45647cdd1b0fa382437f92e33655715146d4cf6e89467d5 index.json
51
  29825997a250cc8fe292450006b82a392fe90c8e0493a1335f74eea72c1de104 kimi-k3/baseline/README.md
52
  7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081 kimi-k3/baseline/config.json
53
  f87d2dd3b390695f6b420a78e126b504abc3d292180a3443deaead174ec1b464 kimi-k3/baseline/results.csv
 
62
  b96e0cc3352efac6140c33cef30873e522db547891101c3c49c6aca42085becb kimi-k3/export-verification.json
63
  3fe1378a23cfe59a37558521ec5bf47ea299a6ad4903a14684147a1b8066ee03 kimi-k3/paired_comparison.csv
64
  4a0f6dde38c5410a6f379192f30eac6d330403c80fcdbc2b357dbe15427a331b kimi-k3/provenance.json
65
+ 1606c3170bec39e7c074659fd8c8d011e02e98b1e69161338554a9e1c416a363 qwen3.7-max/baseline/README.md
66
+ 9f553440e741863e82f32fec11470affe6f0dbffa03c668c647bd63d8f73227f qwen3.7-max/baseline/config.json
67
+ f88769bb10cbebd47ed1101a7e41c4cc5ba680bf49a768ef010aa3681bfc3845 qwen3.7-max/baseline/results.csv
68
+ 4924e929d67956f9deb6c319605042d09f40b12f7bd2f18baa0464b8f7c9851a qwen3.7-max/baseline/results.jsonl
69
+ e506814d33789673fcd6b8e6869cdcbd52371cd3d0455ce53ae27e1098a52214 qwen3.7-max/baseline/summary.json
70
+ c032c3f9953bbd6e7685028f1b08be3539aee4fc9d186c941b5a248895c2289c qwen3.7-max/deepagents/README.md
71
+ 3862d438d381ca3e57b1bf68921ef01a8c38727902152ada2c6e1fcf7b70cc77 qwen3.7-max/deepagents/config.json
72
+ 09d09b2411c67b5c026e5d82662b016f36bd7373563db541e40b2fa0429998dc qwen3.7-max/deepagents/results.csv
73
+ e009f8c5c88b05896836330822afbda33af463c1f264f9adb9bbf5b094e61dfa qwen3.7-max/deepagents/results.jsonl
74
+ a5dcb2297ccc31a921e3f4eeb025ed81a764e16556a12068733cfe3bfe0bdbe3 qwen3.7-max/deepagents/summary.json
75
+ af5df4bd67b09d31325e71769cdedf0d2d8c4ed41ee7ce3a6eb10a2919737589 qwen3.7-max/diagnostics.json
76
+ 7b2d6aa947e21941e412e38f91aa01b75ff91b0668fea9071b5a9787f3a77347 qwen3.7-max/export-verification.json
77
+ 2600a5ed895e9cf217f96ba7cf2bf61762012886cb8f0918693888ab4bb084d3 qwen3.7-max/paired_comparison.csv
78
+ 0b4d4df9d186cb3ab5abd0d0a52bfcf89bad5efb4a135039277b282797d9580e qwen3.7-max/provenance.json
79
  06ed669caaba3ca5e252d12af3de32223019d50aefd9c46751ca3cb35110fce2 verify_kimi_gemini.py
80
+ 9c6d1d6b44998a459141969065e4e23fb0126b3556074ce35803f970ce8c0869 verify_qwen_deepseek.py
results/qwen3.7-max/baseline/README.md ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Qwen3.7-Max / baseline
2
+
3
+ 固定 1,000 题:成功 483、确定失败 517、任务 unknown 0、invalid evidence 0。已封存 1000 条;无待开始或在途记录。原生完整复现率 19.8%。
4
+
5
+ `results.jsonl` 与 `results.csv` 保存逐题状态、指标、用量和结果哈希;`summary.json` 分别保存本地 HF 四维和 local-paper-v1 五维;`config.json` 保存任务清单、协议、源代码版本和接续来源。上一级含配对表、来源与导出核验。
6
+
7
+ `resolved` 仅表示有效封存,`definite_outcome` 表示成功或确定失败。任务 unknown、无效证据与指标 null 分开保留,均不从 1,000 的分母移除。无效记录仅有 `unverified_result_sha256`,不可当作有效执行结果。后续阶段 pending 可能表示前序失败后没有执行到,不表示任务仍待派发。
8
+
9
+ 原始预测三元组不在源评分快照中,prediction 字段为空,并有明确导出状态;不能据此认为所有模型都没有生成预测。配对比较含历史协议修订,非严格等预算因果实验。本地评分尚未核实官方一致性。
10
+
11
+ 本次仅归档已接受的评分和状态,没有重新执行模型或程序。不包含 benchmark 输入、隐藏答案、原始提示词/API 轨迹、凭据或机器路径。
results/qwen3.7-max/baseline/config.json ADDED
@@ -0,0 +1,1199 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "qwen3.7-max",
4
+ "arm": "baseline",
5
+ "harness": "none",
6
+ "archived_at_utc": "2026-09-16T17:18:47.373930+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "engineering_pilot": false,
14
+ "expected_tasks": 1000,
15
+ "task_ids": [
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4,
20
+ 5,
21
+ 6,
22
+ 7,
23
+ 8,
24
+ 9,
25
+ 10,
26
+ 11,
27
+ 12,
28
+ 13,
29
+ 14,
30
+ 15,
31
+ 16,
32
+ 17,
33
+ 18,
34
+ 19,
35
+ 20,
36
+ 21,
37
+ 22,
38
+ 23,
39
+ 24,
40
+ 25,
41
+ 26,
42
+ 27,
43
+ 28,
44
+ 29,
45
+ 30,
46
+ 31,
47
+ 32,
48
+ 33,
49
+ 34,
50
+ 35,
51
+ 36,
52
+ 37,
53
+ 38,
54
+ 39,
55
+ 40,
56
+ 41,
57
+ 42,
58
+ 43,
59
+ 44,
60
+ 45,
61
+ 46,
62
+ 47,
63
+ 48,
64
+ 49,
65
+ 50,
66
+ 51,
67
+ 52,
68
+ 53,
69
+ 54,
70
+ 55,
71
+ 56,
72
+ 57,
73
+ 58,
74
+ 59,
75
+ 60,
76
+ 61,
77
+ 62,
78
+ 63,
79
+ 64,
80
+ 65,
81
+ 66,
82
+ 67,
83
+ 68,
84
+ 69,
85
+ 70,
86
+ 71,
87
+ 72,
88
+ 73,
89
+ 74,
90
+ 75,
91
+ 76,
92
+ 77,
93
+ 78,
94
+ 79,
95
+ 80,
96
+ 81,
97
+ 82,
98
+ 83,
99
+ 84,
100
+ 85,
101
+ 86,
102
+ 87,
103
+ 88,
104
+ 89,
105
+ 90,
106
+ 91,
107
+ 92,
108
+ 93,
109
+ 94,
110
+ 95,
111
+ 96,
112
+ 97,
113
+ 98,
114
+ 100,
115
+ 101,
116
+ 102,
117
+ 103,
118
+ 104,
119
+ 105,
120
+ 106,
121
+ 107,
122
+ 108,
123
+ 109,
124
+ 110,
125
+ 111,
126
+ 112,
127
+ 113,
128
+ 114,
129
+ 115,
130
+ 116,
131
+ 117,
132
+ 118,
133
+ 119,
134
+ 120,
135
+ 121,
136
+ 122,
137
+ 123,
138
+ 124,
139
+ 125,
140
+ 126,
141
+ 127,
142
+ 128,
143
+ 129,
144
+ 130,
145
+ 131,
146
+ 132,
147
+ 133,
148
+ 134,
149
+ 135,
150
+ 136,
151
+ 137,
152
+ 138,
153
+ 139,
154
+ 140,
155
+ 141,
156
+ 142,
157
+ 143,
158
+ 144,
159
+ 145,
160
+ 146,
161
+ 147,
162
+ 148,
163
+ 149,
164
+ 150,
165
+ 151,
166
+ 152,
167
+ 153,
168
+ 154,
169
+ 155,
170
+ 156,
171
+ 157,
172
+ 158,
173
+ 159,
174
+ 160,
175
+ 161,
176
+ 162,
177
+ 163,
178
+ 164,
179
+ 165,
180
+ 166,
181
+ 167,
182
+ 168,
183
+ 169,
184
+ 170,
185
+ 171,
186
+ 172,
187
+ 173,
188
+ 174,
189
+ 175,
190
+ 176,
191
+ 177,
192
+ 178,
193
+ 179,
194
+ 180,
195
+ 181,
196
+ 182,
197
+ 183,
198
+ 184,
199
+ 185,
200
+ 186,
201
+ 187,
202
+ 188,
203
+ 189,
204
+ 190,
205
+ 191,
206
+ 192,
207
+ 193,
208
+ 194,
209
+ 195,
210
+ 196,
211
+ 197,
212
+ 198,
213
+ 199,
214
+ 200,
215
+ 201,
216
+ 202,
217
+ 203,
218
+ 204,
219
+ 205,
220
+ 206,
221
+ 207,
222
+ 208,
223
+ 209,
224
+ 210,
225
+ 211,
226
+ 212,
227
+ 213,
228
+ 214,
229
+ 215,
230
+ 216,
231
+ 217,
232
+ 218,
233
+ 219,
234
+ 220,
235
+ 221,
236
+ 222,
237
+ 223,
238
+ 224,
239
+ 225,
240
+ 226,
241
+ 227,
242
+ 228,
243
+ 229,
244
+ 230,
245
+ 231,
246
+ 232,
247
+ 233,
248
+ 234,
249
+ 235,
250
+ 236,
251
+ 237,
252
+ 238,
253
+ 239,
254
+ 240,
255
+ 241,
256
+ 242,
257
+ 248,
258
+ 249,
259
+ 250,
260
+ 251,
261
+ 252,
262
+ 253,
263
+ 254,
264
+ 255,
265
+ 256,
266
+ 257,
267
+ 258,
268
+ 259,
269
+ 261,
270
+ 262,
271
+ 263,
272
+ 264,
273
+ 265,
274
+ 266,
275
+ 290,
276
+ 291,
277
+ 292,
278
+ 293,
279
+ 294,
280
+ 295,
281
+ 302,
282
+ 303,
283
+ 304,
284
+ 305,
285
+ 306,
286
+ 307,
287
+ 308,
288
+ 309,
289
+ 310,
290
+ 311,
291
+ 312,
292
+ 313,
293
+ 314,
294
+ 315,
295
+ 316,
296
+ 317,
297
+ 318,
298
+ 319,
299
+ 320,
300
+ 321,
301
+ 322,
302
+ 323,
303
+ 324,
304
+ 325,
305
+ 326,
306
+ 327,
307
+ 328,
308
+ 329,
309
+ 330,
310
+ 331,
311
+ 332,
312
+ 333,
313
+ 334,
314
+ 335,
315
+ 336,
316
+ 337,
317
+ 338,
318
+ 339,
319
+ 340,
320
+ 341,
321
+ 342,
322
+ 343,
323
+ 344,
324
+ 345,
325
+ 346,
326
+ 347,
327
+ 348,
328
+ 349,
329
+ 350,
330
+ 351,
331
+ 352,
332
+ 353,
333
+ 354,
334
+ 355,
335
+ 356,
336
+ 357,
337
+ 358,
338
+ 359,
339
+ 360,
340
+ 361,
341
+ 362,
342
+ 363,
343
+ 364,
344
+ 365,
345
+ 366,
346
+ 367,
347
+ 368,
348
+ 369,
349
+ 370,
350
+ 371,
351
+ 372,
352
+ 373,
353
+ 374,
354
+ 375,
355
+ 376,
356
+ 377,
357
+ 378,
358
+ 379,
359
+ 380,
360
+ 381,
361
+ 382,
362
+ 383,
363
+ 384,
364
+ 385,
365
+ 386,
366
+ 387,
367
+ 388,
368
+ 389,
369
+ 390,
370
+ 391,
371
+ 392,
372
+ 393,
373
+ 394,
374
+ 395,
375
+ 396,
376
+ 397,
377
+ 398,
378
+ 399,
379
+ 400,
380
+ 401,
381
+ 402,
382
+ 403,
383
+ 404,
384
+ 405,
385
+ 406,
386
+ 407,
387
+ 408,
388
+ 409,
389
+ 410,
390
+ 411,
391
+ 412,
392
+ 413,
393
+ 414,
394
+ 415,
395
+ 416,
396
+ 417,
397
+ 418,
398
+ 419,
399
+ 420,
400
+ 421,
401
+ 422,
402
+ 423,
403
+ 424,
404
+ 425,
405
+ 426,
406
+ 427,
407
+ 428,
408
+ 429,
409
+ 430,
410
+ 431,
411
+ 432,
412
+ 433,
413
+ 434,
414
+ 435,
415
+ 436,
416
+ 437,
417
+ 438,
418
+ 439,
419
+ 440,
420
+ 441,
421
+ 442,
422
+ 443,
423
+ 444,
424
+ 445,
425
+ 446,
426
+ 447,
427
+ 448,
428
+ 449,
429
+ 450,
430
+ 451,
431
+ 452,
432
+ 453,
433
+ 454,
434
+ 455,
435
+ 456,
436
+ 457,
437
+ 458,
438
+ 459,
439
+ 460,
440
+ 461,
441
+ 462,
442
+ 463,
443
+ 464,
444
+ 465,
445
+ 466,
446
+ 467,
447
+ 468,
448
+ 469,
449
+ 470,
450
+ 471,
451
+ 472,
452
+ 473,
453
+ 474,
454
+ 475,
455
+ 476,
456
+ 477,
457
+ 478,
458
+ 479,
459
+ 480,
460
+ 481,
461
+ 482,
462
+ 483,
463
+ 484,
464
+ 485,
465
+ 486,
466
+ 487,
467
+ 488,
468
+ 489,
469
+ 490,
470
+ 491,
471
+ 492,
472
+ 493,
473
+ 494,
474
+ 495,
475
+ 496,
476
+ 497,
477
+ 498,
478
+ 499,
479
+ 500,
480
+ 501,
481
+ 502,
482
+ 503,
483
+ 504,
484
+ 505,
485
+ 506,
486
+ 507,
487
+ 508,
488
+ 509,
489
+ 510,
490
+ 511,
491
+ 512,
492
+ 513,
493
+ 514,
494
+ 515,
495
+ 516,
496
+ 517,
497
+ 518,
498
+ 519,
499
+ 520,
500
+ 521,
501
+ 522,
502
+ 523,
503
+ 524,
504
+ 525,
505
+ 526,
506
+ 527,
507
+ 528,
508
+ 529,
509
+ 530,
510
+ 531,
511
+ 532,
512
+ 533,
513
+ 534,
514
+ 535,
515
+ 536,
516
+ 537,
517
+ 538,
518
+ 539,
519
+ 540,
520
+ 541,
521
+ 542,
522
+ 543,
523
+ 544,
524
+ 545,
525
+ 546,
526
+ 547,
527
+ 548,
528
+ 549,
529
+ 550,
530
+ 551,
531
+ 552,
532
+ 553,
533
+ 554,
534
+ 555,
535
+ 556,
536
+ 557,
537
+ 558,
538
+ 559,
539
+ 560,
540
+ 561,
541
+ 562,
542
+ 563,
543
+ 565,
544
+ 566,
545
+ 567,
546
+ 568,
547
+ 570,
548
+ 571,
549
+ 572,
550
+ 573,
551
+ 574,
552
+ 575,
553
+ 576,
554
+ 577,
555
+ 578,
556
+ 579,
557
+ 580,
558
+ 581,
559
+ 582,
560
+ 583,
561
+ 584,
562
+ 585,
563
+ 586,
564
+ 587,
565
+ 588,
566
+ 589,
567
+ 590,
568
+ 591,
569
+ 592,
570
+ 593,
571
+ 594,
572
+ 595,
573
+ 596,
574
+ 597,
575
+ 598,
576
+ 601,
577
+ 602,
578
+ 603,
579
+ 604,
580
+ 605,
581
+ 606,
582
+ 607,
583
+ 608,
584
+ 609,
585
+ 610,
586
+ 611,
587
+ 612,
588
+ 613,
589
+ 614,
590
+ 615,
591
+ 616,
592
+ 617,
593
+ 618,
594
+ 619,
595
+ 620,
596
+ 621,
597
+ 622,
598
+ 623,
599
+ 624,
600
+ 664,
601
+ 665,
602
+ 666,
603
+ 667,
604
+ 668,
605
+ 669,
606
+ 670,
607
+ 671,
608
+ 672,
609
+ 673,
610
+ 674,
611
+ 675,
612
+ 676,
613
+ 677,
614
+ 678,
615
+ 679,
616
+ 680,
617
+ 681,
618
+ 682,
619
+ 683,
620
+ 684,
621
+ 685,
622
+ 686,
623
+ 687,
624
+ 688,
625
+ 689,
626
+ 690,
627
+ 691,
628
+ 692,
629
+ 693,
630
+ 694,
631
+ 695,
632
+ 696,
633
+ 697,
634
+ 698,
635
+ 699,
636
+ 700,
637
+ 701,
638
+ 702,
639
+ 703,
640
+ 704,
641
+ 705,
642
+ 706,
643
+ 707,
644
+ 708,
645
+ 709,
646
+ 710,
647
+ 711,
648
+ 712,
649
+ 713,
650
+ 714,
651
+ 715,
652
+ 716,
653
+ 717,
654
+ 718,
655
+ 719,
656
+ 720,
657
+ 721,
658
+ 722,
659
+ 723,
660
+ 724,
661
+ 725,
662
+ 726,
663
+ 727,
664
+ 728,
665
+ 729,
666
+ 730,
667
+ 731,
668
+ 732,
669
+ 733,
670
+ 734,
671
+ 735,
672
+ 736,
673
+ 737,
674
+ 739,
675
+ 740,
676
+ 741,
677
+ 742,
678
+ 743,
679
+ 744,
680
+ 745,
681
+ 746,
682
+ 747,
683
+ 748,
684
+ 749,
685
+ 750,
686
+ 751,
687
+ 752,
688
+ 753,
689
+ 754,
690
+ 755,
691
+ 756,
692
+ 757,
693
+ 758,
694
+ 759,
695
+ 760,
696
+ 761,
697
+ 762,
698
+ 763,
699
+ 764,
700
+ 765,
701
+ 766,
702
+ 767,
703
+ 768,
704
+ 772,
705
+ 773,
706
+ 774,
707
+ 775,
708
+ 776,
709
+ 777,
710
+ 778,
711
+ 779,
712
+ 780,
713
+ 781,
714
+ 782,
715
+ 783,
716
+ 784,
717
+ 785,
718
+ 786,
719
+ 787,
720
+ 788,
721
+ 789,
722
+ 790,
723
+ 791,
724
+ 793,
725
+ 794,
726
+ 795,
727
+ 796,
728
+ 797,
729
+ 798,
730
+ 799,
731
+ 800,
732
+ 801,
733
+ 802,
734
+ 803,
735
+ 804,
736
+ 805,
737
+ 806,
738
+ 807,
739
+ 808,
740
+ 825,
741
+ 826,
742
+ 827,
743
+ 828,
744
+ 829,
745
+ 830,
746
+ 831,
747
+ 832,
748
+ 833,
749
+ 834,
750
+ 835,
751
+ 836,
752
+ 837,
753
+ 838,
754
+ 839,
755
+ 840,
756
+ 841,
757
+ 842,
758
+ 843,
759
+ 844,
760
+ 845,
761
+ 846,
762
+ 847,
763
+ 848,
764
+ 849,
765
+ 850,
766
+ 851,
767
+ 852,
768
+ 853,
769
+ 854,
770
+ 855,
771
+ 856,
772
+ 857,
773
+ 858,
774
+ 859,
775
+ 860,
776
+ 861,
777
+ 862,
778
+ 863,
779
+ 864,
780
+ 865,
781
+ 866,
782
+ 867,
783
+ 868,
784
+ 869,
785
+ 870,
786
+ 871,
787
+ 872,
788
+ 873,
789
+ 874,
790
+ 875,
791
+ 876,
792
+ 877,
793
+ 878,
794
+ 879,
795
+ 880,
796
+ 881,
797
+ 882,
798
+ 883,
799
+ 884,
800
+ 885,
801
+ 886,
802
+ 887,
803
+ 888,
804
+ 889,
805
+ 890,
806
+ 891,
807
+ 892,
808
+ 893,
809
+ 894,
810
+ 895,
811
+ 896,
812
+ 897,
813
+ 898,
814
+ 899,
815
+ 900,
816
+ 901,
817
+ 902,
818
+ 903,
819
+ 906,
820
+ 907,
821
+ 913,
822
+ 914,
823
+ 915,
824
+ 917,
825
+ 920,
826
+ 921,
827
+ 922,
828
+ 923,
829
+ 924,
830
+ 925,
831
+ 926,
832
+ 927,
833
+ 928,
834
+ 929,
835
+ 930,
836
+ 931,
837
+ 932,
838
+ 933,
839
+ 934,
840
+ 935,
841
+ 936,
842
+ 937,
843
+ 938,
844
+ 939,
845
+ 940,
846
+ 941,
847
+ 942,
848
+ 943,
849
+ 944,
850
+ 945,
851
+ 946,
852
+ 947,
853
+ 948,
854
+ 949,
855
+ 950,
856
+ 951,
857
+ 952,
858
+ 953,
859
+ 954,
860
+ 955,
861
+ 956,
862
+ 957,
863
+ 958,
864
+ 959,
865
+ 960,
866
+ 961,
867
+ 962,
868
+ 963,
869
+ 964,
870
+ 965,
871
+ 966,
872
+ 967,
873
+ 968,
874
+ 969,
875
+ 970,
876
+ 971,
877
+ 972,
878
+ 973,
879
+ 974,
880
+ 975,
881
+ 976,
882
+ 977,
883
+ 978,
884
+ 979,
885
+ 980,
886
+ 981,
887
+ 982,
888
+ 983,
889
+ 984,
890
+ 985,
891
+ 1001,
892
+ 1002,
893
+ 1003,
894
+ 1004,
895
+ 1005,
896
+ 1006,
897
+ 1007,
898
+ 1008,
899
+ 1009,
900
+ 1010,
901
+ 1011,
902
+ 1012,
903
+ 1013,
904
+ 1014,
905
+ 1015,
906
+ 1016,
907
+ 1017,
908
+ 1018,
909
+ 1019,
910
+ 1020,
911
+ 1021,
912
+ 1022,
913
+ 1023,
914
+ 1024,
915
+ 1025,
916
+ 1026,
917
+ 1027,
918
+ 1028,
919
+ 1029,
920
+ 1030,
921
+ 1031,
922
+ 1032,
923
+ 1033,
924
+ 1034,
925
+ 1035,
926
+ 1036,
927
+ 1037,
928
+ 1038,
929
+ 1039,
930
+ 1040,
931
+ 1041,
932
+ 1042,
933
+ 1043,
934
+ 1044,
935
+ 1045,
936
+ 1046,
937
+ 1047,
938
+ 1048,
939
+ 1049,
940
+ 1050,
941
+ 1051,
942
+ 1052,
943
+ 1053,
944
+ 1054,
945
+ 1055,
946
+ 1056,
947
+ 1057,
948
+ 1058,
949
+ 1059,
950
+ 1060,
951
+ 1061,
952
+ 1062,
953
+ 1063,
954
+ 1064,
955
+ 1065,
956
+ 1066,
957
+ 1067,
958
+ 1068,
959
+ 1069,
960
+ 1070,
961
+ 1071,
962
+ 1072,
963
+ 1073,
964
+ 1074,
965
+ 1075,
966
+ 1076,
967
+ 1077,
968
+ 1078,
969
+ 1079,
970
+ 1080,
971
+ 1081,
972
+ 1082,
973
+ 1083,
974
+ 1084,
975
+ 1085,
976
+ 1086,
977
+ 1087,
978
+ 1088,
979
+ 1089,
980
+ 1090,
981
+ 1091,
982
+ 1092,
983
+ 1093,
984
+ 1094,
985
+ 1095,
986
+ 1096,
987
+ 1097,
988
+ 1098,
989
+ 1099,
990
+ 1100,
991
+ 1101,
992
+ 1102,
993
+ 1103,
994
+ 1104,
995
+ 1105,
996
+ 1106,
997
+ 1107,
998
+ 1108,
999
+ 1109,
1000
+ 1110,
1001
+ 1111,
1002
+ 1112,
1003
+ 1113,
1004
+ 1114,
1005
+ 1115,
1006
+ 1116,
1007
+ 1117,
1008
+ 1118,
1009
+ 1119,
1010
+ 1120,
1011
+ 1121,
1012
+ 1122,
1013
+ 1123,
1014
+ 1124,
1015
+ 1125
1016
+ ]
1017
+ },
1018
+ "generation": {
1019
+ "model": "qwen3.7-max",
1020
+ "provider": "chat-completions",
1021
+ "reasoning_effort": null,
1022
+ "stream": true,
1023
+ "timeout": 1800
1024
+ },
1025
+ "effective_model_call_cap": 1,
1026
+ "recorded_protocols": [
1027
+ {
1028
+ "core_sha256": {
1029
+ "protocol.json": "7cc76ea33174d6afd9d4039e8bbe090935bfc8625ec4dfb648657e56fe721948",
1030
+ "inputs.manifest.json": "cd35a8f02e25511da81f4b1966b17c2309f5641cd3db381cc1b7d48eada193ac",
1031
+ "schedule.json": "1c61c21d79fa441ec4b6b46fa2e548e98bf9a2b8e92f1629c86b809a76bc1fd7",
1032
+ "study.json": "ad423d7cd139f779654a19d4e15c3483c8d2c8ca267275b5f1922ea6ada463ca",
1033
+ "environment.json": "04f141c80a13f3620f971c3217dccd28c2ee58f22a3d1fc5f7db2bfd1f4cc2ef"
1034
+ },
1035
+ "population_sha256": "50bac93cf8286c76821d0869fd9011f1ddb7aec0457422786f9149ada9894f58",
1036
+ "task_count": 993,
1037
+ "task_input_blobs_rehashed": false,
1038
+ "source_and_dependencies_match": true,
1039
+ "limits": {
1040
+ "output_tokens": 32768,
1041
+ "wall_seconds": 1800,
1042
+ "model_calls": 6,
1043
+ "tool_calls": 4,
1044
+ "tool_timeout": 90
1045
+ },
1046
+ "generation": {
1047
+ "model": "qwen3.7-max",
1048
+ "provider": "chat-completions",
1049
+ "reasoning_effort": null,
1050
+ "stream": true,
1051
+ "timeout": 1800
1052
+ },
1053
+ "dataset_provenance": {
1054
+ "repo_id": "CamoAiLab/InferenceNet",
1055
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1056
+ "task_directory": "Selected_1000",
1057
+ "task_list": "Selected_1000/1000_new.csv",
1058
+ "prior_results_imported": false,
1059
+ "engineering_pilot": false
1060
+ },
1061
+ "execution_identity": {
1062
+ "backend": "dsw-bwrap-v3",
1063
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1064
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1065
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1066
+ "engine_memory_bytes": 38654705664,
1067
+ "engine_cpus": 12,
1068
+ "cpu_slots": [
1069
+ [
1070
+ 4,
1071
+ 5
1072
+ ],
1073
+ [
1074
+ 6,
1075
+ 7
1076
+ ],
1077
+ [
1078
+ 8,
1079
+ 9
1080
+ ],
1081
+ [
1082
+ 10,
1083
+ 11
1084
+ ],
1085
+ [
1086
+ 12,
1087
+ 13
1088
+ ],
1089
+ [
1090
+ 14,
1091
+ 15
1092
+ ]
1093
+ ],
1094
+ "resources": {
1095
+ "cpus": 2.0,
1096
+ "memory": "6g",
1097
+ "tmpfs": "256m"
1098
+ },
1099
+ "memory_enforcement": "RLIMIT_AS",
1100
+ "single_payload_process": true,
1101
+ "network": "none",
1102
+ "uid": 10001,
1103
+ "capacity_profile": "independent-six-slots-20260916",
1104
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1105
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1106
+ },
1107
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1108
+ "scorer_sha256": {
1109
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1110
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1111
+ },
1112
+ "origin": "child"
1113
+ },
1114
+ {
1115
+ "core_sha256": {
1116
+ "protocol.json": "82df8a37d69e6d815b440a40fb4ebcaa04597fa52bdc54f6ea0112d2096565d9",
1117
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1118
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1119
+ "study.json": "ab523f28178c00ecec133dd6773da6f2cbbe8d26630cf38671f7d8745406919c",
1120
+ "environment.json": "f0e702a8a418d683763e0a0d35acd4377d246e006d1c7533c9bf6cac38fd70bf"
1121
+ },
1122
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1123
+ "task_count": 1000,
1124
+ "task_input_blobs_rehashed": false,
1125
+ "source_and_dependencies_match": true,
1126
+ "limits": {
1127
+ "output_tokens": 32768,
1128
+ "wall_seconds": 1800,
1129
+ "model_calls": 6,
1130
+ "tool_calls": 4,
1131
+ "tool_timeout": 90
1132
+ },
1133
+ "generation": {
1134
+ "model": "qwen3.7-max",
1135
+ "provider": "chat-completions",
1136
+ "reasoning_effort": null,
1137
+ "stream": true,
1138
+ "timeout": 1800
1139
+ },
1140
+ "dataset_provenance": {
1141
+ "repo_id": "CamoAiLab/InferenceNet",
1142
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1143
+ "task_directory": "Selected_1000",
1144
+ "task_list": "Selected_1000/1000_new.csv",
1145
+ "prior_results_imported": false,
1146
+ "engineering_pilot": false
1147
+ },
1148
+ "execution_identity": {
1149
+ "backend": "dsw-bwrap-v3",
1150
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1151
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1152
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1153
+ "engine_memory_bytes": 12884901888,
1154
+ "engine_cpus": 4,
1155
+ "cpu_slots": [
1156
+ [
1157
+ 16,
1158
+ 17
1159
+ ],
1160
+ [
1161
+ 18,
1162
+ 19
1163
+ ]
1164
+ ],
1165
+ "resources": {
1166
+ "memory": "6g",
1167
+ "cpus": 2.0,
1168
+ "tmpfs": "256m"
1169
+ },
1170
+ "memory_enforcement": "RLIMIT_AS",
1171
+ "single_payload_process": true,
1172
+ "network": "none",
1173
+ "uid": 10001
1174
+ },
1175
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1176
+ "scorer_sha256": {
1177
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1178
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1179
+ },
1180
+ "origin": "parent"
1181
+ }
1182
+ ],
1183
+ "selected_origin_counts": {
1184
+ "child": 993,
1185
+ "parent": 7
1186
+ },
1187
+ "source_commits": [
1188
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1189
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1190
+ ],
1191
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1192
+ "official_parity_verified": false,
1193
+ "notes": [
1194
+ "Historical disjoint parent/child continuation; original source and protocol hashes retained.",
1195
+ "Different baseline/harness call budgets; not an equal-cost causal comparison.",
1196
+ "Unknown and invalid evidence remain in the denominator; no replacement or rescoring.",
1197
+ "Invalid evidence is not sealed; its unverified hash is distinct from a verified result hash."
1198
+ ]
1199
+ }
results/qwen3.7-max/baseline/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/qwen3.7-max/baseline/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/qwen3.7-max/baseline/summary.json ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
5
+ "model": "qwen3.7-max",
6
+ "arm": "baseline",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": true,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 486,
37
+ "denominator": 1000,
38
+ "rate": 0.486,
39
+ "score": 48.6,
40
+ "unknown_count": 0,
41
+ "failure_count": 514,
42
+ "assessable": 1000,
43
+ "coverage_percent": 100.0
44
+ },
45
+ "partial_replication": {
46
+ "count": 315,
47
+ "denominator": 1000,
48
+ "rate": 0.315,
49
+ "score": 31.5,
50
+ "unknown_count": 518,
51
+ "failure_count": 167,
52
+ "assessable": 482,
53
+ "coverage_percent": 48.2
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 454,
57
+ "denominator": 1000,
58
+ "rate": 0.454,
59
+ "score": 45.4,
60
+ "unknown_count": 518,
61
+ "failure_count": 28,
62
+ "assessable": 482,
63
+ "coverage_percent": 48.2
64
+ },
65
+ "significance_level": {
66
+ "count": 396,
67
+ "denominator": 1000,
68
+ "rate": 0.396,
69
+ "score": 39.6,
70
+ "unknown_count": 518,
71
+ "failure_count": 86,
72
+ "assessable": 482,
73
+ "coverage_percent": 48.2
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "qwen3.7-max baseline",
78
+ "Compilation Success": 48.6,
79
+ "Partial Replication": 31.5,
80
+ "Correct Coefficient Direction": 45.4,
81
+ "Significant Level Correctness": 39.6
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 1000,
94
+ "task_unknown_count": 0,
95
+ "invalid_evidence_count": 0,
96
+ "valid_prediction_count": 483,
97
+ "task_counts": {
98
+ "succeeded": 483,
99
+ "failed": 517,
100
+ "unknown": 0,
101
+ "unsealed": 0,
102
+ "not_started": 0,
103
+ "invalid_evidence": 0,
104
+ "completed": 1000,
105
+ "planned": 1000
106
+ }
107
+ },
108
+ "legacy_local_paper": {
109
+ "profile": "local-paper-v1",
110
+ "definitions": {
111
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
112
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
113
+ "coefficient_only": "coefficient relative error <= .05",
114
+ "direction": "matching strictly positive or strictly negative coefficient signs",
115
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
116
+ },
117
+ "metrics": {
118
+ "perfect": {
119
+ "count": 198,
120
+ "denominator": 1000,
121
+ "rate": 0.198,
122
+ "score": 19.8,
123
+ "unknown_count": 519,
124
+ "failure_count": 283,
125
+ "assessable": 481,
126
+ "coverage_percent": 48.1
127
+ },
128
+ "partial": {
129
+ "count": 238,
130
+ "denominator": 1000,
131
+ "rate": 0.238,
132
+ "score": 23.8,
133
+ "unknown_count": 519,
134
+ "failure_count": 243,
135
+ "assessable": 481,
136
+ "coverage_percent": 48.1
137
+ },
138
+ "coefficient_only": {
139
+ "count": 315,
140
+ "denominator": 1000,
141
+ "rate": 0.315,
142
+ "score": 31.5,
143
+ "unknown_count": 519,
144
+ "failure_count": 166,
145
+ "assessable": 481,
146
+ "coverage_percent": 48.1
147
+ },
148
+ "direction": {
149
+ "count": 453,
150
+ "denominator": 1000,
151
+ "rate": 0.453,
152
+ "score": 45.3,
153
+ "unknown_count": 519,
154
+ "failure_count": 28,
155
+ "assessable": 481,
156
+ "coverage_percent": 48.1
157
+ },
158
+ "significance": {
159
+ "count": 380,
160
+ "denominator": 1000,
161
+ "rate": 0.38,
162
+ "score": 38.0,
163
+ "unknown_count": 519,
164
+ "failure_count": 101,
165
+ "assessable": 481,
166
+ "coverage_percent": 48.1
167
+ }
168
+ }
169
+ },
170
+ "verification_scope": "Offline equality to accepted flags, hashes and states; raw predictions unavailable; not a fresh native audit or official scorer parity check.",
171
+ "usage": {
172
+ "sealed_model_calls": 1000,
173
+ "unknown_usage_calls": 0,
174
+ "unsealed_episodes_excluded": 0,
175
+ "invalid_evidence_episodes_excluded": 0,
176
+ "known_token_subtotal": {
177
+ "input_tokens": 448041,
178
+ "output_tokens": 6100871,
179
+ "total_tokens": 6548912
180
+ },
181
+ "unknown_by_field": {
182
+ "input_tokens": 0,
183
+ "output_tokens": 0,
184
+ "total_tokens": 0
185
+ },
186
+ "total_tokens_complete": true,
187
+ "usd": null,
188
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
189
+ "summed_episode_seconds": 123626.58500000021
190
+ }
191
+ }
results/qwen3.7-max/deepagents/README.md ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Qwen3.7-Max / deepagents
2
+
3
+ 固定 1,000 题:成功 519、确定失败 474、任务 unknown 6、invalid evidence 1。已封存 999 条;无待开始或在途记录。原生完整复现率 24.8%。
4
+
5
+ `results.jsonl` 与 `results.csv` 保存逐题状态、指标、用量和结果哈希;`summary.json` 分别保存本地 HF 四维和 local-paper-v1 五维;`config.json` 保存任务清单、协议、源代码版本和接续来源。上一级含配对表、来源与导出核验。
6
+
7
+ `resolved` 仅表示有效封存,`definite_outcome` 表示成功或确定失败。任务 unknown、无效证据与指标 null 分开保留,均不从 1,000 的分母移除。无效记录仅有 `unverified_result_sha256`,不可当作有效执行结果。后续阶段 pending 可能表示前序失败后没有执行到,不表示任务仍待派发。
8
+
9
+ 原始预测三元组不在源评分快照中,prediction 字段为空,并有明确导出状态;不能据此认为所有模型都没有生成预测。配对比较含历史协议修订,非严格等预算因果实验。本地评分尚未核实官方一致性。
10
+
11
+ 本次仅归档已接受的评分和状态,没有重新执行模型或程序。不包含 benchmark 输入、隐藏答案、原始提示词/API 轨迹、凭据或机器路径。
results/qwen3.7-max/deepagents/config.json ADDED
@@ -0,0 +1,1199 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "qwen3.7-max",
4
+ "arm": "deepagents",
5
+ "harness": "DeepAgents 0.7.13",
6
+ "archived_at_utc": "2026-09-16T17:18:47.373930+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "engineering_pilot": false,
14
+ "expected_tasks": 1000,
15
+ "task_ids": [
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4,
20
+ 5,
21
+ 6,
22
+ 7,
23
+ 8,
24
+ 9,
25
+ 10,
26
+ 11,
27
+ 12,
28
+ 13,
29
+ 14,
30
+ 15,
31
+ 16,
32
+ 17,
33
+ 18,
34
+ 19,
35
+ 20,
36
+ 21,
37
+ 22,
38
+ 23,
39
+ 24,
40
+ 25,
41
+ 26,
42
+ 27,
43
+ 28,
44
+ 29,
45
+ 30,
46
+ 31,
47
+ 32,
48
+ 33,
49
+ 34,
50
+ 35,
51
+ 36,
52
+ 37,
53
+ 38,
54
+ 39,
55
+ 40,
56
+ 41,
57
+ 42,
58
+ 43,
59
+ 44,
60
+ 45,
61
+ 46,
62
+ 47,
63
+ 48,
64
+ 49,
65
+ 50,
66
+ 51,
67
+ 52,
68
+ 53,
69
+ 54,
70
+ 55,
71
+ 56,
72
+ 57,
73
+ 58,
74
+ 59,
75
+ 60,
76
+ 61,
77
+ 62,
78
+ 63,
79
+ 64,
80
+ 65,
81
+ 66,
82
+ 67,
83
+ 68,
84
+ 69,
85
+ 70,
86
+ 71,
87
+ 72,
88
+ 73,
89
+ 74,
90
+ 75,
91
+ 76,
92
+ 77,
93
+ 78,
94
+ 79,
95
+ 80,
96
+ 81,
97
+ 82,
98
+ 83,
99
+ 84,
100
+ 85,
101
+ 86,
102
+ 87,
103
+ 88,
104
+ 89,
105
+ 90,
106
+ 91,
107
+ 92,
108
+ 93,
109
+ 94,
110
+ 95,
111
+ 96,
112
+ 97,
113
+ 98,
114
+ 100,
115
+ 101,
116
+ 102,
117
+ 103,
118
+ 104,
119
+ 105,
120
+ 106,
121
+ 107,
122
+ 108,
123
+ 109,
124
+ 110,
125
+ 111,
126
+ 112,
127
+ 113,
128
+ 114,
129
+ 115,
130
+ 116,
131
+ 117,
132
+ 118,
133
+ 119,
134
+ 120,
135
+ 121,
136
+ 122,
137
+ 123,
138
+ 124,
139
+ 125,
140
+ 126,
141
+ 127,
142
+ 128,
143
+ 129,
144
+ 130,
145
+ 131,
146
+ 132,
147
+ 133,
148
+ 134,
149
+ 135,
150
+ 136,
151
+ 137,
152
+ 138,
153
+ 139,
154
+ 140,
155
+ 141,
156
+ 142,
157
+ 143,
158
+ 144,
159
+ 145,
160
+ 146,
161
+ 147,
162
+ 148,
163
+ 149,
164
+ 150,
165
+ 151,
166
+ 152,
167
+ 153,
168
+ 154,
169
+ 155,
170
+ 156,
171
+ 157,
172
+ 158,
173
+ 159,
174
+ 160,
175
+ 161,
176
+ 162,
177
+ 163,
178
+ 164,
179
+ 165,
180
+ 166,
181
+ 167,
182
+ 168,
183
+ 169,
184
+ 170,
185
+ 171,
186
+ 172,
187
+ 173,
188
+ 174,
189
+ 175,
190
+ 176,
191
+ 177,
192
+ 178,
193
+ 179,
194
+ 180,
195
+ 181,
196
+ 182,
197
+ 183,
198
+ 184,
199
+ 185,
200
+ 186,
201
+ 187,
202
+ 188,
203
+ 189,
204
+ 190,
205
+ 191,
206
+ 192,
207
+ 193,
208
+ 194,
209
+ 195,
210
+ 196,
211
+ 197,
212
+ 198,
213
+ 199,
214
+ 200,
215
+ 201,
216
+ 202,
217
+ 203,
218
+ 204,
219
+ 205,
220
+ 206,
221
+ 207,
222
+ 208,
223
+ 209,
224
+ 210,
225
+ 211,
226
+ 212,
227
+ 213,
228
+ 214,
229
+ 215,
230
+ 216,
231
+ 217,
232
+ 218,
233
+ 219,
234
+ 220,
235
+ 221,
236
+ 222,
237
+ 223,
238
+ 224,
239
+ 225,
240
+ 226,
241
+ 227,
242
+ 228,
243
+ 229,
244
+ 230,
245
+ 231,
246
+ 232,
247
+ 233,
248
+ 234,
249
+ 235,
250
+ 236,
251
+ 237,
252
+ 238,
253
+ 239,
254
+ 240,
255
+ 241,
256
+ 242,
257
+ 248,
258
+ 249,
259
+ 250,
260
+ 251,
261
+ 252,
262
+ 253,
263
+ 254,
264
+ 255,
265
+ 256,
266
+ 257,
267
+ 258,
268
+ 259,
269
+ 261,
270
+ 262,
271
+ 263,
272
+ 264,
273
+ 265,
274
+ 266,
275
+ 290,
276
+ 291,
277
+ 292,
278
+ 293,
279
+ 294,
280
+ 295,
281
+ 302,
282
+ 303,
283
+ 304,
284
+ 305,
285
+ 306,
286
+ 307,
287
+ 308,
288
+ 309,
289
+ 310,
290
+ 311,
291
+ 312,
292
+ 313,
293
+ 314,
294
+ 315,
295
+ 316,
296
+ 317,
297
+ 318,
298
+ 319,
299
+ 320,
300
+ 321,
301
+ 322,
302
+ 323,
303
+ 324,
304
+ 325,
305
+ 326,
306
+ 327,
307
+ 328,
308
+ 329,
309
+ 330,
310
+ 331,
311
+ 332,
312
+ 333,
313
+ 334,
314
+ 335,
315
+ 336,
316
+ 337,
317
+ 338,
318
+ 339,
319
+ 340,
320
+ 341,
321
+ 342,
322
+ 343,
323
+ 344,
324
+ 345,
325
+ 346,
326
+ 347,
327
+ 348,
328
+ 349,
329
+ 350,
330
+ 351,
331
+ 352,
332
+ 353,
333
+ 354,
334
+ 355,
335
+ 356,
336
+ 357,
337
+ 358,
338
+ 359,
339
+ 360,
340
+ 361,
341
+ 362,
342
+ 363,
343
+ 364,
344
+ 365,
345
+ 366,
346
+ 367,
347
+ 368,
348
+ 369,
349
+ 370,
350
+ 371,
351
+ 372,
352
+ 373,
353
+ 374,
354
+ 375,
355
+ 376,
356
+ 377,
357
+ 378,
358
+ 379,
359
+ 380,
360
+ 381,
361
+ 382,
362
+ 383,
363
+ 384,
364
+ 385,
365
+ 386,
366
+ 387,
367
+ 388,
368
+ 389,
369
+ 390,
370
+ 391,
371
+ 392,
372
+ 393,
373
+ 394,
374
+ 395,
375
+ 396,
376
+ 397,
377
+ 398,
378
+ 399,
379
+ 400,
380
+ 401,
381
+ 402,
382
+ 403,
383
+ 404,
384
+ 405,
385
+ 406,
386
+ 407,
387
+ 408,
388
+ 409,
389
+ 410,
390
+ 411,
391
+ 412,
392
+ 413,
393
+ 414,
394
+ 415,
395
+ 416,
396
+ 417,
397
+ 418,
398
+ 419,
399
+ 420,
400
+ 421,
401
+ 422,
402
+ 423,
403
+ 424,
404
+ 425,
405
+ 426,
406
+ 427,
407
+ 428,
408
+ 429,
409
+ 430,
410
+ 431,
411
+ 432,
412
+ 433,
413
+ 434,
414
+ 435,
415
+ 436,
416
+ 437,
417
+ 438,
418
+ 439,
419
+ 440,
420
+ 441,
421
+ 442,
422
+ 443,
423
+ 444,
424
+ 445,
425
+ 446,
426
+ 447,
427
+ 448,
428
+ 449,
429
+ 450,
430
+ 451,
431
+ 452,
432
+ 453,
433
+ 454,
434
+ 455,
435
+ 456,
436
+ 457,
437
+ 458,
438
+ 459,
439
+ 460,
440
+ 461,
441
+ 462,
442
+ 463,
443
+ 464,
444
+ 465,
445
+ 466,
446
+ 467,
447
+ 468,
448
+ 469,
449
+ 470,
450
+ 471,
451
+ 472,
452
+ 473,
453
+ 474,
454
+ 475,
455
+ 476,
456
+ 477,
457
+ 478,
458
+ 479,
459
+ 480,
460
+ 481,
461
+ 482,
462
+ 483,
463
+ 484,
464
+ 485,
465
+ 486,
466
+ 487,
467
+ 488,
468
+ 489,
469
+ 490,
470
+ 491,
471
+ 492,
472
+ 493,
473
+ 494,
474
+ 495,
475
+ 496,
476
+ 497,
477
+ 498,
478
+ 499,
479
+ 500,
480
+ 501,
481
+ 502,
482
+ 503,
483
+ 504,
484
+ 505,
485
+ 506,
486
+ 507,
487
+ 508,
488
+ 509,
489
+ 510,
490
+ 511,
491
+ 512,
492
+ 513,
493
+ 514,
494
+ 515,
495
+ 516,
496
+ 517,
497
+ 518,
498
+ 519,
499
+ 520,
500
+ 521,
501
+ 522,
502
+ 523,
503
+ 524,
504
+ 525,
505
+ 526,
506
+ 527,
507
+ 528,
508
+ 529,
509
+ 530,
510
+ 531,
511
+ 532,
512
+ 533,
513
+ 534,
514
+ 535,
515
+ 536,
516
+ 537,
517
+ 538,
518
+ 539,
519
+ 540,
520
+ 541,
521
+ 542,
522
+ 543,
523
+ 544,
524
+ 545,
525
+ 546,
526
+ 547,
527
+ 548,
528
+ 549,
529
+ 550,
530
+ 551,
531
+ 552,
532
+ 553,
533
+ 554,
534
+ 555,
535
+ 556,
536
+ 557,
537
+ 558,
538
+ 559,
539
+ 560,
540
+ 561,
541
+ 562,
542
+ 563,
543
+ 565,
544
+ 566,
545
+ 567,
546
+ 568,
547
+ 570,
548
+ 571,
549
+ 572,
550
+ 573,
551
+ 574,
552
+ 575,
553
+ 576,
554
+ 577,
555
+ 578,
556
+ 579,
557
+ 580,
558
+ 581,
559
+ 582,
560
+ 583,
561
+ 584,
562
+ 585,
563
+ 586,
564
+ 587,
565
+ 588,
566
+ 589,
567
+ 590,
568
+ 591,
569
+ 592,
570
+ 593,
571
+ 594,
572
+ 595,
573
+ 596,
574
+ 597,
575
+ 598,
576
+ 601,
577
+ 602,
578
+ 603,
579
+ 604,
580
+ 605,
581
+ 606,
582
+ 607,
583
+ 608,
584
+ 609,
585
+ 610,
586
+ 611,
587
+ 612,
588
+ 613,
589
+ 614,
590
+ 615,
591
+ 616,
592
+ 617,
593
+ 618,
594
+ 619,
595
+ 620,
596
+ 621,
597
+ 622,
598
+ 623,
599
+ 624,
600
+ 664,
601
+ 665,
602
+ 666,
603
+ 667,
604
+ 668,
605
+ 669,
606
+ 670,
607
+ 671,
608
+ 672,
609
+ 673,
610
+ 674,
611
+ 675,
612
+ 676,
613
+ 677,
614
+ 678,
615
+ 679,
616
+ 680,
617
+ 681,
618
+ 682,
619
+ 683,
620
+ 684,
621
+ 685,
622
+ 686,
623
+ 687,
624
+ 688,
625
+ 689,
626
+ 690,
627
+ 691,
628
+ 692,
629
+ 693,
630
+ 694,
631
+ 695,
632
+ 696,
633
+ 697,
634
+ 698,
635
+ 699,
636
+ 700,
637
+ 701,
638
+ 702,
639
+ 703,
640
+ 704,
641
+ 705,
642
+ 706,
643
+ 707,
644
+ 708,
645
+ 709,
646
+ 710,
647
+ 711,
648
+ 712,
649
+ 713,
650
+ 714,
651
+ 715,
652
+ 716,
653
+ 717,
654
+ 718,
655
+ 719,
656
+ 720,
657
+ 721,
658
+ 722,
659
+ 723,
660
+ 724,
661
+ 725,
662
+ 726,
663
+ 727,
664
+ 728,
665
+ 729,
666
+ 730,
667
+ 731,
668
+ 732,
669
+ 733,
670
+ 734,
671
+ 735,
672
+ 736,
673
+ 737,
674
+ 739,
675
+ 740,
676
+ 741,
677
+ 742,
678
+ 743,
679
+ 744,
680
+ 745,
681
+ 746,
682
+ 747,
683
+ 748,
684
+ 749,
685
+ 750,
686
+ 751,
687
+ 752,
688
+ 753,
689
+ 754,
690
+ 755,
691
+ 756,
692
+ 757,
693
+ 758,
694
+ 759,
695
+ 760,
696
+ 761,
697
+ 762,
698
+ 763,
699
+ 764,
700
+ 765,
701
+ 766,
702
+ 767,
703
+ 768,
704
+ 772,
705
+ 773,
706
+ 774,
707
+ 775,
708
+ 776,
709
+ 777,
710
+ 778,
711
+ 779,
712
+ 780,
713
+ 781,
714
+ 782,
715
+ 783,
716
+ 784,
717
+ 785,
718
+ 786,
719
+ 787,
720
+ 788,
721
+ 789,
722
+ 790,
723
+ 791,
724
+ 793,
725
+ 794,
726
+ 795,
727
+ 796,
728
+ 797,
729
+ 798,
730
+ 799,
731
+ 800,
732
+ 801,
733
+ 802,
734
+ 803,
735
+ 804,
736
+ 805,
737
+ 806,
738
+ 807,
739
+ 808,
740
+ 825,
741
+ 826,
742
+ 827,
743
+ 828,
744
+ 829,
745
+ 830,
746
+ 831,
747
+ 832,
748
+ 833,
749
+ 834,
750
+ 835,
751
+ 836,
752
+ 837,
753
+ 838,
754
+ 839,
755
+ 840,
756
+ 841,
757
+ 842,
758
+ 843,
759
+ 844,
760
+ 845,
761
+ 846,
762
+ 847,
763
+ 848,
764
+ 849,
765
+ 850,
766
+ 851,
767
+ 852,
768
+ 853,
769
+ 854,
770
+ 855,
771
+ 856,
772
+ 857,
773
+ 858,
774
+ 859,
775
+ 860,
776
+ 861,
777
+ 862,
778
+ 863,
779
+ 864,
780
+ 865,
781
+ 866,
782
+ 867,
783
+ 868,
784
+ 869,
785
+ 870,
786
+ 871,
787
+ 872,
788
+ 873,
789
+ 874,
790
+ 875,
791
+ 876,
792
+ 877,
793
+ 878,
794
+ 879,
795
+ 880,
796
+ 881,
797
+ 882,
798
+ 883,
799
+ 884,
800
+ 885,
801
+ 886,
802
+ 887,
803
+ 888,
804
+ 889,
805
+ 890,
806
+ 891,
807
+ 892,
808
+ 893,
809
+ 894,
810
+ 895,
811
+ 896,
812
+ 897,
813
+ 898,
814
+ 899,
815
+ 900,
816
+ 901,
817
+ 902,
818
+ 903,
819
+ 906,
820
+ 907,
821
+ 913,
822
+ 914,
823
+ 915,
824
+ 917,
825
+ 920,
826
+ 921,
827
+ 922,
828
+ 923,
829
+ 924,
830
+ 925,
831
+ 926,
832
+ 927,
833
+ 928,
834
+ 929,
835
+ 930,
836
+ 931,
837
+ 932,
838
+ 933,
839
+ 934,
840
+ 935,
841
+ 936,
842
+ 937,
843
+ 938,
844
+ 939,
845
+ 940,
846
+ 941,
847
+ 942,
848
+ 943,
849
+ 944,
850
+ 945,
851
+ 946,
852
+ 947,
853
+ 948,
854
+ 949,
855
+ 950,
856
+ 951,
857
+ 952,
858
+ 953,
859
+ 954,
860
+ 955,
861
+ 956,
862
+ 957,
863
+ 958,
864
+ 959,
865
+ 960,
866
+ 961,
867
+ 962,
868
+ 963,
869
+ 964,
870
+ 965,
871
+ 966,
872
+ 967,
873
+ 968,
874
+ 969,
875
+ 970,
876
+ 971,
877
+ 972,
878
+ 973,
879
+ 974,
880
+ 975,
881
+ 976,
882
+ 977,
883
+ 978,
884
+ 979,
885
+ 980,
886
+ 981,
887
+ 982,
888
+ 983,
889
+ 984,
890
+ 985,
891
+ 1001,
892
+ 1002,
893
+ 1003,
894
+ 1004,
895
+ 1005,
896
+ 1006,
897
+ 1007,
898
+ 1008,
899
+ 1009,
900
+ 1010,
901
+ 1011,
902
+ 1012,
903
+ 1013,
904
+ 1014,
905
+ 1015,
906
+ 1016,
907
+ 1017,
908
+ 1018,
909
+ 1019,
910
+ 1020,
911
+ 1021,
912
+ 1022,
913
+ 1023,
914
+ 1024,
915
+ 1025,
916
+ 1026,
917
+ 1027,
918
+ 1028,
919
+ 1029,
920
+ 1030,
921
+ 1031,
922
+ 1032,
923
+ 1033,
924
+ 1034,
925
+ 1035,
926
+ 1036,
927
+ 1037,
928
+ 1038,
929
+ 1039,
930
+ 1040,
931
+ 1041,
932
+ 1042,
933
+ 1043,
934
+ 1044,
935
+ 1045,
936
+ 1046,
937
+ 1047,
938
+ 1048,
939
+ 1049,
940
+ 1050,
941
+ 1051,
942
+ 1052,
943
+ 1053,
944
+ 1054,
945
+ 1055,
946
+ 1056,
947
+ 1057,
948
+ 1058,
949
+ 1059,
950
+ 1060,
951
+ 1061,
952
+ 1062,
953
+ 1063,
954
+ 1064,
955
+ 1065,
956
+ 1066,
957
+ 1067,
958
+ 1068,
959
+ 1069,
960
+ 1070,
961
+ 1071,
962
+ 1072,
963
+ 1073,
964
+ 1074,
965
+ 1075,
966
+ 1076,
967
+ 1077,
968
+ 1078,
969
+ 1079,
970
+ 1080,
971
+ 1081,
972
+ 1082,
973
+ 1083,
974
+ 1084,
975
+ 1085,
976
+ 1086,
977
+ 1087,
978
+ 1088,
979
+ 1089,
980
+ 1090,
981
+ 1091,
982
+ 1092,
983
+ 1093,
984
+ 1094,
985
+ 1095,
986
+ 1096,
987
+ 1097,
988
+ 1098,
989
+ 1099,
990
+ 1100,
991
+ 1101,
992
+ 1102,
993
+ 1103,
994
+ 1104,
995
+ 1105,
996
+ 1106,
997
+ 1107,
998
+ 1108,
999
+ 1109,
1000
+ 1110,
1001
+ 1111,
1002
+ 1112,
1003
+ 1113,
1004
+ 1114,
1005
+ 1115,
1006
+ 1116,
1007
+ 1117,
1008
+ 1118,
1009
+ 1119,
1010
+ 1120,
1011
+ 1121,
1012
+ 1122,
1013
+ 1123,
1014
+ 1124,
1015
+ 1125
1016
+ ]
1017
+ },
1018
+ "generation": {
1019
+ "model": "qwen3.7-max",
1020
+ "provider": "chat-completions",
1021
+ "reasoning_effort": null,
1022
+ "stream": true,
1023
+ "timeout": 1800
1024
+ },
1025
+ "effective_model_call_cap": 6,
1026
+ "recorded_protocols": [
1027
+ {
1028
+ "core_sha256": {
1029
+ "protocol.json": "7cc76ea33174d6afd9d4039e8bbe090935bfc8625ec4dfb648657e56fe721948",
1030
+ "inputs.manifest.json": "cd35a8f02e25511da81f4b1966b17c2309f5641cd3db381cc1b7d48eada193ac",
1031
+ "schedule.json": "1c61c21d79fa441ec4b6b46fa2e548e98bf9a2b8e92f1629c86b809a76bc1fd7",
1032
+ "study.json": "ad423d7cd139f779654a19d4e15c3483c8d2c8ca267275b5f1922ea6ada463ca",
1033
+ "environment.json": "04f141c80a13f3620f971c3217dccd28c2ee58f22a3d1fc5f7db2bfd1f4cc2ef"
1034
+ },
1035
+ "population_sha256": "50bac93cf8286c76821d0869fd9011f1ddb7aec0457422786f9149ada9894f58",
1036
+ "task_count": 993,
1037
+ "task_input_blobs_rehashed": false,
1038
+ "source_and_dependencies_match": true,
1039
+ "limits": {
1040
+ "output_tokens": 32768,
1041
+ "wall_seconds": 1800,
1042
+ "model_calls": 6,
1043
+ "tool_calls": 4,
1044
+ "tool_timeout": 90
1045
+ },
1046
+ "generation": {
1047
+ "model": "qwen3.7-max",
1048
+ "provider": "chat-completions",
1049
+ "reasoning_effort": null,
1050
+ "stream": true,
1051
+ "timeout": 1800
1052
+ },
1053
+ "dataset_provenance": {
1054
+ "repo_id": "CamoAiLab/InferenceNet",
1055
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1056
+ "task_directory": "Selected_1000",
1057
+ "task_list": "Selected_1000/1000_new.csv",
1058
+ "prior_results_imported": false,
1059
+ "engineering_pilot": false
1060
+ },
1061
+ "execution_identity": {
1062
+ "backend": "dsw-bwrap-v3",
1063
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1064
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1065
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1066
+ "engine_memory_bytes": 38654705664,
1067
+ "engine_cpus": 12,
1068
+ "cpu_slots": [
1069
+ [
1070
+ 4,
1071
+ 5
1072
+ ],
1073
+ [
1074
+ 6,
1075
+ 7
1076
+ ],
1077
+ [
1078
+ 8,
1079
+ 9
1080
+ ],
1081
+ [
1082
+ 10,
1083
+ 11
1084
+ ],
1085
+ [
1086
+ 12,
1087
+ 13
1088
+ ],
1089
+ [
1090
+ 14,
1091
+ 15
1092
+ ]
1093
+ ],
1094
+ "resources": {
1095
+ "cpus": 2.0,
1096
+ "memory": "6g",
1097
+ "tmpfs": "256m"
1098
+ },
1099
+ "memory_enforcement": "RLIMIT_AS",
1100
+ "single_payload_process": true,
1101
+ "network": "none",
1102
+ "uid": 10001,
1103
+ "capacity_profile": "independent-six-slots-20260916",
1104
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1105
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1106
+ },
1107
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1108
+ "scorer_sha256": {
1109
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1110
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1111
+ },
1112
+ "origin": "child"
1113
+ },
1114
+ {
1115
+ "core_sha256": {
1116
+ "protocol.json": "82df8a37d69e6d815b440a40fb4ebcaa04597fa52bdc54f6ea0112d2096565d9",
1117
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1118
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1119
+ "study.json": "ab523f28178c00ecec133dd6773da6f2cbbe8d26630cf38671f7d8745406919c",
1120
+ "environment.json": "f0e702a8a418d683763e0a0d35acd4377d246e006d1c7533c9bf6cac38fd70bf"
1121
+ },
1122
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1123
+ "task_count": 1000,
1124
+ "task_input_blobs_rehashed": false,
1125
+ "source_and_dependencies_match": true,
1126
+ "limits": {
1127
+ "output_tokens": 32768,
1128
+ "wall_seconds": 1800,
1129
+ "model_calls": 6,
1130
+ "tool_calls": 4,
1131
+ "tool_timeout": 90
1132
+ },
1133
+ "generation": {
1134
+ "model": "qwen3.7-max",
1135
+ "provider": "chat-completions",
1136
+ "reasoning_effort": null,
1137
+ "stream": true,
1138
+ "timeout": 1800
1139
+ },
1140
+ "dataset_provenance": {
1141
+ "repo_id": "CamoAiLab/InferenceNet",
1142
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1143
+ "task_directory": "Selected_1000",
1144
+ "task_list": "Selected_1000/1000_new.csv",
1145
+ "prior_results_imported": false,
1146
+ "engineering_pilot": false
1147
+ },
1148
+ "execution_identity": {
1149
+ "backend": "dsw-bwrap-v3",
1150
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1151
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1152
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1153
+ "engine_memory_bytes": 12884901888,
1154
+ "engine_cpus": 4,
1155
+ "cpu_slots": [
1156
+ [
1157
+ 16,
1158
+ 17
1159
+ ],
1160
+ [
1161
+ 18,
1162
+ 19
1163
+ ]
1164
+ ],
1165
+ "resources": {
1166
+ "memory": "6g",
1167
+ "cpus": 2.0,
1168
+ "tmpfs": "256m"
1169
+ },
1170
+ "memory_enforcement": "RLIMIT_AS",
1171
+ "single_payload_process": true,
1172
+ "network": "none",
1173
+ "uid": 10001
1174
+ },
1175
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1176
+ "scorer_sha256": {
1177
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1178
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1179
+ },
1180
+ "origin": "parent"
1181
+ }
1182
+ ],
1183
+ "selected_origin_counts": {
1184
+ "child": 993,
1185
+ "parent": 7
1186
+ },
1187
+ "source_commits": [
1188
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1189
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1190
+ ],
1191
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1192
+ "official_parity_verified": false,
1193
+ "notes": [
1194
+ "Historical disjoint parent/child continuation; original source and protocol hashes retained.",
1195
+ "Different baseline/harness call budgets; not an equal-cost causal comparison.",
1196
+ "Unknown and invalid evidence remain in the denominator; no replacement or rescoring.",
1197
+ "Invalid evidence is not sealed; its unverified hash is distinct from a verified result hash."
1198
+ ]
1199
+ }
results/qwen3.7-max/deepagents/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/qwen3.7-max/deepagents/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/qwen3.7-max/deepagents/summary.json ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T17:18:47.373930+00:00",
5
+ "model": "qwen3.7-max",
6
+ "arm": "deepagents",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": false,
29
+ "all_slots_sealed": false,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 523,
37
+ "denominator": 1000,
38
+ "rate": 0.523,
39
+ "score": 52.3,
40
+ "unknown_count": 7,
41
+ "failure_count": 470,
42
+ "assessable": 993,
43
+ "coverage_percent": 99.3
44
+ },
45
+ "partial_replication": {
46
+ "count": 379,
47
+ "denominator": 1000,
48
+ "rate": 0.379,
49
+ "score": 37.9,
50
+ "unknown_count": 481,
51
+ "failure_count": 140,
52
+ "assessable": 519,
53
+ "coverage_percent": 51.9
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 488,
57
+ "denominator": 1000,
58
+ "rate": 0.488,
59
+ "score": 48.8,
60
+ "unknown_count": 481,
61
+ "failure_count": 31,
62
+ "assessable": 519,
63
+ "coverage_percent": 51.9
64
+ },
65
+ "significance_level": {
66
+ "count": 433,
67
+ "denominator": 1000,
68
+ "rate": 0.433,
69
+ "score": 43.3,
70
+ "unknown_count": 481,
71
+ "failure_count": 86,
72
+ "assessable": 519,
73
+ "coverage_percent": 51.9
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "qwen3.7-max deepagents",
78
+ "Compilation Success": 52.3,
79
+ "Partial Replication": 37.9,
80
+ "Correct Coefficient Direction": 48.8,
81
+ "Significant Level Correctness": 43.3
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 999,
92
+ "sealed_count": 999,
93
+ "definite_outcome_count": 993,
94
+ "task_unknown_count": 6,
95
+ "invalid_evidence_count": 1,
96
+ "valid_prediction_count": 519,
97
+ "task_counts": {
98
+ "succeeded": 519,
99
+ "failed": 474,
100
+ "unknown": 6,
101
+ "unsealed": 0,
102
+ "not_started": 0,
103
+ "invalid_evidence": 1,
104
+ "completed": 993,
105
+ "planned": 1000
106
+ }
107
+ },
108
+ "legacy_local_paper": {
109
+ "profile": "local-paper-v1",
110
+ "definitions": {
111
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
112
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
113
+ "coefficient_only": "coefficient relative error <= .05",
114
+ "direction": "matching strictly positive or strictly negative coefficient signs",
115
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
116
+ },
117
+ "metrics": {
118
+ "perfect": {
119
+ "count": 248,
120
+ "denominator": 1000,
121
+ "rate": 0.248,
122
+ "score": 24.8,
123
+ "unknown_count": 481,
124
+ "failure_count": 271,
125
+ "assessable": 519,
126
+ "coverage_percent": 51.9
127
+ },
128
+ "partial": {
129
+ "count": 311,
130
+ "denominator": 1000,
131
+ "rate": 0.311,
132
+ "score": 31.1,
133
+ "unknown_count": 481,
134
+ "failure_count": 208,
135
+ "assessable": 519,
136
+ "coverage_percent": 51.9
137
+ },
138
+ "coefficient_only": {
139
+ "count": 379,
140
+ "denominator": 1000,
141
+ "rate": 0.379,
142
+ "score": 37.9,
143
+ "unknown_count": 481,
144
+ "failure_count": 140,
145
+ "assessable": 519,
146
+ "coverage_percent": 51.9
147
+ },
148
+ "direction": {
149
+ "count": 488,
150
+ "denominator": 1000,
151
+ "rate": 0.488,
152
+ "score": 48.8,
153
+ "unknown_count": 481,
154
+ "failure_count": 31,
155
+ "assessable": 519,
156
+ "coverage_percent": 51.9
157
+ },
158
+ "significance": {
159
+ "count": 416,
160
+ "denominator": 1000,
161
+ "rate": 0.416,
162
+ "score": 41.6,
163
+ "unknown_count": 481,
164
+ "failure_count": 103,
165
+ "assessable": 519,
166
+ "coverage_percent": 51.9
167
+ }
168
+ }
169
+ },
170
+ "verification_scope": "Offline equality to accepted flags, hashes and states; raw predictions unavailable; not a fresh native audit or official scorer parity check.",
171
+ "usage": {
172
+ "sealed_model_calls": 5447,
173
+ "unknown_usage_calls": 6,
174
+ "unsealed_episodes_excluded": 0,
175
+ "invalid_evidence_episodes_excluded": 1,
176
+ "known_token_subtotal": {
177
+ "input_tokens": 28634568,
178
+ "output_tokens": 2738654,
179
+ "total_tokens": 31373222
180
+ },
181
+ "unknown_by_field": {
182
+ "input_tokens": 6,
183
+ "output_tokens": 6,
184
+ "total_tokens": 6
185
+ },
186
+ "total_tokens_complete": false,
187
+ "usd": null,
188
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
189
+ "summed_episode_seconds": 228788.56599999996
190
+ }
191
+ }
results/qwen3.7-max/diagnostics.json ADDED
@@ -0,0 +1,74 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage_counts": {
3
+ "baseline": [
4
+ {
5
+ "generation": "succeeded",
6
+ "execution": "failed",
7
+ "scoring": "pending",
8
+ "count": 500
9
+ },
10
+ {
11
+ "generation": "succeeded",
12
+ "execution": "succeeded",
13
+ "scoring": "scored",
14
+ "count": 483
15
+ },
16
+ {
17
+ "generation": "failed",
18
+ "execution": "pending",
19
+ "scoring": "pending",
20
+ "count": 17
21
+ }
22
+ ],
23
+ "deepagents": [
24
+ {
25
+ "generation": "failed",
26
+ "execution": "pending",
27
+ "scoring": "pending",
28
+ "count": 401
29
+ },
30
+ {
31
+ "generation": "succeeded",
32
+ "execution": "succeeded",
33
+ "scoring": "scored",
34
+ "count": 519
35
+ },
36
+ {
37
+ "generation": "succeeded",
38
+ "execution": "failed",
39
+ "scoring": "pending",
40
+ "count": 73
41
+ },
42
+ {
43
+ "generation": "uncertain",
44
+ "execution": "pending",
45
+ "scoring": "pending",
46
+ "count": 6
47
+ },
48
+ {
49
+ "generation": null,
50
+ "execution": null,
51
+ "scoring": null,
52
+ "count": 1
53
+ }
54
+ ]
55
+ },
56
+ "task_unknown_ids": {
57
+ "baseline": [],
58
+ "deepagents": [
59
+ 86,
60
+ 440,
61
+ 597,
62
+ 609,
63
+ 718,
64
+ 882
65
+ ]
66
+ },
67
+ "invalid_evidence_ids": {
68
+ "baseline": [],
69
+ "deepagents": [
70
+ 453
71
+ ]
72
+ },
73
+ "raw_error_reasons_verified": false
74
+ }
results/qwen3.7-max/export-verification.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "passed_offline",
3
+ "per_arm": {
4
+ "baseline": {
5
+ "records": 1000,
6
+ "unique_task_ids": 1000,
7
+ "counts": {
8
+ "succeeded": 483,
9
+ "failed": 517,
10
+ "unknown": 0,
11
+ "unsealed": 0,
12
+ "not_started": 0,
13
+ "invalid_evidence": 0,
14
+ "completed": 1000,
15
+ "planned": 1000
16
+ },
17
+ "sealed_count": 1000,
18
+ "all_slots_sealed": true,
19
+ "metrics_recomputed": 9,
20
+ "matches_accepted_flags_states_hashes": true
21
+ },
22
+ "deepagents": {
23
+ "records": 1000,
24
+ "unique_task_ids": 1000,
25
+ "counts": {
26
+ "succeeded": 519,
27
+ "failed": 474,
28
+ "unknown": 6,
29
+ "unsealed": 0,
30
+ "not_started": 0,
31
+ "invalid_evidence": 1,
32
+ "completed": 993,
33
+ "planned": 1000
34
+ },
35
+ "sealed_count": 999,
36
+ "all_slots_sealed": false,
37
+ "metrics_recomputed": 9,
38
+ "matches_accepted_flags_states_hashes": true
39
+ }
40
+ },
41
+ "paired_rows": 1000,
42
+ "metrics_recomputed": 18,
43
+ "official_parity_verified": false
44
+ }
results/qwen3.7-max/paired_comparison.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/qwen3.7-max/provenance.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "qwen3.7-max",
3
+ "archived_at_utc": "2026-09-16T17:18:47.373930+00:00",
4
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
5
+ "accepted_results_sha256": "dd72f352be1895b11914dd2aea1fdceb21c453932877c588ecaf81ec09861105",
6
+ "accepted_task_records_sha256": "32a797b4da8dfe90d44b37d7715fe726a29c5b4f93a1b24d7d5ffbffb81e9749",
7
+ "record_count": 2000,
8
+ "selected_task_count": 1000,
9
+ "source_commits": [
10
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
11
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
12
+ ],
13
+ "historical_native_audit": {
14
+ "mode": "fresh_in_accepted_snapshot",
15
+ "collection_interval": {
16
+ "started_at": "2026-09-16T06:46:01.177744+00:00",
17
+ "finished_at": "2026-09-16T07:43:33.877909+00:00"
18
+ }
19
+ },
20
+ "official_parity_verified": false,
21
+ "new_model_calls": 0,
22
+ "new_benchmark_executions": 0,
23
+ "verification_scope": "Offline equality and metric reaggregation of accepted score exports. No fresh input/native audit; no raw prediction rescoring."
24
+ }
results/verify_qwen_deepseek.py ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Offline validation for the four added Qwen/DeepSeek result groups."""
2
+ import csv
3
+ import hashlib
4
+ import json
5
+ from collections import Counter
6
+ from pathlib import Path
7
+
8
+ root=Path(__file__).resolve().parent
9
+ for line in (root/'manifest.sha256').read_text().splitlines():
10
+ expected,name=line.split(' ',1)
11
+ p=(root/name).resolve()
12
+ assert p.is_relative_to(root) and p.is_file(),name
13
+ assert hashlib.sha256(p.read_bytes()).hexdigest()==expected,name
14
+ rows_by_model={}
15
+ checked=0
16
+ for model in ('qwen3.7-max','deepseek-v4-pro'):
17
+ rows_by_model[model]={}
18
+ for arm in ('baseline','deepagents'):
19
+ p=root/model/arm
20
+ rows=[json.loads(line) for line in (p/'results.jsonl').read_text().splitlines()]
21
+ config=json.loads((p/'config.json').read_text())
22
+ summary=json.loads((p/'summary.json').read_text())
23
+ assert len(rows)==1000 and len({r['task_id'] for r in rows})==1000
24
+ assert [r['task_id'] for r in rows]==config['dataset']['task_ids']
25
+ assert all(r['model']==model and r['arm']==arm for r in rows)
26
+ assert all(r['resolved']==(r['state']!='invalid_evidence') and r['definite_outcome']==(r['state'] in ('succeeded','failed')) for r in rows)
27
+ assert all(r['prediction_export_status']==('invalid_evidence_quarantined' if r['state']=='invalid_evidence' else 'not_available_in_accepted_score_snapshot') for r in rows)
28
+ assert all(r[k] is None for r in rows for k in ('prediction_coefficient','prediction_standard_error','prediction_p_value'))
29
+ with (p/'results.csv').open() as f:csv_rows=list(csv.DictReader(f))
30
+ assert len(csv_rows)==1000
31
+ for j,c in zip(rows,csv_rows):
32
+ assert set(j)==set(c)
33
+ assert all(c[k]==('' if v is None else str(v)) for k,v in j.items())
34
+ assert summary['official_parity_verified'] is False and summary['included_in_displayed_leaderboard'] is False
35
+ counts=summary['result']['task_counts']
36
+ assert Counter(r['state'] for r in rows)=={k:counts[k] for k in ('succeeded','failed','unknown','invalid_evidence') if counts[k]}
37
+ assert sum(r['state']=='unknown' for r in rows)==summary['result']['task_unknown_count']
38
+ assert summary['result']['invalid_evidence_count']==sum(r['state']=='invalid_evidence' for r in rows)
39
+ assert summary['result']['sealed_count']==sum(r['resolved'] for r in rows)
40
+ assert summary['all_slots_sealed']==all(r['resolved'] for r in rows)
41
+ for r in rows:
42
+ if r['state']=='invalid_evidence':
43
+ assert r['selected_result_sha256'] is None and r['unverified_result_sha256']
44
+ assert all(r[k] is None for k in r if k.startswith('local_') or k in summary['result']['metrics'])
45
+ for metrics,prefix in [(summary['result']['metrics'],''),(summary['legacy_local_paper']['metrics'],'local_')]:
46
+ for name,m in metrics.items():
47
+ flags=[r[prefix+name] for r in rows]
48
+ assert all(v is None or type(v) is bool for v in flags)
49
+ good,bad,unknown=sum(v is True for v in flags),sum(v is False for v in flags),sum(v is None for v in flags)
50
+ assert (m['count'],m['failure_count'],m['unknown_count'],m['denominator'])==(good,bad,unknown,1000)
51
+ assert m['score']==round(good/10,1) and m['rate']==good/1000
52
+ assert m['assessable']==good+bad and m['coverage_percent']==round((good+bad)/10,1)
53
+ checked+=1
54
+ rows_by_model[model][arm]=rows
55
+ b,d=rows_by_model[model]['baseline'],rows_by_model[model]['deepagents']
56
+ assert [r['task_id'] for r in b]==[r['task_id'] for r in d]
57
+ with (root/model/'paired_comparison.csv').open() as f:pairs=list(csv.DictReader(f))
58
+ assert len(pairs)==1000
59
+ for i,pair in enumerate(pairs):
60
+ assert int(pair['task_id'])==b[i]['task_id']
61
+ for arm in ('baseline','deepagents'):
62
+ r=rows_by_model[model][arm][i]
63
+ assert pair[arm+'_result_sha256']==(r['selected_result_sha256'] or '')
64
+ assert pair[arm+'_unverified_result_sha256']==(r['unverified_result_sha256'] or '')
65
+ assert pair[arm+'_state']==r['state']
66
+ for field,value in r.items():
67
+ key=arm+'_'+field if field.startswith('local_') else arm+'_hf_'+field
68
+ if key in pair:assert pair[key]==('' if value is None else str(value))
69
+ assert checked==36
70
+ print('PASS: archive manifest; 4 x 1000 unique task records; paired rows; 36 metric aggregates; unknowns and invalid evidence retained.')
71
+ print('Offline export verification only; not fresh raw prediction scoring or official scorer parity.')