Archive Kimi K3 and Gemini 3.1 Pro paired evaluation results

#2
by YICHEN013 - opened
Files changed (32) hide show
  1. results/README.md +10 -0
  2. results/gemini-3.1-pro-preview/baseline/README.md +16 -0
  3. results/gemini-3.1-pro-preview/baseline/config.json +1205 -0
  4. results/gemini-3.1-pro-preview/baseline/results.csv +0 -0
  5. results/gemini-3.1-pro-preview/baseline/results.jsonl +0 -0
  6. results/gemini-3.1-pro-preview/baseline/summary.json +190 -0
  7. results/gemini-3.1-pro-preview/deepagents/README.md +16 -0
  8. results/gemini-3.1-pro-preview/deepagents/config.json +1205 -0
  9. results/gemini-3.1-pro-preview/deepagents/results.csv +0 -0
  10. results/gemini-3.1-pro-preview/deepagents/results.jsonl +0 -0
  11. results/gemini-3.1-pro-preview/deepagents/summary.json +190 -0
  12. results/gemini-3.1-pro-preview/diagnostics.json +50 -0
  13. results/gemini-3.1-pro-preview/export-verification.json +47 -0
  14. results/gemini-3.1-pro-preview/paired_comparison.csv +0 -0
  15. results/gemini-3.1-pro-preview/provenance.json +61 -0
  16. results/index.json +26 -1
  17. results/kimi-k3/baseline/README.md +16 -0
  18. results/kimi-k3/baseline/config.json +1276 -0
  19. results/kimi-k3/baseline/results.csv +0 -0
  20. results/kimi-k3/baseline/results.jsonl +0 -0
  21. results/kimi-k3/baseline/summary.json +190 -0
  22. results/kimi-k3/deepagents/README.md +16 -0
  23. results/kimi-k3/deepagents/config.json +1276 -0
  24. results/kimi-k3/deepagents/results.csv +0 -0
  25. results/kimi-k3/deepagents/results.jsonl +0 -0
  26. results/kimi-k3/deepagents/summary.json +190 -0
  27. results/kimi-k3/diagnostics.json +90 -0
  28. results/kimi-k3/export-verification.json +49 -0
  29. results/kimi-k3/paired_comparison.csv +0 -0
  30. results/kimi-k3/provenance.json +45 -0
  31. results/manifest.sha256 +31 -2
  32. results/verify_kimi_gemini.py +63 -0
results/README.md CHANGED
@@ -6,9 +6,19 @@
6
  |---|---|---|
7
  | GPT-5.6 Sol | [1,000 题](gpt-5.6-sol/baseline/) | [1,000 题](gpt-5.6-sol/deepagents/) |
8
  | Claude Opus 4.8 | [1,000 题](claude-opus-4-8/baseline/) | [1,000 题](claude-opus-4-8/deepagents/) |
 
 
9
 
10
  这里先保存完成的实验产物。页面仍只读取根目录 `/results.csv`,本次没有更新它或修改页面代码。等各模型结果齐备,再核对数据、评分和协议后统一整理正式榜单。
11
 
12
  每组均保留失败、无预测与未知记录,固定分母为 1,000。历史恢复与补充协议写入元数据。四维离线导出与旧 local-paper-v1 指标分开保存,目前不声称已与官方 scorer 完全对齐。
13
 
14
  只归档选定最终预测、状态、指标及配置;没有重新调用模型或执行实验。原始数据、隐藏答案、API 请求/响应和凭据不在上传范围内。
 
 
 
 
 
 
 
 
 
6
  |---|---|---|
7
  | GPT-5.6 Sol | [1,000 题](gpt-5.6-sol/baseline/) | [1,000 题](gpt-5.6-sol/deepagents/) |
8
  | Claude Opus 4.8 | [1,000 题](claude-opus-4-8/baseline/) | [1,000 题](claude-opus-4-8/deepagents/) |
9
+ | Kimi K3 | [1,000 题](kimi-k3/baseline/) | [1,000 题](kimi-k3/deepagents/) |
10
+ | Gemini 3.1 Pro Preview | [1,000 题](gemini-3.1-pro-preview/baseline/) | [1,000 题](gemini-3.1-pro-preview/deepagents/) |
11
 
12
  这里先保存完成的实验产物。页面仍只读取根目录 `/results.csv`,本次没有更新它或修改页面代码。等各模型结果齐备,再核对数据、评分和协议后统一整理正式榜单。
13
 
14
  每组均保留失败、无预测与未知记录,固定分母为 1,000。历史恢复与补充协议写入元数据。四维离线导出与旧 local-paper-v1 指标分开保存,目前不声称已与官方 scorer 完全对齐。
15
 
16
  只归档选定最终预测、状态、指标及配置;没有重新调用模型或执行实验。原始数据、隐藏答案、API 请求/响应和凭据不在上传范围内。
17
+
18
+ ## Kimi K3 与 Gemini 3.1 Pro Preview
19
+
20
+ 新增两模型的单次生成 / DeepAgents 共四组、4,000 条记录。原生完整复现率分别为 K3 23.0% / 39.0%、Gemini 24.6% / 39.1%;这些与 HF 四维指标定义不同。
21
+
22
+ K3 保留两组 5 / 4 条任务 unknown,Gemini 无任务 unknown;每组固定 1,000 题分母。指标级 null 与任务级 unknown 分别记账。新增记录的原始预测不在源评分快照中,prediction 字段因此留空并带有导出状态标记;不把导出缺失解释为模型未生成预测。详细协议、源版本、阶段诊断和核验范围见各模型目录。
23
+
24
+ 运行 `python3 -B results/verify_kimi_gemini.py` 可检查归档清单、4,000 条新增逐题记录及 36 项指标汇总;这是离线导出核验,不是重新运行实验或官方 scorer 一致性认证。
results/gemini-3.1-pro-preview/baseline/README.md ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Gemini 3.1 Pro Preview / baseline
2
+
3
+ 1,000 个已封存槽位;成功 540,确定失败 460,任务 unknown 0。固定分母为 1,000;完整复现率 24.6%。
4
+
5
+ - `results.jsonl` / `results.csv`:相同的逐题状态、指标、调用用量和结果哈希。
6
+ - `summary.json`:本地 HF 四维与 local-paper-v1 五维分别保存,包含各自 unknown 和可判定覆盖率。
7
+ - `config.json`:固定任务清单、模型接口、预算、父/子接续来源及哈希。
8
+ - 上一级的 `paired_comparison.csv`、`provenance.json`、`diagnostics.json` 与 `export-verification.json` 保留双组配对与原导出核验边界。
9
+
10
+ `resolved=true` 表示槽位已封存;`definite_outcome` 才表示结果确定。未知和失败都保留在分母中。阶段 pending 可能表示前序失败后未执行到该阶段,不代表还在运行。
11
+
12
+ **此批次只有评分/状态导出,原始预测三元组未包含在源快照中。** 因此 prediction 字段保留 null,并用 `prediction_export_status=not_available_in_accepted_score_snapshot` 区分“导出不可用”与“模型没有有效预测”;不能据此把所有任务判为无预测。`summary.json` 的 valid_prediction_count 来自已核验的执行/评分状态。
13
+
14
+ 四维指标为本地实现,尚未验证与官方 scorer 完全一致。本次为历史 amended comparison 的归档,不自动更新正式榜单,没有重新调用模型或执行程序。用量字段是已知小计,须结合各字段 unknown_usage_calls 阅读。
15
+
16
+ JSON null / CSV 空单元格表示未知或不可用。原始数据、隐藏答案、完整提示词/API 轨迹、凭据和机器路径未包含在归档中。
results/gemini-3.1-pro-preview/baseline/config.json ADDED
@@ -0,0 +1,1205 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "gemini-3.1-pro-preview",
4
+ "arm": "baseline",
5
+ "harness": "none",
6
+ "archived_at_utc": "2026-09-16T15:28:00.953066+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "engineering_pilot": false,
14
+ "expected_tasks": 1000,
15
+ "task_ids": [
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4,
20
+ 5,
21
+ 6,
22
+ 7,
23
+ 8,
24
+ 9,
25
+ 10,
26
+ 11,
27
+ 12,
28
+ 13,
29
+ 14,
30
+ 15,
31
+ 16,
32
+ 17,
33
+ 18,
34
+ 19,
35
+ 20,
36
+ 21,
37
+ 22,
38
+ 23,
39
+ 24,
40
+ 25,
41
+ 26,
42
+ 27,
43
+ 28,
44
+ 29,
45
+ 30,
46
+ 31,
47
+ 32,
48
+ 33,
49
+ 34,
50
+ 35,
51
+ 36,
52
+ 37,
53
+ 38,
54
+ 39,
55
+ 40,
56
+ 41,
57
+ 42,
58
+ 43,
59
+ 44,
60
+ 45,
61
+ 46,
62
+ 47,
63
+ 48,
64
+ 49,
65
+ 50,
66
+ 51,
67
+ 52,
68
+ 53,
69
+ 54,
70
+ 55,
71
+ 56,
72
+ 57,
73
+ 58,
74
+ 59,
75
+ 60,
76
+ 61,
77
+ 62,
78
+ 63,
79
+ 64,
80
+ 65,
81
+ 66,
82
+ 67,
83
+ 68,
84
+ 69,
85
+ 70,
86
+ 71,
87
+ 72,
88
+ 73,
89
+ 74,
90
+ 75,
91
+ 76,
92
+ 77,
93
+ 78,
94
+ 79,
95
+ 80,
96
+ 81,
97
+ 82,
98
+ 83,
99
+ 84,
100
+ 85,
101
+ 86,
102
+ 87,
103
+ 88,
104
+ 89,
105
+ 90,
106
+ 91,
107
+ 92,
108
+ 93,
109
+ 94,
110
+ 95,
111
+ 96,
112
+ 97,
113
+ 98,
114
+ 100,
115
+ 101,
116
+ 102,
117
+ 103,
118
+ 104,
119
+ 105,
120
+ 106,
121
+ 107,
122
+ 108,
123
+ 109,
124
+ 110,
125
+ 111,
126
+ 112,
127
+ 113,
128
+ 114,
129
+ 115,
130
+ 116,
131
+ 117,
132
+ 118,
133
+ 119,
134
+ 120,
135
+ 121,
136
+ 122,
137
+ 123,
138
+ 124,
139
+ 125,
140
+ 126,
141
+ 127,
142
+ 128,
143
+ 129,
144
+ 130,
145
+ 131,
146
+ 132,
147
+ 133,
148
+ 134,
149
+ 135,
150
+ 136,
151
+ 137,
152
+ 138,
153
+ 139,
154
+ 140,
155
+ 141,
156
+ 142,
157
+ 143,
158
+ 144,
159
+ 145,
160
+ 146,
161
+ 147,
162
+ 148,
163
+ 149,
164
+ 150,
165
+ 151,
166
+ 152,
167
+ 153,
168
+ 154,
169
+ 155,
170
+ 156,
171
+ 157,
172
+ 158,
173
+ 159,
174
+ 160,
175
+ 161,
176
+ 162,
177
+ 163,
178
+ 164,
179
+ 165,
180
+ 166,
181
+ 167,
182
+ 168,
183
+ 169,
184
+ 170,
185
+ 171,
186
+ 172,
187
+ 173,
188
+ 174,
189
+ 175,
190
+ 176,
191
+ 177,
192
+ 178,
193
+ 179,
194
+ 180,
195
+ 181,
196
+ 182,
197
+ 183,
198
+ 184,
199
+ 185,
200
+ 186,
201
+ 187,
202
+ 188,
203
+ 189,
204
+ 190,
205
+ 191,
206
+ 192,
207
+ 193,
208
+ 194,
209
+ 195,
210
+ 196,
211
+ 197,
212
+ 198,
213
+ 199,
214
+ 200,
215
+ 201,
216
+ 202,
217
+ 203,
218
+ 204,
219
+ 205,
220
+ 206,
221
+ 207,
222
+ 208,
223
+ 209,
224
+ 210,
225
+ 211,
226
+ 212,
227
+ 213,
228
+ 214,
229
+ 215,
230
+ 216,
231
+ 217,
232
+ 218,
233
+ 219,
234
+ 220,
235
+ 221,
236
+ 222,
237
+ 223,
238
+ 224,
239
+ 225,
240
+ 226,
241
+ 227,
242
+ 228,
243
+ 229,
244
+ 230,
245
+ 231,
246
+ 232,
247
+ 233,
248
+ 234,
249
+ 235,
250
+ 236,
251
+ 237,
252
+ 238,
253
+ 239,
254
+ 240,
255
+ 241,
256
+ 242,
257
+ 248,
258
+ 249,
259
+ 250,
260
+ 251,
261
+ 252,
262
+ 253,
263
+ 254,
264
+ 255,
265
+ 256,
266
+ 257,
267
+ 258,
268
+ 259,
269
+ 261,
270
+ 262,
271
+ 263,
272
+ 264,
273
+ 265,
274
+ 266,
275
+ 290,
276
+ 291,
277
+ 292,
278
+ 293,
279
+ 294,
280
+ 295,
281
+ 302,
282
+ 303,
283
+ 304,
284
+ 305,
285
+ 306,
286
+ 307,
287
+ 308,
288
+ 309,
289
+ 310,
290
+ 311,
291
+ 312,
292
+ 313,
293
+ 314,
294
+ 315,
295
+ 316,
296
+ 317,
297
+ 318,
298
+ 319,
299
+ 320,
300
+ 321,
301
+ 322,
302
+ 323,
303
+ 324,
304
+ 325,
305
+ 326,
306
+ 327,
307
+ 328,
308
+ 329,
309
+ 330,
310
+ 331,
311
+ 332,
312
+ 333,
313
+ 334,
314
+ 335,
315
+ 336,
316
+ 337,
317
+ 338,
318
+ 339,
319
+ 340,
320
+ 341,
321
+ 342,
322
+ 343,
323
+ 344,
324
+ 345,
325
+ 346,
326
+ 347,
327
+ 348,
328
+ 349,
329
+ 350,
330
+ 351,
331
+ 352,
332
+ 353,
333
+ 354,
334
+ 355,
335
+ 356,
336
+ 357,
337
+ 358,
338
+ 359,
339
+ 360,
340
+ 361,
341
+ 362,
342
+ 363,
343
+ 364,
344
+ 365,
345
+ 366,
346
+ 367,
347
+ 368,
348
+ 369,
349
+ 370,
350
+ 371,
351
+ 372,
352
+ 373,
353
+ 374,
354
+ 375,
355
+ 376,
356
+ 377,
357
+ 378,
358
+ 379,
359
+ 380,
360
+ 381,
361
+ 382,
362
+ 383,
363
+ 384,
364
+ 385,
365
+ 386,
366
+ 387,
367
+ 388,
368
+ 389,
369
+ 390,
370
+ 391,
371
+ 392,
372
+ 393,
373
+ 394,
374
+ 395,
375
+ 396,
376
+ 397,
377
+ 398,
378
+ 399,
379
+ 400,
380
+ 401,
381
+ 402,
382
+ 403,
383
+ 404,
384
+ 405,
385
+ 406,
386
+ 407,
387
+ 408,
388
+ 409,
389
+ 410,
390
+ 411,
391
+ 412,
392
+ 413,
393
+ 414,
394
+ 415,
395
+ 416,
396
+ 417,
397
+ 418,
398
+ 419,
399
+ 420,
400
+ 421,
401
+ 422,
402
+ 423,
403
+ 424,
404
+ 425,
405
+ 426,
406
+ 427,
407
+ 428,
408
+ 429,
409
+ 430,
410
+ 431,
411
+ 432,
412
+ 433,
413
+ 434,
414
+ 435,
415
+ 436,
416
+ 437,
417
+ 438,
418
+ 439,
419
+ 440,
420
+ 441,
421
+ 442,
422
+ 443,
423
+ 444,
424
+ 445,
425
+ 446,
426
+ 447,
427
+ 448,
428
+ 449,
429
+ 450,
430
+ 451,
431
+ 452,
432
+ 453,
433
+ 454,
434
+ 455,
435
+ 456,
436
+ 457,
437
+ 458,
438
+ 459,
439
+ 460,
440
+ 461,
441
+ 462,
442
+ 463,
443
+ 464,
444
+ 465,
445
+ 466,
446
+ 467,
447
+ 468,
448
+ 469,
449
+ 470,
450
+ 471,
451
+ 472,
452
+ 473,
453
+ 474,
454
+ 475,
455
+ 476,
456
+ 477,
457
+ 478,
458
+ 479,
459
+ 480,
460
+ 481,
461
+ 482,
462
+ 483,
463
+ 484,
464
+ 485,
465
+ 486,
466
+ 487,
467
+ 488,
468
+ 489,
469
+ 490,
470
+ 491,
471
+ 492,
472
+ 493,
473
+ 494,
474
+ 495,
475
+ 496,
476
+ 497,
477
+ 498,
478
+ 499,
479
+ 500,
480
+ 501,
481
+ 502,
482
+ 503,
483
+ 504,
484
+ 505,
485
+ 506,
486
+ 507,
487
+ 508,
488
+ 509,
489
+ 510,
490
+ 511,
491
+ 512,
492
+ 513,
493
+ 514,
494
+ 515,
495
+ 516,
496
+ 517,
497
+ 518,
498
+ 519,
499
+ 520,
500
+ 521,
501
+ 522,
502
+ 523,
503
+ 524,
504
+ 525,
505
+ 526,
506
+ 527,
507
+ 528,
508
+ 529,
509
+ 530,
510
+ 531,
511
+ 532,
512
+ 533,
513
+ 534,
514
+ 535,
515
+ 536,
516
+ 537,
517
+ 538,
518
+ 539,
519
+ 540,
520
+ 541,
521
+ 542,
522
+ 543,
523
+ 544,
524
+ 545,
525
+ 546,
526
+ 547,
527
+ 548,
528
+ 549,
529
+ 550,
530
+ 551,
531
+ 552,
532
+ 553,
533
+ 554,
534
+ 555,
535
+ 556,
536
+ 557,
537
+ 558,
538
+ 559,
539
+ 560,
540
+ 561,
541
+ 562,
542
+ 563,
543
+ 565,
544
+ 566,
545
+ 567,
546
+ 568,
547
+ 570,
548
+ 571,
549
+ 572,
550
+ 573,
551
+ 574,
552
+ 575,
553
+ 576,
554
+ 577,
555
+ 578,
556
+ 579,
557
+ 580,
558
+ 581,
559
+ 582,
560
+ 583,
561
+ 584,
562
+ 585,
563
+ 586,
564
+ 587,
565
+ 588,
566
+ 589,
567
+ 590,
568
+ 591,
569
+ 592,
570
+ 593,
571
+ 594,
572
+ 595,
573
+ 596,
574
+ 597,
575
+ 598,
576
+ 601,
577
+ 602,
578
+ 603,
579
+ 604,
580
+ 605,
581
+ 606,
582
+ 607,
583
+ 608,
584
+ 609,
585
+ 610,
586
+ 611,
587
+ 612,
588
+ 613,
589
+ 614,
590
+ 615,
591
+ 616,
592
+ 617,
593
+ 618,
594
+ 619,
595
+ 620,
596
+ 621,
597
+ 622,
598
+ 623,
599
+ 624,
600
+ 664,
601
+ 665,
602
+ 666,
603
+ 667,
604
+ 668,
605
+ 669,
606
+ 670,
607
+ 671,
608
+ 672,
609
+ 673,
610
+ 674,
611
+ 675,
612
+ 676,
613
+ 677,
614
+ 678,
615
+ 679,
616
+ 680,
617
+ 681,
618
+ 682,
619
+ 683,
620
+ 684,
621
+ 685,
622
+ 686,
623
+ 687,
624
+ 688,
625
+ 689,
626
+ 690,
627
+ 691,
628
+ 692,
629
+ 693,
630
+ 694,
631
+ 695,
632
+ 696,
633
+ 697,
634
+ 698,
635
+ 699,
636
+ 700,
637
+ 701,
638
+ 702,
639
+ 703,
640
+ 704,
641
+ 705,
642
+ 706,
643
+ 707,
644
+ 708,
645
+ 709,
646
+ 710,
647
+ 711,
648
+ 712,
649
+ 713,
650
+ 714,
651
+ 715,
652
+ 716,
653
+ 717,
654
+ 718,
655
+ 719,
656
+ 720,
657
+ 721,
658
+ 722,
659
+ 723,
660
+ 724,
661
+ 725,
662
+ 726,
663
+ 727,
664
+ 728,
665
+ 729,
666
+ 730,
667
+ 731,
668
+ 732,
669
+ 733,
670
+ 734,
671
+ 735,
672
+ 736,
673
+ 737,
674
+ 739,
675
+ 740,
676
+ 741,
677
+ 742,
678
+ 743,
679
+ 744,
680
+ 745,
681
+ 746,
682
+ 747,
683
+ 748,
684
+ 749,
685
+ 750,
686
+ 751,
687
+ 752,
688
+ 753,
689
+ 754,
690
+ 755,
691
+ 756,
692
+ 757,
693
+ 758,
694
+ 759,
695
+ 760,
696
+ 761,
697
+ 762,
698
+ 763,
699
+ 764,
700
+ 765,
701
+ 766,
702
+ 767,
703
+ 768,
704
+ 772,
705
+ 773,
706
+ 774,
707
+ 775,
708
+ 776,
709
+ 777,
710
+ 778,
711
+ 779,
712
+ 780,
713
+ 781,
714
+ 782,
715
+ 783,
716
+ 784,
717
+ 785,
718
+ 786,
719
+ 787,
720
+ 788,
721
+ 789,
722
+ 790,
723
+ 791,
724
+ 793,
725
+ 794,
726
+ 795,
727
+ 796,
728
+ 797,
729
+ 798,
730
+ 799,
731
+ 800,
732
+ 801,
733
+ 802,
734
+ 803,
735
+ 804,
736
+ 805,
737
+ 806,
738
+ 807,
739
+ 808,
740
+ 825,
741
+ 826,
742
+ 827,
743
+ 828,
744
+ 829,
745
+ 830,
746
+ 831,
747
+ 832,
748
+ 833,
749
+ 834,
750
+ 835,
751
+ 836,
752
+ 837,
753
+ 838,
754
+ 839,
755
+ 840,
756
+ 841,
757
+ 842,
758
+ 843,
759
+ 844,
760
+ 845,
761
+ 846,
762
+ 847,
763
+ 848,
764
+ 849,
765
+ 850,
766
+ 851,
767
+ 852,
768
+ 853,
769
+ 854,
770
+ 855,
771
+ 856,
772
+ 857,
773
+ 858,
774
+ 859,
775
+ 860,
776
+ 861,
777
+ 862,
778
+ 863,
779
+ 864,
780
+ 865,
781
+ 866,
782
+ 867,
783
+ 868,
784
+ 869,
785
+ 870,
786
+ 871,
787
+ 872,
788
+ 873,
789
+ 874,
790
+ 875,
791
+ 876,
792
+ 877,
793
+ 878,
794
+ 879,
795
+ 880,
796
+ 881,
797
+ 882,
798
+ 883,
799
+ 884,
800
+ 885,
801
+ 886,
802
+ 887,
803
+ 888,
804
+ 889,
805
+ 890,
806
+ 891,
807
+ 892,
808
+ 893,
809
+ 894,
810
+ 895,
811
+ 896,
812
+ 897,
813
+ 898,
814
+ 899,
815
+ 900,
816
+ 901,
817
+ 902,
818
+ 903,
819
+ 906,
820
+ 907,
821
+ 913,
822
+ 914,
823
+ 915,
824
+ 917,
825
+ 920,
826
+ 921,
827
+ 922,
828
+ 923,
829
+ 924,
830
+ 925,
831
+ 926,
832
+ 927,
833
+ 928,
834
+ 929,
835
+ 930,
836
+ 931,
837
+ 932,
838
+ 933,
839
+ 934,
840
+ 935,
841
+ 936,
842
+ 937,
843
+ 938,
844
+ 939,
845
+ 940,
846
+ 941,
847
+ 942,
848
+ 943,
849
+ 944,
850
+ 945,
851
+ 946,
852
+ 947,
853
+ 948,
854
+ 949,
855
+ 950,
856
+ 951,
857
+ 952,
858
+ 953,
859
+ 954,
860
+ 955,
861
+ 956,
862
+ 957,
863
+ 958,
864
+ 959,
865
+ 960,
866
+ 961,
867
+ 962,
868
+ 963,
869
+ 964,
870
+ 965,
871
+ 966,
872
+ 967,
873
+ 968,
874
+ 969,
875
+ 970,
876
+ 971,
877
+ 972,
878
+ 973,
879
+ 974,
880
+ 975,
881
+ 976,
882
+ 977,
883
+ 978,
884
+ 979,
885
+ 980,
886
+ 981,
887
+ 982,
888
+ 983,
889
+ 984,
890
+ 985,
891
+ 1001,
892
+ 1002,
893
+ 1003,
894
+ 1004,
895
+ 1005,
896
+ 1006,
897
+ 1007,
898
+ 1008,
899
+ 1009,
900
+ 1010,
901
+ 1011,
902
+ 1012,
903
+ 1013,
904
+ 1014,
905
+ 1015,
906
+ 1016,
907
+ 1017,
908
+ 1018,
909
+ 1019,
910
+ 1020,
911
+ 1021,
912
+ 1022,
913
+ 1023,
914
+ 1024,
915
+ 1025,
916
+ 1026,
917
+ 1027,
918
+ 1028,
919
+ 1029,
920
+ 1030,
921
+ 1031,
922
+ 1032,
923
+ 1033,
924
+ 1034,
925
+ 1035,
926
+ 1036,
927
+ 1037,
928
+ 1038,
929
+ 1039,
930
+ 1040,
931
+ 1041,
932
+ 1042,
933
+ 1043,
934
+ 1044,
935
+ 1045,
936
+ 1046,
937
+ 1047,
938
+ 1048,
939
+ 1049,
940
+ 1050,
941
+ 1051,
942
+ 1052,
943
+ 1053,
944
+ 1054,
945
+ 1055,
946
+ 1056,
947
+ 1057,
948
+ 1058,
949
+ 1059,
950
+ 1060,
951
+ 1061,
952
+ 1062,
953
+ 1063,
954
+ 1064,
955
+ 1065,
956
+ 1066,
957
+ 1067,
958
+ 1068,
959
+ 1069,
960
+ 1070,
961
+ 1071,
962
+ 1072,
963
+ 1073,
964
+ 1074,
965
+ 1075,
966
+ 1076,
967
+ 1077,
968
+ 1078,
969
+ 1079,
970
+ 1080,
971
+ 1081,
972
+ 1082,
973
+ 1083,
974
+ 1084,
975
+ 1085,
976
+ 1086,
977
+ 1087,
978
+ 1088,
979
+ 1089,
980
+ 1090,
981
+ 1091,
982
+ 1092,
983
+ 1093,
984
+ 1094,
985
+ 1095,
986
+ 1096,
987
+ 1097,
988
+ 1098,
989
+ 1099,
990
+ 1100,
991
+ 1101,
992
+ 1102,
993
+ 1103,
994
+ 1104,
995
+ 1105,
996
+ 1106,
997
+ 1107,
998
+ 1108,
999
+ 1109,
1000
+ 1110,
1001
+ 1111,
1002
+ 1112,
1003
+ 1113,
1004
+ 1114,
1005
+ 1115,
1006
+ 1116,
1007
+ 1117,
1008
+ 1118,
1009
+ 1119,
1010
+ 1120,
1011
+ 1121,
1012
+ 1122,
1013
+ 1123,
1014
+ 1124,
1015
+ 1125
1016
+ ]
1017
+ },
1018
+ "generation": {
1019
+ "model": "gemini-3.1-pro-preview",
1020
+ "provider": "chat-completions",
1021
+ "reasoning_effort": null,
1022
+ "stream": true,
1023
+ "timeout": 1800
1024
+ },
1025
+ "effective_model_call_cap": 1,
1026
+ "recorded_protocols": [
1027
+ {
1028
+ "core_sha256": {
1029
+ "protocol.json": "f7fe5ba2701d96e1f73873ce89cdc6699b4c87e9fd7f670075fb9cdd0c24207c",
1030
+ "inputs.manifest.json": "3b60ffa939b62e02ac963a73aa79998bf39158c6ba68f2a74f98c92acc7fd60a",
1031
+ "schedule.json": "919aee60ee5e8e65a350f7e27ffb6a8558bf2677f2717e8a969ffe3d8593d69e",
1032
+ "study.json": "1a808aeb0cf0153b049743d07598743835d504b1be7fdeb78c1365c263279d67",
1033
+ "environment.json": "04f141c80a13f3620f971c3217dccd28c2ee58f22a3d1fc5f7db2bfd1f4cc2ef"
1034
+ },
1035
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1036
+ "task_count": 1000,
1037
+ "task_input_blobs_rehashed": false,
1038
+ "source_and_dependencies_match": true,
1039
+ "limits": {
1040
+ "output_tokens": 32768,
1041
+ "wall_seconds": 1800,
1042
+ "model_calls": 6,
1043
+ "tool_calls": 4,
1044
+ "tool_timeout": 90
1045
+ },
1046
+ "generation": {
1047
+ "model": "gemini-3.1-pro-preview",
1048
+ "provider": "chat-completions",
1049
+ "reasoning_effort": null,
1050
+ "stream": true,
1051
+ "timeout": 1800
1052
+ },
1053
+ "dataset_provenance": {
1054
+ "repo_id": "CamoAiLab/InferenceNet",
1055
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1056
+ "task_directory": "Selected_1000",
1057
+ "task_list": "Selected_1000/1000_new.csv",
1058
+ "prior_results_imported": false,
1059
+ "engineering_pilot": false
1060
+ },
1061
+ "execution_identity": {
1062
+ "backend": "dsw-bwrap-v3",
1063
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1064
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1065
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1066
+ "engine_memory_bytes": 38654705664,
1067
+ "engine_cpus": 12,
1068
+ "cpu_slots": [
1069
+ [
1070
+ 4,
1071
+ 5
1072
+ ],
1073
+ [
1074
+ 6,
1075
+ 7
1076
+ ],
1077
+ [
1078
+ 8,
1079
+ 9
1080
+ ],
1081
+ [
1082
+ 10,
1083
+ 11
1084
+ ],
1085
+ [
1086
+ 12,
1087
+ 13
1088
+ ],
1089
+ [
1090
+ 14,
1091
+ 15
1092
+ ]
1093
+ ],
1094
+ "resources": {
1095
+ "cpus": 2.0,
1096
+ "memory": "6g",
1097
+ "tmpfs": "256m"
1098
+ },
1099
+ "memory_enforcement": "RLIMIT_AS",
1100
+ "single_payload_process": true,
1101
+ "network": "none",
1102
+ "uid": 10001,
1103
+ "capacity_profile": "independent-six-slots-20260916",
1104
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1105
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1106
+ },
1107
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1108
+ "scorer_sha256": {
1109
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1110
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1111
+ },
1112
+ "arms": [
1113
+ "baseline"
1114
+ ],
1115
+ "origin": "child"
1116
+ },
1117
+ {
1118
+ "core_sha256": {
1119
+ "protocol.json": "29deb8f48389e2b8d34c352ae373bc6443ca56ea3f718979e9148957e136df90",
1120
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1121
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1122
+ "study.json": "f5cb2944526f5328274059fc2a175eebf7f75ecf742263063ba1264af708b0a7",
1123
+ "environment.json": "f0e702a8a418d683763e0a0d35acd4377d246e006d1c7533c9bf6cac38fd70bf"
1124
+ },
1125
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1126
+ "task_count": 1000,
1127
+ "task_input_blobs_rehashed": false,
1128
+ "source_and_dependencies_match": true,
1129
+ "limits": {
1130
+ "output_tokens": 32768,
1131
+ "wall_seconds": 1800,
1132
+ "model_calls": 6,
1133
+ "tool_calls": 4,
1134
+ "tool_timeout": 90
1135
+ },
1136
+ "generation": {
1137
+ "model": "gemini-3.1-pro-preview",
1138
+ "provider": "chat-completions",
1139
+ "reasoning_effort": null,
1140
+ "stream": true,
1141
+ "timeout": 1800
1142
+ },
1143
+ "dataset_provenance": {
1144
+ "repo_id": "CamoAiLab/InferenceNet",
1145
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1146
+ "task_directory": "Selected_1000",
1147
+ "task_list": "Selected_1000/1000_new.csv",
1148
+ "prior_results_imported": false,
1149
+ "engineering_pilot": false
1150
+ },
1151
+ "execution_identity": {
1152
+ "backend": "dsw-bwrap-v3",
1153
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1154
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1155
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1156
+ "engine_memory_bytes": 12884901888,
1157
+ "engine_cpus": 4,
1158
+ "cpu_slots": [
1159
+ [
1160
+ 16,
1161
+ 17
1162
+ ],
1163
+ [
1164
+ 18,
1165
+ 19
1166
+ ]
1167
+ ],
1168
+ "resources": {
1169
+ "memory": "6g",
1170
+ "cpus": 2.0,
1171
+ "tmpfs": "256m"
1172
+ },
1173
+ "memory_enforcement": "RLIMIT_AS",
1174
+ "single_payload_process": true,
1175
+ "network": "none",
1176
+ "uid": 10001
1177
+ },
1178
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1179
+ "scorer_sha256": {
1180
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1181
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1182
+ },
1183
+ "arms": [
1184
+ "baseline"
1185
+ ],
1186
+ "origin": "parent"
1187
+ }
1188
+ ],
1189
+ "selected_origin_counts": {
1190
+ "child": 1000
1191
+ },
1192
+ "source_commits": [
1193
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1194
+ ],
1195
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1196
+ "source_export_manifest_sha256": "1173b8ae18088d6cb9d62a0a747751f34ca5ebb5d91c06813ad0f40d173ac40c",
1197
+ "official_parity_verified": false,
1198
+ "notes": [
1199
+ "Selected parent and child records are disjoint; original protocol hashes and source versions retained.",
1200
+ "Baseline has one model call; DeepAgents has up to six model calls and four tool calls.",
1201
+ "Output-token budget 32768; task wall budget 1800 seconds; execution timeout 90 seconds.",
1202
+ "Final execution is separate from trial tools; unknown results are never dropped or replaced.",
1203
+ "Predictions are unavailable in this score export, even for tasks with successful scored predictions."
1204
+ ]
1205
+ }
results/gemini-3.1-pro-preview/baseline/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/gemini-3.1-pro-preview/baseline/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/gemini-3.1-pro-preview/baseline/summary.json ADDED
@@ -0,0 +1,190 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
5
+ "model": "gemini-3.1-pro-preview",
6
+ "arm": "baseline",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": true,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 543,
37
+ "denominator": 1000,
38
+ "rate": 0.543,
39
+ "score": 54.3,
40
+ "unknown_count": 0,
41
+ "failure_count": 457,
42
+ "assessable": 1000,
43
+ "coverage_percent": 100.0
44
+ },
45
+ "partial_replication": {
46
+ "count": 366,
47
+ "denominator": 1000,
48
+ "rate": 0.366,
49
+ "score": 36.6,
50
+ "unknown_count": 460,
51
+ "failure_count": 174,
52
+ "assessable": 540,
53
+ "coverage_percent": 54.0
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 506,
57
+ "denominator": 1000,
58
+ "rate": 0.506,
59
+ "score": 50.6,
60
+ "unknown_count": 460,
61
+ "failure_count": 34,
62
+ "assessable": 540,
63
+ "coverage_percent": 54.0
64
+ },
65
+ "significance_level": {
66
+ "count": 437,
67
+ "denominator": 1000,
68
+ "rate": 0.437,
69
+ "score": 43.7,
70
+ "unknown_count": 460,
71
+ "failure_count": 103,
72
+ "assessable": 540,
73
+ "coverage_percent": 54.0
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "gemini-3.1-pro-preview baseline",
78
+ "Compilation Success": 54.3,
79
+ "Partial Replication": 36.6,
80
+ "Correct Coefficient Direction": 50.6,
81
+ "Significant Level Correctness": 43.7
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 1000,
94
+ "task_unknown_count": 0,
95
+ "valid_prediction_count": 540,
96
+ "task_counts": {
97
+ "succeeded": 540,
98
+ "failed": 460,
99
+ "unknown": 0,
100
+ "unsealed": 0,
101
+ "not_started": 0,
102
+ "invalid_evidence": 0,
103
+ "completed": 1000,
104
+ "planned": 1000
105
+ }
106
+ },
107
+ "legacy_local_paper": {
108
+ "profile": "local-paper-v1",
109
+ "definitions": {
110
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
111
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
112
+ "coefficient_only": "coefficient relative error <= .05",
113
+ "direction": "matching strictly positive or strictly negative coefficient signs",
114
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
115
+ },
116
+ "metrics": {
117
+ "perfect": {
118
+ "count": 246,
119
+ "denominator": 1000,
120
+ "rate": 0.246,
121
+ "score": 24.6,
122
+ "unknown_count": 462,
123
+ "failure_count": 292,
124
+ "assessable": 538,
125
+ "coverage_percent": 53.8
126
+ },
127
+ "partial": {
128
+ "count": 304,
129
+ "denominator": 1000,
130
+ "rate": 0.304,
131
+ "score": 30.4,
132
+ "unknown_count": 462,
133
+ "failure_count": 234,
134
+ "assessable": 538,
135
+ "coverage_percent": 53.8
136
+ },
137
+ "coefficient_only": {
138
+ "count": 366,
139
+ "denominator": 1000,
140
+ "rate": 0.366,
141
+ "score": 36.6,
142
+ "unknown_count": 462,
143
+ "failure_count": 172,
144
+ "assessable": 538,
145
+ "coverage_percent": 53.8
146
+ },
147
+ "direction": {
148
+ "count": 504,
149
+ "denominator": 1000,
150
+ "rate": 0.504,
151
+ "score": 50.4,
152
+ "unknown_count": 462,
153
+ "failure_count": 34,
154
+ "assessable": 538,
155
+ "coverage_percent": 53.8
156
+ },
157
+ "significance": {
158
+ "count": 422,
159
+ "denominator": 1000,
160
+ "rate": 0.422,
161
+ "score": 42.2,
162
+ "unknown_count": 462,
163
+ "failure_count": 116,
164
+ "assessable": 538,
165
+ "coverage_percent": 53.8
166
+ }
167
+ }
168
+ },
169
+ "verification_scope": "Offline reaggregation of accepted per-task flags; raw predictions are not available in this export; official parity not verified.",
170
+ "usage": {
171
+ "sealed_model_calls": 1000,
172
+ "unknown_usage_calls": 0,
173
+ "unsealed_episodes_excluded": 0,
174
+ "invalid_evidence_episodes_excluded": 0,
175
+ "known_token_subtotal": {
176
+ "input_tokens": 451078,
177
+ "output_tokens": 8586499,
178
+ "total_tokens": 9037577
179
+ },
180
+ "unknown_by_field": {
181
+ "input_tokens": 0,
182
+ "output_tokens": 0,
183
+ "total_tokens": 0
184
+ },
185
+ "total_tokens_complete": true,
186
+ "usd": null,
187
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
188
+ "summed_episode_seconds": 109504.42499999989
189
+ }
190
+ }
results/gemini-3.1-pro-preview/deepagents/README.md ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Gemini 3.1 Pro Preview / deepagents
2
+
3
+ 1,000 个已封存槽位;成功 848,确定失败 152,任务 unknown 0。固定分母为 1,000;完整复现率 39.1%。
4
+
5
+ - `results.jsonl` / `results.csv`:相同的逐题状态、指标、调用用量和结果哈希。
6
+ - `summary.json`:本地 HF 四维与 local-paper-v1 五维分别保存,包含各自 unknown 和可判定覆盖率。
7
+ - `config.json`:固定任务清单、模型接口、预算、父/子接续来源及哈希。
8
+ - 上一级的 `paired_comparison.csv`、`provenance.json`、`diagnostics.json` 与 `export-verification.json` 保留双组配对与原导出核验边界。
9
+
10
+ `resolved=true` 表示槽位已封存;`definite_outcome` 才表示结果确定。未知和失败都保留在分母中。阶段 pending 可能表示前序失败后未执行到该阶段,不代表还在运行。
11
+
12
+ **此批次只有评分/状态导出,原始预测三元组未包含在源快照中。** 因此 prediction 字段保留 null,并用 `prediction_export_status=not_available_in_accepted_score_snapshot` 区分“导出不可用”与“模型没有有效预测”;不能据此把所有任务判为无预测。`summary.json` 的 valid_prediction_count 来自已核验的执行/评分状态。
13
+
14
+ 四维指标为本地实现,尚未验证与官方 scorer 完全一致。本次为历史 amended comparison 的归档,不自动更新正式榜单,没有重新调用模型或执行程序。用量字段是已知小计,须结合各字段 unknown_usage_calls 阅读。
15
+
16
+ JSON null / CSV 空单元格表示未知或不可用。原始数据、隐藏答案、完整提示词/API 轨迹、凭据和机器路径未包含在归档中。
results/gemini-3.1-pro-preview/deepagents/config.json ADDED
@@ -0,0 +1,1205 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "gemini-3.1-pro-preview",
4
+ "arm": "deepagents",
5
+ "harness": "DeepAgents 0.7.13",
6
+ "archived_at_utc": "2026-09-16T15:28:00.953066+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "engineering_pilot": false,
14
+ "expected_tasks": 1000,
15
+ "task_ids": [
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4,
20
+ 5,
21
+ 6,
22
+ 7,
23
+ 8,
24
+ 9,
25
+ 10,
26
+ 11,
27
+ 12,
28
+ 13,
29
+ 14,
30
+ 15,
31
+ 16,
32
+ 17,
33
+ 18,
34
+ 19,
35
+ 20,
36
+ 21,
37
+ 22,
38
+ 23,
39
+ 24,
40
+ 25,
41
+ 26,
42
+ 27,
43
+ 28,
44
+ 29,
45
+ 30,
46
+ 31,
47
+ 32,
48
+ 33,
49
+ 34,
50
+ 35,
51
+ 36,
52
+ 37,
53
+ 38,
54
+ 39,
55
+ 40,
56
+ 41,
57
+ 42,
58
+ 43,
59
+ 44,
60
+ 45,
61
+ 46,
62
+ 47,
63
+ 48,
64
+ 49,
65
+ 50,
66
+ 51,
67
+ 52,
68
+ 53,
69
+ 54,
70
+ 55,
71
+ 56,
72
+ 57,
73
+ 58,
74
+ 59,
75
+ 60,
76
+ 61,
77
+ 62,
78
+ 63,
79
+ 64,
80
+ 65,
81
+ 66,
82
+ 67,
83
+ 68,
84
+ 69,
85
+ 70,
86
+ 71,
87
+ 72,
88
+ 73,
89
+ 74,
90
+ 75,
91
+ 76,
92
+ 77,
93
+ 78,
94
+ 79,
95
+ 80,
96
+ 81,
97
+ 82,
98
+ 83,
99
+ 84,
100
+ 85,
101
+ 86,
102
+ 87,
103
+ 88,
104
+ 89,
105
+ 90,
106
+ 91,
107
+ 92,
108
+ 93,
109
+ 94,
110
+ 95,
111
+ 96,
112
+ 97,
113
+ 98,
114
+ 100,
115
+ 101,
116
+ 102,
117
+ 103,
118
+ 104,
119
+ 105,
120
+ 106,
121
+ 107,
122
+ 108,
123
+ 109,
124
+ 110,
125
+ 111,
126
+ 112,
127
+ 113,
128
+ 114,
129
+ 115,
130
+ 116,
131
+ 117,
132
+ 118,
133
+ 119,
134
+ 120,
135
+ 121,
136
+ 122,
137
+ 123,
138
+ 124,
139
+ 125,
140
+ 126,
141
+ 127,
142
+ 128,
143
+ 129,
144
+ 130,
145
+ 131,
146
+ 132,
147
+ 133,
148
+ 134,
149
+ 135,
150
+ 136,
151
+ 137,
152
+ 138,
153
+ 139,
154
+ 140,
155
+ 141,
156
+ 142,
157
+ 143,
158
+ 144,
159
+ 145,
160
+ 146,
161
+ 147,
162
+ 148,
163
+ 149,
164
+ 150,
165
+ 151,
166
+ 152,
167
+ 153,
168
+ 154,
169
+ 155,
170
+ 156,
171
+ 157,
172
+ 158,
173
+ 159,
174
+ 160,
175
+ 161,
176
+ 162,
177
+ 163,
178
+ 164,
179
+ 165,
180
+ 166,
181
+ 167,
182
+ 168,
183
+ 169,
184
+ 170,
185
+ 171,
186
+ 172,
187
+ 173,
188
+ 174,
189
+ 175,
190
+ 176,
191
+ 177,
192
+ 178,
193
+ 179,
194
+ 180,
195
+ 181,
196
+ 182,
197
+ 183,
198
+ 184,
199
+ 185,
200
+ 186,
201
+ 187,
202
+ 188,
203
+ 189,
204
+ 190,
205
+ 191,
206
+ 192,
207
+ 193,
208
+ 194,
209
+ 195,
210
+ 196,
211
+ 197,
212
+ 198,
213
+ 199,
214
+ 200,
215
+ 201,
216
+ 202,
217
+ 203,
218
+ 204,
219
+ 205,
220
+ 206,
221
+ 207,
222
+ 208,
223
+ 209,
224
+ 210,
225
+ 211,
226
+ 212,
227
+ 213,
228
+ 214,
229
+ 215,
230
+ 216,
231
+ 217,
232
+ 218,
233
+ 219,
234
+ 220,
235
+ 221,
236
+ 222,
237
+ 223,
238
+ 224,
239
+ 225,
240
+ 226,
241
+ 227,
242
+ 228,
243
+ 229,
244
+ 230,
245
+ 231,
246
+ 232,
247
+ 233,
248
+ 234,
249
+ 235,
250
+ 236,
251
+ 237,
252
+ 238,
253
+ 239,
254
+ 240,
255
+ 241,
256
+ 242,
257
+ 248,
258
+ 249,
259
+ 250,
260
+ 251,
261
+ 252,
262
+ 253,
263
+ 254,
264
+ 255,
265
+ 256,
266
+ 257,
267
+ 258,
268
+ 259,
269
+ 261,
270
+ 262,
271
+ 263,
272
+ 264,
273
+ 265,
274
+ 266,
275
+ 290,
276
+ 291,
277
+ 292,
278
+ 293,
279
+ 294,
280
+ 295,
281
+ 302,
282
+ 303,
283
+ 304,
284
+ 305,
285
+ 306,
286
+ 307,
287
+ 308,
288
+ 309,
289
+ 310,
290
+ 311,
291
+ 312,
292
+ 313,
293
+ 314,
294
+ 315,
295
+ 316,
296
+ 317,
297
+ 318,
298
+ 319,
299
+ 320,
300
+ 321,
301
+ 322,
302
+ 323,
303
+ 324,
304
+ 325,
305
+ 326,
306
+ 327,
307
+ 328,
308
+ 329,
309
+ 330,
310
+ 331,
311
+ 332,
312
+ 333,
313
+ 334,
314
+ 335,
315
+ 336,
316
+ 337,
317
+ 338,
318
+ 339,
319
+ 340,
320
+ 341,
321
+ 342,
322
+ 343,
323
+ 344,
324
+ 345,
325
+ 346,
326
+ 347,
327
+ 348,
328
+ 349,
329
+ 350,
330
+ 351,
331
+ 352,
332
+ 353,
333
+ 354,
334
+ 355,
335
+ 356,
336
+ 357,
337
+ 358,
338
+ 359,
339
+ 360,
340
+ 361,
341
+ 362,
342
+ 363,
343
+ 364,
344
+ 365,
345
+ 366,
346
+ 367,
347
+ 368,
348
+ 369,
349
+ 370,
350
+ 371,
351
+ 372,
352
+ 373,
353
+ 374,
354
+ 375,
355
+ 376,
356
+ 377,
357
+ 378,
358
+ 379,
359
+ 380,
360
+ 381,
361
+ 382,
362
+ 383,
363
+ 384,
364
+ 385,
365
+ 386,
366
+ 387,
367
+ 388,
368
+ 389,
369
+ 390,
370
+ 391,
371
+ 392,
372
+ 393,
373
+ 394,
374
+ 395,
375
+ 396,
376
+ 397,
377
+ 398,
378
+ 399,
379
+ 400,
380
+ 401,
381
+ 402,
382
+ 403,
383
+ 404,
384
+ 405,
385
+ 406,
386
+ 407,
387
+ 408,
388
+ 409,
389
+ 410,
390
+ 411,
391
+ 412,
392
+ 413,
393
+ 414,
394
+ 415,
395
+ 416,
396
+ 417,
397
+ 418,
398
+ 419,
399
+ 420,
400
+ 421,
401
+ 422,
402
+ 423,
403
+ 424,
404
+ 425,
405
+ 426,
406
+ 427,
407
+ 428,
408
+ 429,
409
+ 430,
410
+ 431,
411
+ 432,
412
+ 433,
413
+ 434,
414
+ 435,
415
+ 436,
416
+ 437,
417
+ 438,
418
+ 439,
419
+ 440,
420
+ 441,
421
+ 442,
422
+ 443,
423
+ 444,
424
+ 445,
425
+ 446,
426
+ 447,
427
+ 448,
428
+ 449,
429
+ 450,
430
+ 451,
431
+ 452,
432
+ 453,
433
+ 454,
434
+ 455,
435
+ 456,
436
+ 457,
437
+ 458,
438
+ 459,
439
+ 460,
440
+ 461,
441
+ 462,
442
+ 463,
443
+ 464,
444
+ 465,
445
+ 466,
446
+ 467,
447
+ 468,
448
+ 469,
449
+ 470,
450
+ 471,
451
+ 472,
452
+ 473,
453
+ 474,
454
+ 475,
455
+ 476,
456
+ 477,
457
+ 478,
458
+ 479,
459
+ 480,
460
+ 481,
461
+ 482,
462
+ 483,
463
+ 484,
464
+ 485,
465
+ 486,
466
+ 487,
467
+ 488,
468
+ 489,
469
+ 490,
470
+ 491,
471
+ 492,
472
+ 493,
473
+ 494,
474
+ 495,
475
+ 496,
476
+ 497,
477
+ 498,
478
+ 499,
479
+ 500,
480
+ 501,
481
+ 502,
482
+ 503,
483
+ 504,
484
+ 505,
485
+ 506,
486
+ 507,
487
+ 508,
488
+ 509,
489
+ 510,
490
+ 511,
491
+ 512,
492
+ 513,
493
+ 514,
494
+ 515,
495
+ 516,
496
+ 517,
497
+ 518,
498
+ 519,
499
+ 520,
500
+ 521,
501
+ 522,
502
+ 523,
503
+ 524,
504
+ 525,
505
+ 526,
506
+ 527,
507
+ 528,
508
+ 529,
509
+ 530,
510
+ 531,
511
+ 532,
512
+ 533,
513
+ 534,
514
+ 535,
515
+ 536,
516
+ 537,
517
+ 538,
518
+ 539,
519
+ 540,
520
+ 541,
521
+ 542,
522
+ 543,
523
+ 544,
524
+ 545,
525
+ 546,
526
+ 547,
527
+ 548,
528
+ 549,
529
+ 550,
530
+ 551,
531
+ 552,
532
+ 553,
533
+ 554,
534
+ 555,
535
+ 556,
536
+ 557,
537
+ 558,
538
+ 559,
539
+ 560,
540
+ 561,
541
+ 562,
542
+ 563,
543
+ 565,
544
+ 566,
545
+ 567,
546
+ 568,
547
+ 570,
548
+ 571,
549
+ 572,
550
+ 573,
551
+ 574,
552
+ 575,
553
+ 576,
554
+ 577,
555
+ 578,
556
+ 579,
557
+ 580,
558
+ 581,
559
+ 582,
560
+ 583,
561
+ 584,
562
+ 585,
563
+ 586,
564
+ 587,
565
+ 588,
566
+ 589,
567
+ 590,
568
+ 591,
569
+ 592,
570
+ 593,
571
+ 594,
572
+ 595,
573
+ 596,
574
+ 597,
575
+ 598,
576
+ 601,
577
+ 602,
578
+ 603,
579
+ 604,
580
+ 605,
581
+ 606,
582
+ 607,
583
+ 608,
584
+ 609,
585
+ 610,
586
+ 611,
587
+ 612,
588
+ 613,
589
+ 614,
590
+ 615,
591
+ 616,
592
+ 617,
593
+ 618,
594
+ 619,
595
+ 620,
596
+ 621,
597
+ 622,
598
+ 623,
599
+ 624,
600
+ 664,
601
+ 665,
602
+ 666,
603
+ 667,
604
+ 668,
605
+ 669,
606
+ 670,
607
+ 671,
608
+ 672,
609
+ 673,
610
+ 674,
611
+ 675,
612
+ 676,
613
+ 677,
614
+ 678,
615
+ 679,
616
+ 680,
617
+ 681,
618
+ 682,
619
+ 683,
620
+ 684,
621
+ 685,
622
+ 686,
623
+ 687,
624
+ 688,
625
+ 689,
626
+ 690,
627
+ 691,
628
+ 692,
629
+ 693,
630
+ 694,
631
+ 695,
632
+ 696,
633
+ 697,
634
+ 698,
635
+ 699,
636
+ 700,
637
+ 701,
638
+ 702,
639
+ 703,
640
+ 704,
641
+ 705,
642
+ 706,
643
+ 707,
644
+ 708,
645
+ 709,
646
+ 710,
647
+ 711,
648
+ 712,
649
+ 713,
650
+ 714,
651
+ 715,
652
+ 716,
653
+ 717,
654
+ 718,
655
+ 719,
656
+ 720,
657
+ 721,
658
+ 722,
659
+ 723,
660
+ 724,
661
+ 725,
662
+ 726,
663
+ 727,
664
+ 728,
665
+ 729,
666
+ 730,
667
+ 731,
668
+ 732,
669
+ 733,
670
+ 734,
671
+ 735,
672
+ 736,
673
+ 737,
674
+ 739,
675
+ 740,
676
+ 741,
677
+ 742,
678
+ 743,
679
+ 744,
680
+ 745,
681
+ 746,
682
+ 747,
683
+ 748,
684
+ 749,
685
+ 750,
686
+ 751,
687
+ 752,
688
+ 753,
689
+ 754,
690
+ 755,
691
+ 756,
692
+ 757,
693
+ 758,
694
+ 759,
695
+ 760,
696
+ 761,
697
+ 762,
698
+ 763,
699
+ 764,
700
+ 765,
701
+ 766,
702
+ 767,
703
+ 768,
704
+ 772,
705
+ 773,
706
+ 774,
707
+ 775,
708
+ 776,
709
+ 777,
710
+ 778,
711
+ 779,
712
+ 780,
713
+ 781,
714
+ 782,
715
+ 783,
716
+ 784,
717
+ 785,
718
+ 786,
719
+ 787,
720
+ 788,
721
+ 789,
722
+ 790,
723
+ 791,
724
+ 793,
725
+ 794,
726
+ 795,
727
+ 796,
728
+ 797,
729
+ 798,
730
+ 799,
731
+ 800,
732
+ 801,
733
+ 802,
734
+ 803,
735
+ 804,
736
+ 805,
737
+ 806,
738
+ 807,
739
+ 808,
740
+ 825,
741
+ 826,
742
+ 827,
743
+ 828,
744
+ 829,
745
+ 830,
746
+ 831,
747
+ 832,
748
+ 833,
749
+ 834,
750
+ 835,
751
+ 836,
752
+ 837,
753
+ 838,
754
+ 839,
755
+ 840,
756
+ 841,
757
+ 842,
758
+ 843,
759
+ 844,
760
+ 845,
761
+ 846,
762
+ 847,
763
+ 848,
764
+ 849,
765
+ 850,
766
+ 851,
767
+ 852,
768
+ 853,
769
+ 854,
770
+ 855,
771
+ 856,
772
+ 857,
773
+ 858,
774
+ 859,
775
+ 860,
776
+ 861,
777
+ 862,
778
+ 863,
779
+ 864,
780
+ 865,
781
+ 866,
782
+ 867,
783
+ 868,
784
+ 869,
785
+ 870,
786
+ 871,
787
+ 872,
788
+ 873,
789
+ 874,
790
+ 875,
791
+ 876,
792
+ 877,
793
+ 878,
794
+ 879,
795
+ 880,
796
+ 881,
797
+ 882,
798
+ 883,
799
+ 884,
800
+ 885,
801
+ 886,
802
+ 887,
803
+ 888,
804
+ 889,
805
+ 890,
806
+ 891,
807
+ 892,
808
+ 893,
809
+ 894,
810
+ 895,
811
+ 896,
812
+ 897,
813
+ 898,
814
+ 899,
815
+ 900,
816
+ 901,
817
+ 902,
818
+ 903,
819
+ 906,
820
+ 907,
821
+ 913,
822
+ 914,
823
+ 915,
824
+ 917,
825
+ 920,
826
+ 921,
827
+ 922,
828
+ 923,
829
+ 924,
830
+ 925,
831
+ 926,
832
+ 927,
833
+ 928,
834
+ 929,
835
+ 930,
836
+ 931,
837
+ 932,
838
+ 933,
839
+ 934,
840
+ 935,
841
+ 936,
842
+ 937,
843
+ 938,
844
+ 939,
845
+ 940,
846
+ 941,
847
+ 942,
848
+ 943,
849
+ 944,
850
+ 945,
851
+ 946,
852
+ 947,
853
+ 948,
854
+ 949,
855
+ 950,
856
+ 951,
857
+ 952,
858
+ 953,
859
+ 954,
860
+ 955,
861
+ 956,
862
+ 957,
863
+ 958,
864
+ 959,
865
+ 960,
866
+ 961,
867
+ 962,
868
+ 963,
869
+ 964,
870
+ 965,
871
+ 966,
872
+ 967,
873
+ 968,
874
+ 969,
875
+ 970,
876
+ 971,
877
+ 972,
878
+ 973,
879
+ 974,
880
+ 975,
881
+ 976,
882
+ 977,
883
+ 978,
884
+ 979,
885
+ 980,
886
+ 981,
887
+ 982,
888
+ 983,
889
+ 984,
890
+ 985,
891
+ 1001,
892
+ 1002,
893
+ 1003,
894
+ 1004,
895
+ 1005,
896
+ 1006,
897
+ 1007,
898
+ 1008,
899
+ 1009,
900
+ 1010,
901
+ 1011,
902
+ 1012,
903
+ 1013,
904
+ 1014,
905
+ 1015,
906
+ 1016,
907
+ 1017,
908
+ 1018,
909
+ 1019,
910
+ 1020,
911
+ 1021,
912
+ 1022,
913
+ 1023,
914
+ 1024,
915
+ 1025,
916
+ 1026,
917
+ 1027,
918
+ 1028,
919
+ 1029,
920
+ 1030,
921
+ 1031,
922
+ 1032,
923
+ 1033,
924
+ 1034,
925
+ 1035,
926
+ 1036,
927
+ 1037,
928
+ 1038,
929
+ 1039,
930
+ 1040,
931
+ 1041,
932
+ 1042,
933
+ 1043,
934
+ 1044,
935
+ 1045,
936
+ 1046,
937
+ 1047,
938
+ 1048,
939
+ 1049,
940
+ 1050,
941
+ 1051,
942
+ 1052,
943
+ 1053,
944
+ 1054,
945
+ 1055,
946
+ 1056,
947
+ 1057,
948
+ 1058,
949
+ 1059,
950
+ 1060,
951
+ 1061,
952
+ 1062,
953
+ 1063,
954
+ 1064,
955
+ 1065,
956
+ 1066,
957
+ 1067,
958
+ 1068,
959
+ 1069,
960
+ 1070,
961
+ 1071,
962
+ 1072,
963
+ 1073,
964
+ 1074,
965
+ 1075,
966
+ 1076,
967
+ 1077,
968
+ 1078,
969
+ 1079,
970
+ 1080,
971
+ 1081,
972
+ 1082,
973
+ 1083,
974
+ 1084,
975
+ 1085,
976
+ 1086,
977
+ 1087,
978
+ 1088,
979
+ 1089,
980
+ 1090,
981
+ 1091,
982
+ 1092,
983
+ 1093,
984
+ 1094,
985
+ 1095,
986
+ 1096,
987
+ 1097,
988
+ 1098,
989
+ 1099,
990
+ 1100,
991
+ 1101,
992
+ 1102,
993
+ 1103,
994
+ 1104,
995
+ 1105,
996
+ 1106,
997
+ 1107,
998
+ 1108,
999
+ 1109,
1000
+ 1110,
1001
+ 1111,
1002
+ 1112,
1003
+ 1113,
1004
+ 1114,
1005
+ 1115,
1006
+ 1116,
1007
+ 1117,
1008
+ 1118,
1009
+ 1119,
1010
+ 1120,
1011
+ 1121,
1012
+ 1122,
1013
+ 1123,
1014
+ 1124,
1015
+ 1125
1016
+ ]
1017
+ },
1018
+ "generation": {
1019
+ "model": "gemini-3.1-pro-preview",
1020
+ "provider": "chat-completions",
1021
+ "reasoning_effort": null,
1022
+ "stream": true,
1023
+ "timeout": 1800
1024
+ },
1025
+ "effective_model_call_cap": 6,
1026
+ "recorded_protocols": [
1027
+ {
1028
+ "core_sha256": {
1029
+ "protocol.json": "f7fe5ba2701d96e1f73873ce89cdc6699b4c87e9fd7f670075fb9cdd0c24207c",
1030
+ "inputs.manifest.json": "3b60ffa939b62e02ac963a73aa79998bf39158c6ba68f2a74f98c92acc7fd60a",
1031
+ "schedule.json": "919aee60ee5e8e65a350f7e27ffb6a8558bf2677f2717e8a969ffe3d8593d69e",
1032
+ "study.json": "1a808aeb0cf0153b049743d07598743835d504b1be7fdeb78c1365c263279d67",
1033
+ "environment.json": "04f141c80a13f3620f971c3217dccd28c2ee58f22a3d1fc5f7db2bfd1f4cc2ef"
1034
+ },
1035
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1036
+ "task_count": 1000,
1037
+ "task_input_blobs_rehashed": false,
1038
+ "source_and_dependencies_match": true,
1039
+ "limits": {
1040
+ "output_tokens": 32768,
1041
+ "wall_seconds": 1800,
1042
+ "model_calls": 6,
1043
+ "tool_calls": 4,
1044
+ "tool_timeout": 90
1045
+ },
1046
+ "generation": {
1047
+ "model": "gemini-3.1-pro-preview",
1048
+ "provider": "chat-completions",
1049
+ "reasoning_effort": null,
1050
+ "stream": true,
1051
+ "timeout": 1800
1052
+ },
1053
+ "dataset_provenance": {
1054
+ "repo_id": "CamoAiLab/InferenceNet",
1055
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1056
+ "task_directory": "Selected_1000",
1057
+ "task_list": "Selected_1000/1000_new.csv",
1058
+ "prior_results_imported": false,
1059
+ "engineering_pilot": false
1060
+ },
1061
+ "execution_identity": {
1062
+ "backend": "dsw-bwrap-v3",
1063
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1064
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1065
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1066
+ "engine_memory_bytes": 38654705664,
1067
+ "engine_cpus": 12,
1068
+ "cpu_slots": [
1069
+ [
1070
+ 4,
1071
+ 5
1072
+ ],
1073
+ [
1074
+ 6,
1075
+ 7
1076
+ ],
1077
+ [
1078
+ 8,
1079
+ 9
1080
+ ],
1081
+ [
1082
+ 10,
1083
+ 11
1084
+ ],
1085
+ [
1086
+ 12,
1087
+ 13
1088
+ ],
1089
+ [
1090
+ 14,
1091
+ 15
1092
+ ]
1093
+ ],
1094
+ "resources": {
1095
+ "cpus": 2.0,
1096
+ "memory": "6g",
1097
+ "tmpfs": "256m"
1098
+ },
1099
+ "memory_enforcement": "RLIMIT_AS",
1100
+ "single_payload_process": true,
1101
+ "network": "none",
1102
+ "uid": 10001,
1103
+ "capacity_profile": "independent-six-slots-20260916",
1104
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1105
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1106
+ },
1107
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1108
+ "scorer_sha256": {
1109
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1110
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1111
+ },
1112
+ "arms": [
1113
+ "deepagents"
1114
+ ],
1115
+ "origin": "child"
1116
+ },
1117
+ {
1118
+ "core_sha256": {
1119
+ "protocol.json": "29deb8f48389e2b8d34c352ae373bc6443ca56ea3f718979e9148957e136df90",
1120
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1121
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1122
+ "study.json": "f5cb2944526f5328274059fc2a175eebf7f75ecf742263063ba1264af708b0a7",
1123
+ "environment.json": "f0e702a8a418d683763e0a0d35acd4377d246e006d1c7533c9bf6cac38fd70bf"
1124
+ },
1125
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1126
+ "task_count": 1000,
1127
+ "task_input_blobs_rehashed": false,
1128
+ "source_and_dependencies_match": true,
1129
+ "limits": {
1130
+ "output_tokens": 32768,
1131
+ "wall_seconds": 1800,
1132
+ "model_calls": 6,
1133
+ "tool_calls": 4,
1134
+ "tool_timeout": 90
1135
+ },
1136
+ "generation": {
1137
+ "model": "gemini-3.1-pro-preview",
1138
+ "provider": "chat-completions",
1139
+ "reasoning_effort": null,
1140
+ "stream": true,
1141
+ "timeout": 1800
1142
+ },
1143
+ "dataset_provenance": {
1144
+ "repo_id": "CamoAiLab/InferenceNet",
1145
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1146
+ "task_directory": "Selected_1000",
1147
+ "task_list": "Selected_1000/1000_new.csv",
1148
+ "prior_results_imported": false,
1149
+ "engineering_pilot": false
1150
+ },
1151
+ "execution_identity": {
1152
+ "backend": "dsw-bwrap-v3",
1153
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1154
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1155
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1156
+ "engine_memory_bytes": 12884901888,
1157
+ "engine_cpus": 4,
1158
+ "cpu_slots": [
1159
+ [
1160
+ 16,
1161
+ 17
1162
+ ],
1163
+ [
1164
+ 18,
1165
+ 19
1166
+ ]
1167
+ ],
1168
+ "resources": {
1169
+ "memory": "6g",
1170
+ "cpus": 2.0,
1171
+ "tmpfs": "256m"
1172
+ },
1173
+ "memory_enforcement": "RLIMIT_AS",
1174
+ "single_payload_process": true,
1175
+ "network": "none",
1176
+ "uid": 10001
1177
+ },
1178
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1179
+ "scorer_sha256": {
1180
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1181
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1182
+ },
1183
+ "arms": [
1184
+ "deepagents"
1185
+ ],
1186
+ "origin": "parent"
1187
+ }
1188
+ ],
1189
+ "selected_origin_counts": {
1190
+ "child": 1000
1191
+ },
1192
+ "source_commits": [
1193
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1194
+ ],
1195
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1196
+ "source_export_manifest_sha256": "1173b8ae18088d6cb9d62a0a747751f34ca5ebb5d91c06813ad0f40d173ac40c",
1197
+ "official_parity_verified": false,
1198
+ "notes": [
1199
+ "Selected parent and child records are disjoint; original protocol hashes and source versions retained.",
1200
+ "Baseline has one model call; DeepAgents has up to six model calls and four tool calls.",
1201
+ "Output-token budget 32768; task wall budget 1800 seconds; execution timeout 90 seconds.",
1202
+ "Final execution is separate from trial tools; unknown results are never dropped or replaced.",
1203
+ "Predictions are unavailable in this score export, even for tasks with successful scored predictions."
1204
+ ]
1205
+ }
results/gemini-3.1-pro-preview/deepagents/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/gemini-3.1-pro-preview/deepagents/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/gemini-3.1-pro-preview/deepagents/summary.json ADDED
@@ -0,0 +1,190 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
5
+ "model": "gemini-3.1-pro-preview",
6
+ "arm": "deepagents",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": true,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 850,
37
+ "denominator": 1000,
38
+ "rate": 0.85,
39
+ "score": 85.0,
40
+ "unknown_count": 0,
41
+ "failure_count": 150,
42
+ "assessable": 1000,
43
+ "coverage_percent": 100.0
44
+ },
45
+ "partial_replication": {
46
+ "count": 632,
47
+ "denominator": 1000,
48
+ "rate": 0.632,
49
+ "score": 63.2,
50
+ "unknown_count": 153,
51
+ "failure_count": 215,
52
+ "assessable": 847,
53
+ "coverage_percent": 84.7
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 811,
57
+ "denominator": 1000,
58
+ "rate": 0.811,
59
+ "score": 81.1,
60
+ "unknown_count": 153,
61
+ "failure_count": 36,
62
+ "assessable": 847,
63
+ "coverage_percent": 84.7
64
+ },
65
+ "significance_level": {
66
+ "count": 718,
67
+ "denominator": 1000,
68
+ "rate": 0.718,
69
+ "score": 71.8,
70
+ "unknown_count": 154,
71
+ "failure_count": 128,
72
+ "assessable": 846,
73
+ "coverage_percent": 84.6
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "gemini-3.1-pro-preview deepagents",
78
+ "Compilation Success": 85.0,
79
+ "Partial Replication": 63.2,
80
+ "Correct Coefficient Direction": 81.1,
81
+ "Significant Level Correctness": 71.8
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 1000,
94
+ "task_unknown_count": 0,
95
+ "valid_prediction_count": 848,
96
+ "task_counts": {
97
+ "succeeded": 848,
98
+ "failed": 152,
99
+ "unknown": 0,
100
+ "unsealed": 0,
101
+ "not_started": 0,
102
+ "invalid_evidence": 0,
103
+ "completed": 1000,
104
+ "planned": 1000
105
+ }
106
+ },
107
+ "legacy_local_paper": {
108
+ "profile": "local-paper-v1",
109
+ "definitions": {
110
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
111
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
112
+ "coefficient_only": "coefficient relative error <= .05",
113
+ "direction": "matching strictly positive or strictly negative coefficient signs",
114
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
115
+ },
116
+ "metrics": {
117
+ "perfect": {
118
+ "count": 391,
119
+ "denominator": 1000,
120
+ "rate": 0.391,
121
+ "score": 39.1,
122
+ "unknown_count": 155,
123
+ "failure_count": 454,
124
+ "assessable": 845,
125
+ "coverage_percent": 84.5
126
+ },
127
+ "partial": {
128
+ "count": 517,
129
+ "denominator": 1000,
130
+ "rate": 0.517,
131
+ "score": 51.7,
132
+ "unknown_count": 155,
133
+ "failure_count": 328,
134
+ "assessable": 845,
135
+ "coverage_percent": 84.5
136
+ },
137
+ "coefficient_only": {
138
+ "count": 632,
139
+ "denominator": 1000,
140
+ "rate": 0.632,
141
+ "score": 63.2,
142
+ "unknown_count": 155,
143
+ "failure_count": 213,
144
+ "assessable": 845,
145
+ "coverage_percent": 84.5
146
+ },
147
+ "direction": {
148
+ "count": 809,
149
+ "denominator": 1000,
150
+ "rate": 0.809,
151
+ "score": 80.9,
152
+ "unknown_count": 155,
153
+ "failure_count": 36,
154
+ "assessable": 845,
155
+ "coverage_percent": 84.5
156
+ },
157
+ "significance": {
158
+ "count": 697,
159
+ "denominator": 1000,
160
+ "rate": 0.697,
161
+ "score": 69.7,
162
+ "unknown_count": 155,
163
+ "failure_count": 148,
164
+ "assessable": 845,
165
+ "coverage_percent": 84.5
166
+ }
167
+ }
168
+ },
169
+ "verification_scope": "Offline reaggregation of accepted per-task flags; raw predictions are not available in this export; official parity not verified.",
170
+ "usage": {
171
+ "sealed_model_calls": 5702,
172
+ "unknown_usage_calls": 0,
173
+ "unsealed_episodes_excluded": 0,
174
+ "invalid_evidence_episodes_excluded": 0,
175
+ "known_token_subtotal": {
176
+ "input_tokens": 26515986,
177
+ "output_tokens": 6407382,
178
+ "total_tokens": 32923368
179
+ },
180
+ "unknown_by_field": {
181
+ "input_tokens": 0,
182
+ "output_tokens": 0,
183
+ "total_tokens": 0
184
+ },
185
+ "total_tokens_complete": true,
186
+ "usd": null,
187
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
188
+ "summed_episode_seconds": 281168.88599999994
189
+ }
190
+ }
results/gemini-3.1-pro-preview/diagnostics.json ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage_counts": {
3
+ "baseline": [
4
+ {
5
+ "generation": "succeeded",
6
+ "execution": "failed",
7
+ "scoring": "pending",
8
+ "count": 460
9
+ },
10
+ {
11
+ "generation": "succeeded",
12
+ "execution": "succeeded",
13
+ "scoring": "scored",
14
+ "count": 540
15
+ }
16
+ ],
17
+ "deepagents": [
18
+ {
19
+ "generation": "failed",
20
+ "execution": "pending",
21
+ "scoring": "pending",
22
+ "count": 41
23
+ },
24
+ {
25
+ "generation": "succeeded",
26
+ "execution": "failed",
27
+ "scoring": "pending",
28
+ "count": 111
29
+ },
30
+ {
31
+ "generation": "succeeded",
32
+ "execution": "succeeded",
33
+ "scoring": "scored",
34
+ "count": 848
35
+ }
36
+ ]
37
+ },
38
+ "task_unknown_ids": {
39
+ "baseline": [],
40
+ "deepagents": []
41
+ },
42
+ "generation_failed_model_call_distribution": {
43
+ "baseline": {},
44
+ "deepagents": {
45
+ "6": 41
46
+ }
47
+ },
48
+ "raw_error_reasons_verified": false,
49
+ "limitation": "Stage labels and counts are from the accepted score snapshot. Raw errors and full traces were not freshly inspected; no new experiment was run. Task unknowns and metric nulls remain separate and are retained in fixed denominators."
50
+ }
results/gemini-3.1-pro-preview/export-verification.json ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "passed_offline",
3
+ "snapshot_hash_matches": true,
4
+ "task_records_match_accepted_package": true,
5
+ "metrics_recomputed": 18,
6
+ "paired_rows": 1000,
7
+ "per_arm": {
8
+ "baseline": {
9
+ "records": 1000,
10
+ "unique_task_ids": 1000,
11
+ "matches_accepted_export": true,
12
+ "counts": {
13
+ "succeeded": 540,
14
+ "failed": 460,
15
+ "unknown": 0,
16
+ "unsealed": 0,
17
+ "not_started": 0,
18
+ "invalid_evidence": 0,
19
+ "completed": 1000,
20
+ "planned": 1000
21
+ },
22
+ "origin_counts": {
23
+ "child": 1000
24
+ }
25
+ },
26
+ "deepagents": {
27
+ "records": 1000,
28
+ "unique_task_ids": 1000,
29
+ "matches_accepted_export": true,
30
+ "counts": {
31
+ "succeeded": 848,
32
+ "failed": 152,
33
+ "unknown": 0,
34
+ "unsealed": 0,
35
+ "not_started": 0,
36
+ "invalid_evidence": 0,
37
+ "completed": 1000,
38
+ "planned": 1000
39
+ },
40
+ "origin_counts": {
41
+ "child": 1000
42
+ }
43
+ }
44
+ },
45
+ "raw_error_root_cause_verified": false,
46
+ "official_parity_verified": false
47
+ }
results/gemini-3.1-pro-preview/paired_comparison.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/gemini-3.1-pro-preview/provenance.json ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "gemini-3.1-pro-preview",
3
+ "exported_at": "2026-09-16T15:11:55.086784+00:00",
4
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
5
+ "accepted_results_sha256": "dd72f352be1895b11914dd2aea1fdceb21c453932877c588ecaf81ec09861105",
6
+ "accepted_task_records_sha256": "32a797b4da8dfe90d44b37d7715fe726a29c5b4f93a1b24d7d5ffbffb81e9749",
7
+ "source_commits": [
8
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
9
+ ],
10
+ "dataset": {
11
+ "repo_id": "CamoAiLab/InferenceNet",
12
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
13
+ "task_directory": "Selected_1000",
14
+ "task_list": "Selected_1000/1000_new.csv",
15
+ "prior_results_imported": false,
16
+ "engineering_pilot": false
17
+ },
18
+ "selected_task_count": 1000,
19
+ "record_count": 2000,
20
+ "official_parity_verified": false,
21
+ "new_model_calls": 0,
22
+ "new_benchmark_executions": 0,
23
+ "historical_native_audit": {
24
+ "mode": "historical_final_cache",
25
+ "snapshot_sha256": "7de17ac129bbb07ac7d1fa08adb92642eeb7cfadab4d4a68946661308d8adde1",
26
+ "journal_sha256": "d35e92010e0f85de995ef587f3d64616260a70600be6749644c33fb287d1a32b",
27
+ "acceptance_sha256": "8ae98f599a7874a5c5f5c30c2cd735ad074a2790efe0c726c0e06c2686b27b94",
28
+ "collector_sha256": "7963b4fd6ef312307d5471d031e6851429f139b03a3e6d7ffc7ca9bd44d98c89",
29
+ "renderer_sha256": "b2cb2f0ae0523431e3be13ac23a6730c8e6a511e7c645921dea79a1cacffe2fd",
30
+ "started_at": "2026-09-16T01:01:30.976377+00:00",
31
+ "finished_at": "2026-09-16T02:24:03.769359+00:00",
32
+ "model_sha256": "20873b7b10df35162074325c38f9acced79e05c8d7bf0ebe52a1c0b45d215afa",
33
+ "source_checkpoint": {
34
+ "record_sha256": "b28a7da705fa8d659290620fc8dbeea2fd581638de018ffcbd90297cf608c58d",
35
+ "record_bytes": 4630907
36
+ },
37
+ "current_identity_check": {
38
+ "kind": "current_source_core_result_hashes_parent_absence",
39
+ "started_at": "2026-09-16T07:21:44.989133+00:00",
40
+ "finished_at": "2026-09-16T07:22:31.844377+00:00",
41
+ "result_hashes_checked": 2000,
42
+ "absent_parent_slots_checked": 2000,
43
+ "native_artifacts_reaudited": false,
44
+ "input_blobs_rehashed": false
45
+ },
46
+ "native_audit_scope": "Historical native artifact/execution/scoring verification; input blobs were not rehashed by that collector."
47
+ },
48
+ "verification_scope": "Offline hash, task identity, accepted-export equality and reaggregation of sealed per-task flags; not fresh prediction scoring or native artifact audit.",
49
+ "related_results": {
50
+ "gpt_5_6_sol": "https://huggingface.co/datasets/YICHEN013/InferenceNet-GPT-5.6-Sol-results"
51
+ },
52
+ "excluded": [
53
+ "benchmark inputs",
54
+ "reference answers",
55
+ "raw predictions unavailable in score snapshot",
56
+ "full prompts and model/API traces",
57
+ "credentials",
58
+ "internal endpoints",
59
+ "machine filesystem paths"
60
+ ]
61
+ }
results/index.json CHANGED
@@ -27,6 +27,31 @@
27
  "arm": "deepagents",
28
  "rows": 1000,
29
  "path": "results/claude-opus-4-8/deepagents"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
30
  }
31
- ]
 
32
  }
 
27
  "arm": "deepagents",
28
  "rows": 1000,
29
  "path": "results/claude-opus-4-8/deepagents"
30
+ },
31
+ {
32
+ "model": "kimi-k3",
33
+ "arm": "baseline",
34
+ "rows": 1000,
35
+ "path": "results/kimi-k3/baseline"
36
+ },
37
+ {
38
+ "model": "kimi-k3",
39
+ "arm": "deepagents",
40
+ "rows": 1000,
41
+ "path": "results/kimi-k3/deepagents"
42
+ },
43
+ {
44
+ "model": "gemini-3.1-pro-preview",
45
+ "arm": "baseline",
46
+ "rows": 1000,
47
+ "path": "results/gemini-3.1-pro-preview/baseline"
48
+ },
49
+ {
50
+ "model": "gemini-3.1-pro-preview",
51
+ "arm": "deepagents",
52
+ "rows": 1000,
53
+ "path": "results/gemini-3.1-pro-preview/deepagents"
54
  }
55
+ ],
56
+ "updated_at_utc": "2026-09-16T15:28:00.953066+00:00"
57
  }
results/kimi-k3/baseline/README.md ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Kimi K3 / baseline
2
+
3
+ 1,000 个已封存槽位;成功 520,确定失败 475,任务 unknown 5。固定分母为 1,000;完整复现率 23.0%。
4
+
5
+ - `results.jsonl` / `results.csv`:相同的逐题状态、指标、调用用量和结果哈希。
6
+ - `summary.json`:本地 HF 四维与 local-paper-v1 五维分别保存,包含各自 unknown 和可判定覆盖率。
7
+ - `config.json`:固定任务清单、模型接口、预算、父/子接续来源及哈希。
8
+ - 上一级的 `paired_comparison.csv`、`provenance.json`、`diagnostics.json` 与 `export-verification.json` 保留双组配对与原导出核验边界。
9
+
10
+ `resolved=true` 表示槽位已封存;`definite_outcome` 才表示结果确定。未知和失败都保留在分母中。阶段 pending 可能表示前序失败后未执行到该阶段,不代表还在运行。
11
+
12
+ **此批次只有评分/状态导出,原始预测三元组未包含在源快照中。** 因此 prediction 字段保留 null,并用 `prediction_export_status=not_available_in_accepted_score_snapshot` 区分“导出不可用”与“模型没有有效预测”;不能据此把所有任务判为无预测。`summary.json` 的 valid_prediction_count 来自已核验的执行/评分状态。
13
+
14
+ 四维指标为本地实现,尚未验证与官方 scorer 完全一致。本次为历史 amended comparison 的归档,不自动更新正式榜单,没有重新调用模型或执行程序。用量字段是已知小计,须结合各字段 unknown_usage_calls 阅读。
15
+
16
+ JSON null / CSV 空单元格表示未知或不可用。原始数据、隐藏答案、完整提示词/API 轨迹、凭据和机器路径未包含在归档中。
results/kimi-k3/baseline/config.json ADDED
@@ -0,0 +1,1276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "kimi-k3",
4
+ "arm": "baseline",
5
+ "harness": "none",
6
+ "archived_at_utc": "2026-09-16T15:28:00.953066+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "expected_tasks": 1000,
14
+ "task_ids": [
15
+ 1,
16
+ 2,
17
+ 3,
18
+ 4,
19
+ 5,
20
+ 6,
21
+ 7,
22
+ 8,
23
+ 9,
24
+ 10,
25
+ 11,
26
+ 12,
27
+ 13,
28
+ 14,
29
+ 15,
30
+ 16,
31
+ 17,
32
+ 18,
33
+ 19,
34
+ 20,
35
+ 21,
36
+ 22,
37
+ 23,
38
+ 24,
39
+ 25,
40
+ 26,
41
+ 27,
42
+ 28,
43
+ 29,
44
+ 30,
45
+ 31,
46
+ 32,
47
+ 33,
48
+ 34,
49
+ 35,
50
+ 36,
51
+ 37,
52
+ 38,
53
+ 39,
54
+ 40,
55
+ 41,
56
+ 42,
57
+ 43,
58
+ 44,
59
+ 45,
60
+ 46,
61
+ 47,
62
+ 48,
63
+ 49,
64
+ 50,
65
+ 51,
66
+ 52,
67
+ 53,
68
+ 54,
69
+ 55,
70
+ 56,
71
+ 57,
72
+ 58,
73
+ 59,
74
+ 60,
75
+ 61,
76
+ 62,
77
+ 63,
78
+ 64,
79
+ 65,
80
+ 66,
81
+ 67,
82
+ 68,
83
+ 69,
84
+ 70,
85
+ 71,
86
+ 72,
87
+ 73,
88
+ 74,
89
+ 75,
90
+ 76,
91
+ 77,
92
+ 78,
93
+ 79,
94
+ 80,
95
+ 81,
96
+ 82,
97
+ 83,
98
+ 84,
99
+ 85,
100
+ 86,
101
+ 87,
102
+ 88,
103
+ 89,
104
+ 90,
105
+ 91,
106
+ 92,
107
+ 93,
108
+ 94,
109
+ 95,
110
+ 96,
111
+ 97,
112
+ 98,
113
+ 100,
114
+ 101,
115
+ 102,
116
+ 103,
117
+ 104,
118
+ 105,
119
+ 106,
120
+ 107,
121
+ 108,
122
+ 109,
123
+ 110,
124
+ 111,
125
+ 112,
126
+ 113,
127
+ 114,
128
+ 115,
129
+ 116,
130
+ 117,
131
+ 118,
132
+ 119,
133
+ 120,
134
+ 121,
135
+ 122,
136
+ 123,
137
+ 124,
138
+ 125,
139
+ 126,
140
+ 127,
141
+ 128,
142
+ 129,
143
+ 130,
144
+ 131,
145
+ 132,
146
+ 133,
147
+ 134,
148
+ 135,
149
+ 136,
150
+ 137,
151
+ 138,
152
+ 139,
153
+ 140,
154
+ 141,
155
+ 142,
156
+ 143,
157
+ 144,
158
+ 145,
159
+ 146,
160
+ 147,
161
+ 148,
162
+ 149,
163
+ 150,
164
+ 151,
165
+ 152,
166
+ 153,
167
+ 154,
168
+ 155,
169
+ 156,
170
+ 157,
171
+ 158,
172
+ 159,
173
+ 160,
174
+ 161,
175
+ 162,
176
+ 163,
177
+ 164,
178
+ 165,
179
+ 166,
180
+ 167,
181
+ 168,
182
+ 169,
183
+ 170,
184
+ 171,
185
+ 172,
186
+ 173,
187
+ 174,
188
+ 175,
189
+ 176,
190
+ 177,
191
+ 178,
192
+ 179,
193
+ 180,
194
+ 181,
195
+ 182,
196
+ 183,
197
+ 184,
198
+ 185,
199
+ 186,
200
+ 187,
201
+ 188,
202
+ 189,
203
+ 190,
204
+ 191,
205
+ 192,
206
+ 193,
207
+ 194,
208
+ 195,
209
+ 196,
210
+ 197,
211
+ 198,
212
+ 199,
213
+ 200,
214
+ 201,
215
+ 202,
216
+ 203,
217
+ 204,
218
+ 205,
219
+ 206,
220
+ 207,
221
+ 208,
222
+ 209,
223
+ 210,
224
+ 211,
225
+ 212,
226
+ 213,
227
+ 214,
228
+ 215,
229
+ 216,
230
+ 217,
231
+ 218,
232
+ 219,
233
+ 220,
234
+ 221,
235
+ 222,
236
+ 223,
237
+ 224,
238
+ 225,
239
+ 226,
240
+ 227,
241
+ 228,
242
+ 229,
243
+ 230,
244
+ 231,
245
+ 232,
246
+ 233,
247
+ 234,
248
+ 235,
249
+ 236,
250
+ 237,
251
+ 238,
252
+ 239,
253
+ 240,
254
+ 241,
255
+ 242,
256
+ 248,
257
+ 249,
258
+ 250,
259
+ 251,
260
+ 252,
261
+ 253,
262
+ 254,
263
+ 255,
264
+ 256,
265
+ 257,
266
+ 258,
267
+ 259,
268
+ 261,
269
+ 262,
270
+ 263,
271
+ 264,
272
+ 265,
273
+ 266,
274
+ 290,
275
+ 291,
276
+ 292,
277
+ 293,
278
+ 294,
279
+ 295,
280
+ 302,
281
+ 303,
282
+ 304,
283
+ 305,
284
+ 306,
285
+ 307,
286
+ 308,
287
+ 309,
288
+ 310,
289
+ 311,
290
+ 312,
291
+ 313,
292
+ 314,
293
+ 315,
294
+ 316,
295
+ 317,
296
+ 318,
297
+ 319,
298
+ 320,
299
+ 321,
300
+ 322,
301
+ 323,
302
+ 324,
303
+ 325,
304
+ 326,
305
+ 327,
306
+ 328,
307
+ 329,
308
+ 330,
309
+ 331,
310
+ 332,
311
+ 333,
312
+ 334,
313
+ 335,
314
+ 336,
315
+ 337,
316
+ 338,
317
+ 339,
318
+ 340,
319
+ 341,
320
+ 342,
321
+ 343,
322
+ 344,
323
+ 345,
324
+ 346,
325
+ 347,
326
+ 348,
327
+ 349,
328
+ 350,
329
+ 351,
330
+ 352,
331
+ 353,
332
+ 354,
333
+ 355,
334
+ 356,
335
+ 357,
336
+ 358,
337
+ 359,
338
+ 360,
339
+ 361,
340
+ 362,
341
+ 363,
342
+ 364,
343
+ 365,
344
+ 366,
345
+ 367,
346
+ 368,
347
+ 369,
348
+ 370,
349
+ 371,
350
+ 372,
351
+ 373,
352
+ 374,
353
+ 375,
354
+ 376,
355
+ 377,
356
+ 378,
357
+ 379,
358
+ 380,
359
+ 381,
360
+ 382,
361
+ 383,
362
+ 384,
363
+ 385,
364
+ 386,
365
+ 387,
366
+ 388,
367
+ 389,
368
+ 390,
369
+ 391,
370
+ 392,
371
+ 393,
372
+ 394,
373
+ 395,
374
+ 396,
375
+ 397,
376
+ 398,
377
+ 399,
378
+ 400,
379
+ 401,
380
+ 402,
381
+ 403,
382
+ 404,
383
+ 405,
384
+ 406,
385
+ 407,
386
+ 408,
387
+ 409,
388
+ 410,
389
+ 411,
390
+ 412,
391
+ 413,
392
+ 414,
393
+ 415,
394
+ 416,
395
+ 417,
396
+ 418,
397
+ 419,
398
+ 420,
399
+ 421,
400
+ 422,
401
+ 423,
402
+ 424,
403
+ 425,
404
+ 426,
405
+ 427,
406
+ 428,
407
+ 429,
408
+ 430,
409
+ 431,
410
+ 432,
411
+ 433,
412
+ 434,
413
+ 435,
414
+ 436,
415
+ 437,
416
+ 438,
417
+ 439,
418
+ 440,
419
+ 441,
420
+ 442,
421
+ 443,
422
+ 444,
423
+ 445,
424
+ 446,
425
+ 447,
426
+ 448,
427
+ 449,
428
+ 450,
429
+ 451,
430
+ 452,
431
+ 453,
432
+ 454,
433
+ 455,
434
+ 456,
435
+ 457,
436
+ 458,
437
+ 459,
438
+ 460,
439
+ 461,
440
+ 462,
441
+ 463,
442
+ 464,
443
+ 465,
444
+ 466,
445
+ 467,
446
+ 468,
447
+ 469,
448
+ 470,
449
+ 471,
450
+ 472,
451
+ 473,
452
+ 474,
453
+ 475,
454
+ 476,
455
+ 477,
456
+ 478,
457
+ 479,
458
+ 480,
459
+ 481,
460
+ 482,
461
+ 483,
462
+ 484,
463
+ 485,
464
+ 486,
465
+ 487,
466
+ 488,
467
+ 489,
468
+ 490,
469
+ 491,
470
+ 492,
471
+ 493,
472
+ 494,
473
+ 495,
474
+ 496,
475
+ 497,
476
+ 498,
477
+ 499,
478
+ 500,
479
+ 501,
480
+ 502,
481
+ 503,
482
+ 504,
483
+ 505,
484
+ 506,
485
+ 507,
486
+ 508,
487
+ 509,
488
+ 510,
489
+ 511,
490
+ 512,
491
+ 513,
492
+ 514,
493
+ 515,
494
+ 516,
495
+ 517,
496
+ 518,
497
+ 519,
498
+ 520,
499
+ 521,
500
+ 522,
501
+ 523,
502
+ 524,
503
+ 525,
504
+ 526,
505
+ 527,
506
+ 528,
507
+ 529,
508
+ 530,
509
+ 531,
510
+ 532,
511
+ 533,
512
+ 534,
513
+ 535,
514
+ 536,
515
+ 537,
516
+ 538,
517
+ 539,
518
+ 540,
519
+ 541,
520
+ 542,
521
+ 543,
522
+ 544,
523
+ 545,
524
+ 546,
525
+ 547,
526
+ 548,
527
+ 549,
528
+ 550,
529
+ 551,
530
+ 552,
531
+ 553,
532
+ 554,
533
+ 555,
534
+ 556,
535
+ 557,
536
+ 558,
537
+ 559,
538
+ 560,
539
+ 561,
540
+ 562,
541
+ 563,
542
+ 565,
543
+ 566,
544
+ 567,
545
+ 568,
546
+ 570,
547
+ 571,
548
+ 572,
549
+ 573,
550
+ 574,
551
+ 575,
552
+ 576,
553
+ 577,
554
+ 578,
555
+ 579,
556
+ 580,
557
+ 581,
558
+ 582,
559
+ 583,
560
+ 584,
561
+ 585,
562
+ 586,
563
+ 587,
564
+ 588,
565
+ 589,
566
+ 590,
567
+ 591,
568
+ 592,
569
+ 593,
570
+ 594,
571
+ 595,
572
+ 596,
573
+ 597,
574
+ 598,
575
+ 601,
576
+ 602,
577
+ 603,
578
+ 604,
579
+ 605,
580
+ 606,
581
+ 607,
582
+ 608,
583
+ 609,
584
+ 610,
585
+ 611,
586
+ 612,
587
+ 613,
588
+ 614,
589
+ 615,
590
+ 616,
591
+ 617,
592
+ 618,
593
+ 619,
594
+ 620,
595
+ 621,
596
+ 622,
597
+ 623,
598
+ 624,
599
+ 664,
600
+ 665,
601
+ 666,
602
+ 667,
603
+ 668,
604
+ 669,
605
+ 670,
606
+ 671,
607
+ 672,
608
+ 673,
609
+ 674,
610
+ 675,
611
+ 676,
612
+ 677,
613
+ 678,
614
+ 679,
615
+ 680,
616
+ 681,
617
+ 682,
618
+ 683,
619
+ 684,
620
+ 685,
621
+ 686,
622
+ 687,
623
+ 688,
624
+ 689,
625
+ 690,
626
+ 691,
627
+ 692,
628
+ 693,
629
+ 694,
630
+ 695,
631
+ 696,
632
+ 697,
633
+ 698,
634
+ 699,
635
+ 700,
636
+ 701,
637
+ 702,
638
+ 703,
639
+ 704,
640
+ 705,
641
+ 706,
642
+ 707,
643
+ 708,
644
+ 709,
645
+ 710,
646
+ 711,
647
+ 712,
648
+ 713,
649
+ 714,
650
+ 715,
651
+ 716,
652
+ 717,
653
+ 718,
654
+ 719,
655
+ 720,
656
+ 721,
657
+ 722,
658
+ 723,
659
+ 724,
660
+ 725,
661
+ 726,
662
+ 727,
663
+ 728,
664
+ 729,
665
+ 730,
666
+ 731,
667
+ 732,
668
+ 733,
669
+ 734,
670
+ 735,
671
+ 736,
672
+ 737,
673
+ 739,
674
+ 740,
675
+ 741,
676
+ 742,
677
+ 743,
678
+ 744,
679
+ 745,
680
+ 746,
681
+ 747,
682
+ 748,
683
+ 749,
684
+ 750,
685
+ 751,
686
+ 752,
687
+ 753,
688
+ 754,
689
+ 755,
690
+ 756,
691
+ 757,
692
+ 758,
693
+ 759,
694
+ 760,
695
+ 761,
696
+ 762,
697
+ 763,
698
+ 764,
699
+ 765,
700
+ 766,
701
+ 767,
702
+ 768,
703
+ 772,
704
+ 773,
705
+ 774,
706
+ 775,
707
+ 776,
708
+ 777,
709
+ 778,
710
+ 779,
711
+ 780,
712
+ 781,
713
+ 782,
714
+ 783,
715
+ 784,
716
+ 785,
717
+ 786,
718
+ 787,
719
+ 788,
720
+ 789,
721
+ 790,
722
+ 791,
723
+ 793,
724
+ 794,
725
+ 795,
726
+ 796,
727
+ 797,
728
+ 798,
729
+ 799,
730
+ 800,
731
+ 801,
732
+ 802,
733
+ 803,
734
+ 804,
735
+ 805,
736
+ 806,
737
+ 807,
738
+ 808,
739
+ 825,
740
+ 826,
741
+ 827,
742
+ 828,
743
+ 829,
744
+ 830,
745
+ 831,
746
+ 832,
747
+ 833,
748
+ 834,
749
+ 835,
750
+ 836,
751
+ 837,
752
+ 838,
753
+ 839,
754
+ 840,
755
+ 841,
756
+ 842,
757
+ 843,
758
+ 844,
759
+ 845,
760
+ 846,
761
+ 847,
762
+ 848,
763
+ 849,
764
+ 850,
765
+ 851,
766
+ 852,
767
+ 853,
768
+ 854,
769
+ 855,
770
+ 856,
771
+ 857,
772
+ 858,
773
+ 859,
774
+ 860,
775
+ 861,
776
+ 862,
777
+ 863,
778
+ 864,
779
+ 865,
780
+ 866,
781
+ 867,
782
+ 868,
783
+ 869,
784
+ 870,
785
+ 871,
786
+ 872,
787
+ 873,
788
+ 874,
789
+ 875,
790
+ 876,
791
+ 877,
792
+ 878,
793
+ 879,
794
+ 880,
795
+ 881,
796
+ 882,
797
+ 883,
798
+ 884,
799
+ 885,
800
+ 886,
801
+ 887,
802
+ 888,
803
+ 889,
804
+ 890,
805
+ 891,
806
+ 892,
807
+ 893,
808
+ 894,
809
+ 895,
810
+ 896,
811
+ 897,
812
+ 898,
813
+ 899,
814
+ 900,
815
+ 901,
816
+ 902,
817
+ 903,
818
+ 906,
819
+ 907,
820
+ 913,
821
+ 914,
822
+ 915,
823
+ 917,
824
+ 920,
825
+ 921,
826
+ 922,
827
+ 923,
828
+ 924,
829
+ 925,
830
+ 926,
831
+ 927,
832
+ 928,
833
+ 929,
834
+ 930,
835
+ 931,
836
+ 932,
837
+ 933,
838
+ 934,
839
+ 935,
840
+ 936,
841
+ 937,
842
+ 938,
843
+ 939,
844
+ 940,
845
+ 941,
846
+ 942,
847
+ 943,
848
+ 944,
849
+ 945,
850
+ 946,
851
+ 947,
852
+ 948,
853
+ 949,
854
+ 950,
855
+ 951,
856
+ 952,
857
+ 953,
858
+ 954,
859
+ 955,
860
+ 956,
861
+ 957,
862
+ 958,
863
+ 959,
864
+ 960,
865
+ 961,
866
+ 962,
867
+ 963,
868
+ 964,
869
+ 965,
870
+ 966,
871
+ 967,
872
+ 968,
873
+ 969,
874
+ 970,
875
+ 971,
876
+ 972,
877
+ 973,
878
+ 974,
879
+ 975,
880
+ 976,
881
+ 977,
882
+ 978,
883
+ 979,
884
+ 980,
885
+ 981,
886
+ 982,
887
+ 983,
888
+ 984,
889
+ 985,
890
+ 1001,
891
+ 1002,
892
+ 1003,
893
+ 1004,
894
+ 1005,
895
+ 1006,
896
+ 1007,
897
+ 1008,
898
+ 1009,
899
+ 1010,
900
+ 1011,
901
+ 1012,
902
+ 1013,
903
+ 1014,
904
+ 1015,
905
+ 1016,
906
+ 1017,
907
+ 1018,
908
+ 1019,
909
+ 1020,
910
+ 1021,
911
+ 1022,
912
+ 1023,
913
+ 1024,
914
+ 1025,
915
+ 1026,
916
+ 1027,
917
+ 1028,
918
+ 1029,
919
+ 1030,
920
+ 1031,
921
+ 1032,
922
+ 1033,
923
+ 1034,
924
+ 1035,
925
+ 1036,
926
+ 1037,
927
+ 1038,
928
+ 1039,
929
+ 1040,
930
+ 1041,
931
+ 1042,
932
+ 1043,
933
+ 1044,
934
+ 1045,
935
+ 1046,
936
+ 1047,
937
+ 1048,
938
+ 1049,
939
+ 1050,
940
+ 1051,
941
+ 1052,
942
+ 1053,
943
+ 1054,
944
+ 1055,
945
+ 1056,
946
+ 1057,
947
+ 1058,
948
+ 1059,
949
+ 1060,
950
+ 1061,
951
+ 1062,
952
+ 1063,
953
+ 1064,
954
+ 1065,
955
+ 1066,
956
+ 1067,
957
+ 1068,
958
+ 1069,
959
+ 1070,
960
+ 1071,
961
+ 1072,
962
+ 1073,
963
+ 1074,
964
+ 1075,
965
+ 1076,
966
+ 1077,
967
+ 1078,
968
+ 1079,
969
+ 1080,
970
+ 1081,
971
+ 1082,
972
+ 1083,
973
+ 1084,
974
+ 1085,
975
+ 1086,
976
+ 1087,
977
+ 1088,
978
+ 1089,
979
+ 1090,
980
+ 1091,
981
+ 1092,
982
+ 1093,
983
+ 1094,
984
+ 1095,
985
+ 1096,
986
+ 1097,
987
+ 1098,
988
+ 1099,
989
+ 1100,
990
+ 1101,
991
+ 1102,
992
+ 1103,
993
+ 1104,
994
+ 1105,
995
+ 1106,
996
+ 1107,
997
+ 1108,
998
+ 1109,
999
+ 1110,
1000
+ 1111,
1001
+ 1112,
1002
+ 1113,
1003
+ 1114,
1004
+ 1115,
1005
+ 1116,
1006
+ 1117,
1007
+ 1118,
1008
+ 1119,
1009
+ 1120,
1010
+ 1121,
1011
+ 1122,
1012
+ 1123,
1013
+ 1124,
1014
+ 1125
1015
+ ]
1016
+ },
1017
+ "generation": {
1018
+ "model": "kimi-k3",
1019
+ "provider": "chat-completions",
1020
+ "reasoning_effort": "max",
1021
+ "stream": true,
1022
+ "timeout": 1800
1023
+ },
1024
+ "effective_model_call_cap": 1,
1025
+ "recorded_protocols": [
1026
+ {
1027
+ "core_sha256": {
1028
+ "protocol.json": "ae255e4239169c9ce077042813e44767190a955295a60bb080de8c358249d284",
1029
+ "inputs.manifest.json": "79b6f3637a4ba4db14be364eb20d7cc59342cb19fc43b45ec450f2cb30a6c372",
1030
+ "schedule.json": "f305e2f927ae588d0b5abaa7119e0abff505b8ee3d502259c0296e0d3206cc90",
1031
+ "study.json": "11a1e87ab0d57e46a36bfef1f5d1cc13cd7ca64dede4c3591f90821d34ffa3d0",
1032
+ "environment.json": "72c02ee46c17c1f8397e97d058f8a7d182c1674ddf451fbde1810048c473611d"
1033
+ },
1034
+ "population_sha256": "7890e3f2652bd76b78a1ea1b31fee0239eaf062d3bc4b33f5187471baacf2f68",
1035
+ "task_count": 956,
1036
+ "task_input_blobs_rehashed": false,
1037
+ "source_and_dependencies_match": true,
1038
+ "limits": {
1039
+ "output_tokens": 32768,
1040
+ "wall_seconds": 1800,
1041
+ "model_calls": 6,
1042
+ "tool_calls": 4,
1043
+ "tool_timeout": 90
1044
+ },
1045
+ "generation": {
1046
+ "model": "kimi-k3",
1047
+ "provider": "chat-completions",
1048
+ "reasoning_effort": "max",
1049
+ "stream": true,
1050
+ "timeout": 1800
1051
+ },
1052
+ "dataset_provenance": {
1053
+ "repo_id": "CamoAiLab/InferenceNet",
1054
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1055
+ "task_directory": "Selected_1000",
1056
+ "task_list": "Selected_1000/1000_new.csv",
1057
+ "prior_results_imported": false
1058
+ },
1059
+ "execution_identity": {
1060
+ "backend": "dsw-bwrap-v3",
1061
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1062
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1063
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1064
+ "engine_memory_bytes": 38654705664,
1065
+ "engine_cpus": 12,
1066
+ "cpu_slots": [
1067
+ [
1068
+ 4,
1069
+ 5
1070
+ ],
1071
+ [
1072
+ 6,
1073
+ 7
1074
+ ],
1075
+ [
1076
+ 8,
1077
+ 9
1078
+ ],
1079
+ [
1080
+ 10,
1081
+ 11
1082
+ ],
1083
+ [
1084
+ 12,
1085
+ 13
1086
+ ],
1087
+ [
1088
+ 14,
1089
+ 15
1090
+ ]
1091
+ ],
1092
+ "resources": {
1093
+ "cpus": 2.0,
1094
+ "memory": "6g",
1095
+ "tmpfs": "256m"
1096
+ },
1097
+ "memory_enforcement": "RLIMIT_AS",
1098
+ "single_payload_process": true,
1099
+ "network": "none",
1100
+ "uid": 10001,
1101
+ "capacity_profile": "independent-six-slots-20260916",
1102
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1103
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1104
+ },
1105
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1106
+ "scorer_sha256": {
1107
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1108
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1109
+ },
1110
+ "arms": [
1111
+ "baseline"
1112
+ ],
1113
+ "origin": "child"
1114
+ },
1115
+ {
1116
+ "core_sha256": {
1117
+ "protocol.json": "7f72f6ce4b86c3446111c200715c7c3e23787a011ffa1f2bfaeca461ae2d004b",
1118
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1119
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1120
+ "study.json": "860384e455b6cb4d3b982a99cb763f76d037cca17e0f4445eba4f5d8537e938b",
1121
+ "environment.json": "fb803d5b8d9c5e49f552841f119343fca52a314d6c26d86a7ccf8c615409aa79"
1122
+ },
1123
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1124
+ "task_count": 1000,
1125
+ "task_input_blobs_rehashed": false,
1126
+ "source_and_dependencies_match": true,
1127
+ "limits": {
1128
+ "output_tokens": 32768,
1129
+ "wall_seconds": 1800,
1130
+ "model_calls": 6,
1131
+ "tool_calls": 4,
1132
+ "tool_timeout": 90
1133
+ },
1134
+ "generation": {
1135
+ "model": "kimi-k3",
1136
+ "provider": "chat-completions",
1137
+ "reasoning_effort": "max",
1138
+ "stream": true,
1139
+ "timeout": 1800
1140
+ },
1141
+ "dataset_provenance": {
1142
+ "repo_id": "CamoAiLab/InferenceNet",
1143
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1144
+ "task_directory": "Selected_1000",
1145
+ "task_list": "Selected_1000/1000_new.csv",
1146
+ "prior_results_imported": false
1147
+ },
1148
+ "execution_identity": {
1149
+ "backend": "dsw-bwrap-v3",
1150
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1151
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1152
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1153
+ "engine_memory_bytes": 12884901888,
1154
+ "engine_cpus": 4,
1155
+ "cpu_slots": [
1156
+ [
1157
+ 16,
1158
+ 17
1159
+ ],
1160
+ [
1161
+ 18,
1162
+ 19
1163
+ ]
1164
+ ],
1165
+ "resources": {
1166
+ "memory": "6g",
1167
+ "cpus": 2.0,
1168
+ "tmpfs": "256m"
1169
+ },
1170
+ "memory_enforcement": "RLIMIT_AS",
1171
+ "single_payload_process": true,
1172
+ "network": "none",
1173
+ "uid": 10001
1174
+ },
1175
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1176
+ "scorer_sha256": {
1177
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1178
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1179
+ },
1180
+ "arms": [
1181
+ "baseline",
1182
+ "deepagents"
1183
+ ],
1184
+ "origin": "parent"
1185
+ },
1186
+ {
1187
+ "core_sha256": {
1188
+ "protocol.json": "7f72f6ce4b86c3446111c200715c7c3e23787a011ffa1f2bfaeca461ae2d004b",
1189
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1190
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1191
+ "study.json": "860384e455b6cb4d3b982a99cb763f76d037cca17e0f4445eba4f5d8537e938b",
1192
+ "environment.json": "fb803d5b8d9c5e49f552841f119343fca52a314d6c26d86a7ccf8c615409aa79"
1193
+ },
1194
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1195
+ "task_count": 1000,
1196
+ "task_input_blobs_rehashed": false,
1197
+ "source_and_dependencies_match": true,
1198
+ "limits": {
1199
+ "output_tokens": 32768,
1200
+ "wall_seconds": 1800,
1201
+ "model_calls": 6,
1202
+ "tool_calls": 4,
1203
+ "tool_timeout": 90
1204
+ },
1205
+ "generation": {
1206
+ "model": "kimi-k3",
1207
+ "provider": "chat-completions",
1208
+ "reasoning_effort": "max",
1209
+ "stream": true,
1210
+ "timeout": 1800
1211
+ },
1212
+ "dataset_provenance": {
1213
+ "repo_id": "CamoAiLab/InferenceNet",
1214
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1215
+ "task_directory": "Selected_1000",
1216
+ "task_list": "Selected_1000/1000_new.csv",
1217
+ "prior_results_imported": false
1218
+ },
1219
+ "execution_identity": {
1220
+ "backend": "dsw-bwrap-v3",
1221
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1222
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1223
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1224
+ "engine_memory_bytes": 12884901888,
1225
+ "engine_cpus": 4,
1226
+ "cpu_slots": [
1227
+ [
1228
+ 16,
1229
+ 17
1230
+ ],
1231
+ [
1232
+ 18,
1233
+ 19
1234
+ ]
1235
+ ],
1236
+ "resources": {
1237
+ "memory": "6g",
1238
+ "cpus": 2.0,
1239
+ "tmpfs": "256m"
1240
+ },
1241
+ "memory_enforcement": "RLIMIT_AS",
1242
+ "single_payload_process": true,
1243
+ "network": "none",
1244
+ "uid": 10001
1245
+ },
1246
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1247
+ "scorer_sha256": {
1248
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1249
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1250
+ },
1251
+ "arms": [
1252
+ "baseline",
1253
+ "deepagents"
1254
+ ],
1255
+ "origin": "parent"
1256
+ }
1257
+ ],
1258
+ "selected_origin_counts": {
1259
+ "child": 956,
1260
+ "parent": 44
1261
+ },
1262
+ "source_commits": [
1263
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1264
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1265
+ ],
1266
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1267
+ "source_export_manifest_sha256": "c974affeed4068cfe707b109c1240c5016b4e453d5c3e7cf37e938632ab812c3",
1268
+ "official_parity_verified": false,
1269
+ "notes": [
1270
+ "Selected parent and child records are disjoint; original protocol hashes and source versions retained.",
1271
+ "Baseline has one model call; DeepAgents has up to six model calls and four tool calls.",
1272
+ "Output-token budget 32768; task wall budget 1800 seconds; execution timeout 90 seconds.",
1273
+ "Final execution is separate from trial tools; unknown results are never dropped or replaced.",
1274
+ "Predictions are unavailable in this score export, even for tasks with successful scored predictions."
1275
+ ]
1276
+ }
results/kimi-k3/baseline/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/kimi-k3/baseline/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/kimi-k3/baseline/summary.json ADDED
@@ -0,0 +1,190 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
5
+ "model": "kimi-k3",
6
+ "arm": "baseline",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": false,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 526,
37
+ "denominator": 1000,
38
+ "rate": 0.526,
39
+ "score": 52.6,
40
+ "unknown_count": 5,
41
+ "failure_count": 469,
42
+ "assessable": 995,
43
+ "coverage_percent": 99.5
44
+ },
45
+ "partial_replication": {
46
+ "count": 366,
47
+ "denominator": 1000,
48
+ "rate": 0.366,
49
+ "score": 36.6,
50
+ "unknown_count": 481,
51
+ "failure_count": 153,
52
+ "assessable": 519,
53
+ "coverage_percent": 51.9
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 494,
57
+ "denominator": 1000,
58
+ "rate": 0.494,
59
+ "score": 49.4,
60
+ "unknown_count": 481,
61
+ "failure_count": 25,
62
+ "assessable": 519,
63
+ "coverage_percent": 51.9
64
+ },
65
+ "significance_level": {
66
+ "count": 431,
67
+ "denominator": 1000,
68
+ "rate": 0.431,
69
+ "score": 43.1,
70
+ "unknown_count": 481,
71
+ "failure_count": 88,
72
+ "assessable": 519,
73
+ "coverage_percent": 51.9
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "kimi-k3 baseline",
78
+ "Compilation Success": 52.6,
79
+ "Partial Replication": 36.6,
80
+ "Correct Coefficient Direction": 49.4,
81
+ "Significant Level Correctness": 43.1
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 995,
94
+ "task_unknown_count": 5,
95
+ "valid_prediction_count": 520,
96
+ "task_counts": {
97
+ "succeeded": 520,
98
+ "failed": 475,
99
+ "unknown": 5,
100
+ "unsealed": 0,
101
+ "not_started": 0,
102
+ "invalid_evidence": 0,
103
+ "completed": 995,
104
+ "planned": 1000
105
+ }
106
+ },
107
+ "legacy_local_paper": {
108
+ "profile": "local-paper-v1",
109
+ "definitions": {
110
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
111
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
112
+ "coefficient_only": "coefficient relative error <= .05",
113
+ "direction": "matching strictly positive or strictly negative coefficient signs",
114
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
115
+ },
116
+ "metrics": {
117
+ "perfect": {
118
+ "count": 230,
119
+ "denominator": 1000,
120
+ "rate": 0.23,
121
+ "score": 23.0,
122
+ "unknown_count": 481,
123
+ "failure_count": 289,
124
+ "assessable": 519,
125
+ "coverage_percent": 51.9
126
+ },
127
+ "partial": {
128
+ "count": 291,
129
+ "denominator": 1000,
130
+ "rate": 0.291,
131
+ "score": 29.1,
132
+ "unknown_count": 481,
133
+ "failure_count": 228,
134
+ "assessable": 519,
135
+ "coverage_percent": 51.9
136
+ },
137
+ "coefficient_only": {
138
+ "count": 366,
139
+ "denominator": 1000,
140
+ "rate": 0.366,
141
+ "score": 36.6,
142
+ "unknown_count": 481,
143
+ "failure_count": 153,
144
+ "assessable": 519,
145
+ "coverage_percent": 51.9
146
+ },
147
+ "direction": {
148
+ "count": 494,
149
+ "denominator": 1000,
150
+ "rate": 0.494,
151
+ "score": 49.4,
152
+ "unknown_count": 481,
153
+ "failure_count": 25,
154
+ "assessable": 519,
155
+ "coverage_percent": 51.9
156
+ },
157
+ "significance": {
158
+ "count": 418,
159
+ "denominator": 1000,
160
+ "rate": 0.418,
161
+ "score": 41.8,
162
+ "unknown_count": 481,
163
+ "failure_count": 101,
164
+ "assessable": 519,
165
+ "coverage_percent": 51.9
166
+ }
167
+ }
168
+ },
169
+ "verification_scope": "Offline reaggregation of accepted per-task flags; raw predictions are not available in this export; official parity not verified.",
170
+ "usage": {
171
+ "sealed_model_calls": 1000,
172
+ "unknown_usage_calls": 1,
173
+ "unsealed_episodes_excluded": 0,
174
+ "invalid_evidence_episodes_excluded": 0,
175
+ "known_token_subtotal": {
176
+ "input_tokens": 504306,
177
+ "output_tokens": 7409222,
178
+ "total_tokens": 7913528
179
+ },
180
+ "unknown_by_field": {
181
+ "input_tokens": 1,
182
+ "output_tokens": 1,
183
+ "total_tokens": 1
184
+ },
185
+ "total_tokens_complete": false,
186
+ "usd": null,
187
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
188
+ "summed_episode_seconds": 231469.68299999973
189
+ }
190
+ }
results/kimi-k3/deepagents/README.md ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Kimi K3 / deepagents
2
+
3
+ 1,000 个已封存槽位;成功 845,确定失败 151,任务 unknown 4。固定分母为 1,000;完整复现率 39.0%。
4
+
5
+ - `results.jsonl` / `results.csv`:相同的逐题状态、指标、调用用量和结果哈希。
6
+ - `summary.json`:本地 HF 四维与 local-paper-v1 五维分别保存,包含各自 unknown 和可判定覆盖率。
7
+ - `config.json`:固定任务清单、模型接口、预算、父/子接续来源及哈希。
8
+ - 上一级的 `paired_comparison.csv`、`provenance.json`、`diagnostics.json` 与 `export-verification.json` 保留双组配对与原导出核验边界。
9
+
10
+ `resolved=true` 表示槽位已封存;`definite_outcome` 才表示结果确定。未知和失败都保留在分母中。阶段 pending 可能表示前序失败后未执行到该阶段,不代表还在运行。
11
+
12
+ **此批次只有评分/状态导出,原始预测三元组未包含在源快照中。** 因此 prediction 字段保留 null,并用 `prediction_export_status=not_available_in_accepted_score_snapshot` 区分“导出不可用”与“模型没有有效预测”;不能据此把所有任务判为无预测。`summary.json` 的 valid_prediction_count 来自已核验的执行/评分状态。
13
+
14
+ 四维指标为本地实现,尚未验证与官方 scorer 完全一致。本次为历史 amended comparison 的归档,不自动更新正式榜单,没有重新调用模型或执行程序。用量字段是已知小计,须结合各字段 unknown_usage_calls 阅读。
15
+
16
+ JSON null / CSV 空单元格表示未知或不可用。原始数据、隐藏答案、完整提示词/API 轨迹、凭据和机器路径未包含在归档中。
results/kimi-k3/deepagents/config.json ADDED
@@ -0,0 +1,1276 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "model": "kimi-k3",
4
+ "arm": "deepagents",
5
+ "harness": "DeepAgents 0.7.13",
6
+ "archived_at_utc": "2026-09-16T15:28:00.953066+00:00",
7
+ "dataset": {
8
+ "repo_id": "CamoAiLab/InferenceNet",
9
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
10
+ "task_directory": "Selected_1000",
11
+ "task_list": "Selected_1000/1000_new.csv",
12
+ "prior_results_imported": false,
13
+ "expected_tasks": 1000,
14
+ "task_ids": [
15
+ 1,
16
+ 2,
17
+ 3,
18
+ 4,
19
+ 5,
20
+ 6,
21
+ 7,
22
+ 8,
23
+ 9,
24
+ 10,
25
+ 11,
26
+ 12,
27
+ 13,
28
+ 14,
29
+ 15,
30
+ 16,
31
+ 17,
32
+ 18,
33
+ 19,
34
+ 20,
35
+ 21,
36
+ 22,
37
+ 23,
38
+ 24,
39
+ 25,
40
+ 26,
41
+ 27,
42
+ 28,
43
+ 29,
44
+ 30,
45
+ 31,
46
+ 32,
47
+ 33,
48
+ 34,
49
+ 35,
50
+ 36,
51
+ 37,
52
+ 38,
53
+ 39,
54
+ 40,
55
+ 41,
56
+ 42,
57
+ 43,
58
+ 44,
59
+ 45,
60
+ 46,
61
+ 47,
62
+ 48,
63
+ 49,
64
+ 50,
65
+ 51,
66
+ 52,
67
+ 53,
68
+ 54,
69
+ 55,
70
+ 56,
71
+ 57,
72
+ 58,
73
+ 59,
74
+ 60,
75
+ 61,
76
+ 62,
77
+ 63,
78
+ 64,
79
+ 65,
80
+ 66,
81
+ 67,
82
+ 68,
83
+ 69,
84
+ 70,
85
+ 71,
86
+ 72,
87
+ 73,
88
+ 74,
89
+ 75,
90
+ 76,
91
+ 77,
92
+ 78,
93
+ 79,
94
+ 80,
95
+ 81,
96
+ 82,
97
+ 83,
98
+ 84,
99
+ 85,
100
+ 86,
101
+ 87,
102
+ 88,
103
+ 89,
104
+ 90,
105
+ 91,
106
+ 92,
107
+ 93,
108
+ 94,
109
+ 95,
110
+ 96,
111
+ 97,
112
+ 98,
113
+ 100,
114
+ 101,
115
+ 102,
116
+ 103,
117
+ 104,
118
+ 105,
119
+ 106,
120
+ 107,
121
+ 108,
122
+ 109,
123
+ 110,
124
+ 111,
125
+ 112,
126
+ 113,
127
+ 114,
128
+ 115,
129
+ 116,
130
+ 117,
131
+ 118,
132
+ 119,
133
+ 120,
134
+ 121,
135
+ 122,
136
+ 123,
137
+ 124,
138
+ 125,
139
+ 126,
140
+ 127,
141
+ 128,
142
+ 129,
143
+ 130,
144
+ 131,
145
+ 132,
146
+ 133,
147
+ 134,
148
+ 135,
149
+ 136,
150
+ 137,
151
+ 138,
152
+ 139,
153
+ 140,
154
+ 141,
155
+ 142,
156
+ 143,
157
+ 144,
158
+ 145,
159
+ 146,
160
+ 147,
161
+ 148,
162
+ 149,
163
+ 150,
164
+ 151,
165
+ 152,
166
+ 153,
167
+ 154,
168
+ 155,
169
+ 156,
170
+ 157,
171
+ 158,
172
+ 159,
173
+ 160,
174
+ 161,
175
+ 162,
176
+ 163,
177
+ 164,
178
+ 165,
179
+ 166,
180
+ 167,
181
+ 168,
182
+ 169,
183
+ 170,
184
+ 171,
185
+ 172,
186
+ 173,
187
+ 174,
188
+ 175,
189
+ 176,
190
+ 177,
191
+ 178,
192
+ 179,
193
+ 180,
194
+ 181,
195
+ 182,
196
+ 183,
197
+ 184,
198
+ 185,
199
+ 186,
200
+ 187,
201
+ 188,
202
+ 189,
203
+ 190,
204
+ 191,
205
+ 192,
206
+ 193,
207
+ 194,
208
+ 195,
209
+ 196,
210
+ 197,
211
+ 198,
212
+ 199,
213
+ 200,
214
+ 201,
215
+ 202,
216
+ 203,
217
+ 204,
218
+ 205,
219
+ 206,
220
+ 207,
221
+ 208,
222
+ 209,
223
+ 210,
224
+ 211,
225
+ 212,
226
+ 213,
227
+ 214,
228
+ 215,
229
+ 216,
230
+ 217,
231
+ 218,
232
+ 219,
233
+ 220,
234
+ 221,
235
+ 222,
236
+ 223,
237
+ 224,
238
+ 225,
239
+ 226,
240
+ 227,
241
+ 228,
242
+ 229,
243
+ 230,
244
+ 231,
245
+ 232,
246
+ 233,
247
+ 234,
248
+ 235,
249
+ 236,
250
+ 237,
251
+ 238,
252
+ 239,
253
+ 240,
254
+ 241,
255
+ 242,
256
+ 248,
257
+ 249,
258
+ 250,
259
+ 251,
260
+ 252,
261
+ 253,
262
+ 254,
263
+ 255,
264
+ 256,
265
+ 257,
266
+ 258,
267
+ 259,
268
+ 261,
269
+ 262,
270
+ 263,
271
+ 264,
272
+ 265,
273
+ 266,
274
+ 290,
275
+ 291,
276
+ 292,
277
+ 293,
278
+ 294,
279
+ 295,
280
+ 302,
281
+ 303,
282
+ 304,
283
+ 305,
284
+ 306,
285
+ 307,
286
+ 308,
287
+ 309,
288
+ 310,
289
+ 311,
290
+ 312,
291
+ 313,
292
+ 314,
293
+ 315,
294
+ 316,
295
+ 317,
296
+ 318,
297
+ 319,
298
+ 320,
299
+ 321,
300
+ 322,
301
+ 323,
302
+ 324,
303
+ 325,
304
+ 326,
305
+ 327,
306
+ 328,
307
+ 329,
308
+ 330,
309
+ 331,
310
+ 332,
311
+ 333,
312
+ 334,
313
+ 335,
314
+ 336,
315
+ 337,
316
+ 338,
317
+ 339,
318
+ 340,
319
+ 341,
320
+ 342,
321
+ 343,
322
+ 344,
323
+ 345,
324
+ 346,
325
+ 347,
326
+ 348,
327
+ 349,
328
+ 350,
329
+ 351,
330
+ 352,
331
+ 353,
332
+ 354,
333
+ 355,
334
+ 356,
335
+ 357,
336
+ 358,
337
+ 359,
338
+ 360,
339
+ 361,
340
+ 362,
341
+ 363,
342
+ 364,
343
+ 365,
344
+ 366,
345
+ 367,
346
+ 368,
347
+ 369,
348
+ 370,
349
+ 371,
350
+ 372,
351
+ 373,
352
+ 374,
353
+ 375,
354
+ 376,
355
+ 377,
356
+ 378,
357
+ 379,
358
+ 380,
359
+ 381,
360
+ 382,
361
+ 383,
362
+ 384,
363
+ 385,
364
+ 386,
365
+ 387,
366
+ 388,
367
+ 389,
368
+ 390,
369
+ 391,
370
+ 392,
371
+ 393,
372
+ 394,
373
+ 395,
374
+ 396,
375
+ 397,
376
+ 398,
377
+ 399,
378
+ 400,
379
+ 401,
380
+ 402,
381
+ 403,
382
+ 404,
383
+ 405,
384
+ 406,
385
+ 407,
386
+ 408,
387
+ 409,
388
+ 410,
389
+ 411,
390
+ 412,
391
+ 413,
392
+ 414,
393
+ 415,
394
+ 416,
395
+ 417,
396
+ 418,
397
+ 419,
398
+ 420,
399
+ 421,
400
+ 422,
401
+ 423,
402
+ 424,
403
+ 425,
404
+ 426,
405
+ 427,
406
+ 428,
407
+ 429,
408
+ 430,
409
+ 431,
410
+ 432,
411
+ 433,
412
+ 434,
413
+ 435,
414
+ 436,
415
+ 437,
416
+ 438,
417
+ 439,
418
+ 440,
419
+ 441,
420
+ 442,
421
+ 443,
422
+ 444,
423
+ 445,
424
+ 446,
425
+ 447,
426
+ 448,
427
+ 449,
428
+ 450,
429
+ 451,
430
+ 452,
431
+ 453,
432
+ 454,
433
+ 455,
434
+ 456,
435
+ 457,
436
+ 458,
437
+ 459,
438
+ 460,
439
+ 461,
440
+ 462,
441
+ 463,
442
+ 464,
443
+ 465,
444
+ 466,
445
+ 467,
446
+ 468,
447
+ 469,
448
+ 470,
449
+ 471,
450
+ 472,
451
+ 473,
452
+ 474,
453
+ 475,
454
+ 476,
455
+ 477,
456
+ 478,
457
+ 479,
458
+ 480,
459
+ 481,
460
+ 482,
461
+ 483,
462
+ 484,
463
+ 485,
464
+ 486,
465
+ 487,
466
+ 488,
467
+ 489,
468
+ 490,
469
+ 491,
470
+ 492,
471
+ 493,
472
+ 494,
473
+ 495,
474
+ 496,
475
+ 497,
476
+ 498,
477
+ 499,
478
+ 500,
479
+ 501,
480
+ 502,
481
+ 503,
482
+ 504,
483
+ 505,
484
+ 506,
485
+ 507,
486
+ 508,
487
+ 509,
488
+ 510,
489
+ 511,
490
+ 512,
491
+ 513,
492
+ 514,
493
+ 515,
494
+ 516,
495
+ 517,
496
+ 518,
497
+ 519,
498
+ 520,
499
+ 521,
500
+ 522,
501
+ 523,
502
+ 524,
503
+ 525,
504
+ 526,
505
+ 527,
506
+ 528,
507
+ 529,
508
+ 530,
509
+ 531,
510
+ 532,
511
+ 533,
512
+ 534,
513
+ 535,
514
+ 536,
515
+ 537,
516
+ 538,
517
+ 539,
518
+ 540,
519
+ 541,
520
+ 542,
521
+ 543,
522
+ 544,
523
+ 545,
524
+ 546,
525
+ 547,
526
+ 548,
527
+ 549,
528
+ 550,
529
+ 551,
530
+ 552,
531
+ 553,
532
+ 554,
533
+ 555,
534
+ 556,
535
+ 557,
536
+ 558,
537
+ 559,
538
+ 560,
539
+ 561,
540
+ 562,
541
+ 563,
542
+ 565,
543
+ 566,
544
+ 567,
545
+ 568,
546
+ 570,
547
+ 571,
548
+ 572,
549
+ 573,
550
+ 574,
551
+ 575,
552
+ 576,
553
+ 577,
554
+ 578,
555
+ 579,
556
+ 580,
557
+ 581,
558
+ 582,
559
+ 583,
560
+ 584,
561
+ 585,
562
+ 586,
563
+ 587,
564
+ 588,
565
+ 589,
566
+ 590,
567
+ 591,
568
+ 592,
569
+ 593,
570
+ 594,
571
+ 595,
572
+ 596,
573
+ 597,
574
+ 598,
575
+ 601,
576
+ 602,
577
+ 603,
578
+ 604,
579
+ 605,
580
+ 606,
581
+ 607,
582
+ 608,
583
+ 609,
584
+ 610,
585
+ 611,
586
+ 612,
587
+ 613,
588
+ 614,
589
+ 615,
590
+ 616,
591
+ 617,
592
+ 618,
593
+ 619,
594
+ 620,
595
+ 621,
596
+ 622,
597
+ 623,
598
+ 624,
599
+ 664,
600
+ 665,
601
+ 666,
602
+ 667,
603
+ 668,
604
+ 669,
605
+ 670,
606
+ 671,
607
+ 672,
608
+ 673,
609
+ 674,
610
+ 675,
611
+ 676,
612
+ 677,
613
+ 678,
614
+ 679,
615
+ 680,
616
+ 681,
617
+ 682,
618
+ 683,
619
+ 684,
620
+ 685,
621
+ 686,
622
+ 687,
623
+ 688,
624
+ 689,
625
+ 690,
626
+ 691,
627
+ 692,
628
+ 693,
629
+ 694,
630
+ 695,
631
+ 696,
632
+ 697,
633
+ 698,
634
+ 699,
635
+ 700,
636
+ 701,
637
+ 702,
638
+ 703,
639
+ 704,
640
+ 705,
641
+ 706,
642
+ 707,
643
+ 708,
644
+ 709,
645
+ 710,
646
+ 711,
647
+ 712,
648
+ 713,
649
+ 714,
650
+ 715,
651
+ 716,
652
+ 717,
653
+ 718,
654
+ 719,
655
+ 720,
656
+ 721,
657
+ 722,
658
+ 723,
659
+ 724,
660
+ 725,
661
+ 726,
662
+ 727,
663
+ 728,
664
+ 729,
665
+ 730,
666
+ 731,
667
+ 732,
668
+ 733,
669
+ 734,
670
+ 735,
671
+ 736,
672
+ 737,
673
+ 739,
674
+ 740,
675
+ 741,
676
+ 742,
677
+ 743,
678
+ 744,
679
+ 745,
680
+ 746,
681
+ 747,
682
+ 748,
683
+ 749,
684
+ 750,
685
+ 751,
686
+ 752,
687
+ 753,
688
+ 754,
689
+ 755,
690
+ 756,
691
+ 757,
692
+ 758,
693
+ 759,
694
+ 760,
695
+ 761,
696
+ 762,
697
+ 763,
698
+ 764,
699
+ 765,
700
+ 766,
701
+ 767,
702
+ 768,
703
+ 772,
704
+ 773,
705
+ 774,
706
+ 775,
707
+ 776,
708
+ 777,
709
+ 778,
710
+ 779,
711
+ 780,
712
+ 781,
713
+ 782,
714
+ 783,
715
+ 784,
716
+ 785,
717
+ 786,
718
+ 787,
719
+ 788,
720
+ 789,
721
+ 790,
722
+ 791,
723
+ 793,
724
+ 794,
725
+ 795,
726
+ 796,
727
+ 797,
728
+ 798,
729
+ 799,
730
+ 800,
731
+ 801,
732
+ 802,
733
+ 803,
734
+ 804,
735
+ 805,
736
+ 806,
737
+ 807,
738
+ 808,
739
+ 825,
740
+ 826,
741
+ 827,
742
+ 828,
743
+ 829,
744
+ 830,
745
+ 831,
746
+ 832,
747
+ 833,
748
+ 834,
749
+ 835,
750
+ 836,
751
+ 837,
752
+ 838,
753
+ 839,
754
+ 840,
755
+ 841,
756
+ 842,
757
+ 843,
758
+ 844,
759
+ 845,
760
+ 846,
761
+ 847,
762
+ 848,
763
+ 849,
764
+ 850,
765
+ 851,
766
+ 852,
767
+ 853,
768
+ 854,
769
+ 855,
770
+ 856,
771
+ 857,
772
+ 858,
773
+ 859,
774
+ 860,
775
+ 861,
776
+ 862,
777
+ 863,
778
+ 864,
779
+ 865,
780
+ 866,
781
+ 867,
782
+ 868,
783
+ 869,
784
+ 870,
785
+ 871,
786
+ 872,
787
+ 873,
788
+ 874,
789
+ 875,
790
+ 876,
791
+ 877,
792
+ 878,
793
+ 879,
794
+ 880,
795
+ 881,
796
+ 882,
797
+ 883,
798
+ 884,
799
+ 885,
800
+ 886,
801
+ 887,
802
+ 888,
803
+ 889,
804
+ 890,
805
+ 891,
806
+ 892,
807
+ 893,
808
+ 894,
809
+ 895,
810
+ 896,
811
+ 897,
812
+ 898,
813
+ 899,
814
+ 900,
815
+ 901,
816
+ 902,
817
+ 903,
818
+ 906,
819
+ 907,
820
+ 913,
821
+ 914,
822
+ 915,
823
+ 917,
824
+ 920,
825
+ 921,
826
+ 922,
827
+ 923,
828
+ 924,
829
+ 925,
830
+ 926,
831
+ 927,
832
+ 928,
833
+ 929,
834
+ 930,
835
+ 931,
836
+ 932,
837
+ 933,
838
+ 934,
839
+ 935,
840
+ 936,
841
+ 937,
842
+ 938,
843
+ 939,
844
+ 940,
845
+ 941,
846
+ 942,
847
+ 943,
848
+ 944,
849
+ 945,
850
+ 946,
851
+ 947,
852
+ 948,
853
+ 949,
854
+ 950,
855
+ 951,
856
+ 952,
857
+ 953,
858
+ 954,
859
+ 955,
860
+ 956,
861
+ 957,
862
+ 958,
863
+ 959,
864
+ 960,
865
+ 961,
866
+ 962,
867
+ 963,
868
+ 964,
869
+ 965,
870
+ 966,
871
+ 967,
872
+ 968,
873
+ 969,
874
+ 970,
875
+ 971,
876
+ 972,
877
+ 973,
878
+ 974,
879
+ 975,
880
+ 976,
881
+ 977,
882
+ 978,
883
+ 979,
884
+ 980,
885
+ 981,
886
+ 982,
887
+ 983,
888
+ 984,
889
+ 985,
890
+ 1001,
891
+ 1002,
892
+ 1003,
893
+ 1004,
894
+ 1005,
895
+ 1006,
896
+ 1007,
897
+ 1008,
898
+ 1009,
899
+ 1010,
900
+ 1011,
901
+ 1012,
902
+ 1013,
903
+ 1014,
904
+ 1015,
905
+ 1016,
906
+ 1017,
907
+ 1018,
908
+ 1019,
909
+ 1020,
910
+ 1021,
911
+ 1022,
912
+ 1023,
913
+ 1024,
914
+ 1025,
915
+ 1026,
916
+ 1027,
917
+ 1028,
918
+ 1029,
919
+ 1030,
920
+ 1031,
921
+ 1032,
922
+ 1033,
923
+ 1034,
924
+ 1035,
925
+ 1036,
926
+ 1037,
927
+ 1038,
928
+ 1039,
929
+ 1040,
930
+ 1041,
931
+ 1042,
932
+ 1043,
933
+ 1044,
934
+ 1045,
935
+ 1046,
936
+ 1047,
937
+ 1048,
938
+ 1049,
939
+ 1050,
940
+ 1051,
941
+ 1052,
942
+ 1053,
943
+ 1054,
944
+ 1055,
945
+ 1056,
946
+ 1057,
947
+ 1058,
948
+ 1059,
949
+ 1060,
950
+ 1061,
951
+ 1062,
952
+ 1063,
953
+ 1064,
954
+ 1065,
955
+ 1066,
956
+ 1067,
957
+ 1068,
958
+ 1069,
959
+ 1070,
960
+ 1071,
961
+ 1072,
962
+ 1073,
963
+ 1074,
964
+ 1075,
965
+ 1076,
966
+ 1077,
967
+ 1078,
968
+ 1079,
969
+ 1080,
970
+ 1081,
971
+ 1082,
972
+ 1083,
973
+ 1084,
974
+ 1085,
975
+ 1086,
976
+ 1087,
977
+ 1088,
978
+ 1089,
979
+ 1090,
980
+ 1091,
981
+ 1092,
982
+ 1093,
983
+ 1094,
984
+ 1095,
985
+ 1096,
986
+ 1097,
987
+ 1098,
988
+ 1099,
989
+ 1100,
990
+ 1101,
991
+ 1102,
992
+ 1103,
993
+ 1104,
994
+ 1105,
995
+ 1106,
996
+ 1107,
997
+ 1108,
998
+ 1109,
999
+ 1110,
1000
+ 1111,
1001
+ 1112,
1002
+ 1113,
1003
+ 1114,
1004
+ 1115,
1005
+ 1116,
1006
+ 1117,
1007
+ 1118,
1008
+ 1119,
1009
+ 1120,
1010
+ 1121,
1011
+ 1122,
1012
+ 1123,
1013
+ 1124,
1014
+ 1125
1015
+ ]
1016
+ },
1017
+ "generation": {
1018
+ "model": "kimi-k3",
1019
+ "provider": "chat-completions",
1020
+ "reasoning_effort": "max",
1021
+ "stream": true,
1022
+ "timeout": 1800
1023
+ },
1024
+ "effective_model_call_cap": 6,
1025
+ "recorded_protocols": [
1026
+ {
1027
+ "core_sha256": {
1028
+ "protocol.json": "7f72f6ce4b86c3446111c200715c7c3e23787a011ffa1f2bfaeca461ae2d004b",
1029
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1030
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1031
+ "study.json": "860384e455b6cb4d3b982a99cb763f76d037cca17e0f4445eba4f5d8537e938b",
1032
+ "environment.json": "fb803d5b8d9c5e49f552841f119343fca52a314d6c26d86a7ccf8c615409aa79"
1033
+ },
1034
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1035
+ "task_count": 1000,
1036
+ "task_input_blobs_rehashed": false,
1037
+ "source_and_dependencies_match": true,
1038
+ "limits": {
1039
+ "output_tokens": 32768,
1040
+ "wall_seconds": 1800,
1041
+ "model_calls": 6,
1042
+ "tool_calls": 4,
1043
+ "tool_timeout": 90
1044
+ },
1045
+ "generation": {
1046
+ "model": "kimi-k3",
1047
+ "provider": "chat-completions",
1048
+ "reasoning_effort": "max",
1049
+ "stream": true,
1050
+ "timeout": 1800
1051
+ },
1052
+ "dataset_provenance": {
1053
+ "repo_id": "CamoAiLab/InferenceNet",
1054
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1055
+ "task_directory": "Selected_1000",
1056
+ "task_list": "Selected_1000/1000_new.csv",
1057
+ "prior_results_imported": false
1058
+ },
1059
+ "execution_identity": {
1060
+ "backend": "dsw-bwrap-v3",
1061
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1062
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1063
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1064
+ "engine_memory_bytes": 12884901888,
1065
+ "engine_cpus": 4,
1066
+ "cpu_slots": [
1067
+ [
1068
+ 16,
1069
+ 17
1070
+ ],
1071
+ [
1072
+ 18,
1073
+ 19
1074
+ ]
1075
+ ],
1076
+ "resources": {
1077
+ "memory": "6g",
1078
+ "cpus": 2.0,
1079
+ "tmpfs": "256m"
1080
+ },
1081
+ "memory_enforcement": "RLIMIT_AS",
1082
+ "single_payload_process": true,
1083
+ "network": "none",
1084
+ "uid": 10001
1085
+ },
1086
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1087
+ "scorer_sha256": {
1088
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1089
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1090
+ },
1091
+ "arms": [
1092
+ "baseline",
1093
+ "deepagents"
1094
+ ],
1095
+ "origin": "parent"
1096
+ },
1097
+ {
1098
+ "core_sha256": {
1099
+ "protocol.json": "13a075289c4a5bd79ed48ba6988d34e094193d05ae6d8e175f43424262f352fe",
1100
+ "inputs.manifest.json": "1c2dab1d0d6dfec969b10eed1300e082162c70d4c9f2a7421dd548f69bb28063",
1101
+ "schedule.json": "a1a9643af25a49cfa0eeb0589eafd685fa00aefa2570012cdbb2b4acdd734f8e",
1102
+ "study.json": "daf4ddb5c1cd0aac658968b1b6f81e52bc68e91c13ca544835022216e30927e1",
1103
+ "environment.json": "72c02ee46c17c1f8397e97d058f8a7d182c1674ddf451fbde1810048c473611d"
1104
+ },
1105
+ "population_sha256": "0c4c8b28d3d2150175eaa8d720f61a8497e04c0e95cdea107f205c366b860b91",
1106
+ "task_count": 960,
1107
+ "task_input_blobs_rehashed": false,
1108
+ "source_and_dependencies_match": true,
1109
+ "limits": {
1110
+ "output_tokens": 32768,
1111
+ "wall_seconds": 1800,
1112
+ "model_calls": 6,
1113
+ "tool_calls": 4,
1114
+ "tool_timeout": 90
1115
+ },
1116
+ "generation": {
1117
+ "model": "kimi-k3",
1118
+ "provider": "chat-completions",
1119
+ "reasoning_effort": "max",
1120
+ "stream": true,
1121
+ "timeout": 1800
1122
+ },
1123
+ "dataset_provenance": {
1124
+ "repo_id": "CamoAiLab/InferenceNet",
1125
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1126
+ "task_directory": "Selected_1000",
1127
+ "task_list": "Selected_1000/1000_new.csv",
1128
+ "prior_results_imported": false
1129
+ },
1130
+ "execution_identity": {
1131
+ "backend": "dsw-bwrap-v3",
1132
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1133
+ "deployment_sha256": "d63e59fd51475a724124091cd655d1b21ad3060ee91180b4c6b57c2ec90b0b03",
1134
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1135
+ "engine_memory_bytes": 38654705664,
1136
+ "engine_cpus": 12,
1137
+ "cpu_slots": [
1138
+ [
1139
+ 4,
1140
+ 5
1141
+ ],
1142
+ [
1143
+ 6,
1144
+ 7
1145
+ ],
1146
+ [
1147
+ 8,
1148
+ 9
1149
+ ],
1150
+ [
1151
+ 10,
1152
+ 11
1153
+ ],
1154
+ [
1155
+ 12,
1156
+ 13
1157
+ ],
1158
+ [
1159
+ 14,
1160
+ 15
1161
+ ]
1162
+ ],
1163
+ "resources": {
1164
+ "cpus": 2.0,
1165
+ "memory": "6g",
1166
+ "tmpfs": "256m"
1167
+ },
1168
+ "memory_enforcement": "RLIMIT_AS",
1169
+ "single_payload_process": true,
1170
+ "network": "none",
1171
+ "uid": 10001,
1172
+ "capacity_profile": "independent-six-slots-20260916",
1173
+ "executor_source_sha256": "5321c07038e4dfb5b827700e2d4cb9df35f21f3d9fcf13c831f09df3e9d3589f",
1174
+ "parent_deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee"
1175
+ },
1176
+ "source_commit": "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980",
1177
+ "scorer_sha256": {
1178
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1179
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1180
+ },
1181
+ "arms": [
1182
+ "deepagents"
1183
+ ],
1184
+ "origin": "child"
1185
+ },
1186
+ {
1187
+ "core_sha256": {
1188
+ "protocol.json": "7f72f6ce4b86c3446111c200715c7c3e23787a011ffa1f2bfaeca461ae2d004b",
1189
+ "inputs.manifest.json": "d794cf9ce233f23f2ec35f0843501c2d3c64dd483cc1e41f3696ec1166d0f11e",
1190
+ "schedule.json": "567db67bdb0d515ade90398142136b94b027ff0f3d2924566b603a1844e74979",
1191
+ "study.json": "860384e455b6cb4d3b982a99cb763f76d037cca17e0f4445eba4f5d8537e938b",
1192
+ "environment.json": "fb803d5b8d9c5e49f552841f119343fca52a314d6c26d86a7ccf8c615409aa79"
1193
+ },
1194
+ "population_sha256": "57e33e1883cc6e1a9e35d068614e28832c87b7cb1c22258cd5d9b3ad1e5ff901",
1195
+ "task_count": 1000,
1196
+ "task_input_blobs_rehashed": false,
1197
+ "source_and_dependencies_match": true,
1198
+ "limits": {
1199
+ "output_tokens": 32768,
1200
+ "wall_seconds": 1800,
1201
+ "model_calls": 6,
1202
+ "tool_calls": 4,
1203
+ "tool_timeout": 90
1204
+ },
1205
+ "generation": {
1206
+ "model": "kimi-k3",
1207
+ "provider": "chat-completions",
1208
+ "reasoning_effort": "max",
1209
+ "stream": true,
1210
+ "timeout": 1800
1211
+ },
1212
+ "dataset_provenance": {
1213
+ "repo_id": "CamoAiLab/InferenceNet",
1214
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
1215
+ "task_directory": "Selected_1000",
1216
+ "task_list": "Selected_1000/1000_new.csv",
1217
+ "prior_results_imported": false
1218
+ },
1219
+ "execution_identity": {
1220
+ "backend": "dsw-bwrap-v3",
1221
+ "protocol": "inferencenet-dsw-amd64-single-process-v3",
1222
+ "deployment_sha256": "6c7d0c001925f436f884ebb737c430c4644b54fc17dcc681320133cca34a4bee",
1223
+ "rootfs_inventory_sha256": "9ac222dbc0ea58880e790c36f5afddddb727f948e23e73cf0978de73ae498e59",
1224
+ "engine_memory_bytes": 12884901888,
1225
+ "engine_cpus": 4,
1226
+ "cpu_slots": [
1227
+ [
1228
+ 16,
1229
+ 17
1230
+ ],
1231
+ [
1232
+ 18,
1233
+ 19
1234
+ ]
1235
+ ],
1236
+ "resources": {
1237
+ "memory": "6g",
1238
+ "cpus": 2.0,
1239
+ "tmpfs": "256m"
1240
+ },
1241
+ "memory_enforcement": "RLIMIT_AS",
1242
+ "single_payload_process": true,
1243
+ "network": "none",
1244
+ "uid": 10001
1245
+ },
1246
+ "source_commit": "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1247
+ "scorer_sha256": {
1248
+ "eval/evaluation/metrics.py": "52d46989ab6dc8b4726edfc281a0e27c9f57dce3a958b7de73143aec47d81a61",
1249
+ "eval/evaluation/leaderboard.py": "f6a6190e97f8bf5c499e97de74f0b882ebb8322dea92e0543711d523ffe6cba3"
1250
+ },
1251
+ "arms": [
1252
+ "baseline",
1253
+ "deepagents"
1254
+ ],
1255
+ "origin": "parent"
1256
+ }
1257
+ ],
1258
+ "selected_origin_counts": {
1259
+ "child": 960,
1260
+ "parent": 40
1261
+ },
1262
+ "source_commits": [
1263
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
1264
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
1265
+ ],
1266
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
1267
+ "source_export_manifest_sha256": "c974affeed4068cfe707b109c1240c5016b4e453d5c3e7cf37e938632ab812c3",
1268
+ "official_parity_verified": false,
1269
+ "notes": [
1270
+ "Selected parent and child records are disjoint; original protocol hashes and source versions retained.",
1271
+ "Baseline has one model call; DeepAgents has up to six model calls and four tool calls.",
1272
+ "Output-token budget 32768; task wall budget 1800 seconds; execution timeout 90 seconds.",
1273
+ "Final execution is separate from trial tools; unknown results are never dropped or replaced.",
1274
+ "Predictions are unavailable in this score export, even for tasks with successful scored predictions."
1275
+ ]
1276
+ }
results/kimi-k3/deepagents/results.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/kimi-k3/deepagents/results.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
results/kimi-k3/deepagents/summary.json ADDED
@@ -0,0 +1,190 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema_version": 2,
3
+ "kind": "inferencenet-offline-leaderboard-export",
4
+ "generated_at": "2026-09-16T15:28:00.953066+00:00",
5
+ "model": "kimi-k3",
6
+ "arm": "deepagents",
7
+ "metric_profile": "hf-leaderboard-v1",
8
+ "definitions": {
9
+ "compilation_success": "Trusted evidence of complete code execution; no prediction JSON gate.",
10
+ "partial_replication": "Coefficient relative error <= .05; no standard-error or p-value gate.",
11
+ "coefficient_direction": "Matching strictly positive or strictly negative coefficient signs.",
12
+ "significance_level": "Equal p-value category: p < .01, p < .05, p < .1, otherwise; no direction gate."
13
+ },
14
+ "assumptions": {
15
+ "official_rules": "Webpage interpretation only; official scorer parity is not verified.",
16
+ "compilation_evidence": "Caller supplies True, False, or None from verified execution evidence, including exit status, timeout, termination confirmation, and execution errors. Missing or uncertain evidence maps to None.",
17
+ "prediction_evidence": "Only execution_status=succeeded and scoring_status=scored rows are evaluated. The current caller requires a complete valid prediction triplet; invalid outputs are not recovered here.",
18
+ "coefficient_boundary": "5% includes equality. Numeric decimal representations are compared exactly without epsilon.",
19
+ "zero_reference_coefficient": "A zero reference coefficient is unknown for relative error and positive/negative direction; it does not block significance. A zero prediction against a nonzero reference is incorrect.",
20
+ "significance_categories": "Four categories use strict .01, .05, and .1 boundaries. Both p-values must be numeric and within [0, 1]. These boundaries and the absent direction gate are explicit local choices.",
21
+ "field_validation": "Each metric requires only its own finite numeric fields. Booleans, numeric strings, and inequalities are not coerced. Standard errors do not affect any of these three formulas.",
22
+ "denominator": "All planned tasks remain in every denominator. Confirmed compilation failures are False; missing or uncertain execution evidence is unknown. For the three result metrics, failed execution, missing predictions, unscored rows, and unsupported reference fields are unknown. Unknowns contribute no successes and are counted separately.",
23
+ "attempt_selection": "The caller supplies one row per task under its frozen attempt policy."
24
+ },
25
+ "official_parity_verified": false,
26
+ "evidence_class": "amended_comparison",
27
+ "original_protocol_complete": false,
28
+ "complete": false,
29
+ "all_slots_sealed": true,
30
+ "dataset_revision": "59f9512a38e594528807744214a60ee00367434e",
31
+ "included_in_displayed_leaderboard": false,
32
+ "publication_kind": "results_archive",
33
+ "result": {
34
+ "metrics": {
35
+ "compilation_success": {
36
+ "count": 867,
37
+ "denominator": 1000,
38
+ "rate": 0.867,
39
+ "score": 86.7,
40
+ "unknown_count": 4,
41
+ "failure_count": 129,
42
+ "assessable": 996,
43
+ "coverage_percent": 99.6
44
+ },
45
+ "partial_replication": {
46
+ "count": 635,
47
+ "denominator": 1000,
48
+ "rate": 0.635,
49
+ "score": 63.5,
50
+ "unknown_count": 156,
51
+ "failure_count": 209,
52
+ "assessable": 844,
53
+ "coverage_percent": 84.4
54
+ },
55
+ "coefficient_direction": {
56
+ "count": 804,
57
+ "denominator": 1000,
58
+ "rate": 0.804,
59
+ "score": 80.4,
60
+ "unknown_count": 156,
61
+ "failure_count": 40,
62
+ "assessable": 844,
63
+ "coverage_percent": 84.4
64
+ },
65
+ "significance_level": {
66
+ "count": 709,
67
+ "denominator": 1000,
68
+ "rate": 0.709,
69
+ "score": 70.9,
70
+ "unknown_count": 157,
71
+ "failure_count": 134,
72
+ "assessable": 843,
73
+ "coverage_percent": 84.3
74
+ }
75
+ },
76
+ "row": {
77
+ "Model ID": "kimi-k3 deepagents",
78
+ "Compilation Success": 86.7,
79
+ "Partial Replication": 63.5,
80
+ "Correct Coefficient Direction": 80.4,
81
+ "Significant Level Correctness": 70.9
82
+ },
83
+ "columns": [
84
+ "Model ID",
85
+ "Compilation Success",
86
+ "Partial Replication",
87
+ "Correct Coefficient Direction",
88
+ "Significant Level Correctness"
89
+ ],
90
+ "expected_count": 1000,
91
+ "resolved_count": 1000,
92
+ "sealed_count": 1000,
93
+ "definite_outcome_count": 996,
94
+ "task_unknown_count": 4,
95
+ "valid_prediction_count": 845,
96
+ "task_counts": {
97
+ "succeeded": 845,
98
+ "failed": 151,
99
+ "unknown": 4,
100
+ "unsealed": 0,
101
+ "not_started": 0,
102
+ "invalid_evidence": 0,
103
+ "completed": 996,
104
+ "planned": 1000
105
+ }
106
+ },
107
+ "legacy_local_paper": {
108
+ "profile": "local-paper-v1",
109
+ "definitions": {
110
+ "perfect": "coefficient and SE relative errors <= .01 and p absolute error <= .01",
111
+ "partial": "coefficient and SE relative errors < .05; no p gate; includes perfect",
112
+ "coefficient_only": "coefficient relative error <= .05",
113
+ "direction": "matching strictly positive or strictly negative coefficient signs",
114
+ "significance": "direction and equal category: p < .01, p < .05, p < .1, otherwise"
115
+ },
116
+ "metrics": {
117
+ "perfect": {
118
+ "count": 390,
119
+ "denominator": 1000,
120
+ "rate": 0.39,
121
+ "score": 39.0,
122
+ "unknown_count": 159,
123
+ "failure_count": 451,
124
+ "assessable": 841,
125
+ "coverage_percent": 84.1
126
+ },
127
+ "partial": {
128
+ "count": 513,
129
+ "denominator": 1000,
130
+ "rate": 0.513,
131
+ "score": 51.3,
132
+ "unknown_count": 159,
133
+ "failure_count": 328,
134
+ "assessable": 841,
135
+ "coverage_percent": 84.1
136
+ },
137
+ "coefficient_only": {
138
+ "count": 634,
139
+ "denominator": 1000,
140
+ "rate": 0.634,
141
+ "score": 63.4,
142
+ "unknown_count": 159,
143
+ "failure_count": 207,
144
+ "assessable": 841,
145
+ "coverage_percent": 84.1
146
+ },
147
+ "direction": {
148
+ "count": 802,
149
+ "denominator": 1000,
150
+ "rate": 0.802,
151
+ "score": 80.2,
152
+ "unknown_count": 159,
153
+ "failure_count": 39,
154
+ "assessable": 841,
155
+ "coverage_percent": 84.1
156
+ },
157
+ "significance": {
158
+ "count": 688,
159
+ "denominator": 1000,
160
+ "rate": 0.688,
161
+ "score": 68.8,
162
+ "unknown_count": 159,
163
+ "failure_count": 153,
164
+ "assessable": 841,
165
+ "coverage_percent": 84.1
166
+ }
167
+ }
168
+ },
169
+ "verification_scope": "Offline reaggregation of accepted per-task flags; raw predictions are not available in this export; official parity not verified.",
170
+ "usage": {
171
+ "sealed_model_calls": 5095,
172
+ "unknown_usage_calls": 1,
173
+ "unsealed_episodes_excluded": 0,
174
+ "invalid_evidence_episodes_excluded": 0,
175
+ "known_token_subtotal": {
176
+ "input_tokens": 31689278,
177
+ "output_tokens": 5765376,
178
+ "total_tokens": 37454654
179
+ },
180
+ "unknown_by_field": {
181
+ "input_tokens": 1,
182
+ "output_tokens": 1,
183
+ "total_tokens": 1
184
+ },
185
+ "total_tokens_complete": false,
186
+ "usd": null,
187
+ "usd_status": "unknown_no_verified_price_or_billing_receipt",
188
+ "summed_episode_seconds": 369394.95999999985
189
+ }
190
+ }
results/kimi-k3/diagnostics.json ADDED
@@ -0,0 +1,90 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage_counts": {
3
+ "baseline": [
4
+ {
5
+ "generation": "failed",
6
+ "execution": "pending",
7
+ "scoring": "pending",
8
+ "count": 25
9
+ },
10
+ {
11
+ "generation": "succeeded",
12
+ "execution": "failed",
13
+ "scoring": "pending",
14
+ "count": 450
15
+ },
16
+ {
17
+ "generation": "succeeded",
18
+ "execution": "pending",
19
+ "scoring": "pending",
20
+ "count": 4
21
+ },
22
+ {
23
+ "generation": "succeeded",
24
+ "execution": "succeeded",
25
+ "scoring": "scored",
26
+ "count": 520
27
+ },
28
+ {
29
+ "generation": "uncertain",
30
+ "execution": "pending",
31
+ "scoring": "pending",
32
+ "count": 1
33
+ }
34
+ ],
35
+ "deepagents": [
36
+ {
37
+ "generation": "failed",
38
+ "execution": "pending",
39
+ "scoring": "pending",
40
+ "count": 50
41
+ },
42
+ {
43
+ "generation": "succeeded",
44
+ "execution": "failed",
45
+ "scoring": "pending",
46
+ "count": 101
47
+ },
48
+ {
49
+ "generation": "succeeded",
50
+ "execution": "succeeded",
51
+ "scoring": "scored",
52
+ "count": 845
53
+ },
54
+ {
55
+ "generation": "uncertain",
56
+ "execution": "pending",
57
+ "scoring": "pending",
58
+ "count": 4
59
+ }
60
+ ]
61
+ },
62
+ "task_unknown_ids": {
63
+ "baseline": [
64
+ 73,
65
+ 113,
66
+ 233,
67
+ 409,
68
+ 981
69
+ ],
70
+ "deepagents": [
71
+ 234,
72
+ 417,
73
+ 495,
74
+ 674
75
+ ]
76
+ },
77
+ "generation_failed_model_call_distribution": {
78
+ "baseline": {
79
+ "1": 25
80
+ },
81
+ "deepagents": {
82
+ "6": 43,
83
+ "1": 3,
84
+ "5": 3,
85
+ "4": 1
86
+ }
87
+ },
88
+ "raw_error_reasons_verified": false,
89
+ "limitation": "Stage labels and counts are from the accepted score snapshot. Raw errors and full traces were not freshly inspected; no new experiment was run. Task unknowns and metric nulls remain separate and are retained in fixed denominators."
90
+ }
results/kimi-k3/export-verification.json ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "status": "passed_offline",
3
+ "snapshot_hash_matches": true,
4
+ "task_records_match_accepted_package": true,
5
+ "metrics_recomputed": 18,
6
+ "paired_rows": 1000,
7
+ "per_arm": {
8
+ "baseline": {
9
+ "records": 1000,
10
+ "unique_task_ids": 1000,
11
+ "matches_accepted_export": true,
12
+ "counts": {
13
+ "succeeded": 520,
14
+ "failed": 475,
15
+ "unknown": 5,
16
+ "unsealed": 0,
17
+ "not_started": 0,
18
+ "invalid_evidence": 0,
19
+ "completed": 995,
20
+ "planned": 1000
21
+ },
22
+ "origin_counts": {
23
+ "child": 956,
24
+ "parent": 44
25
+ }
26
+ },
27
+ "deepagents": {
28
+ "records": 1000,
29
+ "unique_task_ids": 1000,
30
+ "matches_accepted_export": true,
31
+ "counts": {
32
+ "succeeded": 845,
33
+ "failed": 151,
34
+ "unknown": 4,
35
+ "unsealed": 0,
36
+ "not_started": 0,
37
+ "invalid_evidence": 0,
38
+ "completed": 996,
39
+ "planned": 1000
40
+ },
41
+ "origin_counts": {
42
+ "child": 960,
43
+ "parent": 40
44
+ }
45
+ }
46
+ },
47
+ "raw_error_root_cause_verified": false,
48
+ "official_parity_verified": false
49
+ }
results/kimi-k3/paired_comparison.csv ADDED
The diff for this file is too large to render. See raw diff
 
results/kimi-k3/provenance.json ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "kimi-k3",
3
+ "exported_at": "2026-09-16T15:11:54.404624+00:00",
4
+ "result_snapshot_sha256": "eff050fe7977423fdffc731fc62b9ea45652743d2f3184931ae70c3b9365d176",
5
+ "accepted_results_sha256": "dd72f352be1895b11914dd2aea1fdceb21c453932877c588ecaf81ec09861105",
6
+ "accepted_task_records_sha256": "32a797b4da8dfe90d44b37d7715fe726a29c5b4f93a1b24d7d5ffbffb81e9749",
7
+ "source_commits": [
8
+ "4afeaefac6183dc7111c7c02a056c9bf1e40b462",
9
+ "9d6a89eb6c4e42fab8c14b768d1bc4e40f15f980"
10
+ ],
11
+ "dataset": {
12
+ "repo_id": "CamoAiLab/InferenceNet",
13
+ "revision": "59f9512a38e594528807744214a60ee00367434e",
14
+ "task_directory": "Selected_1000",
15
+ "task_list": "Selected_1000/1000_new.csv",
16
+ "prior_results_imported": false
17
+ },
18
+ "selected_task_count": 1000,
19
+ "record_count": 2000,
20
+ "official_parity_verified": false,
21
+ "new_model_calls": 0,
22
+ "new_benchmark_executions": 0,
23
+ "historical_native_audit": {
24
+ "mode": "fresh_in_accepted_snapshot",
25
+ "collection_interval": {
26
+ "started_at": "2026-09-16T06:46:01.177744+00:00",
27
+ "finished_at": "2026-09-16T07:43:33.877909+00:00"
28
+ },
29
+ "input_blobs_rehashed": false,
30
+ "note": "This release only rechecks the accepted local snapshot; it is not a fresh native audit."
31
+ },
32
+ "verification_scope": "Offline hash, task identity, accepted-export equality and reaggregation of sealed per-task flags; not fresh prediction scoring or native artifact audit.",
33
+ "related_results": {
34
+ "gpt_5_6_sol": "https://huggingface.co/datasets/YICHEN013/InferenceNet-GPT-5.6-Sol-results"
35
+ },
36
+ "excluded": [
37
+ "benchmark inputs",
38
+ "reference answers",
39
+ "raw predictions unavailable in score snapshot",
40
+ "full prompts and model/API traces",
41
+ "credentials",
42
+ "internal endpoints",
43
+ "machine filesystem paths"
44
+ ]
45
+ }
results/manifest.sha256 CHANGED
@@ -1,4 +1,4 @@
1
- a81742c1a2a8c7667409f699b7aac78c4318141bcf6b3d33b3e3980d2976a61e README.md
2
  2ad22f4778525b1460f1857e1095f2eea5620063fba1ffa6314ca7fd50f14703 claude-opus-4-8/baseline/README.md
3
  8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56 claude-opus-4-8/baseline/config.json
4
  4d5f859e472b79195c9be6b1f0e9c5feecf59a81dfe32ac08eb4c1d4967ed98e claude-opus-4-8/baseline/results.csv
@@ -9,6 +9,20 @@ dfb00967c2a2ff320d5a6d8c142cbc6ef37d4fb144abcc7a451673515c2b4c50 claude-opus-4-
9
  bafc0a72cd4d8f53c1da33030250171f36ce441488bc8456cba8c3f52b80e6bd claude-opus-4-8/deepagents/results.csv
10
  3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a claude-opus-4-8/deepagents/results.jsonl
11
  f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614 claude-opus-4-8/deepagents/summary.json
 
 
 
 
 
 
 
 
 
 
 
 
 
 
12
  ed22b41edf2afe7a500f4caf1b61f0b4c6e6c73759522c543fb86ce478eb4d04 gpt-5.6-sol/baseline/README.md
13
  90ec4e0781b57498fef52de31eacf5c8ec9424549643865a7140a1ccdb04c1d4 gpt-5.6-sol/baseline/config.json
14
  a54171a3d3e7ecda3e61311414a84131e74a680bfa1342e8627cfaa32fc7c873 gpt-5.6-sol/baseline/results.csv
@@ -19,4 +33,19 @@ a3e24175f2395cffbbcd6d9d3c89689bd3a3b22dd4ce61bb78d5da33c6b6cc1b gpt-5.6-sol/de
19
  27abe57119d4be49c56b29c318a1bc2757a97f44427764d7fc8d188be4d10671 gpt-5.6-sol/deepagents/results.csv
20
  378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be gpt-5.6-sol/deepagents/results.jsonl
21
  821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5 gpt-5.6-sol/deepagents/summary.json
22
- 72096fb6d8744ef0cbc32e6b169627a7a0f936e2902dad9ad42857ec8f3f1ec3 index.json
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ f09b15f887f461339ba94542280056bbc9c519eeb3128c64cbab5a00390e7e2e README.md
2
  2ad22f4778525b1460f1857e1095f2eea5620063fba1ffa6314ca7fd50f14703 claude-opus-4-8/baseline/README.md
3
  8058afcd14c34d666277c0f32ac5962398b10618df5b1aa0e13cd9862344dc56 claude-opus-4-8/baseline/config.json
4
  4d5f859e472b79195c9be6b1f0e9c5feecf59a81dfe32ac08eb4c1d4967ed98e claude-opus-4-8/baseline/results.csv
 
9
  bafc0a72cd4d8f53c1da33030250171f36ce441488bc8456cba8c3f52b80e6bd claude-opus-4-8/deepagents/results.csv
10
  3bef3b9850f24b9cb0b5806efd8425c4e4cc650735a6b28b45befdd55a6ec28a claude-opus-4-8/deepagents/results.jsonl
11
  f86da9c76f8becb75b4970c486160013209a6aa4239923018b217e55f1b0a614 claude-opus-4-8/deepagents/summary.json
12
+ ac7253c934e43a88d0a2f6b105d1bb13424ef004e6838d2a632d5e9c366c72f2 gemini-3.1-pro-preview/baseline/README.md
13
+ ff9884d2095695a1f40f5fc573982b9a49fd53d41dca2a6fc36d04dc81f10998 gemini-3.1-pro-preview/baseline/config.json
14
+ fc708f327a76b22ebe6f453206b0c42e4cf46176afb8cb12eda2c1b71fe90bde gemini-3.1-pro-preview/baseline/results.csv
15
+ 2f1e2019157d105efee7ba311a49eba99abe4a7b84f0a4b4e0ac174e2d56cdba gemini-3.1-pro-preview/baseline/results.jsonl
16
+ a8ba5100a28f1f0aeae6949eb3669c7717c844c4e2eee6fa42ded351aef6b5c7 gemini-3.1-pro-preview/baseline/summary.json
17
+ 49ff37264a63f3b5615632213bb9248d2b6b8feb66b4ec989ed655b45e5faed2 gemini-3.1-pro-preview/deepagents/README.md
18
+ f09d9d99d26328ae2f54e9173b22f69d68f7d7bb50c3cc5e352c996b98dd57f9 gemini-3.1-pro-preview/deepagents/config.json
19
+ af16c74fdc1163682a9a98f597ac0d101bc0de64d60c6a8582f69312948c8069 gemini-3.1-pro-preview/deepagents/results.csv
20
+ c2824f8277b283a6dd60118b6e6a0cb0419e4c8b9193de15a7bcf7aaaad24181 gemini-3.1-pro-preview/deepagents/results.jsonl
21
+ bf9ce721c94103f91b0ca576f717f4f5dca7cc9965601658a65fc6d2a15b7045 gemini-3.1-pro-preview/deepagents/summary.json
22
+ 67b24d52eeabbcd84baf00c7395e72d4efe7cfcc82d7de1b0f8ee30f8018a1f7 gemini-3.1-pro-preview/diagnostics.json
23
+ 885b4de7e83334c416f33ea1535a72f1f70dfe61834d1b622e01ddb3cccfaed4 gemini-3.1-pro-preview/export-verification.json
24
+ 930a9ae67dc0dd4210e7e0c96098cab24899c765e4b3d041f75455f58fc072fe gemini-3.1-pro-preview/paired_comparison.csv
25
+ 84378a81ca5533439b21dc1cf683925f5accc23532d392d140a4477e59c4edfb gemini-3.1-pro-preview/provenance.json
26
  ed22b41edf2afe7a500f4caf1b61f0b4c6e6c73759522c543fb86ce478eb4d04 gpt-5.6-sol/baseline/README.md
27
  90ec4e0781b57498fef52de31eacf5c8ec9424549643865a7140a1ccdb04c1d4 gpt-5.6-sol/baseline/config.json
28
  a54171a3d3e7ecda3e61311414a84131e74a680bfa1342e8627cfaa32fc7c873 gpt-5.6-sol/baseline/results.csv
 
33
  27abe57119d4be49c56b29c318a1bc2757a97f44427764d7fc8d188be4d10671 gpt-5.6-sol/deepagents/results.csv
34
  378cdec8e7a880f091ba442649d1c034997acefdc623898b60a063baad29d7be gpt-5.6-sol/deepagents/results.jsonl
35
  821ca6e55a66006c7fb958951a7859a9dccde29d87de8415d5c3f07f523563c5 gpt-5.6-sol/deepagents/summary.json
36
+ 2b6c21661ef74f692d64095f1e3fd64f9e696cd9876b5186810ed340aff0c24e index.json
37
+ 29825997a250cc8fe292450006b82a392fe90c8e0493a1335f74eea72c1de104 kimi-k3/baseline/README.md
38
+ 7191c406e8695a5cb760ee8fa99c4fb0dd71bca59347acc4d6d33fe0630be081 kimi-k3/baseline/config.json
39
+ f87d2dd3b390695f6b420a78e126b504abc3d292180a3443deaead174ec1b464 kimi-k3/baseline/results.csv
40
+ 0502ae8892dd6c66ee92890a83c678527cff6c3396aa4c13faaadcf41e1857c1 kimi-k3/baseline/results.jsonl
41
+ 2edd2e96f3b448ceceb6189edb7f23c737b245c00c9820728c69f5237feb34b7 kimi-k3/baseline/summary.json
42
+ d6fe2ba624876d49373d4a0c180e0c27067405501b247e59c2d8ca5c74427ff7 kimi-k3/deepagents/README.md
43
+ 312c9260c01d5988878e4f2b88da55a16362ee994e8217c63b15ba772da3857c kimi-k3/deepagents/config.json
44
+ 572c3270a3ec255544b587aada9d8bd56cc2555d5246ce26feea2a3d35eb8f94 kimi-k3/deepagents/results.csv
45
+ bdab99f8f1d81885f74326cf9c8c8e57b9bb8a8d616bfbb1b5a2f6bf32b737ca kimi-k3/deepagents/results.jsonl
46
+ 81c8cabaa58acc65ef6afba792241b70951b666ab90c3f7c25194b61b0d9a228 kimi-k3/deepagents/summary.json
47
+ bfaac06563fc53d5f564197fdfb69e24ddbdff622b85d67bbf2a13fb7fb5d9ba kimi-k3/diagnostics.json
48
+ b96e0cc3352efac6140c33cef30873e522db547891101c3c49c6aca42085becb kimi-k3/export-verification.json
49
+ 3fe1378a23cfe59a37558521ec5bf47ea299a6ad4903a14684147a1b8066ee03 kimi-k3/paired_comparison.csv
50
+ 4a0f6dde38c5410a6f379192f30eac6d330403c80fcdbc2b357dbe15427a331b kimi-k3/provenance.json
51
+ 06ed669caaba3ca5e252d12af3de32223019d50aefd9c46751ca3cb35110fce2 verify_kimi_gemini.py
results/verify_kimi_gemini.py ADDED
@@ -0,0 +1,63 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Offline validation for the four added Kimi/Gemini result groups."""
2
+ import csv
3
+ import hashlib
4
+ import json
5
+ from collections import Counter
6
+ from pathlib import Path
7
+
8
+ root=Path(__file__).resolve().parent
9
+ for line in (root/'manifest.sha256').read_text().splitlines():
10
+ expected,name=line.split(' ',1)
11
+ p=(root/name).resolve()
12
+ assert p.is_relative_to(root) and p.is_file(),name
13
+ assert hashlib.sha256(p.read_bytes()).hexdigest()==expected,name
14
+ rows_by_model={}
15
+ checked=0
16
+ for model in ('kimi-k3','gemini-3.1-pro-preview'):
17
+ rows_by_model[model]={}
18
+ for arm in ('baseline','deepagents'):
19
+ p=root/model/arm
20
+ rows=[json.loads(line) for line in (p/'results.jsonl').read_text().splitlines()]
21
+ config=json.loads((p/'config.json').read_text())
22
+ summary=json.loads((p/'summary.json').read_text())
23
+ assert len(rows)==1000 and len({r['task_id'] for r in rows})==1000
24
+ assert [r['task_id'] for r in rows]==config['dataset']['task_ids']
25
+ assert all(r['model']==model and r['arm']==arm for r in rows)
26
+ assert all(r['resolved'] is True and r['definite_outcome']==(r['state'] in ('succeeded','failed')) for r in rows)
27
+ assert all(r['prediction_export_status']=='not_available_in_accepted_score_snapshot' for r in rows)
28
+ assert all(r[k] is None for r in rows for k in ('prediction_coefficient','prediction_standard_error','prediction_p_value'))
29
+ with (p/'results.csv').open() as f:csv_rows=list(csv.DictReader(f))
30
+ assert len(csv_rows)==1000
31
+ for j,c in zip(rows,csv_rows):
32
+ assert set(j)==set(c)
33
+ assert all(c[k]==('' if v is None else str(v)) for k,v in j.items())
34
+ assert summary['official_parity_verified'] is False and summary['included_in_displayed_leaderboard'] is False
35
+ counts=summary['result']['task_counts']
36
+ assert Counter(r['state'] for r in rows)=={k:counts[k] for k in ('succeeded','failed','unknown') if counts[k]}
37
+ assert sum(r['state']=='unknown' for r in rows)==summary['result']['task_unknown_count']
38
+ for metrics,prefix in [(summary['result']['metrics'],''),(summary['legacy_local_paper']['metrics'],'local_')]:
39
+ for name,m in metrics.items():
40
+ flags=[r[prefix+name] for r in rows]
41
+ assert all(v is None or type(v) is bool for v in flags)
42
+ good,bad,unknown=sum(v is True for v in flags),sum(v is False for v in flags),sum(v is None for v in flags)
43
+ assert (m['count'],m['failure_count'],m['unknown_count'],m['denominator'])==(good,bad,unknown,1000)
44
+ assert m['score']==round(good/10,1) and m['rate']==good/1000
45
+ assert m['assessable']==good+bad and m['coverage_percent']==round((good+bad)/10,1)
46
+ checked+=1
47
+ rows_by_model[model][arm]=rows
48
+ b,d=rows_by_model[model]['baseline'],rows_by_model[model]['deepagents']
49
+ assert [r['task_id'] for r in b]==[r['task_id'] for r in d]
50
+ with (root/model/'paired_comparison.csv').open() as f:pairs=list(csv.DictReader(f))
51
+ assert len(pairs)==1000
52
+ for i,pair in enumerate(pairs):
53
+ assert int(pair['task_id'])==b[i]['task_id']
54
+ for arm in ('baseline','deepagents'):
55
+ r=rows_by_model[model][arm][i]
56
+ assert pair[arm+'_result_sha256']==r['selected_result_sha256']
57
+ assert pair[arm+'_state']==r['state']
58
+ for field,value in r.items():
59
+ key=arm+'_'+field if field.startswith('local_') else arm+'_hf_'+field
60
+ if key in pair:assert pair[key]==('' if value is None else str(value))
61
+ assert checked==36
62
+ print('PASS: archive manifest; 4 x 1000 unique task records; paired rows; 36 metric aggregates; unknowns retained.')
63
+ print('Offline export verification only; not fresh raw prediction scoring or official scorer parity.')