Add DeepSeek API + DSH results and independent leaderboard entry

#6
by lry520111 - opened

Adds the completed 1,000-task DeepSeek API + DeepSeek Harness (DSH) research run and its leaderboard entry. The existing twelve baseline / DeepAgents archives and six paired comparisons are preserved. The Harness view gains a seventh, independently labelled DSH entry; merging this PR updates both the results archive and the displayed leaderboard.

Metric Passed / 1,000 Score Unknown
Execution success 699 69.9% 30
Partial replication โ€” coefficient only 556 55.6% 302
Coefficient direction 683 68.3% 302
Significance category 598 59.8% 304
Full replication โ€” separate local-paper-v1 metric 372 37.2% 305

All rates retain the full 1,000-task denominator. The selected outcomes are 699 scored, 271 failed and 30 uncertain. The four HF-style metrics and the local full-replication metric have distinct definitions; 37.2% is not an average of the other four.

The fixed selection contains 913 original attempts and 87 preselected technical replacements, selected regardless of outcome. The replacements include 60 infrastructure cases and 27 input-schema cases; the latter change the input-profile prompt. All 1,087 attempts are retained in cost accounting, including 38 attempts with unknown usage. No best-of selection or additional model calls were performed for publication.

The request alias and observed API response label are deepseek-v4-pro. The underlying weights are unverified because the provider's model/routing descriptions conflict. Official scorer parity and complete historical environment/reasoning parity are also unverified. DSH is therefore excluded from the baseline / DeepAgents paired-lift calculation. The archive contains the actual SDK/runtime pins, code commits, budgets, task population, amendments, per-task predictions, status flags and source hashes.

The archive is committed first; every new page evidence link is pinned to that real commit:

Validation: all 1,000 selected rows matched sealed source evidence; 4,000 HF-style flags, 5,000 local flags and exported prediction values were compared. All 36 existing result/configuration/summary files stayed byte-identical. Independent public-package and cost reviews passed; seven builder tests, seven interface-logic tests and fifty desktop/mobile bilingual browser checks passed. Hidden reference answers, input datasets, raw API conversations, credentials and local machine paths are excluded.

lry520111 changed pull request title from Archive DeepSeek API + DSH full-1000 results and recovery provenance to Add DeepSeek API + DSH results and independent leaderboard entry
YICHEN013 changed pull request status to merged

Sign up or log in to comment