Add DeepSeek API + DSH results and independent leaderboard entry
Adds the completed 1,000-task DeepSeek API + DeepSeek Harness (DSH) research run and its leaderboard entry. The existing twelve baseline / DeepAgents archives and six paired comparisons are preserved. The Harness view gains a seventh, independently labelled DSH entry; merging this PR updates both the results archive and the displayed leaderboard.
| Metric | Passed / 1,000 | Score | Unknown |
|---|---|---|---|
| Execution success | 699 | 69.9% | 30 |
| Partial replication โ coefficient only | 556 | 55.6% | 302 |
| Coefficient direction | 683 | 68.3% | 302 |
| Significance category | 598 | 59.8% | 304 |
| Full replication โ separate local-paper-v1 metric | 372 | 37.2% | 305 |
All rates retain the full 1,000-task denominator. The selected outcomes are 699 scored, 271 failed and 30 uncertain. The four HF-style metrics and the local full-replication metric have distinct definitions; 37.2% is not an average of the other four.
The fixed selection contains 913 original attempts and 87 preselected technical replacements, selected regardless of outcome. The replacements include 60 infrastructure cases and 27 input-schema cases; the latter change the input-profile prompt. All 1,087 attempts are retained in cost accounting, including 38 attempts with unknown usage. No best-of selection or additional model calls were performed for publication.
The request alias and observed API response label are deepseek-v4-pro. The underlying weights are unverified because the provider's model/routing descriptions conflict. Official scorer parity and complete historical environment/reasoning parity are also unverified. DSH is therefore excluded from the baseline / DeepAgents paired-lift calculation. The archive contains the actual SDK/runtime pins, code commits, budgets, task population, amendments, per-task predictions, status flags and source hashes.
The archive is committed first; every new page evidence link is pinned to that real commit:
Validation: all 1,000 selected rows matched sealed source evidence; 4,000 HF-style flags, 5,000 local flags and exported prediction values were compared. All 36 existing result/configuration/summary files stayed byte-identical. Independent public-package and cost reviews passed; seven builder tests, seven interface-logic tests and fifty desktop/mobile bilingual browser checks passed. Hidden reference answers, input datasets, raw API conversations, credentials and local machine paths are excluded.