Refresh leaderboard with single-pass Agent and DeepAgents comparisons
The Space now displays six archived models in paired, single-pass Agent, and Agent + Harness views (12 groups / 12,000 task records). It adds English/Chinese switching, separate scoring profiles, sorting, model search, CSV export and pinned evidence links.
Results remain explicitly provisional: official scorer parity is unverified, budgets differ, and unknown/invalid evidence stays in the 1,000-task denominator. All archived experiment files and the original results.csv stay unchanged; the original display is preserved at legacy.html.
Validation: 60 displayed metrics checked against accepted summaries; 56 also recomputed from archived task flags. All 15 view/metric combinations, three sorts, search, language persistence, mobile layout, load failure/retry and malformed-data rejection verified in-browser. No inference or rescoring performed.