Refresh leaderboard with single-pass Agent and DeepAgents comparisons

#4
by YICHEN013 - opened
Lab for AI Agents in Business and Economics of HKU CAMO org Hugging Face CLI

The Space now displays six archived models in paired, single-pass Agent, and Agent + Harness views (12 groups / 12,000 task records). It adds English/Chinese switching, separate scoring profiles, sorting, model search, CSV export and pinned evidence links.

Results remain explicitly provisional: official scorer parity is unverified, budgets differ, and unknown/invalid evidence stays in the 1,000-task denominator. All archived experiment files and the original results.csv stay unchanged; the original display is preserved at legacy.html.

Validation: 60 displayed metrics checked against accepted summaries; 56 also recomputed from archived task flags. All 15 view/metric combinations, three sorts, search, language persistence, mobile layout, load failure/retry and malformed-data rejection verified in-browser. No inference or rescoring performed.

YICHEN013 changed pull request status to merged
Lab for AI Agents in Business and Economics of HKU CAMO org

Sign up or log in to comment