Archive Qwen3.7-Max and DeepSeek V4 Pro paired results

#3
by YICHEN013 - opened
Lab for AI Agents in Business and Economics of HKU CAMO org Hugging Face CLI

Add baseline and DeepAgents archives for Qwen3.7-Max and DeepSeek V4 Pro: four groups of 1,000 tasks, completing six models and 12 groups in the result archive. Include task JSONL/CSV, separate HF four-dimensional and local-paper summaries, protocol/source provenance, paired tables and an offline verifier.

Preserve Qwen DeepAgents task-level unknowns (6) and invalid evidence (1, task 453), without replacement or denominator changes. DeepSeek has 1,000 definite outcomes in each arm. Invalid evidence is explicitly not a sealed or trusted result; original predictions are unavailable in the accepted score snapshot. All 4,000 rows match accepted flags, states and result hashes; 36 metric aggregates were rechecked.

This archives historical amended comparisons, with unequal model-call budgets and unverified official scorer parity. GPT-5.5 is excluded. Existing model archives, root results.csv and page code are unchanged. No benchmark inputs, hidden answers, full model traces, credentials or machine paths are included.

YICHEN013 changed pull request status to merged

Sign up or log in to comment