Archive Qwen3.7-Max and DeepSeek V4 Pro paired results
Add baseline and DeepAgents archives for Qwen3.7-Max and DeepSeek V4 Pro: four groups of 1,000 tasks, completing six models and 12 groups in the result archive. Include task JSONL/CSV, separate HF four-dimensional and local-paper summaries, protocol/source provenance, paired tables and an offline verifier.
Preserve Qwen DeepAgents task-level unknowns (6) and invalid evidence (1, task 453), without replacement or denominator changes. DeepSeek has 1,000 definite outcomes in each arm. Invalid evidence is explicitly not a sealed or trusted result; original predictions are unavailable in the accepted score snapshot. All 4,000 rows match accepted flags, states and result hashes; 36 metric aggregates were rechecked.
This archives historical amended comparisons, with unequal model-call budgets and unverified official scorer parity. GPT-5.5 is excluded. Existing model archives, root results.csv and page code are unchanged. No benchmark inputs, hidden answers, full model traces, credentials or machine paths are included.