File size: 12,651 Bytes
03c1286 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 | <!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width,initial-scale=1">
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 32 32'%3E%3Crect width='32' height='32' rx='6' fill='%230c1219'/%3E%3Ctext x='5' y='23' font-family='sans-serif' font-size='21' fill='%2378dbc0'%3EiN%3C/text%3E%3C/svg%3E">
<meta name="color-scheme" content="light dark">
<meta name="description" content="How single-pass execution, DeepAgents and DeepSeek Harness approach an InferenceNet task.">
<title>Agent / Harness — InferenceNet Challenge</title>
<link rel="stylesheet" href="./leaderboard.css">
<link rel="stylesheet" href="./portal.css">
<script type="module" src="./portal.js">
</script>
</head>
<body>
<a class="skip" href="#main">
<span data-en="Skip to content" data-zh="跳转到正文" >Skip to content</span>
</a>
<div class="page">
<nav class="topbar" aria-label="Main navigation">
<a class="brand" href="./index.html">InferenceNet <small>CHALLENGE</small>
</a>
<div class="nav-links">
<a href="./index.html">
<span data-en="Home" data-zh="主页" >Home</span>
</a>
<a href="./data.html">
<span data-en="Data" data-zh="数据" >Data</span>
</a>
<a href="./leaderboard.html">
<span data-en="Leaderboard" data-zh="榜单" >Leaderboard</span>
</a>
<a href="./agent.html" aria-current="page">
<span data-en="Agent / Harness" data-zh="Agent / Harness" >Agent / Harness</span>
</a>
<button id="language" type="button" aria-label="Switch to Chinese">中文</button>
<button id="theme" type="button" aria-pressed="false">
<span class="theme-icon" aria-hidden="true">
</span>
<span id="theme-label">Light</span>
</button>
</div>
</nav>
<main id="main">
<header class="portal-hero">
<p class="eyebrow">AGENT / HARNESS</p>
<h1 data-en="From one attempt to an execution loop." data-zh="从单次生成,到执行反馈循环。" >From one attempt to an execution loop.</h1>
<p data-en="The model proposes the analysis. The harness manages tools, feedback, budgets and final artifacts. InferenceNet makes the resulting process and output inspectable." data-zh="模型提出分析方案;Harness 管理工具、反馈、预算与最终产物。InferenceNet 让执行过程和结果都可以被检查。" class="portal-lead">The model proposes the analysis. The harness manages tools, feedback, budgets and final artifacts. InferenceNet makes the resulting process and output inspectable.</p>
</header>
<section class="portal-section">
<div class="portal-heading">
<p data-en="01 / EXECUTION MODES" data-zh=" 01 / EXECUTION MODES" class="eyebrow">01 / EXECUTION MODES</p>
<h2 data-en="The harness is part of the experiment." data-zh="Harness 是实验的一部分。" >The harness is part of the experiment.</h2>
</div>
<div class="profile-grid">
<article class="profile">
<p class="eyebrow">A / BASELINE</p>
<h3>Single-pass Agent</h3>
<div class="profile-budget">1 <span data-en="model call" data-zh="次模型调用" >model call</span>
</div>
<p data-en="Generate one program, then execute it. There is no iterative tool-feedback loop." data-zh="生成一个程序后执行,不进行多轮工具反馈循环。" >Generate one program, then execute it. There is no iterative tool-feedback loop.</p>
<a href="./leaderboard.html" class="text-link">
<span data-en="See published results →" data-zh="查看已发布结果 →" >See published results →</span>
</a>
</article>
<article class="profile">
<p class="eyebrow">B / AGENT LOOP</p>
<h3>DeepAgents</h3>
<div class="profile-budget">≤ 6 <span data-en="model calls" data-zh="次模型调用" >model calls</span>
</div>
<p data-en="Plan, inspect data, execute trial programs and revise. Up to four trial tool calls, followed by final execution." data-zh="规划、检查数据、试运行与修订;最多四次试运行工具调用,然后执行最终程序。" >Plan, inspect data, execute trial programs and revise. Up to four trial tool calls, followed by final execution.</p>
<a href="https://easonai-5589.github.io/inferencenet-challenge/agent.html" class="text-link" target="_blank" rel="noopener noreferrer">
<span data-en="Read the full walkthrough ↗" data-zh="阅读完整执行记录 ↗" >Read the full walkthrough ↗</span>
</a>
</article>
<article class="profile">
<p class="eyebrow">C / INDEPENDENT</p>
<h3>DeepSeek Harness · DSH</h3>
<div class="profile-budget">≤ 6 <span data-en="calls / attempt" data-zh="次调用 / 尝试" >calls / attempt</span>
</div>
<p data-en="Official DSH SDK. The published run uses technical recovery and stays separate from the baseline / DeepAgents paired comparison." data-zh="使用官方 DSH SDK;已发布实验包含技术补测,独立于单次生成 / DeepAgents 配对比较展示。" >Official DSH SDK. The published run uses technical recovery and stays separate from the baseline / DeepAgents paired comparison.</p>
<a href="https://huggingface.co/spaces/CamoAiLab/InferenceNet-Leaderboard/blob/6ea7c5684664bfe5c6e3d282a5e28f68244acd9e/results/deepseek-v4-pro/dsh/README.md" class="text-link" target="_blank" rel="noopener noreferrer">
<span data-en="Read the run protocol ↗" data-zh="查看运行协议 ↗" >Read the run protocol ↗</span>
</a>
</article>
</div>
<p data-en="These are the published run settings, not universal framework limits. Unequal total budgets and differing protocols do not support an equal-cost causal claim." data-zh="以上为已发布实验的设定,并非框架本身的限制。各组总预算与协议存在差异,不能据此作等成本的因果判断。" class="source-note">These are the published run settings, not universal framework limits. Unequal total budgets and differing protocols do not support an equal-cost causal claim.</p>
</section>
<section class="portal-section">
<div class="portal-heading">
<p data-en="02 / FEEDBACK" data-zh=" 02 / FEEDBACK" class="eyebrow">02 / FEEDBACK</p>
<h2 data-en="A loop with an explicit stopping point." data-zh="有明确停止条件的反馈循环。" >A loop with an explicit stopping point.</h2>
</div>
<div class="agent-loop">
<ol class="workflow">
<li>
<span class="step-number">01</span>
<h3 data-en="Read the task" data-zh="读取任务" >Read the task</h3>
<p data-en="A research specification and its source data." data-zh="研究要求与对应的数据文件。" >A research specification and its source data.</p>
</li>
<li>
<span class="step-number">02</span>
<h3 data-en="Write & run code" data-zh="编写并执行代码" >Write & run code</h3>
<p data-en="Fit the requested econometric model." data-zh="按照要求估计计量模型。" >Fit the requested econometric model.</p>
</li>
<li>
<span class="step-number">03</span>
<h3 data-en="Inspect & revise" data-zh="检查并修订" >Inspect & revise</h3>
<p data-en="A harness can use execution feedback." data-zh="Harness 可根据执行反馈修订程序。" >A harness can use execution feedback.</p>
</li>
<li>
<span class="step-number">04</span>
<h3 data-en="Recover the finding" data-zh="复现研究结果" >Recover the finding</h3>
<p data-en="Return the coefficient, standard error and p-value." data-zh="输出系数、标准误和 p 值。" >Return the coefficient, standard error and p-value.</p>
</li>
</ol>
<div class="loop-feedback">↶ <span data-en="Execution feedback returns to the model until it finalizes or reaches its budget." data-zh="执行反馈返回模型,直到模型提交最终程序或达到预算上限。" >Execution feedback returns to the model until it finalizes or reaches its budget.</span>
</div>
</div>
</section>
<section class="portal-section" id="case">
<div class="portal-heading">
<p data-en="03 / RECORDED CASE" data-zh=" 03 / RECORDED CASE" class="eyebrow">03 / RECORDED CASE</p>
<h2 data-en="Same task. Two execution paths." data-zh="同一道任务,两种执行路径。" >Same task. Two execution paths.</h2>
</div>
<p class="portal-lead">
<span data-en="Task 0011 asks for a monthly-clustered regression on Spain, 1905–1945. This recorded GPT-5.5 example shows why inspecting the data can change an outcome." data-zh="任务 0011 要求对 1905–1945 年的西班牙样本进行按月聚类的回归。这个已记录的 GPT-5.5 案例展示了检查数据如何影响执行结果。" >Task 0011 asks for a monthly-clustered regression on Spain, 1905–1945. This recorded GPT-5.5 example shows why inspecting the data can change an outcome.</span>
</p>
<div class="case-pair">
<article>
<p class="eyebrow">SINGLE-PASS</p>
<h3 data-en="A guessed schema stops execution." data-zh="对数据结构的猜测导致执行失败。" >A guessed schema stops execution.</h3>
<p data-en="The generated program searches for common month/date column names. It misses the dataset’s tid column and exits before producing a result." data-zh="生成的程序只搜索常见的月份 / 日期列名,遗漏数据中的 tid 列,因此未能输出结果。" >The generated program searches for common month/date column names. It misses the dataset’s tid column and exits before producing a result.</p>
<div class="case-outcome">1 <span data-en="model call · no valid result JSON" data-zh="次模型调用 · 未生成有效结果 JSON" >model call · no valid result JSON</span>
</div>
</article>
<article>
<p class="eyebrow">DEEPAGENTS</p>
<h3 data-en="Inspect, estimate, then finalize." data-zh="先检查,再估计,最后提交。" >Inspect, estimate, then finalize.</h3>
<p data-en="The agent inspects 38 columns, checks the time structure, then fits OLS with tid-clustered errors on 977 observations. Its final program recovers all three statistics." data-zh="Agent 检查了 38 列数据和时间结构,然后对 977 条观测进行按 tid 聚类的 OLS 估计。最终程序复现了三项统计量。" >The agent inspects 38 columns, checks the time structure, then fits OLS with tid-clustered errors on 977 observations. Its final program recovers all three statistics.</p>
<div class="case-outcome">4 <span data-en="model calls · 3 tool calls" data-zh="次模型调用 · 3 次工具调用" >model calls · 3 tool calls</span>
</div>
</article>
</div>
<div class="artifact">
<div>
<h3 data-en="The final artifact" data-zh="最终产物" >The final artifact</h3>
<p data-en="A program writes these three numeric fields. Execution logs and the evaluation record are stored separately." data-zh="程序写出这三个数值字段;执行日志与评分记录单独保存。" >A program writes these three numeric fields. Execution logs and the evaluation record are stored separately.</p>
</div>
<pre>
<code>{
"coefficient": 0.27273630753072026,
"standard_error": 0.054224701707830926,
"p_value": 6.899021793750105e-7
}</code>
</pre>
</div>
<p data-en="Illustrative internal run from 13 September 2026, not a result in the current leaderboard. One example does not establish an average harness benefit." data-zh="2026 年 9 月 13 日的内部执行示例,不计入当前榜单。单个案例不代表 Harness 的平均提升。" class="source-note">Illustrative internal run from 13 September 2026, not a result in the current leaderboard. One example does not establish an average harness benefit.</p>
<div class="actions">
<a href="https://easonai-5589.github.io/inferencenet-challenge/agent.html#walkthrough" class="text-link" target="_blank" rel="noopener noreferrer">
<span data-en="Inspect the complete trace ↗" data-zh="查看完整轨迹 ↗" >Inspect the complete trace ↗</span>
</a>
<a href="./leaderboard.html#protocol-details" class="text-link">
<span data-en="Scoring definitions & evidence →" data-zh="评分定义与证据 →" >Scoring definitions & evidence →</span>
</a>
</div>
</section>
</main>
<footer>
<span>InferenceNet <span class="muted">/</span> <span data-en="AI for empirical research." data-zh="面向实证研究的 AI。" >AI for empirical research.</span>
</span>
<div>
<a href="https://huggingface.co/datasets/CamoAiLab/InferenceNet" class="" target="_blank" rel="noopener noreferrer">
<span data-en="Dataset ↗" data-zh="数据集 ↗" >Dataset ↗</span>
</a>
<a href="https://easonai-5589.github.io/inferencenet-challenge/" class="" target="_blank" rel="noopener noreferrer">
<span data-en="Project website ↗" data-zh="项目官网 ↗" >Project website ↗</span>
</a>
<a href="https://easonai-5589.github.io/inferencenet-challenge/teams.html" class="" target="_blank" rel="noopener noreferrer">
<span data-en="Team & contact ↗" data-zh="团队与联系 ↗" >Team & contact ↗</span>
</a>
</div>
</footer>
</div>
</body>
</html>
|