File size: 11,954 Bytes
f8be768
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<script>try{var t=localStorage.getItem("amb-theme")||(matchMedia("(prefers-color-scheme: dark)").matches?"dark":"light");document.documentElement.setAttribute("data-theme",t)}catch(e){}</script>
<title>Submit &amp; reproduce — agent-memory-bench</title>
<meta name="description" content="How to enter a memory product into agent-memory-bench, the rules every arm plays by, and how to re-run and verify the benchmark yourself.">
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 16 16'%3E%3Crect width='16' height='16' fill='%23131311'/%3E%3Crect x='3' y='3' width='4' height='4' fill='%23fbfbf9'/%3E%3C/svg%3E">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Newsreader:ital,opsz,wght@0,6..72,400..700;1,6..72,400..700&family=Spline+Sans+Mono:wght@400..700&display=swap" rel="stylesheet">
<link rel="stylesheet" href="styles.css">
</head>
<body>

<header class="masthead">
  <div class="masthead-inner">
    <a class="wordmark" href="index.html" aria-label="agent-memory-bench home">
      <span class="tick" aria-hidden="true"></span>
      <span><strong class="long">agent-memory-bench</strong><strong class="short">AMB</strong><span class="full"> / AMB</span></span>
    </a>
    <nav aria-label="Primary">
      <a href="index.html">Overview</a>
      <a href="method.html">Method</a>
      <a href="leaderboard.html">Leaderboard</a>
      <a href="submit.html" aria-current="page">Submit</a>
      <a href="https://github.com/GiulioDER/agent-memory-bench">GitHub&nbsp;&#8599;</a>
      <button class="theme-toggle" type="button" aria-label="Switch between light and dark"><span class="sw" aria-hidden="true"></span></button>
    </nav>
  </div>
</header>

<main class="sheet">

  <section class="page-head">
    <span class="kicker reveal">Submit &amp; reproduce</span>
    <h1 class="reveal r2">Same feed. Same gate.<br><em>Your</em> write path.</h1>
    <p class="dek reveal r3">Any memory product with a published Claude Code integration can
    enter. The eight rules below are the ones every existing arm plays by, and the harness
    enforces most of them mechanically.</p>
  </section>

  <section class="pad-y-s" id="rules">
    <div class="section-head">
      <span class="no">01</span>
      <h2>The rules</h2>
      <span class="aside">enforced by the harness where possible</span>
    </div>

    <ol class="rulebook mt-s">
      <li>
        <div>
          <h3>Enter through your published integration</h3>
          <p>A product competes as its own shipped Claude Code integration: a plugin, an MCP
          server, or lifecycle hooks. No bespoke benchmark builds, no unreleased branches. If
          your users cannot install it, the benchmark does not run it.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>One adapter directory, hash-pinned</h3>
          <p>Everything an arm is lives in <code>adapters/&lt;name&gt;/</code>: the adapter
          implementing the <code>MemoryAdapter</code> contract, version pins, and
          <code>config.frozen.json</code> whose hash is recorded before the run. A config change
          after the freeze means a new run, not an amended one.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>Ingest the neutral feed as it is</h3>
          <p>Every adapter receives identical bytes: verbatim recorded session transcripts,
          pinned by a sha256 manifest. Your extraction pipeline decides what to keep, and that
          decision is part of what is measured. No task-specific tuning, no peeking at the task
          suite.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>Prove the treatment or lose the cell</h3>
          <p>The <a href="method.html#gate">admission gate</a> discards any cell where the arm
          cannot prove it was actually applied: MCP tools listed at session init, hooks fired
          with output, sandbox files digest-verified, no cross-arm contamination. Discard counts
          are published per arm.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>Real credentials or an honest refusal</h3>
          <p>Arm credentials come from <code>.env</code>, never from the repository. An absent
          key makes the harness refuse the arm at startup rather than fake it. Local arms run
          via docker compose, with their extraction LLM traffic metered through the harness
          proxy so ingestion tokens are counted.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>Vendor review before the run</h3>
          <p>Every vendor is publicly invited to review their adapter and frozen config before
          any measured run. The invitation, the response, or the documented silence is committed
          in <code>adapters/&lt;name&gt;/VENDOR_REVIEW.md</code>. Silence does not block the
          run; it is simply on the record.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>Preregistration binds the run</h3>
          <p>The question, predictions, endpoints, exclusion rules and sizing are committed
          under <code>preregistration/</code> before the first session. The run scripts refuse
          to start while that directory is dirty. Numbers in a committed preregistration are
          never edited; corrections are appended.</p>
        </div>
      </li>
      <li>
        <div>
          <h3>Results publish in full, win or lose</h3>
          <p>Per-session logs, streams, admission verdicts and the complete cost ledger land in
          <code>results/&lt;run_id&gt;/</code>. There is no private preview and no retraction
          path: a preregistered run that embarrasses an arm, including the authors' own, is
          published like any other.</p>
        </div>
      </li>
    </ol>
  </section>

  <section class="pad-y-s" id="add">
    <div class="section-head">
      <span class="no">02</span>
      <h2>Adding your product</h2>
      <span class="aside">a pull request, not a form</span>
    </div>

    <div class="prose mt-s">
      <p>Submission is a pull request against the repository. It should contain, and review
      will check for, exactly four things:</p>
    </div>

    <div class="table-scroll mt-s">
      <table>
        <thead><tr><th>file</th><th>contents</th></tr></thead>
        <tbody>
          <tr><td><span class="m">adapters/&lt;name&gt;/adapter.py</span></td>
              <td>implements the <code>MemoryAdapter</code> contract in <code>harness/adapters/base.py</code>: ingest the feed, install the integration, tear down cleanly</td></tr>
          <tr><td><span class="m">adapters/&lt;name&gt;/config.frozen.json</span></td>
              <td>the exact configuration the run uses, hash-pinned; defaults your users would get, not a tuned special</td></tr>
          <tr><td><span class="m">adapters/&lt;name&gt;/VENDOR_REVIEW.md</span></td>
              <td>from the template; records the review invitation and its outcome</td></tr>
          <tr><td><span class="m">adapters/&lt;name&gt;/pins</span></td>
              <td>package and image versions, exact; the run must be reconstructible from them</td></tr>
        </tbody>
      </table>
    </div>

  </section>

  <section class="pad-y-s" id="rerun">
    <div class="section-head">
      <span class="no">03</span>
      <h2>Re-run and test</h2>
      <span class="aside">trust nothing, execute everything</span>
    </div>

    <div class="prose mt-s">
      <p>The harness, the tasks, the checkers and both reference solutions are in the open
      repository. Verifying the benchmark's own claims takes one command:</p>
    </div>

<pre><code><span class="c"># clone, then: harness self-tests, task validation,</span>
<span class="c"># and the CI assertion that every naive reference fails</span>
<span class="c"># and every informed reference passes</span>
python -m pytest tests/ -q</code></pre>

    <div class="prose">
      <p>A real measured run additionally needs:</p>
    </div>

    <div class="table-scroll mt-s">
      <table>
        <thead><tr><th>requirement</th><th>why</th></tr></thead>
        <tbody>
          <tr><td><span class="m">Claude Code CLI ≥ 2.1.221</span></td>
              <td>below that, a pending MCP server runs the session without its tools while reporting success; the admission gate exists because that happened</td></tr>
          <tr><td><span class="m">.env credentials</span></td>
              <td>per-arm keys from <code>.env.example</code>; a missing key refuses that arm rather than faking it</td></tr>
          <tr><td><span class="m">clean preregistration/</span></td>
              <td>the run scripts refuse to start while the preregistration directory is dirty</td></tr>
          <tr><td><span class="m">explicit prices</span></td>
              <td><code>--price-in</code>, <code>--price-out</code> and <code>--price-as-of</code> are required for a live run and have no defaults anywhere, because three runners once carried three different ones and none matched the frozen rates. Dry runs need none</td></tr>
          <tr><td><span class="m">docker</span></td>
              <td>for the self-hosted arms. <code>docker/compose.yaml</code> brings up the vector database and the harness image, and starts no memory server, so one-command full reproduction is not there yet</td></tr>
        </tbody>
      </table>
    </div>

    <div class="callout callout-invert mt-l">
      <span class="kicker">Reproduction is the product</span>
      <div class="prose">
        <p>If a published result cannot be regenerated from the pinned versions, the frozen
        configs, the sha256-pinned corpus and the committed preregistration, that is a bug in
        the benchmark and should be filed as one. Disagreement with a result starts with
        <code>results/&lt;run_id&gt;/</code>, which contains every session's logs, admission
        verdicts and costs.</p>
        <p>By that standard the benchmark half fails its own rule, and says so here rather than
        in a footnote. The product arm is now pinned to a released package and its frozen config
        names environment variables instead of one machine's paths, so it no longer has to be
        edited to run elsewhere. What is still missing: the published runs resolved that package
        from a local checkout, there is no <code>versions.lock</code>, and the compose stack
        starts a database but no memory server, so a reader must still supply Postgres, an
        embedding key and a built index. Tracked in
        <a href="https://github.com/GiulioDER/agent-memory-bench/blob/master/docs/STATUS.md">docs/STATUS.md</a>.</p>
      </div>
    </div>
  </section>

</main>

<footer>
  <div class="footer-inner">
    <div>
      <div class="foot-title">agent-memory-bench</div>
      <div class="dim">A preregistered, execution-graded benchmark of pluggable memory layers
      for coding agents. Apache-2.0. Built in the open; results published win or lose.</div>
    </div>
    <div>
      <div class="foot-title">Pages</div>
      <a href="method.html">Method</a><br>
      <a href="leaderboard.html">Leaderboard</a><br>
      <a href="submit.html">Submit &amp; reproduce</a>
    </div>
    <div>
      <div class="foot-title">Source</div>
      <a href="https://github.com/GiulioDER/agent-memory-bench">Repository</a><br>
      <a href="https://github.com/GiulioDER/agent-memory-bench/tree/master/preregistration">Preregistrations</a><br>
      <a href="https://github.com/GiulioDER/agent-memory-bench/blob/master/adapters/VENDOR_REVIEW_TEMPLATE.md">Vendor review</a>
    </div>
  </div>
</footer>

<script src="site.js"></script>
</body>
</html>