Spaces:
Running
Running
File size: 30,256 Bytes
f8be768 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 | <!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<script>try{var t=localStorage.getItem("amb-theme")||(matchMedia("(prefers-color-scheme: dark)").matches?"dark":"light");document.documentElement.setAttribute("data-theme",t)}catch(e){}</script>
<title>Method — agent-memory-bench</title>
<meta name="description" content="How agent-memory-bench works: neutral corpus, per-arm ingestion, sandboxed Claude Code sessions, an admission gate, execution grading, and preregistered analysis.">
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 16 16'%3E%3Crect width='16' height='16' fill='%23131311'/%3E%3Crect x='3' y='3' width='4' height='4' fill='%23fbfbf9'/%3E%3C/svg%3E">
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Newsreader:ital,opsz,wght@0,6..72,400..700;1,6..72,400..700&family=Spline+Sans+Mono:wght@400..700&display=swap" rel="stylesheet">
<link rel="stylesheet" href="styles.css">
</head>
<body>
<header class="masthead">
<div class="masthead-inner">
<a class="wordmark" href="index.html" aria-label="agent-memory-bench home">
<span class="tick" aria-hidden="true"></span>
<span><strong class="long">agent-memory-bench</strong><strong class="short">AMB</strong><span class="full"> / AMB</span></span>
</a>
<nav aria-label="Primary">
<a href="index.html">Overview</a>
<a href="method.html" aria-current="page">Method</a>
<a href="leaderboard.html">Leaderboard</a>
<a href="submit.html">Submit</a>
<a href="https://github.com/GiulioDER/agent-memory-bench">GitHub ↗</a>
<button class="theme-toggle" type="button" aria-label="Switch between light and dark"><span class="sw" aria-hidden="true"></span></button>
</nav>
</div>
</header>
<main class="sheet">
<section class="page-head">
<span class="kicker reveal">Method</span>
<h1 class="reveal r2">One neutral feed in.<br>One <em>executed</em> artifact out.</h1>
<p class="dek reveal r3">Every arm receives the same recorded experience, works the same
tasks in the same sandbox, and is graded by the same checkers. The only variable is the
memory layer. Everything below is enforced by the harness, not by convention.</p>
</section>
<section class="pad-y-s" id="pipeline">
<div class="section-head">
<span class="no">01</span>
<h2>The pipeline</h2>
<span class="aside">harness/</span>
</div>
<div class="pipeline mt-s" role="img" aria-label="Pipeline diagram: corpus feeds per-arm ingestion, then a sandboxed session, then the admission gate which either discards the cell or passes it to execution grading, and results land in the ledger.">
<svg viewBox="0 0 1040 176" xmlns="http://www.w3.org/2000/svg" fill="none">
<style>
.bx { stroke: var(--ink); fill: var(--paper); }
.tt { font: 600 13px "Spline Sans Mono", monospace; fill: var(--ink); }
.ts { font: 400 9.5px "Spline Sans Mono", monospace; fill: var(--ink-faint); }
.ln { stroke: var(--ink); }
.dsh { stroke: var(--ink); stroke-dasharray: 3 4; }
rect.dsh { fill: var(--paper); }
</style>
<rect class="bx" x="4" y="24" width="148" height="52"/>
<rect class="bx" x="180" y="24" width="148" height="52"/>
<rect class="bx" x="356" y="24" width="148" height="52"/>
<rect class="bx" x="532" y="24" width="148" height="52"/>
<rect class="bx" x="708" y="24" width="148" height="52"/>
<rect class="bx" x="884" y="24" width="152" height="52"/>
<text class="tt" x="18" y="46">CORPUS</text> <text class="ts" x="18" y="63">verbatim transcripts</text>
<text class="tt" x="194" y="46">INGEST</text> <text class="ts" x="194" y="63">each arm's write path</text>
<text class="tt" x="370" y="46">SESSION</text> <text class="ts" x="370" y="63">Claude Code, sandboxed</text>
<text class="tt" x="546" y="46">GATE</text> <text class="ts" x="546" y="63">proof of treatment</text>
<text class="tt" x="722" y="46">EXECUTE</text> <text class="ts" x="722" y="63">checker vs oracle</text>
<text class="tt" x="898" y="46">LEDGER</text> <text class="ts" x="898" y="63">score + full costs</text>
<line class="ln" x1="152" y1="50" x2="180" y2="50"/>
<line class="ln" x1="328" y1="50" x2="356" y2="50"/>
<line class="ln" x1="504" y1="50" x2="532" y2="50"/>
<line class="ln" x1="680" y1="50" x2="708" y2="50"/>
<line class="ln" x1="856" y1="50" x2="884" y2="50"/>
<path d="M175 46 L180 50 L175 54" class="ln"/>
<path d="M351 46 L356 50 L351 54" class="ln"/>
<path d="M527 46 L532 50 L527 54" class="ln"/>
<path d="M703 46 L708 50 L703 54" class="ln"/>
<path d="M879 46 L884 50 L879 54" class="ln"/>
<line class="dsh" x1="606" y1="76" x2="606" y2="122"/>
<path d="M602 117 L606 122 L610 117" class="ln"/>
<rect class="dsh" x="516" y="122" width="180" height="42"/>
<text class="tt" x="530" y="140">DISCARDED</text>
<text class="ts" x="530" y="156">counts published per arm</text>
</svg>
</div>
<div class="prose mt-m">
<p>The corpus is fed to every adapter as identical bytes. Each product ingests it through
its own published write path, so what its extraction pipeline keeps, and what it throws
away, is part of what is measured. The agent then works each task in a fresh sandbox with
that arm's integration installed and nothing else.</p>
<p>Before a session is scored it must pass the admission gate. After it, the artifact the
agent left behind is run against an executable checker whose oracle inputs the sandbox
never contained. The result, with every token spent on both ingestion and the session,
lands in one per-arm ledger.</p>
</div>
</section>
<section class="pad-y-s" id="arms">
<div class="section-head">
<span class="no">02</span>
<h2>The arms</h2>
<span class="aside">what each one removes</span>
</div>
<div class="prose mt-s">
<p>An arm is not a contestant. Each one exists to <strong>remove a competing
explanation</strong> for what the product does, so that a difference at the end has one
story left that fits it. Read down the table and the design is the argument: by the last
row, "the memory layer did it" is the only surviving account.</p>
</div>
<div class="table-scroll mt-s">
<table>
<thead>
<tr><th>arm</th><th>what it is</th><th>the explanation it removes</th></tr>
</thead>
<tbody>
<tr>
<td><span class="m">bare</span></td>
<td>no memory layer, no CLAUDE.md, nothing but the repository and the prompt</td>
<td>"the task was solvable anyway." It is the floor, it runs in every cell, and
<strong>damage is defined against it</strong>: an arm damaged a cell when it failed
what <code>bare</code> solved. Without it, harm has no referent</td>
</tr>
<tr>
<td><span class="m">placebo</span></td>
<td>project-shaped prose with no memory content, matched to the baseline bundle on
line count and whitespace tokens</td>
<td>"any extra context would have helped." If <code>placebo</code> moves the number,
the treatment was <em>volume of plausible text</em>, not retrieval. This is the
control most memory benchmarks omit, and omitting it is how context stuffing gets
reported as memory</td>
</tr>
<tr>
<td><span class="m">claude_md</span></td>
<td>a curated static instruction bundle, the file a careful team already maintains</td>
<td>"a memory product beats having nothing." The honest competitor is not nothing, it
is a good README. It costs zero tokens per query and never goes stale mid-session, so
<strong>a memory layer that cannot beat it has not earned its retrieval
budget.</strong> It is the baseline, and every delta is quoted against it</td>
</tr>
<tr>
<td><span class="m">fs_grep</span></td>
<td>the whole corpus on disk, and grep</td>
<td>"a memory product beats searching the transcripts." The cheap answer costs no
index, no embedding and no server. A product is worth its infrastructure only if it
beats this, and it is a <em>non-memory</em> retrieval baseline, so a gap here is
about the memory layer rather than about having the text at all</td>
</tr>
<tr>
<td><span class="m">recall</span></td>
<td>MCP server, 8 read and navigation tools including its reasoning and graph surface</td>
<td>product under test</td>
</tr>
<tr>
<td><span class="m">recall_prefetch</span></td>
<td>the same retrieval, run in the harness with the task prompt already in hand</td>
<td>"the product cannot find it." Query formulation is removed, so this is the
ceiling the live arm is reaching for. It is a reference track, never ranked</td>
</tr>
</tbody>
</table>
</div>
<div class="prose mt-m">
<p><strong>Write tools and write-side hooks are withheld.</strong> The product ships them
and would use them, but the corpus has to be frozen across arms and cells: a session that
wrote would change what the next seed reads, and one that solved a task could write its
answer where the next cell retrieves it. That limitation is stated in the scope line above
every ranking rather than in a footnote: <strong>this measures retrieval, not memory
formation</strong>, and gives no credit for extraction or consolidation at write time.</p>
<p>Each product carries <strong>its own shipped instruction</strong>. Equalising the text
measures a denominator no vendor ships; this measures what a user installs. Per-arm
instruction sizes are published with every run so the asymmetry is visible.</p>
</div>
</section>
<section class="pad-y-s" id="tasks">
<div class="section-head">
<span class="no">03</span>
<h2>Tasks</h2>
<span class="aside">tasks/<id>/ · oracles/<id>/</span>
</div>
<div class="prose mt-s">
<p>A task is a fixture repository, a task spec, and an executable checker. Success depends
on something learned in earlier sessions: a recorded decision, a convention, a constraint
that lives in the corpus and not in the prompt. The suite holds <strong>34 tasks</strong>
(identifiers like <code>ts-tz-utc</code>, <code>ts-stable-sort</code>,
<code>ts-log-mask</code>). Three of them, the <code>xs-</code> set, state their governing
rule across two sessions rather than one, so no single document answers them: a suite where
combining sessions is never necessary cannot detect a product that combines them.</p>
<p>Every task ships two reference solutions, asserted in CI on every commit:</p>
</div>
<div class="table-scroll mt-s">
<table>
<thead>
<tr><th>reference</th><th>encodes</th><th>must</th></tr>
</thead>
<tbody>
<tr>
<td><span class="m">naive/</span></td>
<td>the plausible solution an agent produces <em>without</em> the recorded knowledge</td>
<td><strong>fail</strong> the checker</td>
</tr>
<tr>
<td><span class="m">informed/</span></td>
<td>the solution that uses what the corpus knows</td>
<td><strong>pass</strong> the checker</td>
</tr>
</tbody>
</table>
</div>
<p class="prose mt-s dim">If the naive solution passes, the task is not measuring memory and
is rejected. A do-nothing session scores zero by construction: there is no partial credit
from a judge, because there is no judge.</p>
<div class="prose mt-m">
<p><strong>How one cell actually runs.</strong> A cell is one task at one seed, and every
arm runs it. Take <code>ts-tz-utc</code>, whose governing fact is a timezone convention
this project settled months ago and wrote down nowhere in the code.</p>
<ol>
<li>The corpus is ingested once per condition, before any session, through each
product's own write path. It is then <strong>frozen</strong>: no arm writes to its store
again for the rest of the run.</li>
<li>A fresh sandbox is built from the task's fixture repository. It contains the code and
nothing else: no oracle, no reference solution, no corpus on disk. The sandbox cannot
reach the benchmark repository, because a single <code>cd ..</code> would otherwise reach
the answers.</li>
<li>The agent gets the prompt and that arm's integration, and only that arm's. The
prompt never states the governing fact. It is answerable from the corpus, or from
nothing.</li>
<li>The session ends. Before it can be scored it passes the <a href="#gate">admission
gate</a>, which checks that the arm's tools were actually listed at session init and
that no arm held another arm's tools.</li>
<li>The checker runs the artifact against oracle inputs the sandbox never contained. For
<code>ts-tz-utc</code> that means feeding timestamps whose correct handling depends on
the convention, and comparing output. Pass or fail. That is the score.</li>
</ol>
<p>The two reference solutions are what make the task <em>about memory</em> rather than
about competence. <code>naive/</code> is the good-faith answer an able engineer writes
without the recorded knowledge, and CI asserts it <strong>fails</strong>. <code>informed/</code>
uses the recorded fact and CI asserts it <strong>passes</strong>. If the naive solution ever
starts passing, the task has stopped measuring memory and is rejected rather than kept.</p>
<p>The official run works the subset that carries planted corpus conditions: <strong>73
task-conditions</strong> in all. One task cannot express <code>contradictory</code>
observably and says so in its own directory rather than being quietly dropped.</p>
</div>
</section>
<section class="pad-y-s" id="corpus">
<div class="section-head">
<span class="no">04</span>
<h2>The corpus</h2>
<span class="aside">corpus/ · sha256 manifest</span>
</div>
<div class="prose mt-s">
<p>The experience feed is <strong>verbatim recorded agent session transcripts</strong>,
pinned by a sha256 manifest. It is deliberately not a curated fact list: it carries the
noise, the dead ends and the distractor sessions real agent history carries. No arm gets a
different feed, and the only asymmetry a product can gain is what it extracts.</p>
<p><strong>It is 4,900 documents per condition, and it was made harder on purpose.</strong>
The 196-document feed every earlier run used was saturated: hit@10 was 1.000, so every
memory arm found the governing session every time and the grid could not separate "the
product retrieved badly" from "the agent never searched". On the current corpus
<code>bm25</code> hit@1 is 0.182 against 0.485, and <code>voyage</code> hit@10 is 0.879
against 1.000. Containment is checked against the built corpus rather than trusted: no fact
term of any task appears in any synthetic document, or the <code>absent</code> condition
would be silently broken for that task.</p>
<p class="dim">That change breaks comparability. No number measured on this corpus may be
differenced against a number from any earlier run.</p>
</div>
</section>
<section class="pad-y-s" id="harm">
<div class="section-head">
<span class="no">05</span>
<h2>The harm suite</h2>
<span class="aside">does memory ever make it worse?</span>
</div>
<div class="prose mt-s">
<p>Every task described above places its governing fact <em>in</em> the corpus, so that
suite can only ask whether memory helps. It is structurally incapable of detecting harm,
and a layer that helps 20% of cells while harming 15% publishes the same headline as one
that helps 20% and harms 2%.</p>
<p>So a second suite varies <strong>what the corpus contains</strong>, never what a product
does about it. Every arm ingests identical bytes under each condition. Whether a
system copes through supersession metadata, recency weighting, reranking, a refusal
threshold, or not at all, is the thing being measured rather than the thing assumed. Damage
is defined against the memory-free arm: failing a cell that <code>bare</code> solved.</p>
</div>
<div class="table-scroll mt-s">
<table>
<thead>
<tr><th>condition</th><th>the corpus holds</th><th>correct behaviour</th><th>damage signature</th></tr>
</thead>
<tbody>
<tr><td><span class="m">absent</span></td>
<td>no governing fact for this task</td>
<td>solve from the repository, or say it is unknown</td>
<td>invents a convention and applies it</td></tr>
<tr><td><span class="m">superseded</span></td>
<td>the old fact and the newer one, both dated</td>
<td>apply the current fact</td>
<td>ships the stale convention</td></tr>
<tr><td><span class="m">contradictory</span></td>
<td>two undated memos that disagree, neither marked</td>
<td>surface the conflict rather than choose</td>
<td>chooses silently</td></tr>
<tr><td><span class="m">adjacent</span></td>
<td>a confident, high-similarity memo governing a <em>different</em> subsystem</td>
<td>recognise that it does not apply here</td>
<td>applies the other subsystem's rule</td></tr>
<tr><td><span class="m">present</span></td>
<td>the governing fact, plainly, with nothing done to it</td>
<td>find it and apply it</td>
<td>none; this is the condition memory should win</td></tr>
</tbody>
</table>
</div>
<p class="prose mt-m dim">The fifth condition exists because the other four all vary how the
evidence is <em>bad</em>, which made never searching a dominant strategy: an arm that ignored
its memory entirely forfeited nothing. <code>present</code> is the identity transform, and it
is what a degenerate strategy loses. A plant is measurable only if every reading of it gives a
different observable outcome, which is why one task carries two damage conditions rather than
four: its <code>adjacent</code> damage would be byte-identical to the factless answer. Eleven
tasks carry all four.</p>
</section>
<section class="pad-y-s" id="gate">
<div class="section-head">
<span class="no">06</span>
<h2>The admission gate</h2>
<span class="aside">discard, never score</span>
</div>
<div class="prose mt-s">
<p>Silent failure is the standing hazard of agent benchmarks: an MCP server that never
attached, a hook that never fired, a sandbox missing its files. A session in that state
measures nothing, and scoring it poisons the average in whichever direction luck chooses.</p>
<p>So a grid cell is <strong>discarded, not scored</strong>, unless every arm proves its
treatment was applied:</p>
</div>
<div class="table-scroll mt-s">
<table>
<thead><tr><th>integration</th><th>required proof</th></tr></thead>
<tbody>
<tr><td><span class="m">MCP server</span></td><td>tools listed at session init, in the session's own record</td></tr>
<tr><td><span class="m">lifecycle hooks</span></td><td>hooks demonstrably fired, with output</td></tr>
<tr><td><span class="m">sandbox files</span></td><td>digest-verified against the frozen bundle</td></tr>
<tr><td><span class="m">isolation</span></td><td>no arm holding another arm's tools</td></tr>
</tbody>
</table>
</div>
<p class="prose mt-s dim">Discard counts are published per arm, so a product that only runs
cleanly half the time cannot hide it in a smaller denominator.</p>
<p class="prose mt-s dim">Two consequences the gate cannot state for itself. Only an arm
with a memory surface can fail to wire, so the rule protects one class of arm's worst outcome
and no other's, and every headline is published beside an intention-to-treat column over all
complete cells. And a startup failure is a transient, not an outcome: a preflight speaks the
protocol to the server before a session is paid for, and a session whose treatment failed to
wire is retried under a rule that reads the admission surface and never reads
<code>success</code>, the checker verdict, or anything the model did.</p>
</section>
<section class="pad-y-s" id="scoring">
<div class="section-head">
<span class="no">07</span>
<h2>Scoring and costs</h2>
<span class="aside">harness/costs.py</span>
</div>
<div class="prose mt-s">
<p>The primary endpoint is <strong>task success</strong>: the checker passes or it does
not. Analysis is paired per task, arms are contrasted against the
<code>claude_md</code> baseline, and deltas below the preregistered minimum effect are
reported as noise rather than dressed up as findings.</p>
<p>Costs are end-to-end. Ingestion tokens, session tokens, wall time and negative-transfer
counts land in one per-arm ledger beside the success rate, because a layer that buys two
points for triple the tokens is a different product than its headline suggests. No arm has
been run at a matched budget, so the ledger carries success per million tokens and reports
the asymmetry rather than smoothing it.</p>
<p>Two details that sound like bookkeeping and are not. <strong>Input is not one
price:</strong> fresh input, cache reads and cache creation are metered separately, because
one arm's input can be two-thirds cache reads while a baseline's is under half, and a single
rate then overstates spend unevenly between exactly the two arms being compared.
<strong>Prices are stated, never defaulted:</strong> a live run refuses to start without
them. Compare runs on tokens.</p>
<p>An arm that ingests with a model on the benchmark host reports zero hosted tokens and
names the model, so its zero is never read as zero cost.</p>
</div>
</section>
<section class="pad-y-s" id="diagnostic">
<div class="section-head">
<span class="no">08</span>
<h2>Diagnostic reference tracks</h2>
<span class="aside">unranked, by design</span>
</div>
<div class="prose mt-s">
<p>Task success says whether the memory path worked, not which part failed. A memory arm can
lose three ways: the agent never searched, it searched badly, or the store did not hold the
answer. <code>recall_prefetch</code> separates the first from the rest by running retrieval
in the harness with the exact task prompt and injecting what comes back. It <strong>is</strong>
in the official run, as an upper bound; it is never ranked against the products, because an
arm that cannot lose does not belong in a ranking.</p>
<p class="dim"><code>oracle_memory</code>, which injects the exact evidence and removes
retrieval altogether, is <strong>not</strong> in the run. Its bundles are keyed by task and
carry no corpus condition, so under <code>absent</code> it would supply an answer that
condition is defined not to contain. It returns when its bundles are condition-aware. The
contrasts below are the decomposition it belongs to, stated so its absence is a recorded
choice.</p>
</div>
<div class="mt-m">
<div class="contrast">
<span class="expr">oracle_memory − claude_md</span>
<p>Oracle headroom: how much this task set can reward correct evidence at all.</p>
</div>
<div class="contrast">
<span class="expr">recall − claude_md</span>
<p>Natural memory lift: the product as an agent actually experiences it.</p>
</div>
<div class="contrast">
<span class="expr">recall_prefetch − claude_md</span>
<p>Prefetch memory lift: retrieval quality with query formulation removed.</p>
</div>
<div class="contrast">
<span class="expr">oracle_memory − recall</span>
<p>Access gap: everything lost between perfect evidence and the live memory path.</p>
</div>
<div class="contrast">
<span class="expr">recall_prefetch − recall</span>
<p>Prefetch gap: how much is lost to the agent's own decision to search, and its query.</p>
</div>
</div>
<p class="prose mt-m dim">A gap is evidence about a causal path, not a claim that a product
is first or unique. Diagnostic arms never enter the product ranking.</p>
</section>
<section class="pad-y-s" id="why">
<div class="section-head">
<span class="no">09</span>
<h2>Why this is a test of memory</h2>
<span class="aside">and not of retrieval prose</span>
</div>
<div class="prose mt-s">
<p>Most memory benchmarks ask a model questions about synthetic conversations and score
the answers with another model: a needle someone planted, found under a judge whose failure
modes correlate with the thing being judged. Five properties here are chosen against
that.</p>
<p><strong>The grade is execution.</strong> A checker runs the artifact against oracle
inputs the sandbox never contained. No partial credit, no rubric, no judge in the primary
endpoint, so a fluent wrong answer scores what a silent one does: zero.</p>
<p><strong>The corpus is raw.</strong> Verbatim recorded sessions, dead ends and noise
included. Distilling signal from that stream is what a memory product claims to do, so the
benchmark refuses to do it for anyone. The only asymmetry available is what a product
keeps.</p>
<p><strong>Harm is measured.</strong> A suite where every governing fact sits in the corpus
can only ask whether memory helps, and a layer that helps 20% of cells while harming 15%
publishes the same headline as one that helps 20% and harms 2%. The corpus conditions ask
whether the product notices stale, contradictory and off-subsystem evidence.</p>
<p><strong>The comparison is paired.</strong> Every arm meets the same task, seed and
fixture, so a cell is a within-subject comparison rather than a difference of averages
across arms that met different work. Arm order is randomised per cell and recorded.</p>
<p><strong>Treatment is proved.</strong> A stdio server that fails to start is invisible in
a transcript: no memory tool calls, which looks exactly like a model that chose not to
search. The gate discards any cell where an arm cannot prove its treatment applied.</p>
<p>The result is a benchmark that can return a null and survive it. If memory layers do not
beat a good static instruction file on real coding work, that is the finding, and it is
published as readily as the opposite.</p>
</div>
</section>
<section class="pad-y-s" id="prereg">
<div class="section-head">
<span class="no">10</span>
<h2>Preregistration</h2>
<span class="aside">preregistration/</span>
</div>
<div class="prose mt-s">
<p>Every measured run is preregistered: the question, the predictions, the endpoints, the
contrast families, the exclusion rules and the sizing are committed <strong>before the
first session starts</strong>. The run scripts enforce the mechanical half: they refuse to
start while the preregistration directory is dirty.</p>
</div>
<div class="callout mt-m">
<span class="kicker">The honest half, by convention</span>
<div class="prose">
<p><strong>Never edit a number in a committed preregistration.</strong> Not a
prediction, not a measured value, not a date. Append a correction underneath.</p>
<p><strong>Results are appended below the frozen prediction</strong>, under a marked
line, in the same file, so prediction and outcome are read together.</p>
<p><strong>Falsified predictions stay.</strong> The gap between expected and measured is
the only part of a result that teaches anything.</p>
</div>
</div>
</section>
</main>
<footer>
<div class="footer-inner">
<div>
<div class="foot-title">agent-memory-bench</div>
<div class="dim">A preregistered, execution-graded benchmark of pluggable memory layers
for coding agents. Apache-2.0. Built in the open; results published win or lose.</div>
</div>
<div>
<div class="foot-title">Pages</div>
<a href="method.html">Method</a><br>
<a href="leaderboard.html">Leaderboard</a><br>
<a href="submit.html">Submit & reproduce</a>
</div>
<div>
<div class="foot-title">Source</div>
<a href="https://github.com/GiulioDER/agent-memory-bench">Repository</a><br>
<a href="https://github.com/GiulioDER/agent-memory-bench/tree/master/preregistration">Preregistrations</a><br>
<a href="https://github.com/GiulioDER/agent-memory-bench/blob/master/adapters/VENDOR_REVIEW_TEMPLATE.md">Vendor review</a>
</div>
</div>
</footer>
<script src="site.js"></script>
</body>
</html>
|