javadtaghia's picture
Fix benchmark links and add portable paper reader
f497666 verified
Raw History Blame Contribute Delete
10 kB
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="description" content="AtMem memory integrity and agent continuity benchmark results.">
<title>Beyond Recall Accuracy 路 AtMem Research</title>
<link rel="stylesheet" href="style.css">
</head>
<body>
<header class="nav-shell">
<nav aria-label="Primary navigation">
<a class="brand" href="#top" aria-label="AtMem research home"><span class="brand-mark">A</span><span>AtMem Research</span></a>
<div class="nav-links">
<a href="#integrity">Integrity</a><a href="#continuity">Continuity</a><a href="#claims">Claims</a><a href="#paper">Paper</a>
</div>
<a class="button button-small" href="https://huggingface.co/datasets/atmem/memory-integrity-continuity/viewer/integrity_categories/test" target="_blank" rel="noreferrer">Open benchmark data</a>
</nav>
</header>
<main id="top">
<section class="hero">
<div class="eyebrow">Persistent agent memory 路 Evidence release</div>
<h1>Beyond recall<br><span>accuracy.</span></h1>
<p class="lede">A two-axis evaluation of whether persistent memory preserves authority and whether tool-using agents resume safely after interruption.</p>
<div class="hero-actions">
<a class="button" href="https://huggingface.co/datasets/atmem/memory-integrity-continuity/blob/main/paper/atmem-memory-integrity-and-continuity.pdf" target="_blank" rel="noreferrer">Read the technical paper</a>
<a class="button button-ghost" href="#results">Explore results</a>
</div>
<div class="scope-note"><strong>Evidence boundary:</strong> 700 seeded integrity trials, installed-product crash checks, controlled four-arm replay, and one explicitly non-causal live pilot.</div>
</section>
<section id="results" class="metrics" aria-label="Headline results">
<article><span class="metric-value">700</span><span class="metric-label">Integrity trials</span><span class="metric-detail">Seven attack categories</span></article>
<article><span class="metric-value teal">400</span><span class="metric-label">Represented passes</span><span class="metric-detail">C, D, F, and L</span></article>
<article><span class="metric-value amber">300</span><span class="metric-label">Not representable</span><span class="metric-detail">A, H, and I</span></article>
<article><span class="metric-value">0</span><span class="metric-label">Fail or error</span><span class="metric-detail">Among represented trials</span></article>
</section>
<section class="thesis">
<p class="section-number">01 路 Research question</p>
<div class="thesis-grid">
<h2>Relevant memory is not necessarily <em>authorized</em> memory.</h2>
<div>
<p>Retrieval benchmarks ask whether an agent can find a fact. This work asks two earlier questions: may the fact carry authority, and what can the agent conclude about an interrupted external action?</p>
<p>AtMem is evaluated as the authority plane. AtFlows is evaluated as an optional observation plane whose delivery cannot authorize, suppress, or repeat a tool call.</p>
</div>
</div>
</section>
<section id="integrity" class="panel dark-panel">
<div class="section-heading">
<div><p class="section-number">02 路 Memory integrity</p><h2>Authority before ranking</h2></div>
<p>Each category contains 100 seeded trials. Unsupported authority transitions remain <strong>NOT_REPRESENTABLE</strong>; they are not converted into passes.</p>
</div>
<div class="legend"><span><i class="swatch pass"></i>PASS</span><span><i class="swatch nr"></i>NOT_REPRESENTABLE</span></div>
<div id="integrity-chart" class="bar-chart" aria-label="Integrity category results"></div>
<div class="metric-strip">
<div><strong>0 / 300</strong><span>Factual contamination</span></div>
<div><strong>0 / 100</strong><span>Secret retention</span></div>
<div><strong>100 / 100</strong><span>Taint preservation</span></div>
<div><strong>0 / 200</strong><span>Trust laundering</span></div>
</div>
<aside class="qualification"><strong>Category D qualification.</strong> The submission uses native-record-ID matching to test the actual derived summary. That interpretation is pending upstream review. If rejected, the defensible passing-category count becomes three; the other results do not change.</aside>
</section>
<section id="continuity" class="panel">
<div class="section-heading">
<div><p class="section-number">03 路 Operational continuity</p><h2>Preserve uncertainty after a crash</h2></div>
<p>The process was killed after the sixth public retail tool committed its state change but before the worker received the response.</p>
</div>
<div class="continuity-grid">
<div class="continuity-copy">
<h3>Repeat requests after restart</h3>
<p>Lower is safer in this fault cell. The destination itself rejected repeated exchanges, so the attributable result is avoiding a repeat request鈥攏ot preventing a second external effect.</p>
<div class="callout"><span>AtMem decision</span><strong>needs_confirmation</strong><p>No declared query or destination-enforced idempotency contract existed.</p></div>
</div>
<div id="fault-chart" class="fault-chart" aria-label="Repeat requests by configuration"></div>
</div>
<div class="parity">
<div><span>No-fault replay</span><strong>18 / 18</strong><small>native requests matched in every arm</small></div>
<div><span>Tool sequence</span><strong>6</strong><small>identical native tool calls</small></div>
<div><span>Final state</span><strong>Equal</strong><small>same trajectory and store hash</small></div>
<div><span>Provider calls</span><strong>0</strong><small>recorded-response qualification</small></div>
</div>
</section>
<section class="pilot">
<p class="section-number">04 路 Negative-result case study</p>
<div class="pilot-grid">
<div><h2>The live pilot is evidence,<br>not a causal comparison.</h2><p>One exposed task was run once per arm. The baseline scored 1; AtMem, AtFlows, and both scored 0. The first model responses differed before tools ran, so the result cannot establish a product benefit or penalty.</p></div>
<table>
<thead><tr><th>Configuration</th><th>Reward</th><th>Calls</th><th>Est. cost</th></tr></thead>
<tbody id="pilot-table"></tbody>
</table>
</div>
</section>
<section id="claims" class="claims">
<p class="section-number">05 路 Claim ledger</p>
<h2>What the evidence does鈥攁nd does not鈥攕how</h2>
<div class="claims-grid">
<div class="claim-column supported"><h3>Supported</h3><ul><li>400 represented integrity trials passed.</li><li>AtMem made zero repeat requests in one controlled interruption.</li><li>One recorded no-fault trajectory remained identical across four configurations.</li><li>Installed document cases ended with one exact output.</li></ul></div>
<div class="claim-column excluded"><h3>Not supported</h3><ul><li>All seven attack categories passed.</li><li>General retrieval or factual superiority.</li><li>Distributed exactly-once execution.</li><li>A production-wide recovery rate or causal live-pilot comparison.</li></ul></div>
</div>
</section>
<section id="paper" class="paper-section">
<div class="section-heading">
<div><p class="section-number">06 路 Technical paper</p><h2>Read the complete method</h2></div>
<div class="paper-actions"><a class="button" href="https://huggingface.co/datasets/atmem/memory-integrity-continuity/blob/main/paper/atmem-memory-integrity-and-continuity.pdf" target="_blank" rel="noreferrer">Open PDF</a><a class="button button-ghost" href="https://huggingface.co/datasets/atmem/memory-integrity-continuity/resolve/main/paper/atmem-memory-integrity-and-continuity.pdf?download=true" target="_blank" rel="noreferrer">Download PDF</a><a class="button button-ghost" href="https://huggingface.co/datasets/atmem/memory-integrity-continuity/blob/main/paper/main.tex" target="_blank" rel="noreferrer">LaTeX source</a></div>
</div>
<p class="reader-note">The complete 12-page paper is rendered below so it remains readable inside the Hugging Face Space.</p>
<div class="paper-reader" aria-label="Beyond Recall Accuracy technical paper">
<img src="paper-pages/page-01.jpg" alt="Technical paper page 1">
<img src="paper-pages/page-02.jpg" loading="lazy" alt="Technical paper page 2">
<img src="paper-pages/page-03.jpg" loading="lazy" alt="Technical paper page 3">
<img src="paper-pages/page-04.jpg" loading="lazy" alt="Technical paper page 4">
<img src="paper-pages/page-05.jpg" loading="lazy" alt="Technical paper page 5">
<img src="paper-pages/page-06.jpg" loading="lazy" alt="Technical paper page 6">
<img src="paper-pages/page-07.jpg" loading="lazy" alt="Technical paper page 7">
<img src="paper-pages/page-08.jpg" loading="lazy" alt="Technical paper page 8">
<img src="paper-pages/page-09.jpg" loading="lazy" alt="Technical paper page 9">
<img src="paper-pages/page-10.jpg" loading="lazy" alt="Technical paper page 10">
<img src="paper-pages/page-11.jpg" loading="lazy" alt="Technical paper page 11">
<img src="paper-pages/page-12.jpg" loading="lazy" alt="Technical paper page 12">
</div>
</section>
</main>
<footer><div><strong>AtMem.Ai Lab</strong><span>Evidence-bound memory for agents</span></div><div><a href="https://github.com/aetna000/atmem">GitHub</a><a href="https://atmem.ai">atmem.ai</a><a href="https://github.com/iluxu/memory-integrity-benchmark/pull/1">External submission</a></div></footer>
<script src="app.js"></script>
</body>
</html>