Download index.html from dreddnafious/mini-agi-replication: direct link, hf CLI and curl.
- Browser
- Download file 26.1 kB
-
https://huggingface.co/spaces/dreddnafious/mini-agi-replication/resolve/main/index.html
- Command line
-
hf download hf://spaces/dreddnafious/mini-agi-replication/index.html
-
curl -L -o index.html https://huggingface.co/spaces/dreddnafious/mini-agi-replication/resolve/main/index.html
26.1 kB
| <html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1"> | |
| <title>Does mini-AGI forget? A small-scale replication (revision 2)</title> | |
| <style> | |
| :root { | |
| --bg:#f6f7f9; --card:#fff; --ink:#1a2230; --muted:#6b7686; --line:#e3e7ee; | |
| --accent:#2a78d6; --good:#159a6c; --good-bg:#e8f6f0; --warn:#c67d0a; --warn-bg:#fdf5e6; | |
| --bad:#c0392b; --bad-bg:#fdeeec; --code-bg:#eef1f6; --up:#eb6834; | |
| } | |
| * { box-sizing:border-box; } | |
| @page { size: Letter; margin: 14mm 12mm; } | |
| html, body { -webkit-print-color-adjust:exact; print-color-adjust:exact; } | |
| body { margin:0; background:var(--bg); color:var(--ink); | |
| font:13px/1.5 -apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,Helvetica,Arial,sans-serif; } | |
| .wrap { max-width:880px; margin:0 auto; padding:4px 4px 16px; } | |
| .brand { font-weight:700; letter-spacing:.14em; text-transform:uppercase; font-size:11px; color:var(--accent); } | |
| h1 { font-size:27px; margin:.15em 0 .25em; font-weight:680; line-height:1.2; } | |
| .sub { color:var(--muted); font-size:13px; margin:0; max-width:80ch; } | |
| .block { background:var(--card); border:1px solid var(--line); border-radius:12px; padding:16px 20px; margin:12px 0; } | |
| .block h2 { font-size:17px; margin:0 0 12px; font-weight:660; } | |
| .block h3 { font-size:12.5px; margin:18px 0 8px; color:var(--muted); font-weight:640; text-transform:uppercase; letter-spacing:.05em; } | |
| p { margin:.4em 0 .7em; } | |
| ul { margin:.3em 0 .7em; padding-left:1.25em; } li { margin:.3em 0; } | |
| code { background:var(--code-bg); padding:1px 5px; border-radius:4px; font-size:11.5px; | |
| font-family:ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; } | |
| pre { background:var(--code-bg); border-radius:8px; padding:10px 12px; font-size:10.5px; line-height:1.45; | |
| white-space:pre-wrap; word-break:break-all; font-family:ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; } | |
| .muted { color:var(--muted); } | |
| .tiles { display:grid; grid-template-columns:repeat(4,1fr); gap:10px; margin:6px 0 4px; } | |
| .tile { border:1px solid var(--line); border-radius:10px; padding:12px 12px 10px; background:var(--bg); } | |
| .tile.up { border-color:var(--up); } | |
| .tile.good { border-color:var(--good); background:var(--good-bg); } | |
| .tile.bad { border-color:var(--bad); background:var(--bad-bg); } | |
| .tile-num { font-size:18px; font-weight:700; line-height:1.15; font-variant-numeric:tabular-nums; } | |
| .tile-label { font-size:11.5px; color:var(--ink); margin-top:4px; font-weight:600; } | |
| .tile-sub { font-size:10.5px; color:var(--muted); margin-top:3px; } | |
| .lede { font-size:14px; margin:8px 0 2px; } | |
| table { border-collapse:collapse; width:100%; font-size:12px; margin:6px 0 10px; font-variant-numeric:tabular-nums; } | |
| th { text-align:left; font-weight:640; color:var(--muted); font-size:11px; text-transform:uppercase; letter-spacing:.04em; | |
| border-bottom:1.5px solid var(--line); padding:6px 8px; } | |
| td { border-bottom:1px solid var(--line); padding:6px 8px; vertical-align:top; } | |
| td.n, th.n { text-align:right; white-space:nowrap; } | |
| tr.hl td { background:#f2f7fd; font-weight:600; } | |
| .pill { display:inline-block; font-size:10px; font-weight:700; letter-spacing:.03em; text-transform:uppercase; | |
| border-radius:20px; padding:2px 8px; white-space:nowrap; } | |
| .pill.good { background:var(--good-bg); color:var(--good); border:1px solid var(--good); } | |
| .pill.warn { background:var(--warn-bg); color:var(--warn); border:1px solid var(--warn); } | |
| .pill.bad { background:var(--bad-bg); color:var(--bad); border:1px solid var(--bad); } | |
| figure { margin:8px 0 12px; } | |
| figure img { width:92%; margin:0 auto; border:1px solid var(--line); border-radius:8px; display:block; } | |
| figcaption { font-size:11.5px; color:var(--muted); margin-top:5px; } | |
| .callout { border:1px solid var(--line); border-left:3px solid var(--accent); border-radius:8px; padding:10px 14px; | |
| background:#f2f7fd; margin:10px 0; } | |
| .two { display:grid; grid-template-columns:1fr 1fr; gap:14px; } | |
| figure, .callout, .tiles, tr, li { break-inside:avoid; page-break-inside:avoid; } | |
| h2, h3 { break-after:avoid; page-break-after:avoid; } | |
| @media print { body { background:#fff; } .block { border-color:#d9dee7; } } | |
| footer { color:var(--muted); font-size:11px; text-align:center; margin-top:18px; } | |
| </style></head> | |
| </head> | |
| <body><div class="wrap"> | |
| <header> | |
| <div class="brand">Replication report · continual learning</div> | |
| <h1>Does mini-AGI forget? A small-scale replication</h1> | |
| <p class="sub" style="margin:0 0 8px;color:var(--ink);font-size:13.5px"><b>Jason Matthews</b> · <a href="https://github.com/dreddnafious">github.com/dreddnafious</a> · September 2026 · revision 2</p> | |
| <p class="sub" style="margin:4px 0 10px;font-size:13px"><a href="report.pdf">Download PDF</a> · <a href="https://huggingface.co/spaces/dreddnafious/mini-agi-replication/tree/main/code">Code, configs & raw results</a> · <a href="https://github.com/volotat/mini-AGI">Upstream project</a></p> | |
| <p class="sub">An independent replication of the continual-learning claims in | |
| <a href="https://github.com/volotat/mini-AGI">volotat/mini-AGI</a> (commit <code>201852d</code>), run on one RTX 4080 Super. | |
| Written as a courtesy to the author. The project is careful and unusually honest about its own history, | |
| and this report tries to return the favour.</p> | |
| <p class="sub" style="margin-top:6px;font-size:12px"><b>AI assistance:</b> the experiments, analysis and this write-up were produced with Claude Opus 5.5 (Anthropic), directed and reviewed by Jason Matthews.</p> | |
| </header> | |
| <div class="callout" style="border-left-color:var(--warn);background:var(--warn-bg)"> | |
| <b>Revision note.</b> The first version of this report said the headline forgetting number did not replicate | |
| (50–110× larger). That was our error. Our scaled-down setup took <b>4× more optimizer steps per character</b> than | |
| upstream, and forgetting turns out to be steeply sensitive to that. Once step density is matched, the headline | |
| replicates. This version reruns the key results at matched density. The step-density effect turned out to be | |
| the most instructive result, and it gets its own section below. | |
| </div> | |
| <section class="block"> | |
| <h2>The short version</h2> | |
| <p class="lede">The README's headline <b>replicates</b>: reading one subject with a slowly-updated shared trunk leaves the | |
| other subjects intact. What it demonstrates is a known principle, clearly: <b>forgetting tracks how far the shared | |
| weights move</b>. Here that looks like roughly the square of the displacement, as in Biderman et al.'s | |
| "learns less, forgets less". It also means the protection holds for a bounded amount of reading, not indefinitely.</p> | |
| <div class="tiles"> | |
| <div class="tile up"><div class="tile-num">+0.0067</div><div class="tile-label">README claim</div> | |
| <div class="tile-sub">trunk 0.1×, 524k chars of chess</div></div> | |
| <div class="tile good"><div class="tile-num">+0.030 / +0.022</div><div class="tile-label">Ours, same probe</div> | |
| <div class="tile-sub">at upstream's step density; 99.3% / 99.5% retained</div></div> | |
| <div class="tile"><div class="tile-num">−0.051 / −0.017</div><div class="tile-label">Control</div> | |
| <div class="tile-sub">all subjects read; our noise floor</div></div> | |
| <div class="tile bad"><div class="tile-num">+1.49 / +1.09</div><div class="tile-label">Same probe, 4M chars</div> | |
| <div class="tile-sub">8× longer; 65% / 75% retained</div></div> | |
| </div> | |
| <p class="muted" style="font-size:11.5px;margin-top:6px">Forgetting = mean change in held-out loss on the 7 subjects <i>not</i> being read, nats/char. | |
| Pairs are <b>seed 0 / seed 1</b> throughout.</p> | |
| <ul> | |
| <li><b>The trunk's learning rate is the lever.</b> A 1× trunk forgets ~30× more than a 0.1× trunk; the experts' rate barely matters.</li> | |
| <li><b>The architecture's distinctive parts don't do the protecting.</b> Running every rate at 0.1× forgets as little as the trunk/expert split, and freezing the expert working set changes nothing.</li> | |
| <li><b>Forgetting grows roughly with the square of the trunk's displacement</b> (learning rate × optimizer steps). That's why a 4× difference in step density produced a ~25× difference in forgetting.</li> | |
| <li><b>The floor is horizon-bound.</b> Forgetting stays at the control level for about 0.5M characters, then climbs steadily, while chess itself improves by only ~0.04 nats.</li> | |
| <li><b>Recovery replicates:</b> about 75% of heavy damage returns within 131k characters of mixed reading, then plateaus at 80–86%.</li> | |
| </ul> | |
| </section> | |
| <section class="block"> | |
| <h2>Claim by claim</h2> | |
| <table> | |
| <tr><th style="width:3%">#</th><th style="width:38%">README claim</th><th style="width:15%">Verdict</th><th>What we measured (at upstream's step density unless noted)</th></tr> | |
| <tr><td>1</td><td>524k chars of chess at trunk 0.1× costs the unread subjects <b>+0.0067</b> nats (99.84% retained)</td> | |
| <td><span class="pill good">Replicates</span></td> | |
| <td>+0.030 / +0.022 (99.3% / 99.5% retained); control −0.051 / −0.017.</td></tr> | |
| <tr><td>2</td><td>At trunk LR = expert LR: +2.23 (swapping) / +2.59 (working set frozen)</td> | |
| <td><span class="pill good">Direction</span></td> | |
| <td>+0.87 / +0.80 swapping: ~30× the 0.1× trunk, though smaller than the README's figure.</td></tr> | |
| <tr><td>3</td><td>The expert pool is not the mechanism (freezing the working set explains 13.8%)</td> | |
| <td><span class="pill good">Replicates</span></td> | |
| <td>Frozen vs swapping indistinguishable (sign flips across seeds; measured at 4× density). Every rate at 0.1× gives +0.026 / +0.016, the same as the split.</td></tr> | |
| <tr><td>4</td><td>About ¾ of the damage comes back in 131k chars of mixed reading</td> | |
| <td><span class="pill good">Replicates</span></td> | |
| <td>76% / 73% (measured at 4× density), then a plateau at 80–86%.</td></tr> | |
| <tr><td>5</td><td>Only 54 of 136 experts receive gradient during the probe</td> | |
| <td><span class="pill good">Replicates</span></td> | |
| <td>24 / 132 and 33 / 134, a similar minority of the pool.</td></tr> | |
| <tr><td></td><td><i>Implied:</i> continual reading without catastrophic forgetting</td> | |
| <td><span class="pill warn">Horizon-bound</span></td> | |
| <td>Same probe over 4M characters: +1.49 / +1.09 (65% / 75% retained), while chess gains +0.045 / +0.038.</td></tr> | |
| </table> | |
| </section> | |
| <section class="block"> | |
| <h2>How we tested it</h2> | |
| <div class="two"> | |
| <div> | |
| <h3 style="margin-top:0">Code & data</h3> | |
| <ul> | |
| <li>Upstream <b>vendored unmodified</b>. Bases trained with upstream's own <code>train.py read</code>: its LR controller, growth, pruning, paging and evaluator.</li> | |
| <li>Only logging-side runtime changes (thinned text sampling; a guard on a crashing diagnostic meter).</li> | |
| <li>Upstream's own 8-subject corpus from <code>python -m corpora all</code>, with the <code>chat_hermes</code> held-out set restored (see the notes for the author).</li> | |
| </ul> | |
| <h3>The probe</h3> | |
| <ul> | |
| <li>Upstream's probe tooling isn't published, so it was <b>rebuilt from the README</b>, reusing the same per-chunk training step as <code>read</code>.</li> | |
| <li>Batch 1, 524,288 characters, each arm from the same base copy, at the learning rate the base's own LR controller had reached. The same held-out text (61k chars per subject) is scored every 65k characters.</li> | |
| <li><b>Step density matched to upstream</b> by accumulating gradients over 4 chunks per optimizer step (256 steps per 524k characters, as at upstream's chunk of 2,048).</li> | |
| <li>Where the README is silent: growth and pruning off; LR frozen at the base controller's value; "frozen" = working set chosen once, never re-chosen.</li> | |
| </ul> | |
| </div> | |
| <div> | |
| <h3 style="margin-top:0">Scale</h3> | |
| <table> | |
| <tr><th></th><th class="n">Upstream</th><th class="n">Here</th></tr> | |
| <tr><td>d_model / heads</td><td class="n">512 / 8</td><td class="n">256 / 4</td></tr> | |
| <tr><td>Loop depth (max)</td><td class="n">24</td><td class="n">8</td></tr> | |
| <tr><td>Expert width / top-k / resident</td><td class="n">2048 / 8 / 32</td><td class="n">512 / 4 / 16</td></tr> | |
| <tr><td>Experts at probe time</td><td class="n">136</td><td class="n">132 / 134</td></tr> | |
| <tr><td>Trunk parameters</td><td class="n">8.3M</td><td class="n">2.1M</td></tr> | |
| <tr><td>Chunk / context</td><td class="n">2,048 / 4,096</td><td class="n">512 / 1,024</td></tr> | |
| <tr class="hl"><td>Optimizer steps per 524k chars</td><td class="n">256</td><td class="n">256 (accumulate 4)</td></tr> | |
| <tr><td>Chars read before probe</td><td class="n">not stated</td><td class="n">77M</td></tr> | |
| <tr><td>Unread-subject loss at probe start</td><td class="n">1.12</td><td class="n">1.28 / 1.24</td></tr> | |
| </table> | |
| <p class="muted" style="font-size:11.5px">Ratios preserved: top-k/resident, mean/max halting depth, chunk/context, trunk_lr_mult 0.1. Two independent seeds (init + data order).</p> | |
| <h3>Metrics</h3> | |
| <ul> | |
| <li><b>Forgetting</b>: mean Δ held-out loss over subjects not read.</li> | |
| <li><b>Retained</b>: 1 − forgetting / (ln 265 − start loss), which reproduces the README's figure.</li> | |
| <li><b>Gain</b>: improvement on the subject being read.</li> | |
| </ul> | |
| </div> | |
| </div> | |
| </section> | |
| <section class="block"> | |
| <h2>1 · The headline probe replicates</h2> | |
| <figure><img src="figs/1_headline.png" alt="headline probe at matched step density"> | |
| <figcaption>Forgetting on the unread subjects while reading 524k characters of chess, at upstream's step density. Solid = seed 0, dashed = seed 1.</figcaption></figure> | |
| <table> | |
| <tr><th>Arm (524k chars of chess, matched density)</th><th class="n">Upstream</th><th class="n">Seed 0</th><th class="n">Seed 1</th><th class="n">Chess gain</th></tr> | |
| <tr><td>Trunk 1×, swapping</td><td class="n">+2.23</td><td class="n">+0.87</td><td class="n">+0.80</td><td class="n">−0.011 / −0.006</td></tr> | |
| <tr class="hl"><td>Trunk 0.1×, swapping</td><td class="n">+0.0067</td><td class="n">+0.030</td><td class="n">+0.022</td><td class="n">+0.020 / +0.020</td></tr> | |
| <tr><td>Every rate at 0.1× (no split)</td><td class="n">n/a</td><td class="n">+0.026</td><td class="n">+0.016</td><td class="n">+0.020 / +0.020</td></tr> | |
| <tr><td>Control, all subjects read</td><td class="n">−0.0077</td><td class="n">−0.051</td><td class="n">−0.017</td><td class="n">n/a</td></tr> | |
| </table> | |
| <p>The 0.1× trunk stays within our noise floor for the whole probe; the control's own spread (−0.05 to +0.02) is the measure of that floor. The contrast with a 1× trunk is ~30×. Removing the trunk/expert split, by slowing the experts too, changes nothing.</p> | |
| </section> | |
| <section class="block"> | |
| <h2>2 · Step density, and why it matters</h2> | |
| <p>Our first version ran every probe at 4× upstream's optimizer steps per character: we shrank the chunk from 2,048 to 512 characters to fit the scaled model, without matching the step count. Holding everything else fixed and varying only the steps:</p> | |
| <table> | |
| <tr><th>Optimizer steps per 524k chars</th><th class="n">Seed 0</th><th class="n">Seed 1</th><th class="n">Chess gain</th></tr> | |
| <tr><td>1,024 (our first version)</td><td class="n">+0.73</td><td class="n">+0.32</td><td class="n">+0.010 / +0.010</td></tr> | |
| <tr><td>512</td><td class="n">+0.25</td><td class="n">+0.066</td><td class="n">+0.017 / +0.018</td></tr> | |
| <tr class="hl"><td>256 (upstream's density)</td><td class="n">+0.030</td><td class="n">+0.022</td><td class="n">+0.020 / +0.020</td></tr> | |
| </table> | |
| <figure><img src="figs/2_displacement.png" alt="forgetting vs displacement"> | |
| <figcaption>Every trunk-0.1× probe, plotted against probe learning rate × optimizer steps. Blue: learning rate varied at fixed steps. Green: steps varied by gradient accumulation at fixed learning rate. Log-log.</figcaption></figure> | |
| <ul> | |
| <li><b>Steps and learning rate are one lever.</b> Adam normalizes each update to roughly the learning rate in size, so how far the trunk moves per character is about <i>learning rate × steps per character</i>. Varying either lands on the same curve, within seed noise.</li> | |
| <li><b>The curve is close to quadratic.</b> At small displacement, the slope on a log-log plot is 2.1 on both seeds, flattening toward 1–1.5 as the damage saturates. That's why a 4× difference in steps produced a ~25× difference in forgetting.</li> | |
| <li><b>More displacement bought no learning.</b> Chess gains ~0.02 at the smallest displacement and ~0.01 at the largest. The extra movement is pure cost.</li> | |
| </ul> | |
| <p>The practical consequence: <b>tokens per optimizer step is a hidden hyperparameter in any streaming-learning claim.</b> A forgetting number without the learning rate, the tokens per step and the horizon isn't comparable across setups. It nearly misled this replication.</p> | |
| </section> | |
| <section class="block"> | |
| <h2>3 · Read for longer</h2> | |
| <figure><img src="figs/3_long.png" alt="4M characters at matched density"> | |
| <figcaption>The headline probe run 8× longer (4M characters) at matched density. The shaded band is the README's probe length. Gray: control, all subjects read (from the first version, at 4× density; it is flat either way).</figcaption></figure> | |
| <p>Forgetting holds at the control level through about 0.5M characters, then climbs steadily to <b>+1.49 / +1.09</b> by 4M: 65% / 75% of the gain over chance retained. Chess improves by just +0.045 / +0.038 over the same stretch. The README's "never leaves the floor" is accurate for the window it measured; the square law explains why the window is finite.</p> | |
| </section> | |
| <section class="block"> | |
| <h2>4 · Recovery</h2> | |
| <figure><img src="figs/4_recover.png" alt="recovery after damage"> | |
| <figcaption>After the trunk-1× probe (measured at 4× step density), share of the damage recovered while reading all subjects in rotation.</figcaption></figure> | |
| <p>Reading all subjects recovers <b>79% / 77% within 65k characters</b>. By then only 2 of the 8 subjects have been read, so recovery is general rather than per-subject relearning, which fits the README's "displacement, not destruction". Recovery then plateaus at 80–86%, with +0.64 / +0.69 nats still missing after 1M characters. (Subjects rotate every 32,768 characters, so the README's 131k window covers 4 of the 8.)</p> | |
| </section> | |
| <section class="block"> | |
| <h2>What this demonstrates</h2> | |
| <p>The result is a clean, cheap demonstration of a principle established in prior work, reproduced in a byte-level, sparse-expert model that learns from a stream.</p> | |
| <ul> | |
| <li><b>Learns less, forgets less.</b> <a href="https://arxiv.org/abs/2405.09673">Biderman et al. (2024)</a> showed that constraining how much a model's weights can change during fine-tuning (there, with low-rank LoRA updates) preserves capabilities outside the target domain, at the cost of learning less of it. mini-AGI constrains the same thing a different way: a slow learning rate on the shared trunk. Our results show the "forgets less" half directly. Forgetting falls with the trunk's displacement, whether that displacement is reduced by the learning rate or by the number of steps. The "learns less" half is invisible here, because the probe subject was already learned (see below).</li> | |
| <li><b>Why it's roughly quadratic.</b> A base trained to a minimum on its subjects has near-zero gradient on them, so a small step costs nothing to first order, and the loss rises only through curvature: ΔL ≈ ½·δᵀHδ. <a href="https://arxiv.org/abs/2006.06958">Mirzadeh et al. (2020)</a> analyse forgetting exactly this way, as a function of displacement and curvature, and show training-regime choices such as learning rate and batch size move it. Our slope of ~2 at small displacement, and the flat-then-rising curve over 4M characters, are what that view predicts.</li> | |
| <li><b>Slowing the shared layers is established practice.</b> Layer-wise learning rates, lower in the shared body than the head, are a standard fine-tuning tool (<a href="https://arxiv.org/abs/1801.06146">Howard & Ruder, 2018</a>). The trunk multiplier is that idea applied to continual reading.</li> | |
| <li><b>The curvature view also points past uniform slowing.</b> <a href="https://arxiv.org/abs/1612.00796">EWC (Kirkpatrick et al., 2017)</a> penalizes movement weighted by each parameter's importance to earlier tasks: the same quadratic, made selective.</li> | |
| </ul> | |
| <p>What's specific to mini-AGI is where the effect lives. The expert pool confines each update to a minority of experts, yet that doesn't protect anything; the dense shared trunk carries the interference. That's a useful data point, given how often sparse or modular architectures are proposed as a continual-learning mechanism.</p> | |
| </section> | |
| <section class="block"> | |
| <h2>Paths that would qualify the claims further</h2> | |
| <p>These are the open questions this replication couldn't settle, offered as directions rather than gaps in the original work:</p> | |
| <ol> | |
| <li><b>A subject the model hasn't learned.</b> The chess probe improves chess by only ~0.02–0.04 nats, so it measures stability, not plasticity. Probing with a subject held out of base training would show the "learns less" side of the trade-off. The quadratic view predicts that learning grows linearly with displacement while forgetting grows quadratically, so slower reading should improve the gain-to-forgetting ratio, at the cost of time.</li> | |
| <li><b>Step count vs learning rate at matched displacement.</b> On one seed, accumulating gradients forgot about 3× less than lowering the learning rate to the same displacement; on the other it didn't. If averaging out batch-1 gradient noise is a real bonus, accumulation is the better knob for streaming learners. More seeds would settle it.</li> | |
| <li><b>Selective vs uniform slowing.</b> An EWC-style or L2-to-anchor penalty on the trunk, compared with the 0.1× trunk at equal learning, would test whether protection can be targeted rather than bought with a uniformly slow trunk.</li> | |
| <li><b>Base maturity.</b> The quadratic picture assumes the base sits at a minimum. On a less-trained base the first-order term returns; a short pilot on a 12-minute base showed large forgetting even at 0.1×.</li> | |
| <li><b>Scale.</b> Our trunk is a quarter of upstream's. A larger trunk may have a wider, flatter basin and hold the floor longer.</li> | |
| </ol> | |
| </section> | |
| <section class="block"> | |
| <h2>Caveats</h2> | |
| <ul> | |
| <li><b>Scale.</b> 2.1M trunk parameters against 8.3M; the bases start the probe at a comparable loss on the unread subjects (1.24–1.28 vs 1.12).</li> | |
| <li><b>Reconstructed probe.</b> Rebuilt from the README's description; where it is silent we took the most literal reading.</li> | |
| <li><b>Mixed densities in supporting results.</b> The headline, trunk-1×, no-split, control and 4M results are at matched density. The frozen-working-set comparison and the recovery curve were measured at 4× density; their conclusions (no effect; ~75% then a plateau) don't depend on the absolute level.</li> | |
| <li><b>Two seeds.</b> They differ in magnitude by up to 2–3× but agree on every verdict.</li> | |
| </ul> | |
| </section> | |
| <section class="block"> | |
| <h2>Notes for the author</h2> | |
| <ol> | |
| <li><b>Tokens per optimizer step is worth reporting with any forgetting number.</b> Forgetting here scales roughly with the square of learning rate × steps per character, so the same 524k-character probe gives +0.03 or +0.73 depending only on the chunk size. We learned this by getting it wrong first.</li> | |
| <li><b>The corpus builder drops the <code>chat_hermes</code> held-out set.</b> Hermes writes its held-out files to <code>data/val/chat</code>, then the synthetic-chat <code>expand</code> step wipes that folder, so a fresh <code>python -m corpora all</code> can't produce the README's eight held-out subjects. Noted on PR #11, which fixes the wipe as a side effect.</li> | |
| <li><b><code>GradSNR.observe</code> crashes when a parameter's grad is <code>None</code> on some steps</b>: a depth-1 step leaves the halting head gradless. Reported in #19; fix proposed in #20.</li> | |
| <li><b>The chess probe barely teaches the model anything</b> (≤0.04 nats even over 4M characters). A probe on an unlearned subject would test the full stability-plasticity trade-off.</li> | |
| <li><b>The probe tooling is referenced but not published</b> (<code>tools/</code>, <code>runs/cl/</code>, <code>runs/results/</code>). Publishing it would make the headline directly checkable.</li> | |
| </ol> | |
| </section> | |
| <section class="block"> | |
| <h2>Reproduce</h2> | |
| <pre>python -m corpora all # upstream corpus, then restore chat_hermes (see code/README.md) | |
| python harness/run_upstream.py configs/rep_s.yaml read upstream/mini-AGI/data/train --save \ | |
| --weights-dir runs/rep_s_seed1/weights --held-out upstream/mini-AGI/data/val \ | |
| --sample-every 0.67 --minutes 240 --seed 1 --shuffle-seed 1 | |
| harness/run_accum4.sh # headline arms at upstream's step density (--accum 4) | |
| harness/run_stepdensity.sh # the same probe at 1, 2 and 4 chunks per step | |
| harness/run_matrix.sh runs/bases/s_seed1 runs/probes/s_seed1 0 # first-version arms (4x density) | |
| python harness/summarize.py runs/probes/*/*/probe.json && python harness/figures.py v2</pre> | |
| <p class="muted" style="font-size:11.5px">Code, configs, every probe's raw results (<code>probe.json</code>) and the full deviation and run log (<code>NOTES.md</code>) are published at <a href="https://huggingface.co/spaces/dreddnafious/mini-agi-replication/tree/main/code">huggingface.co/spaces/dreddnafious/mini-agi-replication</a>. The two summary commands above regenerate every table and figure from them without a GPU.</p> | |
| </section> | |
| <footer>mini-AGI replication · revision 2 · September 2026</footer> | |
| </div></body></html> | |