dreddnafious's picture
Revision 2: step density matched; headline replicates
0662d8e verified
|
Raw
History Blame Contribute Delete
5.45 kB

Replicating mini-AGI's continual-learning claims

Upstream: https://github.com/volotat/mini-AGI at commit 201852d, vendored unmodified in upstream/mini-AGI/ (only its data/ layout was changed, see below).

Claims under test (README, "How continual learning works")

  1. Reading 524k characters of chess only, trunk LR at 0.1x the experts' leaves the other seven subjects at +0.0067 nats (99.84% retained vs chance). At trunk LR = expert LR: +2.23 (swapping) / +2.59 (working set frozen). Control, all subjects read: -0.0077.
  2. The expert pool is not the mechanism: freezing the working set explains only 13.8% of the effect.
  3. Displacement, not destruction: ~3/4 of the damage comes back in 131k characters of mixed reading.
  4. During the 0.1x probe only 54 of 136 experts received any gradient.

Scale

Upstream: 540M params (8.3M trunk), 409M characters, RTX 3070 laptop, ~713 c/s. Here: configs/rep_s.yaml, 14.7M params at start (2.1M trunk, 32 experts of 0.39M, 16 resident), ~8k c/s per run with two runs sharing the 4080 Super. Every scaled value is annotated in the config with its upstream value; ratios (top_k/resident, halting mean/max depth, chunk/context) are preserved.

Deviations and assumptions

What Upstream Here Why
Held-out chat_hermes scored separately restored to data/val/chat_hermes upstream's builder overwrites it: expand wipes data/val/chat after Hermes writes there
Training lanes 8 (inferred) 8, 1:1 with held-out subjects the builder puts Hermes under train/chat/hermes, merging it into the synthetic chat lane
Context grows 2048 -> 4096 fixed 1024 context growth is not part of any claim
Optimiser steps per character 1 per 2,048 chars (chunk 2048) revision 1: 1 per 512 (4x, unnoticed); revision 2: 1 per 2,048 via --accum 4 forgetting scales ~quadratically with LR x steps; this deviation produced revision 1's error
Sample text generation every eval every 15th eval (SAMPLE_GEN_EVERY) logging only; held-out eval still runs every time and drives the LR controller
GradSNR meter crashes if a grad is None missing grad counted as zeros latent upstream bug: depth 1 leaves the halting head gradless; ~1.4%/step at our depth, ~3e-6 at upstream's
Probe code unpublished (tools/, runs/cl/ missing) harness/probe.py rebuilt from README text
Probe LR unstated base LR x the base checkpoint's plasticity scale, frozen for the probe every arm reads at the same rate
Growth/pruning in probe unstated off a dry read does neither
"Frozen working set" unstated chosen once by peeking the first chess chunk, restored after each eval most literal reading
Control arm "all seven subjects read" all 8 lanes round-robin, trunk 0.1x the normal training regime

Observation on claim 3: with passage = 32,768 characters, 131k characters of mixed reading visits only 4 of 8 subjects. Recovery here runs 1M characters (4 full cycles) and is reported per subject.

Run log

  • 2026-09-22 18:27: seed 0 base died at 12 min with a PyTorch internal assert (CUDAGuardImpl.h:28) in a plain matmul in PagedPool._key_of, during evaluation. Not reproduced since; treated as transient. Its probe matrix ran on the 12-minute base and is quarantined in runs/probes_INVALID_seed0_12min_base (a pilot only). Seed 0 was restarted from scratch under harness/supervise.py, which resumes after crashes and stops at seed 1's final character count.

  • 2026-09-23 00:03: both bases done (77.6M / 77.3M chars; plasticity scale 0.775 / 0.614). Full matrix ran on both, no failures.

  • 2026-09-23: follow-ups harness/run_lrscale.sh (R3 at fixed probe LR scale 0.1/0.3/0.6/1.0) and L_swap_t0.1_scale0.1_4M (low-rate 4M probe). Key result: the README's +0.0067 reappears at probe LR 0.1x, but forgetting still climbs to +0.90 / +0.63 by 4M characters. Write-up in REPORT.md, figures from harness/figures.py into report_figs/.

  • 2026-09-23: correction — step density. Our chunk of 512 (vs upstream 2048) meant 4x more optimiser steps per character. harness/probe.py --accum N (gradient accumulation) matches upstream at N=4. harness/run_stepdensity.sh (accum 1/2/4 on the headline arm) and harness/run_accum4.sh (trunk 1x, all-0.1x, control, 4M probe at accum 4). At matched density the headline replicates (+0.030 / +0.022). The first public report (Space + issue #21) overstated forgetting; the Space was taken down and revision 2 (report/report.html, figures report_figs_v2/ via harness/figures.py v2) replaces it. Revision 1 kept locally as report/report_v1.*.

Metrics

  • forgetting = mean over unread subjects of (loss_end - loss_start), same held-out text every eval
  • retained = 1 - forgetting / (ln 265 - mean start loss of unread subjects) - reproduces upstream's figure
  • gain = start - end loss on the subjects read. Upstream reports chess +0.013 at 0.1x; the question is whether low forgetting is bought with low learning.

Experiment matrix (harness/run_matrix.sh)

R1-R4 replicate the four arms. E1 (everything at 0.1x) tests whether the trunk/pool split matters or just a lower rate. E2-E5 sweep the trunk multiplier for the stability-plasticity frontier. E6/E7 run 8x longer. Two base seeds (runs/bases/s_seed{0,1}), probes land in runs/probes/.