Code and raw results
Everything needed to check or rerun the replication. Upstream mini-AGI is not included; the harness
imports it unmodified from upstream/mini-AGI/.
Regenerate the tables and figures (no GPU needed)
pip install numpy matplotlib
python harness/summarize.py runs/probes/*/*/probe.json # every arm's forgetting / retained / gain
python harness/figures.py v2 # revision-2 figures -> report_figs_v2/*.png
runs/probes/s_seed{0,1}/<arm>/probe.json holds the full evaluation trajectory of each probe: per-subject
held-out loss at every evaluation, the probe's effective LR scale, and how many experts received gradient.
Arm names: A4_* and S_*_accum4 are at upstream's optimiser step density (--accum 4) and back revision 2's
headline results; S_* vary the step density. R*, E* and L* are revision 1's arms at 4x step density
(chunk 512, one step per chunk) and support the step-density analysis. The harness scripts define each arm's flags.
Rerun from scratch (one 8 GB+ CUDA GPU)
git clone https://github.com/volotat/mini-AGI upstream/mini-AGI && git -C upstream/mini-AGI checkout 201852d
python -m venv .venv && .venv/bin/pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
.venv/bin/pip install numpy==1.26.4 pyyaml matplotlib scipy chess zstandard datasets
(cd upstream/mini-AGI && ../../.venv/bin/python -m corpora all)
Then restore the chat_hermes held-out set that upstream's builder overwrites, and split Hermes into its
own training lane (see "Deviations" in NOTES.md):
cd upstream/mini-AGI
../../.venv/bin/python -m corpora fetch --dataset teknium/OpenHermes-2.5 --kind chat --limit 20000 \
--hold 400 --out /tmp/hermes_scratch --held-out data/val/chat_hermes
mkdir -p data/train/chat_hermes && mv data/train/chat/hermes/* data/train/chat_hermes/ && rmdir data/train/chat/hermes
cd ../..
Train a base (about 4 h on an RTX 4080 Super), then run the probes:
SAMPLE_GEN_EVERY=15 .venv/bin/python harness/run_upstream.py configs/rep_s.yaml read upstream/mini-AGI/data/train \
--save --weights-dir runs/rep_s_seed1/weights --held-out upstream/mini-AGI/data/val \
--sample-every 0.67 --sample-log runs/rep_s_seed1/samples.txt --minutes 240 --no-plots --seed 1 --shuffle-seed 1
mkdir -p runs/bases && cp -r runs/rep_s_seed1/weights runs/bases/s_seed1
harness/run_matrix.sh runs/bases/s_seed1 runs/probes/s_seed1 0
harness/run_lrscale.sh # expects both runs/bases/s_seed0 and s_seed1
harness/run_accum4.sh # revision-2 arms at upstream step density (both seeds)
harness/run_stepdensity.sh # the headline arm at 1, 2 and 4 chunks per optimiser step
harness/supervise.py runs a base to a character target and resumes after crashes; we used it to match
seed 0's length to seed 1's.
License: MIT, the same as upstream.