dreddnafious's picture
Revision 2: step density matched; headline replicates
0662d8e verified
|
Raw
History Blame Contribute Delete
2.9 kB

Code and raw results

Everything needed to check or rerun the replication. Upstream mini-AGI is not included; the harness imports it unmodified from upstream/mini-AGI/.

Regenerate the tables and figures (no GPU needed)

pip install numpy matplotlib
python harness/summarize.py runs/probes/*/*/probe.json   # every arm's forgetting / retained / gain
python harness/figures.py v2                             # revision-2 figures -> report_figs_v2/*.png

runs/probes/s_seed{0,1}/<arm>/probe.json holds the full evaluation trajectory of each probe: per-subject held-out loss at every evaluation, the probe's effective LR scale, and how many experts received gradient. Arm names: A4_* and S_*_accum4 are at upstream's optimiser step density (--accum 4) and back revision 2's headline results; S_* vary the step density. R*, E* and L* are revision 1's arms at 4x step density (chunk 512, one step per chunk) and support the step-density analysis. The harness scripts define each arm's flags.

Rerun from scratch (one 8 GB+ CUDA GPU)

git clone https://github.com/volotat/mini-AGI upstream/mini-AGI && git -C upstream/mini-AGI checkout 201852d
python -m venv .venv && .venv/bin/pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
.venv/bin/pip install numpy==1.26.4 pyyaml matplotlib scipy chess zstandard datasets
(cd upstream/mini-AGI && ../../.venv/bin/python -m corpora all)

Then restore the chat_hermes held-out set that upstream's builder overwrites, and split Hermes into its own training lane (see "Deviations" in NOTES.md):

cd upstream/mini-AGI
../../.venv/bin/python -m corpora fetch --dataset teknium/OpenHermes-2.5 --kind chat --limit 20000 \
  --hold 400 --out /tmp/hermes_scratch --held-out data/val/chat_hermes
mkdir -p data/train/chat_hermes && mv data/train/chat/hermes/* data/train/chat_hermes/ && rmdir data/train/chat/hermes
cd ../..

Train a base (about 4 h on an RTX 4080 Super), then run the probes:

SAMPLE_GEN_EVERY=15 .venv/bin/python harness/run_upstream.py configs/rep_s.yaml read upstream/mini-AGI/data/train \
  --save --weights-dir runs/rep_s_seed1/weights --held-out upstream/mini-AGI/data/val \
  --sample-every 0.67 --sample-log runs/rep_s_seed1/samples.txt --minutes 240 --no-plots --seed 1 --shuffle-seed 1
mkdir -p runs/bases && cp -r runs/rep_s_seed1/weights runs/bases/s_seed1
harness/run_matrix.sh runs/bases/s_seed1 runs/probes/s_seed1 0
harness/run_lrscale.sh      # expects both runs/bases/s_seed0 and s_seed1
harness/run_accum4.sh       # revision-2 arms at upstream step density (both seeds)
harness/run_stepdensity.sh  # the headline arm at 1, 2 and 4 chunks per optimiser step

harness/supervise.py runs a base to a character target and resumes after crashes; we used it to match seed 0's length to seed 1's.

License: MIT, the same as upstream.