File size: 2,897 Bytes
0662d8e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
# Code and raw results

Everything needed to check or rerun the replication. Upstream mini-AGI is **not** included; the harness
imports it unmodified from `upstream/mini-AGI/`.

## Regenerate the tables and figures (no GPU needed)

```bash
pip install numpy matplotlib
python harness/summarize.py runs/probes/*/*/probe.json   # every arm's forgetting / retained / gain
python harness/figures.py v2                             # revision-2 figures -> report_figs_v2/*.png
```

`runs/probes/s_seed{0,1}/<arm>/probe.json` holds the full evaluation trajectory of each probe: per-subject
held-out loss at every evaluation, the probe's effective LR scale, and how many experts received gradient.
Arm names: `A4_*` and `S_*_accum4` are at upstream's optimiser step density (`--accum 4`) and back revision 2's
headline results; `S_*` vary the step density. `R*`, `E*` and `L*` are revision 1's arms at 4x step density
(chunk 512, one step per chunk) and support the step-density analysis. The harness scripts define each arm's flags.

## Rerun from scratch (one 8 GB+ CUDA GPU)

```bash
git clone https://github.com/volotat/mini-AGI upstream/mini-AGI && git -C upstream/mini-AGI checkout 201852d
python -m venv .venv && .venv/bin/pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
.venv/bin/pip install numpy==1.26.4 pyyaml matplotlib scipy chess zstandard datasets
(cd upstream/mini-AGI && ../../.venv/bin/python -m corpora all)
```

Then restore the `chat_hermes` held-out set that upstream's builder overwrites, and split Hermes into its
own training lane (see "Deviations" in `NOTES.md`):

```bash
cd upstream/mini-AGI
../../.venv/bin/python -m corpora fetch --dataset teknium/OpenHermes-2.5 --kind chat --limit 20000 \
  --hold 400 --out /tmp/hermes_scratch --held-out data/val/chat_hermes
mkdir -p data/train/chat_hermes && mv data/train/chat/hermes/* data/train/chat_hermes/ && rmdir data/train/chat/hermes
cd ../..
```

Train a base (about 4 h on an RTX 4080 Super), then run the probes:

```bash
SAMPLE_GEN_EVERY=15 .venv/bin/python harness/run_upstream.py configs/rep_s.yaml read upstream/mini-AGI/data/train \
  --save --weights-dir runs/rep_s_seed1/weights --held-out upstream/mini-AGI/data/val \
  --sample-every 0.67 --sample-log runs/rep_s_seed1/samples.txt --minutes 240 --no-plots --seed 1 --shuffle-seed 1
mkdir -p runs/bases && cp -r runs/rep_s_seed1/weights runs/bases/s_seed1
harness/run_matrix.sh runs/bases/s_seed1 runs/probes/s_seed1 0
harness/run_lrscale.sh      # expects both runs/bases/s_seed0 and s_seed1
harness/run_accum4.sh       # revision-2 arms at upstream step density (both seeds)
harness/run_stepdensity.sh  # the headline arm at 1, 2 and 4 chunks per optimiser step
```

`harness/supervise.py` runs a base to a character target and resumes after crashes; we used it to match
seed 0's length to seed 1's.

License: MIT, the same as upstream.