YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
OLens ablation round β reconstructors, inverters, and the rulers that disagree
Every arm here differs from the recipe of record by exactly one flag. Data, pool, crop seed, layer set, learning rate, batch and seeds are held fixed across arms so the only free variable is the one named.
The headline
The whitened-FVE ruler that the program optimises and reports anti-correlates with lens quality. The AR trained without whitening scores far worse on the ruler and makes a materially better lens, and running both to saturation does not rescue it.
Lens quality (val CE on the 594-crop u64 carve, lower is better)
| arm | cum steps | examples | span tokens | val CE |
|---|---|---|---|---|
base |
42,899 | 43,928,576 | 782M | 1.1996 |
unitcos |
28,351 | 29,031,424 | 515M | 1.2323 |
rawcos_c |
29,351 | 30,055,424 | 533M | 1.1952 |
rawcos |
42,899 | 43,928,576 | 782M | 1.1341 |
tf.scaled |
15,516 | 15,888,384 | 281M | 1.3734 |
tf.raw |
15,516 | 15,888,384 | 281M | 1.3547 |
Compare arms only at MATCHED cumulative steps. Arms stopped early by a budget cut
have lower totals, and reading their endpoints against a fully-trained arm inverts the
ordering: unitcos looks worse than base on endpoints but is identical to it at
matched steps.
Reconstructor quality (frozen 8192-row ruler, identical rows, two whitener bases)
| AR | FVE (chat basis) | FVE (pooled basis) | whitened-cos loss (chat) |
|---|---|---|---|
ar.asst.ptag.chat.k2.s3_final |
0.1830 | 0.1957 | 1.2388 |
ar.ab.k2.unitcos_final |
0.1656 | 0.1783 | 1.2771 |
ar.ab.k2.base_final |
0.1652 | 0.1779 | 1.2779 |
ar.asst.ptag.chat.k2.s0_ex11404800 |
0.1606 | β | 1.2893 |
ar.ab.k2.rawcos_c_final |
0.1102 | 0.1243 | 1.4066 |
ar.ab.k2.rawcos_final |
0.0920 | 0.1022 | 1.5067 |
ar.asst.ptag.chat.k2.s0_ex2880000 |
0.0752 | β | 1.5039 |
ar.xm.4b.p1_ex11405184 |
0.0526 | β | 1.5763 |
ar.xm.4b.p1_ex2880384 |
0.0335 | β | 1.6622 |
ar.xm.27b.p1.affmap_ex11405184 |
0.0088 | β | 1.8424 |
ar.xm.27b.p1.affmap.kaiming_ex11405184 |
0.0084 | β | 1.8455 |
ar.xm.27b.p1.affmap_ex2880384 |
0.0066 | β | 1.8663 |
These are the only comparable AR numbers. Each run also logs a val_fve during
training, but every arm carves its own val set from its own wave, so those share no
ruler and must never be compared across arms. Note the ordering here is the REVERSE
of the lens table above: the AR that scores worst on this ruler (rawcos) makes the
best lens.
RL'd arms (rl/)
Each saturated lens was lightly distilled into a 4-bullet student (dictionary labels: K=4,
length penalty mu=0.0002, 4,000 rows/layer x 21 layers, 1 epoch, warm-started from the lens
endpoint; the dictionary embedded by that arm's OWN AR), then GRPO'd for 300 steps with that
arm's own AR as the reward. Everything else -- activations labelled, labelling whitener, phrase
dictionary, student recipe, RL data/config, and the fixed 231-item ruler scored through the AR of
record -- is identical across arms. Ruler numbers, when scored, are in results/eval_table.md.
| checkpoint | ruler mean FVE | median | L20-32 / L36-48 / L52-60 | no-bullet | bullets/gen | % 3 / 4 / 5 / 6+ | bullet tok | train reward@step (FVE, KL, bullets, len, cut) |
|---|---|---|---|---|---|---|---|---|
RL rl.ab4.unitcos.g0 step 100 |
0.2782 | 0.2666 | 0.2447 / 0.2602 / 0.3414 | 0.0% | 4.009 | 1.7 / 95.7 / 2.6 / 0.0 | 9.78 | 0.178 (0.178, 0.3685, 4.00, 73, 0%) |
RL rl.ab4.unitcos.g0 step 200 |
0.3003 | 0.2865 | 0.2589 / 0.2773 / 0.3792 | 0.0% | 3.981 | 4.1 / 93.3 / 2.2 / 0.2 | 10.14 | 0.246 (0.246, 0.3689, 3.99, 86, 0%) |
RL rl.ab4.unitcos.g0 step 300 |
0.3183 | 0.3109 | 0.276 / 0.2974 / 0.3956 | 0.0% | 4.05 | 6.7 / 82.0 / 10.0 / 1.1 | 11.43 | 0.301 (0.301, 0.4070, 3.98, 100, 0%) |
What each arm changed
Reconstructors (ars/)
| arm | change | flag |
|---|---|---|
ar.ab.k2.base |
the recipe of record: whitened cosine on centred vectors | --loss-space whiten |
ar.ab.k2.unitcos |
unit-norm the AR output before the loss, pinning the free scale beta | --head-output unitnorm |
ar.ab.k2.rawcos_c |
drop whitening, keep mean-centring: cos(x-mu, T-mu) | --loss-space rawcos_c |
ar.ab.k2.rawcos |
drop both: cos(x, T), no whitening and no centring | --loss-space rawcos |
Inverters (inverters/)
| run | what it is | final val CE |
|---|---|---|
ao.ab.k2.ds.base |
lens reading the base AR (parent segment) | 1.4724 |
ao.ab.k2.ds.base.x |
base lens run to saturation on the full w5+6 pool (teacher for the 4-bullet student) | 1.1996 |
ao.ab.k2.ds.unitcos |
lens reading the unitcos AR | 1.4827 |
ao.ab.k2.ds.unitcos.x |
unitcos lens, saturation extension (teacher) | 1.2323 |
ao.ab.k2.ds.rawcos_c |
lens reading the rawcos_c AR | 1.4209 |
ao.ab.k2.ds.rawcos_c.x |
rawcos_c lens, saturation extension (teacher) | 1.1952 |
ao.ab.k2.ds.rawcos |
lens reading the rawcos AR | 1.4160 |
ao.ab.k2.ds.rawcos.x |
rawcos lens run to saturation (teacher) | 1.1341 |
ao.ab.k2.scaled |
injection transform: scale*x (the recipe) | 1.5109 |
ao.ab.k2.raw |
injection transform: x, unscaled | 1.5020 |
ao.ab.k2.unit |
injection transform: alpha*x/ | |
ao.ab.k2.centered |
injection transform: scale*(x-mu) | 1.5125 |
ao.ab.k2.wht |
injection transform: scale*W(x-mu) | 1.5282 |
ao.ab.k2.scramble |
control: vectors shuffled across rows (the no-information floor) | 3.4192 |
ao.ab.k2.w56.gt |
inject TRUE residuals instead of AR reconstructions | 1.5400 |
ao.ab.k2.w56.scaled |
the matched partner of the gt arm: AR reconstructions, same crops | 1.2777 |
ao.ab4.base.sft |
4-bullet student: light distill (K=4, mu=0.0002, 4k rows/layer, 1 epoch) from the base lens endpoint; dictionary embedded by the base AR | β |
ao.ab4.unitcos.sft |
4-bullet student from the unitcos lens endpoint; dictionary embedded by the unitcos AR | β |
ao.ab4.rawcos_c.sft |
4-bullet student from the rawcos_c lens endpoint; dictionary embedded by the rawcos_c AR | β |
ao.ab4.rawcos.sft |
4-bullet student from the rawcos lens endpoint; dictionary embedded by the rawcos AR | β |
How to compare
- Lens CE across arms: only at matched cumulative steps, and only against the same
594-crop
val_pool_iolens_u64carve. The scramble arm is the no-information floor. - AR FVE across arms: only from
results/ruler_*.json, never from training logs. - Never infer lens quality from AR FVE. That is the finding of this round.
How to run
# a reconstructor arm (one flag differs per arm; see the table above)
olens/ab/ar_arm.sh # ARM=base|unitcos|rawcos_c|rawcos
# the lens that reads it
ARM=<arm> bash olens/ab/ao_arm.sh
# extend a lens to saturation on fresh crops
ARM=<arm> TRAIN_AROUT_SHARED=1 bash olens/ab/ao_ext.sh
# score every AR on the frozen ruler, both bases
bash olens/ab/rescore_ars.sh
TRAIN_AROUT_SHARED=1 matters: without it the reconstruction bank is written to
node-local NVMe and dies with the pod, costing ~5h of recompute on resume.