YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

OLens ablation round β€” reconstructors, inverters, and the rulers that disagree

Every arm here differs from the recipe of record by exactly one flag. Data, pool, crop seed, layer set, learning rate, batch and seeds are held fixed across arms so the only free variable is the one named.

The headline

The whitened-FVE ruler that the program optimises and reports anti-correlates with lens quality. The AR trained without whitening scores far worse on the ruler and makes a materially better lens, and running both to saturation does not rescue it.

Lens quality (val CE on the 594-crop u64 carve, lower is better)

arm cum steps examples span tokens val CE
base 42,899 43,928,576 782M 1.1996
unitcos 28,351 29,031,424 515M 1.2323
rawcos_c 29,351 30,055,424 533M 1.1952
rawcos 42,899 43,928,576 782M 1.1341
tf.scaled 15,516 15,888,384 281M 1.3734
tf.raw 15,516 15,888,384 281M 1.3547

Compare arms only at MATCHED cumulative steps. Arms stopped early by a budget cut have lower totals, and reading their endpoints against a fully-trained arm inverts the ordering: unitcos looks worse than base on endpoints but is identical to it at matched steps.

Reconstructor quality (frozen 8192-row ruler, identical rows, two whitener bases)

AR FVE (chat basis) FVE (pooled basis) whitened-cos loss (chat)
ar.asst.ptag.chat.k2.s3_final 0.1830 0.1957 1.2388
ar.ab.k2.unitcos_final 0.1656 0.1783 1.2771
ar.ab.k2.base_final 0.1652 0.1779 1.2779
ar.asst.ptag.chat.k2.s0_ex11404800 0.1606 β€” 1.2893
ar.ab.k2.rawcos_c_final 0.1102 0.1243 1.4066
ar.ab.k2.rawcos_final 0.0920 0.1022 1.5067
ar.asst.ptag.chat.k2.s0_ex2880000 0.0752 β€” 1.5039
ar.xm.4b.p1_ex11405184 0.0526 β€” 1.5763
ar.xm.4b.p1_ex2880384 0.0335 β€” 1.6622
ar.xm.27b.p1.affmap_ex11405184 0.0088 β€” 1.8424
ar.xm.27b.p1.affmap.kaiming_ex11405184 0.0084 β€” 1.8455
ar.xm.27b.p1.affmap_ex2880384 0.0066 β€” 1.8663

These are the only comparable AR numbers. Each run also logs a val_fve during training, but every arm carves its own val set from its own wave, so those share no ruler and must never be compared across arms. Note the ordering here is the REVERSE of the lens table above: the AR that scores worst on this ruler (rawcos) makes the best lens.

RL'd arms (rl/)

Each saturated lens was lightly distilled into a 4-bullet student (dictionary labels: K=4, length penalty mu=0.0002, 4,000 rows/layer x 21 layers, 1 epoch, warm-started from the lens endpoint; the dictionary embedded by that arm's OWN AR), then GRPO'd for 300 steps with that arm's own AR as the reward. Everything else -- activations labelled, labelling whitener, phrase dictionary, student recipe, RL data/config, and the fixed 231-item ruler scored through the AR of record -- is identical across arms. Ruler numbers, when scored, are in results/eval_table.md.

checkpoint ruler mean FVE median L20-32 / L36-48 / L52-60 no-bullet bullets/gen % 3 / 4 / 5 / 6+ bullet tok train reward@step (FVE, KL, bullets, len, cut)
RL rl.ab4.unitcos.g0 step 100 0.2782 0.2666 0.2447 / 0.2602 / 0.3414 0.0% 4.009 1.7 / 95.7 / 2.6 / 0.0 9.78 0.178 (0.178, 0.3685, 4.00, 73, 0%)
RL rl.ab4.unitcos.g0 step 200 0.3003 0.2865 0.2589 / 0.2773 / 0.3792 0.0% 3.981 4.1 / 93.3 / 2.2 / 0.2 10.14 0.246 (0.246, 0.3689, 3.99, 86, 0%)
RL rl.ab4.unitcos.g0 step 300 0.3183 0.3109 0.276 / 0.2974 / 0.3956 0.0% 4.05 6.7 / 82.0 / 10.0 / 1.1 11.43 0.301 (0.301, 0.4070, 3.98, 100, 0%)

What each arm changed

Reconstructors (ars/)

arm change flag
ar.ab.k2.base the recipe of record: whitened cosine on centred vectors --loss-space whiten
ar.ab.k2.unitcos unit-norm the AR output before the loss, pinning the free scale beta --head-output unitnorm
ar.ab.k2.rawcos_c drop whitening, keep mean-centring: cos(x-mu, T-mu) --loss-space rawcos_c
ar.ab.k2.rawcos drop both: cos(x, T), no whitening and no centring --loss-space rawcos

Inverters (inverters/)

run what it is final val CE
ao.ab.k2.ds.base lens reading the base AR (parent segment) 1.4724
ao.ab.k2.ds.base.x base lens run to saturation on the full w5+6 pool (teacher for the 4-bullet student) 1.1996
ao.ab.k2.ds.unitcos lens reading the unitcos AR 1.4827
ao.ab.k2.ds.unitcos.x unitcos lens, saturation extension (teacher) 1.2323
ao.ab.k2.ds.rawcos_c lens reading the rawcos_c AR 1.4209
ao.ab.k2.ds.rawcos_c.x rawcos_c lens, saturation extension (teacher) 1.1952
ao.ab.k2.ds.rawcos lens reading the rawcos AR 1.4160
ao.ab.k2.ds.rawcos.x rawcos lens run to saturation (teacher) 1.1341
ao.ab.k2.scaled injection transform: scale*x (the recipe) 1.5109
ao.ab.k2.raw injection transform: x, unscaled 1.5020
ao.ab.k2.unit injection transform: alpha*x/
ao.ab.k2.centered injection transform: scale*(x-mu) 1.5125
ao.ab.k2.wht injection transform: scale*W(x-mu) 1.5282
ao.ab.k2.scramble control: vectors shuffled across rows (the no-information floor) 3.4192
ao.ab.k2.w56.gt inject TRUE residuals instead of AR reconstructions 1.5400
ao.ab.k2.w56.scaled the matched partner of the gt arm: AR reconstructions, same crops 1.2777
ao.ab4.base.sft 4-bullet student: light distill (K=4, mu=0.0002, 4k rows/layer, 1 epoch) from the base lens endpoint; dictionary embedded by the base AR β€”
ao.ab4.unitcos.sft 4-bullet student from the unitcos lens endpoint; dictionary embedded by the unitcos AR β€”
ao.ab4.rawcos_c.sft 4-bullet student from the rawcos_c lens endpoint; dictionary embedded by the rawcos_c AR β€”
ao.ab4.rawcos.sft 4-bullet student from the rawcos lens endpoint; dictionary embedded by the rawcos AR β€”

How to compare

  1. Lens CE across arms: only at matched cumulative steps, and only against the same 594-crop val_pool_iolens_u64carve. The scramble arm is the no-information floor.
  2. AR FVE across arms: only from results/ruler_*.json, never from training logs.
  3. Never infer lens quality from AR FVE. That is the finding of this round.

How to run

# a reconstructor arm (one flag differs per arm; see the table above)
olens/ab/ar_arm.sh              # ARM=base|unitcos|rawcos_c|rawcos
# the lens that reads it
ARM=<arm> bash olens/ab/ao_arm.sh
# extend a lens to saturation on fresh crops
ARM=<arm> TRAIN_AROUT_SHARED=1 bash olens/ab/ao_ext.sh
# score every AR on the frozen ruler, both bases
bash olens/ab/rescore_ars.sh

TRAIN_AROUT_SHARED=1 matters: without it the reconstruction bank is written to node-local NVMe and dies with the pod, costing ~5h of recompute on resume.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support