Binomial-Ordering SAEs (Pythia & OPT-BabyLM)
Sparse autoencoders over the residual streams of six small language models, built for interpretability research on English binomial ordering preferences ("salt and pepper" vs "pepper and salt"): which internal features bias a model's word-order choice for pairs it has never encountered.
Models
Each folder holds one SAE of the residual-stream input to transformer
block layer (equivalently hidden_states[layer] in Hugging Face terms,
blocks.{layer}.hook_resid_pre in transformer_lens):
| folder | base model | block | d_in | d_sae | k (batch-mean L0) | val EV | dead |
|---|---|---|---|---|---|---|---|
| babylm-125m | znhoughton/opt-babylm-125m-20eps-seed964 | 8 | 768 | 24576 | 64 | 0.897 | 5 |
| babylm-350m | znhoughton/opt-babylm-350m-20eps-seed964 | 16 | 1024 | 32768 | 64 | 0.841 | 2 |
| babylm-1.3b | znhoughton/opt-babylm-1.3B-20eps-seed964 | 16 | 2048 | 65536 | 64 | 0.737 | 2 |
| pythia-160m | EleutherAI/pythia-160m | 8 | 768 | 24576 | 64 | 0.928 | 13 |
| pythia-410m | EleutherAI/pythia-410m | 10 | 1024 | 32768 | 64 | 0.966 | 55 |
| pythia-1.4b | EleutherAI/pythia-1.4b | 11 | 2048 | 65536 | 64 | 0.929 | 36 |
Protocol (identical for all six): sae_lens BatchTopK, dictionary = 32x model
width, k = 64, Adam (lr 1e-4, beta2 0.9999) with cosine decay + warmup,
1024-token context, aux-reconstruction loss for dead-latent revival, trained
on a shared pinned Wikipedia corpus (98M tokens; sha256 cd0c879c...),
stopped at the plateau of held-out reconstruction explained-variance on a
2M-token validation split (patience 3 x 0.05%, floor 64M tokens, cap 400M).
Training was exported weights reproduce sae_lens's encode bit-exactly
(max abs diff 0.0, verified on held-out text for all six).
Using these SAEs
Load from the Hub (or SAE.load_from_disk on a local copy of a folder):
from huggingface_hub import snapshot_download
from sae_lens import SAE
path = snapshot_download( # fetches one folder (~150 MB-1 GB)
"znhoughton/binomial-abspref-saes",
allow_patterns=["pythia-160m/*"], # one of the six folders above
)
sae = SAE.load_from_disk(f"{path}/pythia-160m") # sae_lens >= 6.x
Anonymous downloads work, but the Hub sometimes rate-limits unauthenticated
traffic (HTTP 429) โ set HF_TOKEN if you hit that. This load path was
tested end-to-end with sae_lens 6.51: encode + decode against the
pythia-160m base model gave reconstruction EV 0.99 on binomial sentences.
Encode the residual stream at the block listed in the table:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("EleutherAI/pythia-160m")
model = AutoModelForCausalLM.from_pretrained("EleutherAI/pythia-160m").eval()
enc = tok(["Bread and butter are common breakfast staples."],
return_tensors="pt")
with torch.no_grad():
hidden = model(**enc, output_hidden_states=True).hidden_states[8] # block 8 input
acts = sae.encode(hidden) # (batch, seq, d_sae), thresholded activations
recon = sae.decode(acts) # reconstruction of the residual stream
For causal experiments (e.g. testing whether latent j biases ordering),
ablate or steer it inside the forward pass: at the block input, either
subtract latent j's contribution (acts[..., j] -> 0, then
x + sae.decode(corrected_acts) - x) or add c * w_dec[j] and re-measure
the model's ordering score.
Each folder contains:
sae_weights.safetensors,cfg.jsonโ weights (W_enc,W_dec,b_enc,b_dec,threshold) and the sae_lens configtrain_meta.jsonโ final audit: stop reason (validation plateau), tokens trained, val EV, L0, dead-latent count, corpus hashesval_history.jsonโ the full held-out EV curve used for early stopping
Semantics note: trained with batch-mean top-k (k = 64); inference uses the exported fixed threshold (JumpReLU-style), so the per-token number of active latents fluctuates around k.
Citation
The three OPT-based BabyLM checkpoints used as base models
(znhoughton/opt-babylm-*) were trained in the project behind this paper.
The SAEs in this repo are separate artifacts and are not discussed in it:
@article{houghton2026holistic,
title = {The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models},
author = {Houghton, Zachary Nicholas and Zhou, Yu and Pluth, Dan and Hosier, Jordan and Gurbani, Vijay K.},
journal = {arXiv preprint arXiv:2606.13993},
year = {2026}
}
Model tree for znhoughton/binomial-abspref-saes
Base model
EleutherAI/pythia-1.4b