Binomial-Ordering SAEs (Pythia & OPT-BabyLM)

Sparse autoencoders over the residual streams of six small language models, built for interpretability research on English binomial ordering preferences ("salt and pepper" vs "pepper and salt"): which internal features bias a model's word-order choice for pairs it has never encountered.

Models

Each folder holds one SAE of the residual-stream input to transformer block layer (equivalently hidden_states[layer] in Hugging Face terms, blocks.{layer}.hook_resid_pre in transformer_lens):

folder base model block d_in d_sae k (batch-mean L0) val EV dead
babylm-125m znhoughton/opt-babylm-125m-20eps-seed964 8 768 24576 64 0.897 5
babylm-350m znhoughton/opt-babylm-350m-20eps-seed964 16 1024 32768 64 0.841 2
babylm-1.3b znhoughton/opt-babylm-1.3B-20eps-seed964 16 2048 65536 64 0.737 2
pythia-160m EleutherAI/pythia-160m 8 768 24576 64 0.928 13
pythia-410m EleutherAI/pythia-410m 10 1024 32768 64 0.966 55
pythia-1.4b EleutherAI/pythia-1.4b 11 2048 65536 64 0.929 36

Protocol (identical for all six): sae_lens BatchTopK, dictionary = 32x model width, k = 64, Adam (lr 1e-4, beta2 0.9999) with cosine decay + warmup, 1024-token context, aux-reconstruction loss for dead-latent revival, trained on a shared pinned Wikipedia corpus (98M tokens; sha256 cd0c879c...), stopped at the plateau of held-out reconstruction explained-variance on a 2M-token validation split (patience 3 x 0.05%, floor 64M tokens, cap 400M). Training was exported weights reproduce sae_lens's encode bit-exactly (max abs diff 0.0, verified on held-out text for all six).

Using these SAEs

Load from the Hub (or SAE.load_from_disk on a local copy of a folder):

from huggingface_hub import snapshot_download
from sae_lens import SAE

path = snapshot_download(                       # fetches one folder (~150 MB-1 GB)
    "znhoughton/binomial-abspref-saes",
    allow_patterns=["pythia-160m/*"],           # one of the six folders above
)
sae = SAE.load_from_disk(f"{path}/pythia-160m") # sae_lens >= 6.x

Anonymous downloads work, but the Hub sometimes rate-limits unauthenticated traffic (HTTP 429) โ€” set HF_TOKEN if you hit that. This load path was tested end-to-end with sae_lens 6.51: encode + decode against the pythia-160m base model gave reconstruction EV 0.99 on binomial sentences.

Encode the residual stream at the block listed in the table:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("EleutherAI/pythia-160m")
model = AutoModelForCausalLM.from_pretrained("EleutherAI/pythia-160m").eval()
enc = tok(["Bread and butter are common breakfast staples."],
          return_tensors="pt")
with torch.no_grad():
    hidden = model(**enc, output_hidden_states=True).hidden_states[8]  # block 8 input
acts = sae.encode(hidden)      # (batch, seq, d_sae), thresholded activations
recon = sae.decode(acts)       # reconstruction of the residual stream

For causal experiments (e.g. testing whether latent j biases ordering), ablate or steer it inside the forward pass: at the block input, either subtract latent j's contribution (acts[..., j] -> 0, then x + sae.decode(corrected_acts) - x) or add c * w_dec[j] and re-measure the model's ordering score.

Each folder contains:

  • sae_weights.safetensors, cfg.json โ€” weights (W_enc, W_dec, b_enc, b_dec, threshold) and the sae_lens config
  • train_meta.json โ€” final audit: stop reason (validation plateau), tokens trained, val EV, L0, dead-latent count, corpus hashes
  • val_history.json โ€” the full held-out EV curve used for early stopping

Semantics note: trained with batch-mean top-k (k = 64); inference uses the exported fixed threshold (JumpReLU-style), so the per-token number of active latents fluctuates around k.

Citation

The three OPT-based BabyLM checkpoints used as base models (znhoughton/opt-babylm-*) were trained in the project behind this paper. The SAEs in this repo are separate artifacts and are not discussed in it:

@article{houghton2026holistic,
  title   = {The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models},
  author  = {Houghton, Zachary Nicholas and Zhou, Yu and Pluth, Dan and Hosier, Jordan and Gurbani, Vijay K.},
  journal = {arXiv preprint arXiv:2606.13993},
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for znhoughton/binomial-abspref-saes

Finetuned
(76)
this model

Paper for znhoughton/binomial-abspref-saes