PFLM: A Prior-Fitted Language Model

PFLM is a 300M-parameter byte-level model pretrained only on samples from a synthetic non-linguistic prior. Given the start of a byte sequence, such as a text in a language it has never seen, it infers the source in context and predicts what comes next.

This repository holds the weights and their configuration. The model code and a byte-level API for scoring streams and measuring in-context learning are at github.com/cbl/prior-fitted-language-model.

Quickstart

pip install "pflm1[hf]"

Watch it learn. The model has never seen prime numbers, yet it gets better at predicting them the more it reads.

import pflm1
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("lennartcb/pflm1", dtype="bfloat16").cuda().eval()

def primes(n):                                   # 1 where the integer is prime, else 0
    flags = bytearray([1]) * n
    flags[:2] = b"\0\0"
    for i in range(2, int(n ** 0.5) + 1):
        if flags[i]:
            flags[i * i::i] = bytes(len(flags[i * i::i]))
    return bytes(48 + f for f in flags)

text = primes(10_000)                            # "0011010100..." ten thousand digits
bits = model.bits_per_byte(text)                 # bits per byte, one entry each
print(bits[:500].mean(), bits[-500:].mean())     # the first 500 digits vs. the last 500

Citation

Learning to Learn a Language

@article{carstensbehrens2026learning,
  title   = {Learning to Learn a Language},
  author  = {Carstens-Behrens, Lennart and Fr{\"o}hlich, Holger},
  journal = {arXiv preprint arXiv:2610.05879},
  year    = {2026},
}
Downloads last month
327
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for lennartcb/pflm1