Papers
arxiv:2610.05879

Learning to Learn a Language

Published on Oct 5
· Submitted by
Lennart Carstens-Behrens
on Oct 6
Authors:

Abstract

We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.

Community

Paper author Paper submitter

PFLM is a 300M-parameter byte-level model trained without any natural language. Every training sequence comes from a freshly sampled recurrent causal model, so each one is a new "language" and the only way to predict it is to infer its rules from context.

With frozen weights, it learns to predict any language in context. The same model counts, compares numbers, predicts the primes better than their density alone allows, and compresses code, DNA and speech below gzip and PPMd.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.05879
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.05879 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.05879 in a Space README.md to link it from this page.

Collections including this paper 1