Papers
arxiv:2609.39827

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

Published on Sep 30
· Submitted by
Atsuki Yamaguchi
on Oct 1
Authors:
,
,
,
,
,

Abstract

Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.

Community

We ask whether "pre-pretraining" (PPT) (i.e., training on synthetic non-natural language data briefly prior to pre-training on natural language data) helps at scale. This is the largest ever PPT study.

(a) PPT is claimed to induce structural priors that transfer to natural language grammar. (b) Prior work sits at or below 1B parameters and short training horizons. We span four scales, four PT mixtures, five PPT tasks, and up to 100B PT tokens to test the effectiveness of PPT in a more realistic setting.

The short answer is "yes". We confirm the efficacy of PPT in terms of downstream performance (both zero-shot & few-shot) up to 7B models.

The benefit of PPT is not tied to specific PPT tasks (e.g., k-Shuffle Dyck (see above Fig (a))). We further test three more recent PPT tasks. What we find is that PPT tasks capturing long-range retrieval well often result in better LM capabilities.

We also conduct various pre-training data sweep experiments. For details, please check out our preprint!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.39827
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39827 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39827 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39827 in a Space README.md to link it from this page.

Collections including this paper 1