Text-pretrained baselines for Self-Play Pretraining with Zero Data

These are byte-level language models trained on ordinary web text (DCLM-Baseline). They use exactly the architecture and model sizes of the self-play learners from Self-Play Pretraining with Zero Data (weights). I trained them to answer a question the paper leaves open: does ordinary pretraining produce the same in-context learning that self-play does, at the same size and the same number of training tokens?

Short answer: no. Each gets a different kind. The full comparison, code and figures are at github.com/mihir-s-05/icl-selfplay-vs-text. Every ICL score is in the companion dataset rihim/icl-selfplay-vs-text-results.

What's in the repo

dclm/<size>/seed-<s>/       trained from scratch on DCLM text
sp2dclm/<size>/seed-0/      started from the paper's self-play checkpoint (round 8191), then DCLM
    config.json             model architecture ({"learner": {...}})
    learner_<tokens>M.pth   {"learner_state_dict", "tokens", "round", "arm", "init_from"}
    log.jsonl               training log: loss, LR, grad norm, validation bits/byte
    run.json                full training configuration

This is the same layout, file format and model class as the paper's own checkpoints, so anything that loads theirs loads these.

Checkpoint schedule.

  • DCLM runs save at 0.1, 0.25, 0.5 and 1B tokens, then at 1.6, 3.2, 6.4, 12.9 and 17.7B. The last five are exactly the learner-token counts of the paper's self-play rounds 256, 512, 1024, 2048 and 2816 (1536 programs × 4096 bytes per round). That makes checkpoint n here directly comparable to self-play round n. Round 2816 (17.7B tokens) is the one the paper uses for its ICL results.
  • Warm starts save at 0.05, 0.1, 0.25 and 0.5B tokens of DCLM.

The runs

Size Non-emb. params Arch (d / heads / layers) DCLM seeds LR Val bits/byte @ 17.7B Warm start: val bits/byte @ 0.5B (scratch @ 0.5B)
100k 65,728 64 / 1 / 1 1 2e-3 2.495 2.786 (2.662)
500k 492,160 128 / 2 / 2 1 4e-3 1.882 2.116 (2.205)
1M 984,192 128 / 2 / 4 2 2e-3 1.711, 1.702 1.999 (2.099)
3M 3,016,960 256 / 4 / 4 2 2e-3 1.584, 1.581 1.821 (1.984)
6M 6,033,664 256 / 4 / 8 2 2e-3 1.506, 1.504 1.741 (1.778)
24M 24,253,184 512 / 8 / 8 1 5e-4 1.354 1.560 (1.666)

Validation is on 64 sequences of a DCLM sample disjoint from training. The paper's own self-play learners score about 6.2 to 7.6 bits/byte on the same kind of text zero-shot. Warm starts beat scratch at equal tokens at every size except 100k, consistent with the paper's Fig. 6.

Training.

  • Data: 18.5 GB of DCLM-Baseline 1.0 text, skipping every file the paper's evaluation corpus draws from. Each sequence is the byte O followed by 4095 text bytes, which is the paper's scoring convention.
  • Optimiser and schedule:
    • AdamW (β 0.9/0.95), weight decay 0.1 on all parameters.
    • Batch 64 × 4096 bytes.
    • 2% linear warmup, then constant LR (no decay), so every checkpoint is a plain mid-run model, like the self-play ones.
    • bf16 autocast and torch.compile, on one A100 80GB.
  • LR choice: picked per size from a 3-point sweep of short runs (0.3 to 0.5B tokens).
  • Initialisation: GPT-2 style, normal(0, 0.02) with scaled residual projections.
  • Warm starts: LR 3e-3, the value the paper tuned for its own warm starts.

Loading one

You need the model class from the paper's repo (scoring/src/framework/model.py):

import json, torch
from huggingface_hub import hf_hub_download
from src.framework.model import ProgramLanguageModel   # from nourya-aliz/self_play_pretraining/scoring

repo = "rihim/icl-selfplay-vs-text-checkpoints"
cfg = json.load(open(hf_hub_download(repo, "dclm/6M/seed-0/config.json")))["learner"]
blob = torch.load(hf_hub_download(repo, "dclm/6M/seed-0/learner_17723M.pth"),
                  map_location="cpu", weights_only=False)
model = ProgramLanguageModel(**cfg)
model.load_state_dict(blob["learner_state_dict"], strict=False)   # same call as for the paper's checkpoints
model.eval()

text = b"O" + "The capital of France is".encode()  # every sequence starts with the byte 'O'
ids = torch.tensor([list(text)])
next_byte = model(ids)[0][0, -1].argmax().item()

To get everything: hf download rihim/icl-selfplay-vs-text-checkpoints --local-dir ckpts.

Headline results (in-context learning at 17.7B tokens)

Mean exact-match accuracy over each suite's tasks. Self-play is the paper's checkpoints (4 seeds, 95% CI); DCLM is these models (one number per seed).

Size Printable-text tasks: self-play Printable-text tasks: DCLM Word tasks: self-play Word tasks: DCLM
1M 0.30 ± 0.15 0.17, 0.13 0.31 ± 0.06 0.51, 0.52
3M 0.27 ± 0.16 0.15, 0.15 0.36 ± 0.08 0.63, 0.51
6M 0.44 ± 0.04 0.24, 0.14 0.40 ± 0.08 0.59, 0.60
24M 0.41 ± 0.10 0.15 0.41 ± 0.06 0.69

What this shows:

  • Self-play wins procedural tasks: copying by position, stack tracking, comparisons.
  • These text models win lookup and meaning: key→value recall, and categorising words they haven't been shown.
  • Both gaps widen with size.

Details, caveats and per-task numbers are in the write-up.

Limitations

  • Small models: these are tiny research models, not useful as general language models.
  • Seeds: there are only 1 to 2 seeds per size.
  • Schedule: the LR is constant (no cooldown), so the final checkpoints are not fully converged for their token budget.

Credit

The architecture, model class and self-play method are from Self-Play Pretraining with Zero Data (arXiv 2609.30063). Training data is DCLM-Baseline 1.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train rihim/icl-selfplay-vs-text-checkpoints

Paper for rihim/icl-selfplay-vs-text-checkpoints