Papers
arxiv:2609.32722

Scaling Properties of Same-Family On-Policy Distillation

Published on Sep 26
Β· Submitted by
Yuntai Bao
on Sep 30
Authors:
,
,
,
,
,
,

Abstract

*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, G) rises approximately linearly in d=mathrm{KL(Ο€_ΞΈVert Ο€_{ref})}, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how G_{peak} and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.

Community

Paper author Paper submitter

Author here πŸ‘‹

We put together an interactive project page for this paper β€” you can explore all 25 gold-score trajectories (with a training replay), the full 5Γ—5 peak grid, and a calculator that predicts the peak from student size, teacher size, and teacher score:
🌐 https://colored-dye.github.io/blog/2026/opd-scaling/

A short summary thread is on X:
🐦 https://x.com/colored_dye/status/2105290534969561508

We have also released the training checkpoints (teachers, students, and baselines, 311 models across Qwen2.5 0.5B–14B): https://huggingface.co/colored-dye/OPD-scaling-checkpoints

Paper author Paper submitter
β€’
This comment has been hidden (marked as Resolved)

this is actually interesting

this research is great

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32722
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32722 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32722 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.