Papers
arxiv:2609.37915

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

Published on Sep 29
· Submitted by
Md. Ismail Hossain
on Oct 1
Authors:
,
,

Abstract

On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.

Community

Paper submitter

We study the scaling limits of on-policy self-distillation for LLM reasoning, focusing on the challenge of supervising long, unverified student trajectories.
We find that models can often recover from failed prefixes, and use this property to train primarily on verified on-policy trajectories rather than relying on unverified long rollouts.
Our approach, OASIS, removes the need for a privileged ground-truth reasoning solution as teacher context and instead uses the student's own generated attempt.
This makes on-policy distillation possible using only final-answer verification, while still improving reasoning performance across Qwen3-1.7B, 4B, and 8B.
Training code and checkpoints are planned for release upon publication, and we hope this work helps make on-policy self-distillation more scalable and practical.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.37915
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.37915 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.37915 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.37915 in a Space README.md to link it from this page.

Collections including this paper 1