Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
Abstract
On-policy self-distillation improves reasoning models primarily through context-induced teacher behavior rather than privileged access to reference solutions, as shown by substituting unrelated problem-solution pairs.
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with OP^{2}SD (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, OP^{2}SD improves over the base model, remains competitive with OPSD. The success of OP^{2}SD implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.
Get this paper in your agent:
hf papers read 2608.09228 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper