Papers
arxiv:2609.33791

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Published on Sep 27
· Submitted by
Wenze Lin
on Sep 29
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.

Community

Paper submitter

We revisit a seemingly fundamental assumption in on-policy distillation: do we really need KL divergence?

Our experiments suggest that the precise magnitude of the KL signal may be less important than expected. What matters most is preserving the update direction toward the teacher, particularly on a small subset of tokens with strong teacher–student disagreement.

Building on this finding, we further introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD), where each sample is supervised by all teachers rather than routed to a single one, leading to consistent improvements on both math and code tasks.

Code is available — feedback and discussions are very welcome!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33791
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33791 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33791 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33791 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.