Papers
arxiv:2609.39687

Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

Published on Sep 30
· Submitted by
dingyi
on Oct 2
Authors:
,
,
,
,
,
,
,
,

Abstract

On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.

Community

Paper submitter

Hi everyone! We’re excited to introduce our new work, N-OPSD.
If the same model sees the same reference solution, is the help it can offer a student always the same?
In on-policy self-distillation (OPSD), the student solves problems on its own, while a teacher with access to the reference solution provides supervision along the student’s generated reasoning trajectory. Typically, this supervision comes from a single set of teacher parameters.
We find that small perturbations to the teacher’s parameters produce experts that offer complementary corrections at different positions. Help that the original teacher does not provide may already exist in its parameter neighborhood.
N-OPSD brings these complementary strengths together: select complementary experts offline, then decide online whom to learn from—and how much support to receive—at each step.
🚀 Offline: Building a Team of Complementary Teachers
When choosing a teacher, we ask what it adds to the team.
We evaluate candidate experts at the same positions in reference solutions, measuring how much each expert increases the probability of the correct next token relative to a fixed base model. We then filter these gains and use them for expert selection.
Experts are selected one at a time based on their marginal contribution to the current pool. After each addition, we update the best correction the pool can provide at every position. The next expert receives credit only for improvements beyond what the pool already offers.
If two teachers consistently provide similar help at the same positions, the second adds little. Experts that fill gaps or offer stronger corrections are more valuable.
Our question is: how much more help can the entire team provide when this expert joins?
We compare this approach with random selection, ranking by problem-solving accuracy, prioritizing coverage of correctly solved problems, and ranking by each expert’s total correction gain. Matched experiments show that selecting complementary experts by marginal contribution, while accounting for overlapping corrections, leads to better student performance.
🚀 Online: Choosing the Direction, Then the Strength
The student learns along its own reasoning trajectory. The expert pool stays fixed, but the source of supervision can change at every generation step.
✦ First, choose the direction
MaxPeak identifies the token with the highest probability across all expert predictions and uses it as an anchor. A direction proposed by a single expert can therefore be adopted without majority agreement.
✦ Then, choose the strength
Among experts whose top-choice token matches the anchor, we compare the gap between each expert’s anchor probability and the student’s, then select an expert using a quantile rule.
These experts agree on the direction but differ in how strongly they support it. The student learns from the selected expert’s full next-token distribution, using the same clipped training objective as OPSD.
With the same expert pool, we compare this approach against averaging expert distributions, random routing, majority-vote consensus, and always choosing the expert with the highest peak probability. The results support choosing direction and strength separately.
💡 The expert with the highest peak probability does not necessarily provide the best training target.
Greater teacher confidence does not guarantee more student learning. Once complementary experts have been selected, how their supervision is used matters just as much.
📊 Main Results
We evaluate three Qwen3 model sizes on AIME 2024, AIME 2025, and HMMT February 2025.
N-OPSD outperforms the comparison methods in our paper across all nine model-size–benchmark combinations. Averaged over three independent training runs, the gains over OPSD in the mean score across the three benchmarks range from 1.67 to 2.75 percentage points across model sizes.
During training, N-OPSD draws complementary supervision from nearby experts. At inference time, only the single distilled student is needed.
Our paper is now public. We welcome your thoughts, questions, and critiques!

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39687 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39687 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39687 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.