Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Abstract
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
Community
Training a diffusion LM doesn't need every token, just the few pivots that shape the rest of the generation. ๐ฏ
Excited to share Pivot-SD (EMNLP 2026 Oral), our new approach to efficient self-distillation for masked diffusion LMs!
Instead of training on the full sequence or running expensive online RL, Pivot-SD uses Information Gain to find the "pivots", the critical denoising steps where uncertainty collapses. We reward successful pivots and penalize failed ones, leaving the rest of the trajectory untouched.
The result? With just 200 questions, Pivot-SD outperforms full-sequence SFT and online RL baselines on math & code, using less compute! ๐
Would love for you to check it out:
๐ Paper: https://arxiv.org/abs/2610.03665
๐ป Project Page: https://sunwoohong.github.io/pivot-sd
Get this paper in your agent:
hf papers read 2610.03665 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper