Papers
arxiv:2610.03665

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Published on Oct 2
ยท Submitted by
Seo Hyun Kim
on Oct 5
Authors:
,
,
,

Abstract

Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.

Community

Paper author Paper submitter

Training a diffusion LM doesn't need every token, just the few pivots that shape the rest of the generation. ๐ŸŽฏ

Excited to share Pivot-SD (EMNLP 2026 Oral), our new approach to efficient self-distillation for masked diffusion LMs!

Instead of training on the full sequence or running expensive online RL, Pivot-SD uses Information Gain to find the "pivots", the critical denoising steps where uncertainty collapses. We reward successful pivots and penalize failed ones, leaving the rest of the trajectory untouched.   

The result? With just 200 questions, Pivot-SD outperforms full-sequence SFT and online RL baselines on math & code, using less compute! ๐Ÿš€   

Would love for you to check it out:
๐Ÿ“„ Paper: https://arxiv.org/abs/2610.03665
๐Ÿ’ป Project Page: https://sunwoohong.github.io/pivot-sd

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.03665
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.03665 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.03665 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.