Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
Abstract
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce Decoupled Credit Self-Distillation (DCSD), which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6\% of tokens and yielding a 1.5times reduction in token credit magnitude.
Community
Can we trust the teacher? ๐ค We introduce Decoupled Credit Self-Distillation (DCSD), which separates credit direction from contribution magnitude to provide more reliable teacher supervision for self-distillation. Across 11 benchmarks, DCSD consistently improves reasoning performance while correcting noisy token-level credit assignment.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents (2026)
- When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment (2026)
- Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation (2026)
- Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization (2026)
- Calibrating Teacher--Student Discrepancy for On-Policy Distillation (2026)
- TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning (2026)
- Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34848 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper