iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs Paper • 2609.24646 • Published 13 days ago • 9
Beam Search as Test-Time Self-Distillation via Counterfactual Contexts Paper • 2609.37041 • Published 5 days ago • 7
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 5 days ago • 10
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 6 days ago • 10
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 5 days ago • 10
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks Paper • 2410.05102 • Published Oct 7, 2024 • 1