Beam Search as Test-Time Self-Distillation via Counterfactual Contexts Paper • 2609.37041 • Published 4 days ago • 7
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 4 days ago • 10
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 5 days ago • 10
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 4 days ago • 10