iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs Paper • 2609.24646 • Published 12 days ago • 9
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 4 days ago • 10
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 5 days ago • 10
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 4 days ago • 10
Sleeping Agents First Agent Template ⚡ Chat with an AI that can code, search the web, and create images