Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Paper • 2609.38536 • Published 3 days ago • 9
Composable Decoding on the Probability Simplex: Theory and Implementation Paper • 2609.34992 • Published 4 days ago • 9
The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning Paper • 2610.00332 • Published 3 days ago • 9
iSDFT: Info-Proximal Self-Distillation Fine-Tuning Collection iSDFT Trained Models • 66 items • Updated 10 days ago • 4
The Y-Combinator for LLMs: Solving Long-Context Rot with λ-Calculus Paper • 2603.20105 • Published Mar 20 • 37
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers Paper • 2602.18292 • Published Feb 20 • 14
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening Paper • 2601.21590 • Published Jan 29 • 14
Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing Paper • 2512.04829 • Published Dec 4, 2025 • 12
Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective Paper • 2509.22921 • Published Sep 26, 2025 • 12
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving Paper • 2507.02726 • Published Jul 3, 2025 • 14