ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 12 days ago • 66
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 28 days ago • 97
Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning Paper • 2606.29985 • Published Jun 29 • 19
LiSA: Lifelong Safety Adaptation via Conservative Policy Induction Paper • 2605.14454 • Published May 14 • 5
ReflectCAP: Detailed Image Captioning with Reflective Memory Paper • 2604.12357 • Published Apr 14 • 2
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? Paper • 2603.24472 • Published Mar 25 • 58
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty Paper • 2603.15500 • Published Mar 16 • 12
CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution Paper • 2602.07918 • Published Feb 8 • 6
Drift: Decoding-time Personalized Alignments with Implicit User Preferences Paper • 2502.14289 • Published Feb 20, 2025 • 1
FlowRL: Matching Reward Distributions for LLM Reasoning Paper • 2509.15207 • Published Sep 18, 2025 • 119
Critic-Guided Decoding for Controlled Text Generation Paper • 2212.10938 • Published Dec 21, 2022 • 2