Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning Paper • 2511.04654 • Published Nov 6, 2025
Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models Paper • 2510.23334 • Published Oct 27, 2025
Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models Paper • 2608.25311 • Published Aug 26
Self-Adapting Group of Experts for Multi-Agent Reasoning Paper • 2609.35412 • Published 14 days ago
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation Paper • 2608.09228 • Published Aug 10
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings Paper • 2602.00574 • Published Jan 31