Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 3 days ago • 15
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 8 days ago • 108
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 8 days ago • 178
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 8 days ago • 70
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 11 days ago • 50
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 9 days ago • 62
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 9 days ago • 79