Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 8 days ago • 65
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published 23 days ago • 36
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 27 days ago • 114
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 178
EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory Paper • 2606.21649 • Published Jun 19 • 36
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published Jul 23 • 40