WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models Paper • 2609.23033 • Published 8 days ago • 8
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems Paper • 2609.34060 • Published 6 days ago • 4
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction Paper • 2609.34054 • Published 6 days ago • 4
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference Paper • 2605.19218 • Published May 19 • 2
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Paper • 2605.16839 • Published May 16 • 13
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding Paper • 2602.23881 • Published Feb 27 • 18
RelayGen: Intra-Generation Model Switching for Efficient Reasoning Paper • 2602.06454 • Published Feb 6 • 12
L4Q: Parameter Efficient Quantization-Aware Training on Large Language Models via LoRA-wise LSQ Paper • 2402.04902 • Published Feb 7, 2024 • 5
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents Paper • 2602.01053 • Published Feb 1 • 8
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection Paper • 2602.03216 • Published Feb 3 • 14
Retrospective Sparse Attention for Efficient Long-Context Generation Paper • 2508.09001 • Published Aug 12, 2025 • 4
Kimi Linear: An Expressive, Efficient Attention Architecture Paper • 2510.26692 • Published Oct 30, 2025 • 138
Exploring Conditions for Diffusion models in Robotic Control Paper • 2510.15510 • Published Oct 17, 2025 • 40
LightMem: Lightweight and Efficient Memory-Augmented Generation Paper • 2510.18866 • Published Oct 21, 2025 • 116
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning Paper • 2510.19338 • Published Oct 22, 2025 • 116