WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models Paper • 2609.23033 • Published 4 days ago • 4
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems Paper • 2609.34060 • Published 1 day ago • 3
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction Paper • 2609.34054 • Published 1 day ago • 3
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference Paper • 2605.19218 • Published May 19 • 2
Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference Paper • 2605.19218 • Published May 19 • 2
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Paper • 2605.16839 • Published May 16 • 13
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Paper • 2605.16839 • Published May 16 • 13
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Paper • 2605.16839 • Published May 16 • 13
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding Paper • 2602.23881 • Published Feb 27 • 18
RelayGen: Intra-Generation Model Switching for Efficient Reasoning Paper • 2602.06454 • Published Feb 6 • 12