IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse Paper • 2603.12201 • Published Mar 12 • 64
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 26 days ago • 81
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference Paper • 2605.25475 • Published May 25
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging Paper • 2506.23266 • Published Jun 29, 2025 • 1
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 26 days ago • 81
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models Paper • 2602.02244 • Published Feb 2 • 2
Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models Paper • 2602.02244 • Published Feb 2 • 2
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning Paper • 2606.10968 • Published Jun 9 • 42