SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving Paper • 2609.34117 • Published 10 days ago • 12 • 4
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization Paper • 2402.18096 • Published Feb 28, 2024 • 1
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs Paper • 2512.17970 • Published Dec 19, 2025
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention Paper • 2602.23057 • Published Feb 26
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding Paper • 2603.03333 • Published Feb 11
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models Paper • 2206.09557 • Published Jun 20, 2022
An Investigation of FP8 Across Accelerators for LLM Inference Paper • 2502.01070 • Published Feb 3, 2025 • 3
Faster Inference of LLMs using FP8 on the Intel Gaudi Paper • 2503.09975 • Published Mar 13, 2025 • 1
SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving Paper • 2609.34117 • Published 10 days ago • 12
SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving Paper • 2609.34117 • Published 10 days ago • 12
SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving Paper • 2609.34117 • Published 10 days ago • 12