| license: gpl-3.0 | |
| tags: | |
| - kv-cache | |
| - llm-inference | |
| - optimization | |
| - compression | |
| - mla | |
| - deepseek | |
| language: | |
| - zh | |
| - en | |
| # 📄 KV Cache Compression for Text-Only LLMs | |
| **纯文本大语言模型 KV 缓存压缩 — 完整技术报告** | **Full Technical Reports** | |
| ## Contents | |
| | File | Description | | |
| |------|-------------| | |
| | `mla_absorbed_cache_report.md` | 🚀 MLA 吸收式缓存优化(中文) | | |
| | `mla_absorbed_cache_report_en.md` | 🚀 **Absorbed MLA Cache Optimization (English)** | | |
| | `kv_cache_compression_report.md` | 📊 KV 缓存压缩完整方案(中文) | | |
| | `kv_compress_plan.md` | 🗺 详细实施计划(中文) | | |
| ## 🚀 Highlight: Absorbed MLA Cache (270KB → 8.4KB/token, 32×) | |
| DeepSeek-V2-Lite MLA 优化,在 L40S 实测: | |
| | Approach | KV/token | Compression | Error | | |
| |----------|----------|-------------|-------| | |
| | Standard MHA | 270 KB | 1x | — | | |
| | Absorbed MLA | 30.4 KB | 8.9x | 0 | | |
| | + per-channel INT8 | 15.2 KB | 17.8x | 0.011 | | |
| | + INT4 (extreme) | **8.4 KB** | **32x** | 0.112 | | |
| **Single L40S: 1.24M tokens context.** | |
| ## Key Insight | |
| > Text-only models can be aggressively compressed (quantization + eviction), while reasoning models (R1-class) must be conservative (quantization + sliding window + tiered storage). **Scenario classification matters more than the algorithm itself.** | |
| --- | |
| *Cloud LTE Studio · 2026-08-08 · GPL-3.0 License* | |