--- license: other license_name: modified-mit license_link: https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE base_model: MiniMaxAI/MiniMax-M2.7 base_model_relation: quantized library_name: vllm pipeline_tag: text-generation tags: - vllm - compressed-tensors - awq - w4a16 - 4bit - marlin - minimax - quantization - moe --- # MiniMax-M2.7 AWQ 4bit (W4A16) W4A16 quantization of [`MiniMaxAI/MiniMax-M2.7`](https://huggingface.co/MiniMaxAI/MiniMax-M2.7), produced with [`llm-compressor`](https://github.com/vllm-project/llm-compressor). - **Format:** `compressed-tensors` pack-quantized, int4 weights / fp16 activations - **Group size:** 128, symmetric - **Calibration:** data-free, MSE observer - **Kept in BF16:** MoE routing gates and `lm_head` only — every other `Linear` is quantized, matching the ignore list from [`cyankiwi/MiniMax-M2.5-AWQ-4bit`](https://huggingface.co/cyankiwi/MiniMax-M2.5-AWQ-4bit). ## vLLM ``` vllm serve demon-zombie/MiniMax-M2.7-AWQ-4bit \ --tensor-parallel-size 4 \ --tool-call-parser minimax_m2 \ --reasoning-parser minimax_m2 \ --enable-auto-tool-choice ```