File size: 6,001 Bytes
ec955c2
 
 
 
 
 
 
 
3e692d1
 
 
 
 
 
 
ec955c2
 
e0f45f6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ec955c2
 
72581b9
ec955c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3e692d1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48263b4
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
base_model: MiniMaxAI/MiniMax-M2.7
library_name: mlx
pipeline_tag: text-generation
license: other
license_name: minimax-model-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/LICENSE
tags:
  - minimax
  - m2.7
  - moe
  - quantized
  - rotorquant
  - kv-cache-quantization
  - mlx
---

> [!TIP]
> **KV-cache quantization without any fork (recommended, 2026):** upstream
> llama.cpp/Ollama now cover this natively — use `-ctk q8_0 -ctv q8_0`
> (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
> `-ctk q4_0 -ctv q4_0` (~quarter memory, ≈7.6% perplexity increase). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
> K and V types symmetric to stay on the fast fused Flash-Attention path.
> Since April 2026, mainline llama.cpp also applies Hadamard rotation to
> KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
> which greatly improves low-bit KV quality (opt-out:
> `LLAMA_ATTN_ROT_DISABLE=1`).
>
> The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
> TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
> is unmaintained relative to mainline. It is NOT required to use this model.
<!-- kv-upstream-note -->

# MiniMax-M2.7-RotorQuant-MLX-4bit

**MLX 4-bit quantized variant of [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) with the legacy RotorQuant KV-cache fork (superseded by upstream llama.cpp KV options), optimized for Apple Silicon.**

## Overview

MiniMax-M2.7 is a massive 256-expert Mixture-of-Experts (MoE) model with 8 experts active per token, totaling approximately 456 billion parameters. This variant combines **4-bit MLX weight quantization** with **RotorQuant** KV-cache quantization for deployment on Apple Silicon hardware.

RotorQuant applies a learned Hadamard rotation matrix to keys and values before quantization, smoothing the activation distribution for better quality retention. At 4-bit, RotorQuant's rotation-based approach helps preserve quality where naive quantization would degrade.

| Property | Value |
|---|---|
| Architecture | MoE (256 experts, 8 active/token) |
| Total Parameters | ~456B |
| Layers | 62 |
| Hidden Size | 3072 |
| Attention Heads | 48 |
| Weight Quantization | 4-bit (MLX) |
| KV-Cache Quantization | RotorQuant |
| Estimated Size | ~225 GB |
| Base Model | MiniMaxAI/MiniMax-M2.7 |

## Quickstart

```bash
pip install mlx-lm
```

```python
from mlx_lm import load, generate

model, tokenizer = load("majentik/MiniMax-M2.7-RotorQuant-MLX-4bit")

prompt = "What is a Comprehensive Geriatric Assessment?"
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)

response = generate(
    model,
    tokenizer,
    prompt=text,
    max_tokens=512,
)
print(response)
```

## RotorQuant vs TurboQuant

| Feature | RotorQuant | TurboQuant |
|---|---|---|
| Technique | Rotation-based KV quantization (Hadamard transform) | Asymmetric per-channel KV quantization |
| Throughput | Slightly lower throughput (rotation overhead) | Higher throughput, lower latency |
| Quality | Better quality preservation at low bit-widths | Good quality preservation |
| Best For | Quality-sensitive tasks, research | High-throughput serving, long contexts |

## Memory Estimates (Apple Silicon)

| Variant | Estimated Size | Minimum Unified Memory |
|---|---|---|
| MLX 8-bit | ~456 GB | 512 GB (Mac Studio M2/M3/M4 Ultra) |
| MLX 5-bit | ~280 GB | 384 GB |
| MLX 4-bit | ~225 GB | 256 GB |
| MLX 3-bit | ~170 GB | 192 GB |
| MLX 2-bit | ~110 GB | 128 GB |

> **Note**: 4-bit quantization requires Apple Silicon with 256 GB+ unified memory, such as a Mac Studio with M2/M3/M4 Ultra.

## See Also

- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) -- Base model
- [majentik/MiniMax-M2.7-TurboQuant-MLX-4bit](https://huggingface.co/majentik/MiniMax-M2.7-TurboQuant-MLX-4bit) -- TurboQuant MLX 4-bit
- [majentik/MiniMax-M2.7-RotorQuant-MLX-3bit](https://huggingface.co/majentik/MiniMax-M2.7-RotorQuant-MLX-3bit) -- MLX 3-bit

## Quant trade-off (MLX lane)

| Bits | Approx size | Use case | Recommendation |
|---|---|---|---|
| 2-bit | ~119 GB | Aggressive quantization | Very low-RAM Macs |
| 3-bit | ~164 GB | Lossy but small | Low-RAM Macs |
| **4-bit** | ~192 GB | Balanced default | **Recommended for most Macs** |
| 5-bit | ~228 GB | Higher fidelity | Quality-sensitive |
| 6-bit | ~274 GB | Approaching FP16 quality | High-fidelity |
| 8-bit | ~347 GB | Near-lossless reference | Fidelity-critical work |

(Current variant — **4bit** — is bolded.)

## Variants in this family

(Showing 12 sibling variants under `majentik/minimax-m2.7-*`. The current variant — `RotorQuant-MLX-4bit` — is **bolded**.)

| Variant | Runtime | Approx size | Use case |
|---|---|---|---|
| [RotorQuant-MLX-3bit](https://huggingface.co/majentik/minimax-m2.7-rotorquant-mlx-3bit) | mlx-lm | ~1.2 GB | Apple Silicon, small |
| **RotorQuant-MLX-4bit** | mlx-lm | ~1.7 GB | Apple Silicon balanced |
| [RotorQuant-MLX-5bit](https://huggingface.co/majentik/minimax-m2.7-rotorquant-mlx-5bit) | mlx-lm | ~2.1 GB | Apple Silicon, higher fidelity |
| [TurboQuant-MLX-3bit](https://huggingface.co/majentik/minimax-m2.7-turboquant-mlx-3bit) | mlx-lm | ~1.2 GB | Apple Silicon, small |
| [TurboQuant-MLX-4bit](https://huggingface.co/majentik/minimax-m2.7-turboquant-mlx-4bit) | mlx-lm | ~1.7 GB | Apple Silicon balanced |
| [TurboQuant-MLX-5bit](https://huggingface.co/majentik/minimax-m2.7-turboquant-mlx-5bit) | mlx-lm | ~2.1 GB | Apple Silicon, higher fidelity |

## About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's **release labels**, not distinct
quantization algorithms — for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured.