Text Generation
MLX
Safetensors
minimax_m2
minimax
m2.7
Mixture of Experts
quantized
rotorquant
kv-cache-quantization
conversational
custom_code
3-bit
Instructions to use majentik/MiniMax-M2.7-RotorQuant-MLX-3bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use majentik/MiniMax-M2.7-RotorQuant-MLX-3bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("majentik/MiniMax-M2.7-RotorQuant-MLX-3bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use majentik/MiniMax-M2.7-RotorQuant-MLX-3bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use majentik/MiniMax-M2.7-RotorQuant-MLX-3bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use majentik/MiniMax-M2.7-RotorQuant-MLX-3bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use majentik/MiniMax-M2.7-RotorQuant-MLX-3bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "majentik/MiniMax-M2.7-RotorQuant-MLX-3bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default majentik/MiniMax-M2.7-RotorQuant-MLX-3bit
Run Hermes
hermes
- Atomic Chat
File size: 6,179 Bytes
ecbbd54 fdfe53f ecbbd54 f8146c2 ecbbd54 bb2fe1f ecbbd54 fdfe53f 605904c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | ---
base_model: MiniMaxAI/MiniMax-M2.7
library_name: mlx
pipeline_tag: text-generation
license: other
license_name: minimax-model-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/LICENSE
tags:
- minimax
- m2.7
- moe
- quantized
- rotorquant
- kv-cache-quantization
- mlx
---
> [!TIP]
> **KV-cache quantization without any fork (recommended, 2026):** upstream
> llama.cpp/Ollama now cover this natively — use `-ctk q8_0 -ctv q8_0`
> (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or
> `-ctk q4_0 -ctv q4_0` (~quarter memory, ≈7.6% perplexity increase). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`. Keep
> K and V types symmetric to stay on the fast fused Flash-Attention path.
> Since April 2026, mainline llama.cpp also applies Hadamard rotation to
> KV activations ([PR #21038](https://github.com/ggml-org/llama.cpp/pull/21038)),
> which greatly improves low-bit KV quality (opt-out:
> `LLAMA_ATTN_ROT_DISABLE=1`).
>
> The RotorQuant/TurboQuant fork flow below is **experimental/legacy**: the
> TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork
> is unmaintained relative to mainline. It is NOT required to use this model.
<!-- kv-upstream-note -->
# MiniMax-M2.7-RotorQuant-MLX-3bit
**MLX 3-bit quantized variant of [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) with the legacy RotorQuant KV-cache fork (superseded by upstream llama.cpp KV options), optimized for Apple Silicon.**
## Overview
MiniMax-M2.7 is a massive 256-expert Mixture-of-Experts (MoE) model with 8 experts active per token, totaling approximately 456 billion parameters. This variant combines **3-bit MLX weight quantization** with **RotorQuant** KV-cache quantization for deployment on Apple Silicon hardware.
RotorQuant applies a learned Hadamard rotation matrix to keys and values before quantization, smoothing the activation distribution for better quality retention. At 3-bit, RotorQuant's rotation-based approach is particularly valuable for preserving output quality where naive quantization would noticeably degrade.
| Property | Value |
|---|---|
| Architecture | MoE (256 experts, 8 active/token) |
| Total Parameters | ~456B |
| Layers | 62 |
| Hidden Size | 3072 |
| Attention Heads | 48 |
| Weight Quantization | 3-bit (MLX) |
| KV-Cache Quantization | RotorQuant |
| Estimated Size | ~170 GB |
| Base Model | MiniMaxAI/MiniMax-M2.7 |
## Quickstart
```bash
pip install mlx-lm
```
```python
from mlx_lm import load, generate
model, tokenizer = load("majentik/MiniMax-M2.7-RotorQuant-MLX-3bit")
prompt = "What is a Comprehensive Geriatric Assessment?"
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
response = generate(
model,
tokenizer,
prompt=text,
max_tokens=512,
)
print(response)
```
## RotorQuant vs TurboQuant
| Feature | RotorQuant | TurboQuant |
|---|---|---|
| Technique | Rotation-based KV quantization (Hadamard transform) | Asymmetric per-channel KV quantization |
| Throughput | Slightly lower throughput (rotation overhead) | Higher throughput, lower latency |
| Quality | Better quality preservation at low bit-widths | Good quality preservation |
| Best For | Quality-sensitive tasks, research | High-throughput serving, long contexts |
> At 3-bit quantization, RotorQuant provides meaningfully better quality than TurboQuant due to its rotation-based outlier smoothing.
## Memory Estimates (Apple Silicon)
| Variant | Estimated Size | Minimum Unified Memory |
|---|---|---|
| MLX 8-bit | ~456 GB | 512 GB (Mac Studio M2/M3/M4 Ultra) |
| MLX 5-bit | ~280 GB | 384 GB |
| MLX 4-bit | ~225 GB | 256 GB |
| MLX 3-bit | ~170 GB | 192 GB |
| MLX 2-bit | ~110 GB | 128 GB |
> **Note**: 3-bit quantization requires Apple Silicon with 192 GB+ unified memory, such as a Mac Studio with M2/M3/M4 Ultra.
## See Also
- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) -- Base model
- [majentik/MiniMax-M2.7-TurboQuant-MLX-3bit](https://huggingface.co/majentik/MiniMax-M2.7-TurboQuant-MLX-3bit) -- TurboQuant MLX 3-bit
- [majentik/MiniMax-M2.7-RotorQuant-MLX-4bit](https://huggingface.co/majentik/MiniMax-M2.7-RotorQuant-MLX-4bit) -- MLX 4-bit
## Quant trade-off (MLX lane)
| Bits | Approx size | Use case | Recommendation |
|---|---|---|---|
| 2-bit | ~119 GB | Aggressive quantization | Very low-RAM Macs |
| **3-bit** | ~164 GB | Lossy but small | **Low-RAM Macs** |
| 4-bit | ~192 GB | Balanced default | Recommended for most Macs |
| 5-bit | ~228 GB | Higher fidelity | Quality-sensitive |
| 6-bit | ~274 GB | Approaching FP16 quality | High-fidelity |
| 8-bit | ~347 GB | Near-lossless reference | Fidelity-critical work |
(Current variant — **3bit** — is bolded.)
## Variants in this family
(Showing 12 sibling variants under `majentik/minimax-m2.7-*`. The current variant — `RotorQuant-MLX-3bit` — is **bolded**.)
| Variant | Runtime | Approx size | Use case |
|---|---|---|---|
| **RotorQuant-MLX-3bit** | mlx-lm | ~1.2 GB | Apple Silicon, small |
| [RotorQuant-MLX-4bit](https://huggingface.co/majentik/minimax-m2.7-rotorquant-mlx-4bit) | mlx-lm | ~1.7 GB | Apple Silicon balanced |
| [RotorQuant-MLX-5bit](https://huggingface.co/majentik/minimax-m2.7-rotorquant-mlx-5bit) | mlx-lm | ~2.1 GB | Apple Silicon, higher fidelity |
| [TurboQuant-MLX-3bit](https://huggingface.co/majentik/minimax-m2.7-turboquant-mlx-3bit) | mlx-lm | ~1.2 GB | Apple Silicon, small |
| [TurboQuant-MLX-4bit](https://huggingface.co/majentik/minimax-m2.7-turboquant-mlx-4bit) | mlx-lm | ~1.7 GB | Apple Silicon balanced |
| [TurboQuant-MLX-5bit](https://huggingface.co/majentik/minimax-m2.7-turboquant-mlx-5bit) | mlx-lm | ~2.1 GB | Apple Silicon, higher fidelity |
## About the RotorQuant / TurboQuant labels
RotorQuant and TurboQuant are this project's **release labels**, not distinct
quantization algorithms — for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured.
|