MiniCPM5-1B-Agent (MLX 2-bit)

Aggressively quantized MLX version of Luminia/MiniCPM5-1B-Agent-GGUF. Quantized with mlx_lm.convert using affine mode (group_size=32, 3.0 bits/weight). Smallest variant — ~387 MB, runs on constrained Apple Silicon.

Note: 2-bit quantization significantly impacts output quality. Use the Q4 variant for better results if memory allows.

About the model

MiniCPM5-1B-Agent is a tiny agentic coding agent for CPU: a full fine-tune of openbmb/MiniCPM5-1B specialized to reason in <think>, call a small tool set (bash/read/write/edit/glob/grep), and run → read output → debug → patch → verify.

  • Base model: openbmb/MiniCPM5-1B (RL+OPD checkpoint)
  • Architecture: LlamaForCausalLM — 24 layers, 16 attention heads (GQA), 1536 hidden, 130560 vocab
  • Parameters: 1,080,632,832 (quantized to ~3.0 bits/weight)
  • Quantization: affine, group_size=32, 2 bits
  • License: Apache-2.0

How to use

from mlx_lm import load, generate

model, tokenizer = load("MC7ever/MiniCPM5-1B-Agent-mlx-q2")

messages = [
    {"role": "user", "content": "Write a Python function to check if a number is prime."}
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
response = generate(model, tokenizer, prompt=prompt, max_tokens=1024)
print(response)

Credits

Other variants

Variant Size Bits/Weight Repo
Safetensors (fp16) 2.0 GB 16 MC7ever/MiniCPM5-1B-Agent-safetensors
MLX Q4 580 MB 4.5 MC7ever/MiniCPM5-1B-Agent-mlx-q4
MLX Q2 387 MB 3.0 This repo
Downloads last month
23
Safetensors
Model size
0.1B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MC7ever/MiniCPM5-1B-Agent-mlx-q2

Quantized
(2)
this model

Paper for MC7ever/MiniCPM5-1B-Agent-mlx-q2