Qwen3.8-27B, 4-bit, with its draft head

Qwen3.8-27B quantized to four bits for MLX, carrying the model's own multi-token prediction head in mtp/. Speculative decoding therefore works from this repository alone -- nothing else to fetch, no environment variable pointing somewhere else.

Generation runs 1.8-2.3x faster with the head on, and says the same thing. Every token it proposes is checked by the model itself, so the reply is the model's own either way; the head only saves passes over the weights.

Speed

Measured on a 48 GB MacBook Pro (M4 Pro), greedy decoding, 96 tokens, with gbx_lm. Decode is timed from the first token, so prefill is not in it.

context decode, head off decode, head on draft acceptance
1,024 14.7 tok/s 33.8 tok/s 0.95
4,096 14.3 27.3 0.78
16,384 13.5 24.8 0.78

Requirements

macOS 15.0 or later
chip Apple Silicon (arm64). There is no Intel build.
Python none -- the binary carries what it needs

These weights are held resident, not paged: on a 512 GB Mac Studio the model settles at about 16 GB, and a machine needs room for that much plus the conversation's cache.

Install

curl -fL -o gbx_lm-darwin-arm64.tar.gz 'https://github.com/GreenBitAI/gbx-lm/releases/latest/download/gbx_lm-darwin-arm64.tar.gz' \
  && tar -xzf gbx_lm-darwin-arm64.tar.gz gbx_lm \
  && mkdir -p "$HOME/.local/bin" \
  && mv gbx_lm "$HOME/.local/bin/gbx_lm" \
  && chmod +x "$HOME/.local/bin/gbx_lm"

gbx_lm -h

command not found means $HOME/.local/bin is not on your PATH: add it, or call the binary by its full path. The build is signed with a Developer ID and notarised, so macOS runs it without the usual detour for a downloaded binary.

Run

gbx_lm --model GreenBitAI/Qwen3.8-27B-4bit

# the draft head is off unless asked for, and found in `mtp/` without a path
GBX_QWEN35_MTP=on gbx_lm --model GreenBitAI/Qwen3.8-27B-4bit

That serves an OpenAI-compatible API on port 11688, which is its default. The weights download on first use into ~/.libra/cache/models; set HF_HOME to put them elsewhere, and HF_TOKEN if you meet the Hub's rate limits for anonymous downloads.

curl http://127.0.0.1:11688/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"GreenBitAI/Qwen3.8-27B-4bit","messages":[{"role":"user","content":"Hello"}]}'

Coding agents

The server speaks three wire protocols on the same port, so the tools that expect a hosted API can be pointed at this one:

path for
/v1/chat/completions anything written against the OpenAI API
/v1/responses Codex
/v1/messages Claude Code

Codex -- a provider in ~/.codex/config.toml:

[model_providers.gbx]
name = "gbx-lm"
base_url = "http://127.0.0.1:11688/v1"
wire_api = "responses"

and a profile in ~/.codex/gbx.config.toml:

model_provider = "gbx"
model = "GreenBitAI/Qwen3.8-27B-4bit"
model_context_window = 262144

Claude Code -- ~/.claude/gbx.settings.json:

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://127.0.0.1:11688",
    "ANTHROPIC_AUTH_TOKEN": "local",
    "ANTHROPIC_MODEL": "GreenBitAI/Qwen3.8-27B-4bit",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "GreenBitAI/Qwen3.8-27B-4bit"
  }
}

Both clients ask for a small model for their own background work, so every name in the settings has to be one this server is serving.

What is in here

mtp/mtp.safetensors is built from the draft head Qwen/Qwen3.8-27B ships under mtp., quantized to match these weights. Apache 2.0, as the original is.

Downloads last month
94
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GreenBitAI/Qwen3.8-27B-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(1240)
this model

Collection including GreenBitAI/Qwen3.8-27B-4bit