Instructions to use Cropduster69/GLM-5.3-Flash-oQ2e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Cropduster69/GLM-5.3-Flash-oQ2e-mtp with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Cropduster69/GLM-5.3-Flash-oQ2e-mtp") config = load_config("Cropduster69/GLM-5.3-Flash-oQ2e-mtp") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Cropduster69/GLM-5.3-Flash-oQ2e-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Cropduster69/GLM-5.3-Flash-oQ2e-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Cropduster69/GLM-5.3-Flash-oQ2e-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Cropduster69/GLM-5.3-Flash-oQ2e-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Cropduster69/GLM-5.3-Flash-oQ2e-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Cropduster69/GLM-5.3-Flash-oQ2e-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Cropduster69/GLM-5.3-Flash-oQ2e-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Cropduster69/GLM-5.3-Flash-oQ2e-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Cropduster69/GLM-5.3-Flash-oQ2e-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GLM-5.3-Flash-oQ2e-mtp (MLX)
2-bit imatrix-calibrated oQ quantization (oQ2e) of
zai-org/GLM-5.3-Flash, built locally on a
Mac Studio M4 Max (128 GB) with oMLX's oQ quantizer. The
MTP (nextn) head is preserved and calibrated like the backbone.
| Source | zai-org/GLM-5.3-Flash (FP8, e4m3, 128x128 block scales), 328.37 GB |
| Format | MLX safetensors, model_type: glm5_next, 3057 tensors |
| Size on disk | 110.3 GB (103 GiB), 21 shards |
| Level | oQ level 2, enhanced (oQe) — imatrix-calibrated, target 2.8 bpw / hard cap 3.0 |
| Non-quantized | bfloat16 |
| MTP head | preserved, 59 tensors (language_model.mtp.0.*), with imatrix entries |
| Built | 2026-10-01, 33.7 min wall clock, one calibration round, no calibration proxy |
Quantization layout (actual per-tensor config)
| Bit width | Scope |
|---|---|
2 bit (default, group_size 64, affine) |
trunk routed experts (mlp.switch_mlp.*, 288 experts / top-8, layers 3-44) |
| 4 bit (3 tensors) | the MTP head's routed experts (mtp.0.block.mlp.switch_mlp.{gate,up,down}_proj) |
| 8 bit (554 tensors) | attention projections incl. indexer + embed_q/unembed_out (q/k/v/o_proj, q_a/q_b, kv_a_proj_with_mqa, b_proj, g_a/g_b, forget_gate.*, indexer.*), shared_experts.*, dense layers 0-2, lm_head, embed_tokens |
So the routed experts carry the 2-bit budget (that is where oQ2 spends it) while attention, shared experts and the head stay protected — plus the MTP head's experts are boosted to 4 bit, which is cheap because the head is a single layer.
Calibration (oQe imatrix)
| Dataset | oMLX oqe_code_multilingual (code/en/zh/ja/ko/tool-calling/reasoning) |
| Budget | 128 samples x 512 tokens, micro-batch 12, 1 round (coverage sufficient) |
| Collection | streamed layer-by-layer from the FP8 checkpoint (load_kind: streaming) |
| Entries applied | 682 (oq_imatrix_report.json: 671 entries) |
| Uncalibrated | none (missing: [], mismatched: []) |
| Expert coverage | 38 688 / 38 688 active, 0 zero-count, min_count 35 (required >= 16) |
The imatrix is collected from the real FP8 weights. A uniform 4-bit calibration proxy
(~181 GB resident) does not fit a 128 GB machine (0.75 * capacity = 92.8 GB), which is why
the streaming path was needed — see jundot/omlx#4173
and jundot/omlx#4174 (CI green; it also fixes two
glm5_next MTP-head/imatrix bugs found while verifying: the resident head pass is skipped for
this family, and embed_q capture is lost for prefixed collector installs).
Verification
- loads on a 128 GB M4 Max with MoE expert offload:
wrapped 42 layers at 82.5% residency (expert tables: 95.13 GB total, 78.61 GB resident),actual: 87.43 GB(fully resident without offload: ~103 GB) - MTP active in generation (
MTP[0] ... accept=8/9 (88.9%)on a short probe) - clean output/thinking split (
finish=stop,content='17 x 23 = 391', thinking inreasoning_content),cached_tokens: 0on the first run - no keys in the wrong namespace and no loader shape errors (the
preserve_mtpregressions reported in jundot/omlx#3845 are not present in this build)
No systematic benchmark was run. Short probes on the machine above measured low-to-mid tens of tokens/s depending on prompt and warm-up; prefill is dominated by reading the ~95 GB of expert tables from disk (offload on).
Usage (oMLX)
Place this folder in your oMLX model directory, reload the model list, and set the per-model settings that GLM-5.3-Flash needs on a <=128 GB machine:
moe_expert_offload_enabled = true
moe_expert_offload_resident_fraction = 0.825
mtp_enabled = true
max_context_window = 262144
model_type_override = "vlm"
Without the offload setting the model loads fully resident (~103 GB), leaving no KV headroom on a 128 GB machine.
How it was built
quantize_oq_streaming(
model_path=<zai-org/GLM-5.3-Flash, FP8>,
output_path=.../GLM-5.3-Flash-oQ2e-mtp,
oq_level=2, group_size=64, dtype="bfloat16",
enhanced=True, # oQe imatrix
preserve_mtp=True, # keep + calibrate the nextn head
stream_calibration=True, # layer-by-layer instead of the RAM proxy
imatrix_num_samples=128, imatrix_seq_length=512,
)
An already-quantized MLX/GGUF build cannot be used as an oQ source — validate_quantizable()
accepts only native FP8/MXFP8 (or unquantized) checkpoints, so the official FP8 release is
the source.
Provenance and license
- Base model:
zai-org/GLM-5.3-Flash, MIT — this derivative is published under the same license, with attribution to zai-org. - Quantized with oMLX (
omlx/oq.py); the streaming calibration used here is in jundot/omlx#4174. oq_imatrix_report.jsonis the calibration report of this exact run (entries, coverage, sample budget, load kind).
Weights are quantized, not retrained: quality tracks the base model within the usual 2-bit / target-2.8-bpw limits. Test before production use.
- Downloads last month
- 231
2-bit
Model tree for Cropduster69/GLM-5.3-Flash-oQ2e-mtp
Base model
zai-org/GLM-5.3-Flash