Instructions to use mlx-community/GLM-5.3-Flash-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/GLM-5.3-Flash-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mlx-community/GLM-5.3-Flash-4bit") config = load_config("mlx-community/GLM-5.3-Flash-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/GLM-5.3-Flash-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/GLM-5.3-Flash-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/GLM-5.3-Flash-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use mlx-community/GLM-5.3-Flash-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/GLM-5.3-Flash-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/GLM-5.3-Flash-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mlx-community/GLM-5.3-Flash-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/GLM-5.3-Flash-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/GLM-5.3-Flash-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
mlx-community/GLM-5.3-Flash-4bit
zai-org/GLM-5.3-Flash in MLX affine format, mixed 4/5/6-bit, group size 64.
The weights are the 4-bit variant of orcarouter/GLM-5.3-Flash-MLX
(revision c80f6810b1a95b5be9042761becc6aa78d189782, the 4-bit/ build, which is also the repo root there). The quantisation is orcarouter's.
The 62 safetensors shards are byte-identical to that build (sha256 checked).
Change from the source: config.json now lists the per-module widths of the MTP layer (layer 45) in
quantization and quantization_config. The source config named them for layers 3-44 only. The tensors are not changed.
Size: 62 shards, 203,992,076,296 bytes (190 GiB). It needs a Mac with 256 GB of memory or more.
Widths
Read from each tensor's shapes (bits = 32 * weight_columns / (group_size * scales_columns)), group size 64 throughout.
| Modules | Layers | Bits |
|---|---|---|
Routed experts gate_proj, up_proj |
3-45 | 4 |
Routed experts down_proj |
3-45 | 5 |
Shared expert gate_proj, up_proj, down_proj |
3-45 | 6 |
Dense MLP gate_proj, up_proj |
0-2 | 4 |
Dense MLP down_proj |
0-2 | 5 |
Sparse attention q_a_proj, q_b_proj, kv_a_proj_with_mqa, o_proj |
3, 7, 11, ..., 43, 45 (12 layers) | 4 |
Linear attention layers, kv_b_proj, indexer, router, mHC, norms, embed_tokens, lm_head, MTP eh_proj, vision tower |
all | bf16 |
Layer 45 is the MTP (next-token prediction) layer.
Usage
mlx-lm does not support the glm5_next architecture. mlx-vlm has a glm5_next model;
see the orcarouter model card for its use.
License
MIT, the same as the base model.
- Downloads last month
- 364
4-bit
Model tree for mlx-community/GLM-5.3-Flash-4bit
Base model
zai-org/GLM-5.3-Flash