Instructions to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S # Run inference directly in the terminal: llama cli -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S # Run inference directly in the terminal: llama cli -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S # Run inference directly in the terminal: ./llama-cli -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Use Docker
docker model run hf.co/LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
- LM Studio
- Jan
- Ollama
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with Ollama:
ollama run hf.co/LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
- Unsloth Desktop
- Pi
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with Docker Model Runner:
docker model run hf.co/LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
- Lemonade
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Run and chat with the model
lemonade run user.GLM-4.7-Flash-Coder-GGUF-UD-IQ1_S
List all available models
lemonade list
- Hermes Agent
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use LakoMoor/GLM-4.7-Flash-Coder-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "LakoMoor/GLM-4.7-Flash-Coder-GGUF:UD-IQ1_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GLM-4.7-Flash-Coder โ Dynamic 1-bit / 2-bit GGUF quants
Ultra-low-bit GGUF quantizations of whitecircle/GLM-4.7-Flash-Coder (MIT, a coding-agent finetune of zai-org/GLM-4.7-Flash, 29.9B params, MoE: 47 layers, 64 routed experts top-4 + 1 shared, MLA attention, layer 0 dense), following the logic of Unsloth's Dynamic 3.0 GGUFs: per-tensor mixed precision from an imatrix, not a uniform bit width.
Method
- Per-tensor type plan: read from the public headers of
unsloth/GLM-4.7-Flash-GGUF(their quant of the base model,plan/unsloth_*.json) and re-applied to this finetune viallama-quantize --tensor-type-file. Pattern: norms/router/biases stay F32, MLA attention projections stay Q8_0/Q5_K/Q6_K, the shared expert and embeddings/output stay 4โ6 bit, and routed-expertgate/up/downtensors carry the bulk of the bit-budget cut (down to IQ1_S/IQ1_M in the most compressible layers, IQ2_SโIQ4_XS elsewhere). - Our own imatrix, computed on our own calibration mix (
calib/calib.txt, ~410K tokens): 55% agentic coding traces (whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces, rendered through the model's own chat template with tool definitions), 20% general chat (SmolTalk), 15% general text (Cosmopedia-v2), 10% multilingual (FineWeb-2, 6 languages) โ no Unsloth weights or imatrix data reused. - Quality check: KL-divergence and top-1 agreement of each quant vs. the Q8_0 reference, on a held-out slice (
calib/heldout.txt, ~24.6K tokens, different documents/traces than calibration).
Results (12 chunks ร 2048 tokens, held-out; Q8_0 baseline PPL 13.98)
| variant | size | bits | mean PPL | ฮPPL vs Q8_0 | mean KL divergence | same top-1 token |
|---|---|---|---|---|---|---|
| UD-IQ1_S (1-bit class) | 8.6 GiB | 2.47 BPW | 15.51 | +17.6% | 1.241 | 68.9% |
| UD-IQ2_XXS (2-bit class) | 9.8 GiB | ~2.8 BPW | 14.00 | +6.1% | 1.068 | 74.9% |
| UD-Q2_K_XL (2-bit class) | 11.1 GiB | ~3.1 BPW | 15.22 | +15.4% | 1.085 | 75.8% |
Caveat on these numbers: the held-out set is small (~24.6K tokens) and mixed-domain, so the ยฑ0.5 PPL error bars are non-trivial relative to the differences between quants โ treat this as a directional signal, not a precise ranking. Notably, IQ2_XXS scored best on PPL/ฮPPL despite being smaller than Q2_K_XL, which is not what Unsloth's own (much larger-scale) evaluation found for the base model; the tensor-type plan was derived from the base model and may not transfer perfectly to this coding-agent finetune's expert routing.
Generation sanity check
Both UD-IQ1_S and UD-IQ2_XXS produce coherent, on-topic reasoning for a coding prompt ("write a function that reverses a singly linked list") โ thinking-mode traces stay on task and don't devolve into repetition or gibberish in a short sample. Not a rigorous eval; see Limitations.
| variant | prompt eval | generation | threads |
|---|---|---|---|
| UD-IQ1_S | 43.5 tok/s | 8.3 tok/s | 56 (CPU) |
| UD-IQ2_XXS | 32.0 tok/s | 6.3 tok/s | 56 (CPU) |
Limitations
- 1-bit quantization measurably degrades this model (+17.6% perplexity, only 69% of tokens agree with the fp16-equivalent reference at top-1). Per Unsloth's own published guidance, expect noticeably worse tool-calling and multi-step reasoning at this level โ this is a research/fits-in-less-VRAM option, not a quality-preserving one.
- 2-bit (IQ2_XXS/Q2_K_XL) is the practical floor for this size/architecture if you want the model to still behave close to its original self.
- No agentic tool-calling benchmark (SWE-bench-style) was run โ only a KL-divergence/perplexity proxy and a short generation spot-check. If you depend on tool-calling accuracy, evaluate that directly before relying on these quants.
- The per-tensor type plan is transplanted from the base model's Unsloth quant, not derived from this finetune's own imatrix-guided layer sensitivity analysis โ a plan computed natively for this checkpoint might do better, especially for the routed experts most reshaped by the coding-agent finetuning.
Files
GLM-4.7-Flash-Coder-UD-IQ1_S.ggufโ 8.6 GiB, 1-bit classGLM-4.7-Flash-Coder-UD-IQ2_XXS.ggufโ 9.8 GiB, 2-bit class, best measured quality/size hereGLM-4.7-Flash-Coder-UD-Q2_K_XL.ggufโ 11.1 GiB, 2-bit class, Unsloth's usual flagship 2-bit recipe
License
MIT (inherited from the base model and this finetune). Calibration data license terms are their own (SmolTalk: Apache-2.0; Cosmopedia-v2: Apache-2.0/ODC-By; FineWeb-2: ODC-By; the agentic traces dataset: see its own card) โ none of it is redistributed here, only used to compute the imatrix.
- Downloads last month
- 813
1-bit
2-bit