Instructions to use whitecircle/GLM-4.7-Flash-Coder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use whitecircle/GLM-4.7-Flash-Coder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="whitecircle/GLM-4.7-Flash-Coder") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("whitecircle/GLM-4.7-Flash-Coder") model = AutoModelForCausalLM.from_pretrained("whitecircle/GLM-4.7-Flash-Coder", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use whitecircle/GLM-4.7-Flash-Coder with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "whitecircle/GLM-4.7-Flash-Coder" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "whitecircle/GLM-4.7-Flash-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/whitecircle/GLM-4.7-Flash-Coder
- SGLang
How to use whitecircle/GLM-4.7-Flash-Coder with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "whitecircle/GLM-4.7-Flash-Coder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "whitecircle/GLM-4.7-Flash-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "whitecircle/GLM-4.7-Flash-Coder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "whitecircle/GLM-4.7-Flash-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use whitecircle/GLM-4.7-Flash-Coder with Docker Model Runner:
docker model run hf.co/whitecircle/GLM-4.7-Flash-Coder
Apply editorial pass from Notion v5
Browse files
README.md
CHANGED
|
@@ -13,6 +13,8 @@ We fine-tuned GLM-4.7-Flash to push its agentic coding capabilities further: fro
|
|
| 13 |
|
| 14 |
This 30B Mixture-of-Experts model beats competitors 3.5-4× its size — Qwen3.5-122B-A10B and Devstral-2 — and surpasses Ling-2.6-Flash by a narrow margin.
|
| 15 |
|
|
|
|
|
|
|
| 16 |

|
| 17 |
|
| 18 |
## More reasoning, fewer turns
|
|
@@ -23,17 +25,19 @@ Although GLM-4.7-Flash-Coder does 2.4× more reasoning than the base model, it c
|
|
| 23 |
|
| 24 |
## The key to turn efficiency is concurrent tool calls
|
| 25 |
|
|
|
|
|
|
|
| 26 |

|
| 27 |
|
| 28 |
## SWE-rebench-V2
|
| 29 |
|
| 30 |
-
We trained and evaluated models on [nebius/SWE-rebench-V2](https://huggingface.co/datasets/nebius/SWE-rebench-V2)
|
| 31 |
|
| 32 |

|
| 33 |
|
| 34 |
## Teacher — GLM-5.1
|
| 35 |
|
| 36 |
-
The training data comes from GLM-5.1 acting as a teacher. We ran it on the train split with 4 rollouts per task and kept only trajectories that passed the tests
|
| 37 |
|
| 38 |
We’re releasing the [training dataset](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces) alongside the [train/test splits](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-clean-python-tasks).
|
| 39 |
|
|
@@ -115,21 +119,21 @@ Perhaps the most surprising transfer was the `git stash` usage:
|
|
| 115 |
> git stash && python -m pytest tests/test_markdown.py::test_markdown_table -xvs 2>&1 | tail -20
|
| 116 |
> ```
|
| 117 |
|
| 118 |
-
## Loss decomposition into Reasoning
|
| 119 |
|
| 120 |
-
We only train on assistant responses. Of 177 million tokens in the dataset, only 46 million carry loss — the remaining 131 million (system prompts and tool responses) are masked.
|
| 121 |
|
| 122 |
-
|
| 123 |
|
| 124 |

|
| 125 |
|
| 126 |
-
Although tool
|
| 127 |
|
| 128 |

|
| 129 |
|
| 130 |
-
|
| 131 |
|
| 132 |
-
|
| 133 |
|
| 134 |
## Test scores
|
| 135 |
|
|
@@ -147,15 +151,13 @@ FlashAttention-4 is a fast attention kernel — but it’s incompatible with GLM
|
|
| 147 |
|
| 148 |
> *FlashAttention-4 produces NaN gradients for this model’s head_dim-256 partial-rotary attention; falling back to `attn_implementation='sdpa'`.*
|
| 149 |
|
| 150 |
-
We found that the incompatibility only appears in the backward pass
|
| 151 |
|
| 152 |
*We’ll have a separate post on FA4-fwd-FA2-bwd in [our research blog](https://whitecircle.com/#research); stay tuned! 💚*
|
| 153 |
|
| 154 |
## Kernels
|
| 155 |
|
| 156 |
-
GLM-4.7-Flash is a Mixture-of-Experts model. Tokens are distributed unevenly across experts
|
| 157 |
-
|
| 158 |
-
The Grouped-GEMM kernel batches matrix multiplications of varying shapes into a single kernel launch. Enabling it gave a 1.6× speedup on the FFN step.
|
| 159 |
|
| 160 |
Additionally, with Liger Fused Linear Cross-Entropy, we were able to compute loss without full logit matrix materialization in VRAM, preventing OOM on long sequences.
|
| 161 |
|
|
@@ -165,19 +167,19 @@ Together, these improvements yield up to a 1.63× training speedup over TRL, rai
|
|
| 165 |
|
| 166 |
## Agentic infra challenge
|
| 167 |
|
| 168 |
-
|
| 169 |
|
| 170 |
-
|
| 171 |
|
| 172 |
Kubernetes struggled at high concurrency on a single host as well: it couldn’t set up and tear down pods fast enough, triggering timeouts. The problem was worst on powerful hosts: they could support enormous concurrency, but pod scheduling couldn’t keep up, so we never reached full CPU or RAM utilization.
|
| 173 |
|
| 174 |
## Overlay sandboxes
|
| 175 |
|
| 176 |
-
Our solution is the Overlay Sandbox: instead of a container image, the agent runs on the bare host,
|
| 177 |
|
| 178 |
It has three layers: a base filesystem with the OS and repository, a hidden layer with the verifier suite, and an upper layer holding all edits and pip-installed packages.
|
| 179 |
|
| 180 |
-
Because overlay mounts are lazy (they don’t even index the filesystem!) and
|
| 181 |
|
| 182 |

|
| 183 |
|
|
|
|
| 13 |
|
| 14 |
This 30B Mixture-of-Experts model beats competitors 3.5-4× its size — Qwen3.5-122B-A10B and Devstral-2 — and surpasses Ling-2.6-Flash by a narrow margin.
|
| 15 |
|
| 16 |
+
We trained it with [Halo](https://github.com/whitecircle/halo), the post-training framework we’re open-sourcing today — the full story of the framework [here](https://whitecircle.com/research).
|
| 17 |
+
|
| 18 |

|
| 19 |
|
| 20 |
## More reasoning, fewer turns
|
|
|
|
| 25 |
|
| 26 |
## The key to turn efficiency is concurrent tool calls
|
| 27 |
|
| 28 |
+
The fine-tuned model learned to batch its tool use instead of working one call at a time. Turns with two or more parallel `read` calls jumped from 1.6% to 9.7% of all turns, parallel `bash` from 0.6% to 7.2%, and parallel `edit`, which is almost nonexistent in the base model, grew ~90×. Reading three files in one turn instead of three turns is exactly the kind of habit that compounds into 45% fewer turns overall.
|
| 29 |
+
|
| 30 |

|
| 31 |
|
| 32 |
## SWE-rebench-V2
|
| 33 |
|
| 34 |
+
We trained and evaluated models on [nebius/SWE-rebench-V2](https://huggingface.co/datasets/nebius/SWE-rebench-V2), a dataset of 32k agentic coding tasks. It consists of real issues from public repositories, closely representing real-world SWE challenges.
|
| 35 |
|
| 36 |

|
| 37 |
|
| 38 |
## Teacher — GLM-5.1
|
| 39 |
|
| 40 |
+
The training data comes from GLM-5.1 acting as a teacher. We ran it on the train split with 4 rollouts per task and kept only trajectories that passed the tests: 7,777 successful multi-turn traces after rejection sampling.
|
| 41 |
|
| 42 |
We’re releasing the [training dataset](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces) alongside the [train/test splits](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-clean-python-tasks).
|
| 43 |
|
|
|
|
| 119 |
> git stash && python -m pytest tests/test_markdown.py::test_markdown_table -xvs 2>&1 | tail -20
|
| 120 |
> ```
|
| 121 |
|
| 122 |
+
## Loss decomposition into Reasoning vs. Tool Calls
|
| 123 |
|
| 124 |
+
We only train on assistant responses. Of the 177 million tokens in the dataset, only 46 million carry loss — the remaining 131 million (system prompts and tool responses) are masked.
|
| 125 |
|
| 126 |
+
Within an assistant message there are two very different kinds of tokens, mainly reasoning and tool calls. Tool-call tokens are far lower-entropy:
|
| 127 |
|
| 128 |

|
| 129 |
|
| 130 |
+
Although tool-call tokens make up 43.3% of input tokens, they contribute only 16.1% of the overall loss. To track convergence separately, we decompose the loss into two components: negative log-likelihood on reasoning tokens and on tool-call tokens.
|
| 131 |
|
| 132 |

|
| 133 |
|
| 134 |
+
The training-loss curves show 10–20% cliffs at epoch boundaries — the usual signature of memorizing repeated data. We read it as benign: eval losses keep improving through the second epoch.
|
| 135 |
|
| 136 |
+
It’s possible to set different learning rates for the two components, but on our data both eval losses reach their global minima at the same moment (the end of epoch 2) so we didn’t need to.
|
| 137 |
|
| 138 |
## Test scores
|
| 139 |
|
|
|
|
| 151 |
|
| 152 |
> *FlashAttention-4 produces NaN gradients for this model’s head_dim-256 partial-rotary attention; falling back to `attn_implementation='sdpa'`.*
|
| 153 |
|
| 154 |
+
We found that the incompatibility only appears in the backward pass; the forward computation is correct. So we designed a hybrid scheme: FA4 handles the forward pass, while FA2 takes over for the backward. The hybrid approach gave a 1.36× increase in attention throughput.
|
| 155 |
|
| 156 |
*We’ll have a separate post on FA4-fwd-FA2-bwd in [our research blog](https://whitecircle.com/#research); stay tuned! 💚*
|
| 157 |
|
| 158 |
## Kernels
|
| 159 |
|
| 160 |
+
GLM-4.7-Flash is a Mixture-of-Experts model. Tokens are distributed unevenly across experts, so the MLP blocks receive tensors of varying width. The Grouped-GEMM kernel batches matrix multiplications of varying shapes into a single kernel launch. Enabling it gave a 1.6× speedup on the FFN step.
|
|
|
|
|
|
|
| 161 |
|
| 162 |
Additionally, with Liger Fused Linear Cross-Entropy, we were able to compute loss without full logit matrix materialization in VRAM, preventing OOM on long sequences.
|
| 163 |
|
|
|
|
| 167 |
|
| 168 |
## Agentic infra challenge
|
| 169 |
|
| 170 |
+
Agentic evaluation at scale turned out to be a real infrastructure problem.
|
| 171 |
|
| 172 |
+
With thousands of concurrent agents, Docker brings high startup overhead and heavy disk pressure, and quickly saturates the local network — producing a high rate of environment startup errors.
|
| 173 |
|
| 174 |
Kubernetes struggled at high concurrency on a single host as well: it couldn’t set up and tear down pods fast enough, triggering timeouts. The problem was worst on powerful hosts: they could support enormous concurrency, but pod scheduling couldn’t keep up, so we never reached full CPU or RAM utilization.
|
| 175 |
|
| 176 |
## Overlay sandboxes
|
| 177 |
|
| 178 |
+
Our solution is the Overlay Sandbox: instead of a container image, the agent runs on the bare host, inside an isolated filesystem and process group.
|
| 179 |
|
| 180 |
It has three layers: a base filesystem with the OS and repository, a hidden layer with the verifier suite, and an upper layer holding all edits and pip-installed packages.
|
| 181 |
|
| 182 |
+
Because overlay mounts are lazy (they don’t even index the filesystem!) and track only modified files, sandbox startup is essentially free.
|
| 183 |
|
| 184 |

|
| 185 |
|