Text Generation
Transformers
Safetensors
glm4_moe_lite
conversational
🇪🇺 Region: EU
olegsmirnov commited on
Commit
33f9caf
·
verified ·
1 Parent(s): 8ec7962

Apply editorial pass from Notion v5

Browse files
Files changed (1) hide show
  1. README.md +18 -16
README.md CHANGED
@@ -13,6 +13,8 @@ We fine-tuned GLM-4.7-Flash to push its agentic coding capabilities further: fro
13
 
14
  This 30B Mixture-of-Experts model beats competitors 3.5-4× its size — Qwen3.5-122B-A10B and Devstral-2 — and surpasses Ling-2.6-Flash by a narrow margin.
15
 
 
 
16
  ![Score vs cost-per-task scatter plot on SWE-rebench-V2, 500 Python tasks: GLM-4.7-Flash-Coder at 0.417 sits 26% above GLM-4.7-Flash at 0.332, at a lower cost per task](media/score_vs_cost.png)
17
 
18
  ## More reasoning, fewer turns
@@ -23,17 +25,19 @@ Although GLM-4.7-Flash-Coder does 2.4× more reasoning than the base model, it c
23
 
24
  ## The key to turn efficiency is concurrent tool calls
25
 
 
 
26
  ![Frequency of responses with 2+ read, bash, and edit tool calls: the SFT model issues simultaneous tool calls 6–90× more often than the base model](media/simultaneous_tool_calls.png)
27
 
28
  ## SWE-rebench-V2
29
 
30
- We trained and evaluated models on [nebius/SWE-rebench-V2](https://huggingface.co/datasets/nebius/SWE-rebench-V2) — a dataset of 32k agentic coding tasks. It consists of real issues from public repositories, closely representing real-world SWE challenges.
31
 
32
  ![SWE-rebench-V2 pipeline: a repository and an issue enter an agentic loop in a sandbox, the generated patch is checked by the test-suite verifier, and the reward is 1 if all tests pass](media/SWE-rebench-V2-pipeline.png)
33
 
34
  ## Teacher — GLM-5.1
35
 
36
- The training data comes from GLM-5.1 acting as a teacher. We ran it on the train split with 4 rollouts per task and kept only trajectories that passed the tests — 7,777 successful multi-turn traces after rejection sampling.
37
 
38
  We’re releasing the [training dataset](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces) alongside the [train/test splits](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-clean-python-tasks).
39
 
@@ -115,21 +119,21 @@ Perhaps the most surprising transfer was the `git stash` usage:
115
  > git stash && python -m pytest tests/test_markdown.py::test_markdown_table -xvs 2>&1 | tail -20
116
  > ```
117
 
118
- ## Loss decomposition into Reasoning & Tool Calls
119
 
120
- We only train on assistant responses. Of 177 million tokens in the dataset, only 46 million carry loss — the remaining 131 million (system prompts and tool responses) are masked.
121
 
122
- Assistant messages are split into two parts — reasoning and tool calls. Naturally, tool call tokens have much lower entropy. These parts are very different in nature:
123
 
124
  ![A training trace rendered with per-token probabilities: reasoning tokens in green, tool call tokens in orange, lighter shades meaning higher probability](media/reasoning-tool-calls-logprobs.png)
125
 
126
- Although tool call tokens make up 43.3% of input tokens, they only contribute to 16.1% of overall loss. To track convergence separately, we decompose loss into two components — negative log-likelihood on reasoning and tool calling tokens, respectively.
127
 
128
  ![Train and eval loss curves decomposed into reasoning and tool-call components with epoch boundaries; every eval loss reaches its global minimum at step 300](media/train_loss_decomposition.png)
129
 
130
- Training loss charts show 10-20% epoch cliffs, indicating benign memorization of repeated data.
131
 
132
- Although it’s possible to set different learning rates for two components, we observed that on our data, eval losses reach their global minima simultaneously — at the end of epoch 2.
133
 
134
  ## Test scores
135
 
@@ -147,15 +151,13 @@ FlashAttention-4 is a fast attention kernel — but it’s incompatible with GLM
147
 
148
  > *FlashAttention-4 produces NaN gradients for this model’s head_dim-256 partial-rotary attention; falling back to `attn_implementation='sdpa'`.*
149
 
150
- We found that the incompatibility only appears in the backward pass — the forward computation is correct. So we designed a hybrid scheme: FA4 handles the forward pass, while FA2 takes over for the backward. The hybrid approach gave a 1.36× increase in attention throughput.
151
 
152
  *We’ll have a separate post on FA4-fwd-FA2-bwd in [our research blog](https://whitecircle.com/#research); stay tuned! 💚*
153
 
154
  ## Kernels
155
 
156
- GLM-4.7-Flash is a Mixture-of-Experts model. Tokens are distributed unevenly across experts — that means MLP blocks receive tensors of varying width.
157
-
158
- The Grouped-GEMM kernel batches matrix multiplications of varying shapes into a single kernel launch. Enabling it gave a 1.6× speedup on the FFN step.
159
 
160
  Additionally, with Liger Fused Linear Cross-Entropy, we were able to compute loss without full logit matrix materialization in VRAM, preventing OOM on long sequences.
161
 
@@ -165,19 +167,19 @@ Together, these improvements yield up to a 1.63× training speedup over TRL, rai
165
 
166
  ## Agentic infra challenge
167
 
168
- We faced a real challenge with agentic evaluation infrastructure.
169
 
170
- At a scale of thousands of concurrent agents, Docker introduces high startup overhead and heavy disk pressure, and quickly saturates the local network — resulting in a high rate of environment startup errors.
171
 
172
  Kubernetes struggled at high concurrency on a single host as well: it couldn’t set up and tear down pods fast enough, triggering timeouts. The problem was worst on powerful hosts: they could support enormous concurrency, but pod scheduling couldn’t keep up, so we never reached full CPU or RAM utilization.
173
 
174
  ## Overlay sandboxes
175
 
176
- Our solution is the Overlay Sandbox: instead of a container image, the agent runs on the bare host, in an isolated filesystem and process group.
177
 
178
  It has three layers: a base filesystem with the OS and repository, a hidden layer with the verifier suite, and an upper layer holding all edits and pip-installed packages.
179
 
180
- Because overlay mounts are lazy (they don’t even index the filesystem!) and only track modified files, sandbox startup is essentially free.
181
 
182
  ![Overlay sandbox diagram: a read-only base filesystem, hidden verifier files not visible to the agent, and a writable lazy overlay mount holding edits and new files](media/overlay_sandbox.png)
183
 
 
13
 
14
  This 30B Mixture-of-Experts model beats competitors 3.5-4× its size — Qwen3.5-122B-A10B and Devstral-2 — and surpasses Ling-2.6-Flash by a narrow margin.
15
 
16
+ We trained it with [Halo](https://github.com/whitecircle/halo), the post-training framework we’re open-sourcing today — the full story of the framework [here](https://whitecircle.com/research).
17
+
18
  ![Score vs cost-per-task scatter plot on SWE-rebench-V2, 500 Python tasks: GLM-4.7-Flash-Coder at 0.417 sits 26% above GLM-4.7-Flash at 0.332, at a lower cost per task](media/score_vs_cost.png)
19
 
20
  ## More reasoning, fewer turns
 
25
 
26
  ## The key to turn efficiency is concurrent tool calls
27
 
28
+ The fine-tuned model learned to batch its tool use instead of working one call at a time. Turns with two or more parallel `read` calls jumped from 1.6% to 9.7% of all turns, parallel `bash` from 0.6% to 7.2%, and parallel `edit`, which is almost nonexistent in the base model, grew ~90×. Reading three files in one turn instead of three turns is exactly the kind of habit that compounds into 45% fewer turns overall.
29
+
30
  ![Frequency of responses with 2+ read, bash, and edit tool calls: the SFT model issues simultaneous tool calls 6–90× more often than the base model](media/simultaneous_tool_calls.png)
31
 
32
  ## SWE-rebench-V2
33
 
34
+ We trained and evaluated models on [nebius/SWE-rebench-V2](https://huggingface.co/datasets/nebius/SWE-rebench-V2), a dataset of 32k agentic coding tasks. It consists of real issues from public repositories, closely representing real-world SWE challenges.
35
 
36
  ![SWE-rebench-V2 pipeline: a repository and an issue enter an agentic loop in a sandbox, the generated patch is checked by the test-suite verifier, and the reward is 1 if all tests pass](media/SWE-rebench-V2-pipeline.png)
37
 
38
  ## Teacher — GLM-5.1
39
 
40
+ The training data comes from GLM-5.1 acting as a teacher. We ran it on the train split with 4 rollouts per task and kept only trajectories that passed the tests: 7,777 successful multi-turn traces after rejection sampling.
41
 
42
  We’re releasing the [training dataset](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces) alongside the [train/test splits](https://huggingface.co/datasets/whitecircle/swe-rebench-v2-clean-python-tasks).
43
 
 
119
  > git stash && python -m pytest tests/test_markdown.py::test_markdown_table -xvs 2>&1 | tail -20
120
  > ```
121
 
122
+ ## Loss decomposition into Reasoning vs. Tool Calls
123
 
124
+ We only train on assistant responses. Of the 177 million tokens in the dataset, only 46 million carry loss — the remaining 131 million (system prompts and tool responses) are masked.
125
 
126
+ Within an assistant message there are two very different kinds of tokens, mainly reasoning and tool calls. Tool-call tokens are far lower-entropy:
127
 
128
  ![A training trace rendered with per-token probabilities: reasoning tokens in green, tool call tokens in orange, lighter shades meaning higher probability](media/reasoning-tool-calls-logprobs.png)
129
 
130
+ Although tool-call tokens make up 43.3% of input tokens, they contribute only 16.1% of the overall loss. To track convergence separately, we decompose the loss into two components: negative log-likelihood on reasoning tokens and on tool-call tokens.
131
 
132
  ![Train and eval loss curves decomposed into reasoning and tool-call components with epoch boundaries; every eval loss reaches its global minimum at step 300](media/train_loss_decomposition.png)
133
 
134
+ The training-loss curves show 10–20% cliffs at epoch boundaries — the usual signature of memorizing repeated data. We read it as benign: eval losses keep improving through the second epoch.
135
 
136
+ It’s possible to set different learning rates for the two components, but on our data both eval losses reach their global minima at the same moment (the end of epoch 2) so we didn’t need to.
137
 
138
  ## Test scores
139
 
 
151
 
152
  > *FlashAttention-4 produces NaN gradients for this model’s head_dim-256 partial-rotary attention; falling back to `attn_implementation='sdpa'`.*
153
 
154
+ We found that the incompatibility only appears in the backward pass; the forward computation is correct. So we designed a hybrid scheme: FA4 handles the forward pass, while FA2 takes over for the backward. The hybrid approach gave a 1.36× increase in attention throughput.
155
 
156
  *We’ll have a separate post on FA4-fwd-FA2-bwd in [our research blog](https://whitecircle.com/#research); stay tuned! 💚*
157
 
158
  ## Kernels
159
 
160
+ GLM-4.7-Flash is a Mixture-of-Experts model. Tokens are distributed unevenly across experts, so the MLP blocks receive tensors of varying width. The Grouped-GEMM kernel batches matrix multiplications of varying shapes into a single kernel launch. Enabling it gave a 1.6× speedup on the FFN step.
 
 
161
 
162
  Additionally, with Liger Fused Linear Cross-Entropy, we were able to compute loss without full logit matrix materialization in VRAM, preventing OOM on long sequences.
163
 
 
167
 
168
  ## Agentic infra challenge
169
 
170
+ Agentic evaluation at scale turned out to be a real infrastructure problem.
171
 
172
+ With thousands of concurrent agents, Docker brings high startup overhead and heavy disk pressure, and quickly saturates the local network — producing a high rate of environment startup errors.
173
 
174
  Kubernetes struggled at high concurrency on a single host as well: it couldn’t set up and tear down pods fast enough, triggering timeouts. The problem was worst on powerful hosts: they could support enormous concurrency, but pod scheduling couldn’t keep up, so we never reached full CPU or RAM utilization.
175
 
176
  ## Overlay sandboxes
177
 
178
+ Our solution is the Overlay Sandbox: instead of a container image, the agent runs on the bare host, inside an isolated filesystem and process group.
179
 
180
  It has three layers: a base filesystem with the OS and repository, a hidden layer with the verifier suite, and an upper layer holding all edits and pip-installed packages.
181
 
182
+ Because overlay mounts are lazy (they don’t even index the filesystem!) and track only modified files, sandbox startup is essentially free.
183
 
184
  ![Overlay sandbox diagram: a read-only base filesystem, hidden verifier files not visible to the agent, and a writable lazy overlay mount holding edits and new files](media/overlay_sandbox.png)
185