Text Generation
Transformers
Safetensors
glm4_moe_lite
conversational
🇪🇺 Region: EU
olegsmirnov commited on
Commit
3d68b8d
·
verified ·
1 Parent(s): 4364941

Unfold Halo config section, remove leading br in details blocks

Browse files
Files changed (1) hide show
  1. README.md +1 -7
README.md CHANGED
@@ -50,8 +50,7 @@ We trained Coder with Supervised Fine-Tuning (SFT) using [Halo](https://github.c
50
  - Learning rate scheduling: Cosine, with linear warm-up
51
  - Training time: 3 hours on 16 B300 GPUs
52
 
53
- <details open>
54
- <summary><h3>Halo 😇 is driven entirely by a single YAML config</h3></summary>
55
 
56
  ```yaml
57
  model_name_or_path: zai-org/GLM-4.7-Flash
@@ -98,8 +97,6 @@ run_name: glm-4.7-flash-coder
98
  enable_efficiency_metrics: true
99
  ```
100
 
101
- </details>
102
-
103
  ## Emergent skills
104
 
105
  The model picked up several of the teacher’s skills and patterns during fine-tuning:
@@ -200,7 +197,6 @@ Inference costs are calculated from OpenRouter pricing tables. We report the the
200
 
201
  <details>
202
  <summary>OpenRouter providers per model</summary>
203
- <br>
204
 
205
  | model | provider | input $/Mtok | output $/Mtok | cache_read $/Mtok | cache_write $/Mtok |
206
  | --- | --- | --- | --- | --- | --- |
@@ -231,7 +227,6 @@ Inference costs are calculated from OpenRouter pricing tables. We report the the
231
 
232
  <details>
233
  <summary>Prompt template</summary>
234
- <br>
235
 
236
  ```markdown
237
  You are solving a real GitHub issue in the `{repo}` repository. The repo is already cloned and set up in the current working directory.
@@ -260,7 +255,6 @@ The following interfaces are expected by the test suite. Pay close attention to
260
 
261
  <details>
262
  <summary>Inference with SGLang 0.5.14, CUDA 13</summary>
263
- <br>
264
 
265
  ```bash
266
  python3 -m sglang_router.launch_server \
 
50
  - Learning rate scheduling: Cosine, with linear warm-up
51
  - Training time: 3 hours on 16 B300 GPUs
52
 
53
+ ### Halo 😇 is driven entirely by a single YAML config
 
54
 
55
  ```yaml
56
  model_name_or_path: zai-org/GLM-4.7-Flash
 
97
  enable_efficiency_metrics: true
98
  ```
99
 
 
 
100
  ## Emergent skills
101
 
102
  The model picked up several of the teacher’s skills and patterns during fine-tuning:
 
197
 
198
  <details>
199
  <summary>OpenRouter providers per model</summary>
 
200
 
201
  | model | provider | input $/Mtok | output $/Mtok | cache_read $/Mtok | cache_write $/Mtok |
202
  | --- | --- | --- | --- | --- | --- |
 
227
 
228
  <details>
229
  <summary>Prompt template</summary>
 
230
 
231
  ```markdown
232
  You are solving a real GitHub issue in the `{repo}` repository. The repo is already cloned and set up in the current working directory.
 
255
 
256
  <details>
257
  <summary>Inference with SGLang 0.5.14, CUDA 13</summary>
 
258
 
259
  ```bash
260
  python3 -m sglang_router.launch_server \