Add LoRA adapter, model card, and eval artifacts

#1
by Poojanbuselvan - opened
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ These LoRA adapter weights are released under the MIT License, matching the
2
+ license declared by the base model, zai-org/GLM-4.7-Flash. Note that the base
3
+ model's repository declares `mit` in its metadata but does not itself ship a
4
+ LICENSE file; the standard MIT text is reproduced below and applies to this
5
+ adapter.
6
+
7
+ The training data is derived from b-mc2/sql-create-context, which is licensed
8
+ CC-BY-4.0. That license is not superseded by this one — downstream use of these
9
+ weights should carry its attribution requirement.
10
+
11
+ ---
12
+
13
+ MIT License
14
+
15
+ Copyright (c) 2026 SASVA AI Model Cognition Labs (MCL)
16
+
17
+ Permission is hereby granted, free of charge, to any person obtaining a copy
18
+ of this software and associated documentation files (the "Software"), to deal
19
+ in the Software without restriction, including without limitation the rights
20
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
21
+ copies of the Software, and to permit persons to whom the Software is
22
+ furnished to do so, subject to the following conditions:
23
+
24
+ The above copyright notice and this permission notice shall be included in all
25
+ copies or substantial portions of the Software.
26
+
27
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
28
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
29
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
30
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
31
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
32
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
33
+ SOFTWARE.
README.md CHANGED
@@ -1,3 +1,412 @@
1
  ---
2
- license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ # ---- Identity -------------------------------------------------------------
3
+ base_model: zai-org/GLM-4.7-Flash
4
+ base_model_relation: adapter
5
+ library_name: peft
6
+ pipeline_tag: text-generation
7
+ language:
8
+ - en
9
+ license: mit # verified via the Hub API: zai-org/GLM-4.7-Flash reports `mit`.
10
+ # The training data is CC-BY-4.0 — see Licence below.
11
+
12
+ # ---- Discovery ------------------------------------------------------------
13
+ tags:
14
+ - lora
15
+ - qlora
16
+ - peft
17
+ - sft
18
+ - trl
19
+ - text-to-sql
20
+ - sql
21
+
22
+ datasets:
23
+ - b-mc2/sql-create-context
24
+
25
+ metrics:
26
+ - exact_match
27
+ - bleu
28
+ - rouge
29
+
30
+ # ---- Structured evaluation ------------------------------------------------
31
+ model-index:
32
+ - name: glm-4.7-flash-sql-create-context-lora
33
+ results:
34
+ - task:
35
+ type: text-generation
36
+ name: Text-to-SQL (natural language + CREATE TABLE schema -> SQL)
37
+ dataset:
38
+ type: sql-create-context-val
39
+ name: sql-create-context derived validation split (400 pairs)
40
+ split: validation
41
+ metrics:
42
+ - type: exact_match
43
+ name: Exact match
44
+ value: 0.805
45
+ - type: bleu
46
+ name: BLEU
47
+ value: 0.940082
48
+ - type: rouge
49
+ name: ROUGE-L
50
+ value: 0.986055
51
+ args:
52
+ rouge_type: rougeL
53
  ---
54
+
55
+ # GLM-4.7-Flash Text-to-SQL (LoRA)
56
+
57
+ Given a natural-language question and a `CREATE TABLE` schema, emits exactly one
58
+ SQL query answering that question against that schema. For natural-language
59
+ query interfaces over a known relational schema.
60
+
61
+ This is a **LoRA adapter for**
62
+ [zai-org/GLM-4.7-Flash](https://huggingface.co/zai-org/GLM-4.7-Flash), trained
63
+ with **QLoRA (4-bit NF4 base, bf16 compute)** via
64
+ [TRL](https://github.com/huggingface/trl) SFT.
65
+
66
+ ## Model details
67
+
68
+ | | |
69
+ |---|---|
70
+ | Developed by | SASVA AI Model Cognition Labs (MCL) Team |
71
+ | Base model | [`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash) |
72
+ | Base parameters | 30B total / 3B active (MoE) — 29,943,396,864 in the merged bf16 build |
73
+ | Architecture family | `glm4_moe_lite` (`Glm4MoeLiteForCausalLM`), 47 layers, hidden size 2048, vocab 154,880 |
74
+ | Adaptation | LoRA (`r=64`, `alpha=128`, `dropout=0.05`) |
75
+ | Trainable modules | `q_a_proj`, `q_b_proj`, `kv_a_proj_with_mqa`, `kv_b_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` |
76
+ | Training method | `qlora` (4-bit NF4, double quant, bf16 compute) |
77
+ | Refinement | none |
78
+ | Language | English (questions) / SQL (outputs) |
79
+ | License | MIT (inherited from the base model) |
80
+
81
+ Trainable parameters: **118,140,928** — **0.3945%** of the base. The adapter file
82
+ is 472,671,000 bytes (752 fp32 tensors: a `lora_A` + `lora_B` pair for each of
83
+ the 8 target modules across all 47 layers).
84
+
85
+ GLM-4.7-Flash uses Multi-head Latent Attention, so the attention target modules
86
+ are the MLA projections (`q_a_proj`/`q_b_proj`/`kv_a_proj_with_mqa`/`kv_b_proj`),
87
+ not `q_proj`/`k_proj`/`v_proj`. Targeting the conventional names would silently
88
+ adapt nothing.
89
+
90
+ ## Intended use
91
+
92
+ **Direct use.** Translate one English question plus one `CREATE TABLE` schema
93
+ into one SQL query. The model was trained on a specific prompt shape and that
94
+ shape is part of the contract:
95
+
96
+ - System prompt (verbatim): *"You are a text-to-SQL engine. Given a
97
+ natural-language question and a CREATE TABLE schema, output exactly one SQL
98
+ query that answers the question against the provided schema. Output only the
99
+ raw SQL query on a single line with no explanation, no markdown formatting,
100
+ and no additional text."*
101
+ - User turn: the question, a blank line, then the schema inside a fenced code
102
+ block.
103
+ - Applied through the tokenizer's chat template (`chat_template.jinja`, shipped
104
+ in this repo) with `enable_thinking=False`. Do not concatenate strings by hand.
105
+ - The query is the **first line** of the generation; discard anything after it.
106
+
107
+ **Out of scope.**
108
+ - **Not validated against a live database.** The model is scored on string
109
+ similarity to a reference query, never on execution. A syntactically perfect
110
+ query can still be semantically wrong. Parse and, where you can, dry-run
111
+ against the real schema before trusting output.
112
+ - **Never interpolate output into a privileged connection.** Treat generated SQL
113
+ as untrusted input: run it read-only, with least privilege, on a connection
114
+ that cannot write or drop.
115
+ - Multi-table joins, CTEs, window functions, subqueries, and DDL/DML are largely
116
+ out of distribution — the training data is dominated by single-table
117
+ `SELECT`s. Measured join accuracy is poor (see Limitations).
118
+ - Dialect is not controllable. The model reproduces the source corpus's
119
+ conventions (double-quoted string literals, lower-cased comparison values),
120
+ which are not portable to every engine.
121
+ - Not a general-purpose assistant. It emits a bare query, never prose.
122
+
123
+ ## How to get started
124
+
125
+ ```python
126
+ import torch
127
+ from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
128
+ from peft import PeftModel
129
+
130
+ BASE = "zai-org/GLM-4.7-Flash"
131
+ ADAPTER = "SASVAAI/GLM-4.7-Flash-sql-create-context"
132
+
133
+ # 4-bit NF4 matches the numerics the adapter was trained against. A bf16 base
134
+ # also works and scores the same (see Merged-weights equivalence) but needs
135
+ # ~60 GB rather than ~22 GB.
136
+ bnb = BitsAndBytesConfig(
137
+ load_in_4bit=True,
138
+ bnb_4bit_quant_type="nf4",
139
+ bnb_4bit_compute_dtype=torch.bfloat16,
140
+ bnb_4bit_use_double_quant=True,
141
+ )
142
+
143
+ tokenizer = AutoTokenizer.from_pretrained(ADAPTER)
144
+ model = AutoModelForCausalLM.from_pretrained(
145
+ BASE, quantization_config=bnb, dtype=torch.bfloat16, device_map="auto"
146
+ )
147
+ model = PeftModel.from_pretrained(model, ADAPTER)
148
+ model.eval()
149
+
150
+ SYSTEM = (
151
+ "You are a text-to-SQL engine. Given a natural-language question and a "
152
+ "CREATE TABLE schema, output exactly one SQL query that answers the question "
153
+ "against the provided schema. Output only the raw SQL query on a single line "
154
+ "with no explanation, no markdown formatting, and no additional text."
155
+ )
156
+
157
+ question = "Which kingdom has Suin as its capital?"
158
+ schema = "CREATE TABLE table_name_65 (name_of_kingdom VARCHAR, capital VARCHAR)"
159
+
160
+ messages = [
161
+ {"role": "system", "content": SYSTEM},
162
+ {"role": "user", "content": f"{question}\n\n```\n{schema}\n```"},
163
+ ]
164
+ prompt = tokenizer.apply_chat_template(
165
+ messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
166
+ )
167
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
168
+
169
+ out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
170
+ text = tokenizer.decode(out[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
171
+ print(text.strip().splitlines()[0])
172
+ # -> SELECT name_of_kingdom FROM table_name_65 WHERE capital = "suin"
173
+ ```
174
+
175
+ The base model is ~59 GB in bfloat16, or ~22 GB per GPU under 4-bit NF4.
176
+
177
+ > Decoding matters. This model was evaluated with greedy decoding
178
+ > (`do_sample=False`, `max_new_tokens=128`). Sampling will not reproduce the
179
+ > reported numbers.
180
+
181
+ **Serving note.** These adapter weights do **not** load as a vLLM LoRA on this
182
+ architecture — vLLM's MLA path asserts inside
183
+ `DeepSeekV2FusedQkvAProjLinear` because `q_a_proj` and `kv_a_proj_with_mqa` are
184
+ fused into one module that a LoRA cannot be attached to. A merged build of these
185
+ weights serves under vLLM without complaint. For vLLM deployment, merge first
186
+ (`peft.merge_and_unload()`).
187
+
188
+ ## Training details
189
+
190
+ **Data.** A 4,000-pair subset of
191
+ [`b-mc2/sql-create-context`](https://huggingface.co/datasets/b-mc2/sql-create-context)
192
+ (78,577 pairs, itself derived from WikiSQL and Spider), split 90/10 by this
193
+ project's data-generation stage. Each record is
194
+ `{"instruction": <question>, "input": <CREATE TABLE ...>, "output": <SQL>}`.
195
+ The subset selection is a generated artifact, not a published split — the
196
+ `id`/`instruction`/`input`/`gold` tuples in `predictions.jsonl` are the
197
+ authoritative record of what was evaluated.
198
+
199
+ | | |
200
+ |---|---|
201
+ | Train samples | 3,600 |
202
+ | Validation samples | 400 |
203
+ | Prompt format | chat template + system prompt (see Intended use) |
204
+
205
+ ### Method
206
+
207
+ | | |
208
+ |---|---|
209
+ | SFT method | `qlora` |
210
+ | Base quantisation during training | 4-bit NF4, double quant, bf16 compute |
211
+ | Refinement stage | none |
212
+ | Hardware | 4x NVIDIA H200 (141 GB), `torchrun --nproc_per_node=4` |
213
+
214
+ `qlora` is one of three methods considered for this model, alongside 8-bit LoRA
215
+ and attention-only 4-bit LoRA. All three were tried; `qlora` scored highest
216
+ every time it was compared against the other two on the same data.
217
+
218
+ No refinement stage ran; the published weights are the SFT adapter.
219
+
220
+ ### Final hyperparameters
221
+
222
+ | Hyperparameter | Value | Source |
223
+ |---|---|---|
224
+ | `learning_rate` | 0.0002 | `training_args.bin` |
225
+ | `lr_scheduler_type` | cosine | `training_args.bin` |
226
+ | `num_train_epochs` | 3 | `training_args.bin` |
227
+ | `per_device_train_batch_size` | 1 | `training_args.bin` |
228
+ | `gradient_accumulation_steps` | 4 | `training_args.bin` |
229
+ | `max_length` | 2048 | `training_args.bin` |
230
+ | `warmup_ratio` | 0.05 | `training_args.bin` |
231
+ | `weight_decay` | 0.01 | `training_args.bin` |
232
+ | `optim` / `max_grad_norm` / `seed` | `adamw_torch` / 1.0 / 42 | `training_args.bin` |
233
+ | `bf16` / `gradient_checkpointing` | true / true (`use_reentrant=False`) | `training_args.bin` |
234
+ | `neftune_noise_alpha` / `packing` | None / false | `training_args.bin` |
235
+ | `lora_r` / `lora_alpha` / `lora_dropout` | 64 / 128 / 0.05 | `adapter_config.json` |
236
+ | `target_modules` | the 8 listed in Model details | `adapter_config.json` |
237
+
238
+ **Effective batch size: 16** (`1 x 4 x 4`). Optimizer steps: 675.
239
+
240
+ KD parameters are omitted deliberately — this is a `qlora` run, not
241
+ `bf16_lora_kd`, so `KD_ALPHA`/`KD_BETA`/`KD_TEMPERATURE` carry inert defaults
242
+ that would imply distillation that did not happen.
243
+
244
+ This configuration is **not a unique optimum**. Other configurations reached the
245
+ same `exact_match`; this one was published for being the simplest and cheapest
246
+ of them — fewest epochs, shortest training time — and because it scored
247
+ marginally higher on BLEU and ROUGE-L. Treat the values as a good working point,
248
+ not a tuned maximum.
249
+
250
+ **Observed training metrics.**
251
+
252
+ | | |
253
+ |---|---|
254
+ | Final train loss | 0.4594 |
255
+ | Mean train loss | 0.5646523337894016 |
256
+ | Train runtime | 6392.6844s |
257
+ | Total FLOPs | 2.3441226939026637e+17 |
258
+ | Throughput | 1.689 samples/s, 0.106 steps/s |
259
+
260
+ No eval loss was computed during training; the loop scores on generation, not
261
+ perplexity. Loss falls from 2.3882 at step 10 to 1.1077 at step 20 and 0.5940 by
262
+ step 170, then improves slowly to ~0.45 by step 660. The task is essentially
263
+ learned within the first quarter of epoch 1; epochs 2 and 3 together buy roughly
264
+ 0.14 of training loss.
265
+
266
+ ## Evaluation
267
+
268
+ **Protocol.** All 400 validation pairs, no sampling. Predictions generated
269
+ greedily (`do_sample=False`, `max_new_tokens=128`) through the same chat template
270
+ used in training, against a 4-bit NF4 base to match training numerics. The
271
+ predicted query is the **first line** of the generation, stripped. No constrained
272
+ decoding and no SQL grammar were applied. Exact match is byte equality against
273
+ the reference query; BLEU and ROUGE-L are computed over the same strings.
274
+
275
+ | Metric | Value |
276
+ |---|---|
277
+ | Exact match | 0.805000 (322 / 400) |
278
+ | BLEU | 0.940082 |
279
+ | ROUGE-L | 0.986055 |
280
+ | Samples | 400 |
281
+
282
+ **Baseline for comparison.** **Not measured.** The untuned
283
+ `zai-org/GLM-4.7-Flash` was never scored on this split, so these numbers
284
+ quantify the fine-tuned model's performance but do not establish how much of it
285
+ the fine-tuning is responsible for.
286
+
287
+ **This is a validation split, not a held-out test set.** Hyperparameters were
288
+ selected against it, so expect optimistic bias. A clean estimate needs a third
289
+ split that was never used for selection.
290
+
291
+ ## Limitations and bias
292
+
293
+ **Exact match understates the model; BLEU and ROUGE-L overstate it.** The
294
+ 0.805 / 0.940 / 0.986 spread is the story of this model. Exact match is byte
295
+ equality, so a semantically identical query loses the point on quoting or a
296
+ missing `DISTINCT`:
297
+
298
+ ```
299
+ question: Find the states where have some college students in tryout and their decisions are yes.
300
+ schema: CREATE TABLE tryout (cName VARCHAR, decision VARCHAR);
301
+ CREATE TABLE college (state VARCHAR, cName VARCHAR)
302
+ gold: SELECT DISTINCT T1.state FROM college AS T1 JOIN tryout AS T2 ON T1.cName = T2.cName WHERE T2.decision = 'yes'
303
+ pred: SELECT T1.state FROM college AS T1 JOIN tryout AS T2 ON T1.cName = T2.cName WHERE T2.decision = "yes"
304
+ ```
305
+
306
+ Relaxing to case- and whitespace-insensitive comparison moves exact match from
307
+ 0.805 to **0.8125** (325/400) — so only 3 of the 78 misses are pure formatting.
308
+ The other 75 are real semantic or structural errors. ROUGE-L at 0.986 mostly
309
+ measures that both strings are short SQL over the same table names; it is not
310
+ evidence of correctness.
311
+
312
+ **Output is never degenerate.** All 400 golds and all 400 predictions begin with
313
+ `SELECT`; the model never emitted prose, markdown, or an empty string. Format
314
+ compliance is not the failure mode.
315
+
316
+ **Joins are the failure mode.** Miss rate by gold-query feature:
317
+
318
+ | Gold query contains | n | misses | miss rate |
319
+ |---|---|---|---|
320
+ | `JOIN` | 13 | 10 | **76.9%** |
321
+ | `ORDER BY` | 3 | 1 | 33.3% |
322
+ | aggregate (`COUNT`/`SUM`/`AVG`/`MIN`/`MAX`) | 138 | 37 | 26.8% |
323
+ | `GROUP BY` | 8 | 2 | 25.0% |
324
+ | multi-predicate `WHERE` (`AND`/`OR`) | 123 | 27 | 22.0% |
325
+ | single-predicate `SELECT..WHERE` | 181 | 22 | 12.2% |
326
+
327
+ The model is reliable on the shape it saw constantly (one table, one predicate)
328
+ and unreliable on the shape it barely saw. **Do not deploy this on a
329
+ multi-table schema.** Note the join, `ORDER BY`, and `GROUP BY` rows rest on
330
+ 13, 3, and 8 examples respectively — read them as a strong warning, not a precise
331
+ rate.
332
+
333
+ Complexity tracks length: correct predictions have a mean gold length of 10.6
334
+ tokens, misses 12.8.
335
+
336
+ **Dialect is baked in.** The model emits the source corpus's double-quoted
337
+ string literals and lower-cases comparison values. On engines where `"x"` is an
338
+ identifier rather than a string (PostgreSQL, ANSI mode), output will not run
339
+ unmodified.
340
+
341
+ **No execution or injection safety.** Correctness is measured only as string
342
+ similarity to a reference. Nothing here prevents a generated query from being
343
+ expensive, wrong, or destructive against a real database.
344
+
345
+ **Inherits all biases and limitations of the base model.** This adapter changes
346
+ 0.3945% of the parameters and was not evaluated for social bias, safety, or
347
+ fairness.
348
+
349
+ ## Merged-weights equivalence
350
+
351
+ A merged build of these weights (base + adapter folded into one standalone bf16
352
+ model, `W + (alpha/r) * B @ A`) was evaluated on the identical split:
353
+
354
+ | Metric | Adapter (4-bit base) | Merged (bf16) | Delta |
355
+ |---|---|---|---|
356
+ | Exact match | 0.805000 | 0.805000 | 0.000000 |
357
+ | BLEU | 0.940082 | 0.942079 | +0.001997 |
358
+ | ROUGE-L | 0.986055 | 0.987663 | +0.001608 |
359
+
360
+ **39 of 400 predictions differ** (9.75%) — far more churn than a bf16-trained
361
+ adapter would show, because merging into an unquantised base genuinely changes
362
+ the numerics the adapter was fitted against. The changes cancel exactly:
363
+ **10 predictions flip correct→incorrect and 10 flip incorrect→correct**, leaving
364
+ exact match identical and BLEU/ROUGE-L marginally higher.
365
+
366
+ So a merged distribution is behaviourally equivalent in aggregate but not
367
+ prediction-for-prediction. Merging is also the practical route to vLLM serving
368
+ (see Serving note). MIT permits distributing derivative works, so publishing a
369
+ merged build is allowed.
370
+
371
+ ## Environmental impact
372
+
373
+ | | |
374
+ |---|---|
375
+ | Hardware | 4x NVIDIA H200 (141 GB) |
376
+ | Training time | 106.5 minutes (6392.6844s) |
377
+ | Cloud provider / region | on-premise |
378
+
379
+ Covers the training of these published weights only. It excludes the wider
380
+ hyperparameter search that selected them, which cost substantially more.
381
+
382
+ ## Framework versions
383
+
384
+ - PEFT 0.18.1
385
+ - TRL: 1.0.0
386
+ - Transformers: 5.7.0.dev0
387
+ - Pytorch: 2.5.1+cu121
388
+ - Datasets: 4.8.4
389
+ - Tokenizers: 0.22.2
390
+ - bitsandbytes: 0.49.2
391
+
392
+ `transformers` is a git-main build: GLM-4.7-Flash's `Glm4MoeLite` architecture is
393
+ not in the stable PyPI release.
394
+
395
+ ## Licence
396
+
397
+ Adapter weights: **MIT**, inherited from
398
+ [`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash)
399
+ (verified via the Hub API). Training data:
400
+ [`b-mc2/sql-create-context`](https://huggingface.co/datasets/b-mc2/sql-create-context),
401
+ licensed **CC-BY-4.0** — downstream use should carry that attribution.
402
+
403
+ ## Citation
404
+
405
+ ```bibtex
406
+ @misc{glm47flash_sql_create_context_lora_2026,
407
+ title = {GLM-4.7-Flash Text-to-SQL (LoRA)},
408
+ author = {{SASVA AI Model Cognition Labs (MCL) Team}},
409
+ year = {2026},
410
+ url = {https://huggingface.co/SASVAAI/GLM-4.7-Flash-sql-create-context}
411
+ }
412
+ ```
adapter_config.json ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "zai-org/GLM-4.7-Flash",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "layer_replication": null,
16
+ "layers_pattern": null,
17
+ "layers_to_transform": null,
18
+ "loftq_config": {},
19
+ "lora_alpha": 128,
20
+ "lora_bias": false,
21
+ "lora_dropout": 0.05,
22
+ "megatron_config": null,
23
+ "megatron_core": "megatron.core",
24
+ "modules_to_save": null,
25
+ "peft_type": "LORA",
26
+ "peft_version": "0.18.1",
27
+ "qalora_group_size": 16,
28
+ "r": 64,
29
+ "rank_pattern": {},
30
+ "revision": null,
31
+ "target_modules": [
32
+ "up_proj",
33
+ "down_proj",
34
+ "gate_proj",
35
+ "q_a_proj",
36
+ "kv_a_proj_with_mqa",
37
+ "o_proj",
38
+ "kv_b_proj",
39
+ "q_b_proj"
40
+ ],
41
+ "target_parameters": null,
42
+ "task_type": "CAUSAL_LM",
43
+ "trainable_token_indices": null,
44
+ "use_dora": false,
45
+ "use_qalora": false,
46
+ "use_rslora": false
47
+ }
adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f7c4b6367c7ac51b2700e675946c6e0dffb7d037d46dadf8096fecd3757b9d41
3
+ size 472671000
all_results.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "total_flos": 2.3441226939026637e+17,
3
+ "train_loss": 0.5646523337894016,
4
+ "train_runtime": 6392.6844,
5
+ "train_samples_per_second": 1.689,
6
+ "train_steps_per_second": 0.106
7
+ }
chat_template.jinja ADDED
@@ -0,0 +1,86 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ [gMASK]<sop>
2
+ {%- if tools -%}
3
+ <|system|>
4
+ # Tools
5
+
6
+ You may call one or more functions to assist with the user query.
7
+
8
+ You are provided with function signatures within <tools></tools> XML tags:
9
+ <tools>
10
+ {% for tool in tools %}
11
+ {{ tool | tojson(ensure_ascii=False) }}
12
+ {% endfor %}
13
+ </tools>
14
+
15
+ For each function call, output the function name and arguments within the following XML format:
16
+ <tool_call>{function-name}<arg_key>{arg-key-1}</arg_key><arg_value>{arg-value-1}</arg_value><arg_key>{arg-key-2}</arg_key><arg_value>{arg-value-2}</arg_value>...</tool_call>{%- endif -%}
17
+ {%- macro visible_text(content) -%}
18
+ {%- if content is string -%}
19
+ {{- content }}
20
+ {%- elif content is iterable and content is not mapping -%}
21
+ {%- for item in content -%}
22
+ {%- if item is mapping and item.type == 'text' -%}
23
+ {{- item.text }}
24
+ {%- elif item is string -%}
25
+ {{- item }}
26
+ {%- endif -%}
27
+ {%- endfor -%}
28
+ {%- else -%}
29
+ {{- content }}
30
+ {%- endif -%}
31
+ {%- endmacro -%}
32
+ {%- set ns = namespace(last_user_index=-1) %}
33
+ {%- for m in messages %}
34
+ {%- if m.role == 'user' %}
35
+ {% set ns.last_user_index = loop.index0 -%}
36
+ {%- endif %}
37
+ {%- endfor %}
38
+ {% for m in messages %}
39
+ {%- if m.role == 'user' -%}<|user|>{{ visible_text(m.content) }}
40
+ {%- elif m.role == 'assistant' -%}
41
+ <|assistant|>
42
+ {%- set reasoning_content = '' %}
43
+ {%- set content = visible_text(m.content) %}
44
+ {%- if m.reasoning_content is string %}
45
+ {%- set reasoning_content = m.reasoning_content %}
46
+ {%- else %}
47
+ {%- if '</think>' in content %}
48
+ {%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
49
+ {%- set content = content.split('</think>')[-1].lstrip('\n') %}
50
+ {%- endif %}
51
+ {%- endif %}
52
+ {%- if ((clear_thinking is defined and not clear_thinking) or loop.index0 > ns.last_user_index) and reasoning_content -%}
53
+ {{ '<think>' + reasoning_content.strip() + '</think>'}}
54
+ {%- else -%}
55
+ {{ '</think>' }}
56
+ {%- endif -%}
57
+ {%- if content.strip() -%}
58
+ {{ content.strip() }}
59
+ {%- endif -%}
60
+ {% if m.tool_calls %}
61
+ {% for tc in m.tool_calls %}
62
+ {%- if tc.function %}
63
+ {%- set tc = tc.function %}
64
+ {%- endif %}
65
+ {{- '<tool_call>' + tc.name -}}
66
+ {% set _args = tc.arguments %}{% for k, v in _args.items() %}<arg_key>{{ k }}</arg_key><arg_value>{{ v | tojson(ensure_ascii=False) if v is not string else v }}</arg_value>{% endfor %}</tool_call>{% endfor %}
67
+ {% endif %}
68
+ {%- elif m.role == 'tool' -%}
69
+ {%- if m.content is string -%}
70
+ {%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
71
+ {{- '<|observation|>' }}
72
+ {%- endif %}
73
+ {{- '<tool_response>' }}
74
+ {{- m.content }}
75
+ {{- '</tool_response>' }}
76
+ {%- else -%}
77
+ <|observation|>{% for tr in m.content %}
78
+ <tool_response>{{ tr.output if tr.output is defined else tr }}</tool_response>{% endfor -%}
79
+ {% endif -%}
80
+ {%- elif m.role == 'system' -%}
81
+ <|system|>{{ visible_text(m.content) }}
82
+ {%- endif -%}
83
+ {%- endfor -%}
84
+ {%- if add_generation_prompt -%}
85
+ <|assistant|>{{- '</think>' if (enable_thinking is defined and not enable_thinking) else '<think>' -}}
86
+ {%- endif -%}
eval_results.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metrics": {
3
+ "exact_match": 0.805,
4
+ "bleu": 0.9400823275452048,
5
+ "rouge_l": 0.9860549005631489
6
+ },
7
+ "num_samples": 400
8
+ }
predictions.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:19e773648cb4e65de8660ea6365e10acca112d42a854923df93db4a6f333a82d
3
+ size 20217442
tokenizer_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "clean_up_tokenization_spaces": false,
4
+ "do_lower_case": false,
5
+ "eos_token": "<|endoftext|>",
6
+ "is_local": false,
7
+ "local_files_only": false,
8
+ "model_max_length": 128000,
9
+ "pad_token": "<|endoftext|>",
10
+ "padding_side": "right",
11
+ "remove_space": false,
12
+ "tokenizer_class": "TokenizersBackend"
13
+ }
train_results.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "total_flos": 2.3441226939026637e+17,
3
+ "train_loss": 0.5646523337894016,
4
+ "train_runtime": 6392.6844,
5
+ "train_samples_per_second": 1.689,
6
+ "train_steps_per_second": 0.106
7
+ }
trainer_state.json ADDED
@@ -0,0 +1,724 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": 675,
3
+ "best_metric": 0.5975516438484192,
4
+ "best_model_checkpoint": "/home/svadouser1/pooja/llm_autotuning/checkpoints/exp_20260827_221617/checkpoint-675",
5
+ "epoch": 3.0,
6
+ "eval_steps": 9999,
7
+ "global_step": 675,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "entropy": 1.461071628332138,
14
+ "epoch": 0.044444444444444446,
15
+ "grad_norm": 1.6171875,
16
+ "learning_rate": 5.294117647058824e-05,
17
+ "loss": 2.388203811645508,
18
+ "mean_token_accuracy": 0.6123090386390686,
19
+ "num_tokens": 19486.0,
20
+ "step": 10
21
+ },
22
+ {
23
+ "entropy": 0.9562640145421029,
24
+ "epoch": 0.08888888888888889,
25
+ "grad_norm": 0.6875,
26
+ "learning_rate": 0.00011176470588235294,
27
+ "loss": 1.1077078819274901,
28
+ "mean_token_accuracy": 0.8430786401033401,
29
+ "num_tokens": 38826.0,
30
+ "step": 20
31
+ },
32
+ {
33
+ "entropy": 0.77465368360281,
34
+ "epoch": 0.13333333333333333,
35
+ "grad_norm": 0.455078125,
36
+ "learning_rate": 0.00017058823529411766,
37
+ "loss": 0.8623821258544921,
38
+ "mean_token_accuracy": 0.864073084294796,
39
+ "num_tokens": 58244.0,
40
+ "step": 30
41
+ },
42
+ {
43
+ "entropy": 0.7483784720301628,
44
+ "epoch": 0.17777777777777778,
45
+ "grad_norm": 0.423828125,
46
+ "learning_rate": 0.00019996997576394573,
47
+ "loss": 0.7713714599609375,
48
+ "mean_token_accuracy": 0.8697837173938752,
49
+ "num_tokens": 77711.0,
50
+ "step": 40
51
+ },
52
+ {
53
+ "entropy": 0.7391687169671058,
54
+ "epoch": 0.2222222222222222,
55
+ "grad_norm": 0.376953125,
56
+ "learning_rate": 0.00019972989003925543,
57
+ "loss": 0.7115508556365967,
58
+ "mean_token_accuracy": 0.8751530349254608,
59
+ "num_tokens": 97104.0,
60
+ "step": 50
61
+ },
62
+ {
63
+ "entropy": 0.7042164742946625,
64
+ "epoch": 0.26666666666666666,
65
+ "grad_norm": 0.330078125,
66
+ "learning_rate": 0.00019925029517454195,
67
+ "loss": 0.6672511577606202,
68
+ "mean_token_accuracy": 0.8725973218679428,
69
+ "num_tokens": 116498.0,
70
+ "step": 60
71
+ },
72
+ {
73
+ "entropy": 0.6745826601982117,
74
+ "epoch": 0.3111111111111111,
75
+ "grad_norm": 0.361328125,
76
+ "learning_rate": 0.00019853234295442638,
77
+ "loss": 0.6365277290344238,
78
+ "mean_token_accuracy": 0.872233435511589,
79
+ "num_tokens": 136089.0,
80
+ "step": 70
81
+ },
82
+ {
83
+ "entropy": 0.6878040120005607,
84
+ "epoch": 0.35555555555555557,
85
+ "grad_norm": 0.37109375,
86
+ "learning_rate": 0.00019757775759738272,
87
+ "loss": 0.6395976066589355,
88
+ "mean_token_accuracy": 0.8741081401705741,
89
+ "num_tokens": 155486.0,
90
+ "step": 80
91
+ },
92
+ {
93
+ "entropy": 0.6898319900035859,
94
+ "epoch": 0.4,
95
+ "grad_norm": 0.330078125,
96
+ "learning_rate": 0.00019638883161489224,
97
+ "loss": 0.6368544578552247,
98
+ "mean_token_accuracy": 0.8765213042497635,
99
+ "num_tokens": 175116.0,
100
+ "step": 90
101
+ },
102
+ {
103
+ "entropy": 0.656438185274601,
104
+ "epoch": 0.4444444444444444,
105
+ "grad_norm": 0.33203125,
106
+ "learning_rate": 0.0001949684203057978,
107
+ "loss": 0.6183670043945313,
108
+ "mean_token_accuracy": 0.8779006212949753,
109
+ "num_tokens": 194495.0,
110
+ "step": 100
111
+ },
112
+ {
113
+ "entropy": 0.6711793914437294,
114
+ "epoch": 0.4888888888888889,
115
+ "grad_norm": 0.388671875,
116
+ "learning_rate": 0.00019331993489907977,
117
+ "loss": 0.633416748046875,
118
+ "mean_token_accuracy": 0.8727847129106522,
119
+ "num_tokens": 214311.0,
120
+ "step": 110
121
+ },
122
+ {
123
+ "entropy": 0.6290019288659096,
124
+ "epoch": 0.5333333333333333,
125
+ "grad_norm": 0.33203125,
126
+ "learning_rate": 0.00019144733436152284,
127
+ "loss": 0.5889303684234619,
128
+ "mean_token_accuracy": 0.8814716726541519,
129
+ "num_tokens": 233676.0,
130
+ "step": 120
131
+ },
132
+ {
133
+ "entropy": 0.6458568304777146,
134
+ "epoch": 0.5777777777777777,
135
+ "grad_norm": 0.318359375,
136
+ "learning_rate": 0.00018935511588994715,
137
+ "loss": 0.6043062686920166,
138
+ "mean_token_accuracy": 0.8794952735304833,
139
+ "num_tokens": 253188.0,
140
+ "step": 130
141
+ },
142
+ {
143
+ "entropy": 0.6591948792338371,
144
+ "epoch": 0.6222222222222222,
145
+ "grad_norm": 0.32421875,
146
+ "learning_rate": 0.000187048304110838,
147
+ "loss": 0.6233312129974365,
148
+ "mean_token_accuracy": 0.873698017001152,
149
+ "num_tokens": 273185.0,
150
+ "step": 140
151
+ },
152
+ {
153
+ "entropy": 0.6109217047691345,
154
+ "epoch": 0.6666666666666666,
155
+ "grad_norm": 0.291015625,
156
+ "learning_rate": 0.00018453243901331195,
157
+ "loss": 0.5925732135772706,
158
+ "mean_token_accuracy": 0.884858050942421,
159
+ "num_tokens": 292431.0,
160
+ "step": 150
161
+ },
162
+ {
163
+ "entropy": 0.6377103522419929,
164
+ "epoch": 0.7111111111111111,
165
+ "grad_norm": 0.3671875,
166
+ "learning_rate": 0.00018181356264439905,
167
+ "loss": 0.5884737968444824,
168
+ "mean_token_accuracy": 0.8802034169435501,
169
+ "num_tokens": 311687.0,
170
+ "step": 160
171
+ },
172
+ {
173
+ "entropy": 0.6447647094726563,
174
+ "epoch": 0.7555555555555555,
175
+ "grad_norm": 0.3046875,
176
+ "learning_rate": 0.0001788982045985939,
177
+ "loss": 0.5939778327941895,
178
+ "mean_token_accuracy": 0.8817022785544395,
179
+ "num_tokens": 330813.0,
180
+ "step": 170
181
+ },
182
+ {
183
+ "entropy": 0.6029744669795036,
184
+ "epoch": 0.8,
185
+ "grad_norm": 0.3125,
186
+ "learning_rate": 0.00017579336633652317,
187
+ "loss": 0.5752908706665039,
188
+ "mean_token_accuracy": 0.8855120778083801,
189
+ "num_tokens": 350308.0,
190
+ "step": 180
191
+ },
192
+ {
193
+ "entropy": 0.6441277623176574,
194
+ "epoch": 0.8444444444444444,
195
+ "grad_norm": 0.259765625,
196
+ "learning_rate": 0.00017250650437038964,
197
+ "loss": 0.5958236217498779,
198
+ "mean_token_accuracy": 0.8808144956827164,
199
+ "num_tokens": 370083.0,
200
+ "step": 190
201
+ },
202
+ {
203
+ "entropy": 0.5766214028000831,
204
+ "epoch": 0.8888888888888888,
205
+ "grad_norm": 0.2734375,
206
+ "learning_rate": 0.0001690455123565743,
207
+ "loss": 0.5667627811431885,
208
+ "mean_token_accuracy": 0.8866458177566529,
209
+ "num_tokens": 389516.0,
210
+ "step": 200
211
+ },
212
+ {
213
+ "entropy": 0.6425503984093666,
214
+ "epoch": 0.9333333333333333,
215
+ "grad_norm": 0.3046875,
216
+ "learning_rate": 0.00016541870213840243,
217
+ "loss": 0.6137699127197266,
218
+ "mean_token_accuracy": 0.8819929376244545,
219
+ "num_tokens": 408940.0,
220
+ "step": 210
221
+ },
222
+ {
223
+ "entropy": 0.6159977808594703,
224
+ "epoch": 0.9777777777777777,
225
+ "grad_norm": 0.265625,
226
+ "learning_rate": 0.0001616347837846011,
227
+ "loss": 0.5788507461547852,
228
+ "mean_token_accuracy": 0.8852896198630333,
229
+ "num_tokens": 428166.0,
230
+ "step": 220
231
+ },
232
+ {
233
+ "entropy": 0.6020576372742653,
234
+ "epoch": 1.0222222222222221,
235
+ "grad_norm": 0.2353515625,
236
+ "learning_rate": 0.000157702844671387,
237
+ "loss": 0.5500371932983399,
238
+ "mean_token_accuracy": 0.8901642486453056,
239
+ "num_tokens": 447438.0,
240
+ "step": 230
241
+ },
242
+ {
243
+ "entropy": 0.5630205765366554,
244
+ "epoch": 1.0666666666666667,
245
+ "grad_norm": 0.275390625,
246
+ "learning_rate": 0.00015363232765842089,
247
+ "loss": 0.5382141590118408,
248
+ "mean_token_accuracy": 0.8885036528110504,
249
+ "num_tokens": 467009.0,
250
+ "step": 240
251
+ },
252
+ {
253
+ "entropy": 0.5631943918764591,
254
+ "epoch": 1.1111111111111112,
255
+ "grad_norm": 0.306640625,
256
+ "learning_rate": 0.00014943300841104094,
257
+ "loss": 0.5194426536560058,
258
+ "mean_token_accuracy": 0.8929360061883926,
259
+ "num_tokens": 486591.0,
260
+ "step": 250
261
+ },
262
+ {
263
+ "entropy": 0.5655099518597126,
264
+ "epoch": 1.1555555555555554,
265
+ "grad_norm": 0.259765625,
266
+ "learning_rate": 0.0001451149719232366,
267
+ "loss": 0.5290531158447266,
268
+ "mean_token_accuracy": 0.8912677451968193,
269
+ "num_tokens": 506180.0,
270
+ "step": 260
271
+ },
272
+ {
273
+ "entropy": 0.5660845704376698,
274
+ "epoch": 1.2,
275
+ "grad_norm": 0.326171875,
276
+ "learning_rate": 0.00014068858829774608,
277
+ "loss": 0.5366004943847656,
278
+ "mean_token_accuracy": 0.8909024119377136,
279
+ "num_tokens": 525696.0,
280
+ "step": 270
281
+ },
282
+ {
283
+ "entropy": 0.5818345256149768,
284
+ "epoch": 1.2444444444444445,
285
+ "grad_norm": 0.26171875,
286
+ "learning_rate": 0.0001361644878414428,
287
+ "loss": 0.5424720764160156,
288
+ "mean_token_accuracy": 0.8896975666284561,
289
+ "num_tokens": 545392.0,
290
+ "step": 280
291
+ },
292
+ {
293
+ "entropy": 0.5546154774725437,
294
+ "epoch": 1.2888888888888888,
295
+ "grad_norm": 0.3359375,
296
+ "learning_rate": 0.00013155353553582057,
297
+ "loss": 0.5176132202148438,
298
+ "mean_token_accuracy": 0.8928953528404235,
299
+ "num_tokens": 565036.0,
300
+ "step": 290
301
+ },
302
+ {
303
+ "entropy": 0.5676225990056991,
304
+ "epoch": 1.3333333333333333,
305
+ "grad_norm": 0.275390625,
306
+ "learning_rate": 0.0001268668049438902,
307
+ "loss": 0.5130613803863525,
308
+ "mean_token_accuracy": 0.8928169339895249,
309
+ "num_tokens": 584507.0,
310
+ "step": 300
311
+ },
312
+ {
313
+ "entropy": 0.5529272347688675,
314
+ "epoch": 1.3777777777777778,
315
+ "grad_norm": 0.326171875,
316
+ "learning_rate": 0.0001221155516161506,
317
+ "loss": 0.5317890644073486,
318
+ "mean_token_accuracy": 0.8913554430007935,
319
+ "num_tokens": 604008.0,
320
+ "step": 310
321
+ },
322
+ {
323
+ "entropy": 0.5658974438905716,
324
+ "epoch": 1.4222222222222223,
325
+ "grad_norm": 0.26953125,
326
+ "learning_rate": 0.0001173111860595032,
327
+ "loss": 0.5273444652557373,
328
+ "mean_token_accuracy": 0.8927861481904984,
329
+ "num_tokens": 623480.0,
330
+ "step": 320
331
+ },
332
+ {
333
+ "entropy": 0.5697685681283474,
334
+ "epoch": 1.4666666666666668,
335
+ "grad_norm": 0.322265625,
336
+ "learning_rate": 0.00011246524633402573,
337
+ "loss": 0.5287697792053223,
338
+ "mean_token_accuracy": 0.8912984907627106,
339
+ "num_tokens": 642979.0,
340
+ "step": 330
341
+ },
342
+ {
343
+ "entropy": 0.5488456651568413,
344
+ "epoch": 1.511111111111111,
345
+ "grad_norm": 0.326171875,
346
+ "learning_rate": 0.00010758937034341787,
347
+ "loss": 0.513739824295044,
348
+ "mean_token_accuracy": 0.8952319726347924,
349
+ "num_tokens": 662396.0,
350
+ "step": 340
351
+ },
352
+ {
353
+ "entropy": 0.55315260887146,
354
+ "epoch": 1.5555555555555556,
355
+ "grad_norm": 0.26953125,
356
+ "learning_rate": 0.00010269526788566408,
357
+ "loss": 0.5225533485412598,
358
+ "mean_token_accuracy": 0.8937178790569306,
359
+ "num_tokens": 681644.0,
360
+ "step": 350
361
+ },
362
+ {
363
+ "entropy": 0.5738558314740658,
364
+ "epoch": 1.6,
365
+ "grad_norm": 0.29296875,
366
+ "learning_rate": 9.779469253103684e-05,
367
+ "loss": 0.5239317417144775,
368
+ "mean_token_accuracy": 0.8908430561423302,
369
+ "num_tokens": 701030.0,
370
+ "step": 360
371
+ },
372
+ {
373
+ "entropy": 0.5420261971652508,
374
+ "epoch": 1.6444444444444444,
375
+ "grad_norm": 0.259765625,
376
+ "learning_rate": 9.289941339497719e-05,
377
+ "loss": 0.5274893283843994,
378
+ "mean_token_accuracy": 0.8941561266779899,
379
+ "num_tokens": 720294.0,
380
+ "step": 370
381
+ },
382
+ {
383
+ "entropy": 0.577882957458496,
384
+ "epoch": 1.6888888888888889,
385
+ "grad_norm": 0.33203125,
386
+ "learning_rate": 8.802118687364284e-05,
387
+ "loss": 0.5202207088470459,
388
+ "mean_token_accuracy": 0.8931182771921158,
389
+ "num_tokens": 739607.0,
390
+ "step": 380
391
+ },
392
+ {
393
+ "entropy": 0.5461296208202839,
394
+ "epoch": 1.7333333333333334,
395
+ "grad_norm": 0.30859375,
396
+ "learning_rate": 8.317172841000173e-05,
397
+ "loss": 0.5105932235717774,
398
+ "mean_token_accuracy": 0.8917569547891617,
399
+ "num_tokens": 759370.0,
400
+ "step": 390
401
+ },
402
+ {
403
+ "entropy": 0.5631573006510735,
404
+ "epoch": 1.7777777777777777,
405
+ "grad_norm": 0.298828125,
406
+ "learning_rate": 7.836268435827875e-05,
407
+ "loss": 0.5263056755065918,
408
+ "mean_token_accuracy": 0.8906426534056664,
409
+ "num_tokens": 778969.0,
410
+ "step": 400
411
+ },
412
+ {
413
+ "entropy": 0.5476421102881431,
414
+ "epoch": 1.8222222222222222,
415
+ "grad_norm": 0.255859375,
416
+ "learning_rate": 7.360560401432401e-05,
417
+ "loss": 0.5025547504425049,
418
+ "mean_token_accuracy": 0.8946620702743531,
419
+ "num_tokens": 798105.0,
420
+ "step": 410
421
+ },
422
+ {
423
+ "entropy": 0.5343898519873619,
424
+ "epoch": 1.8666666666666667,
425
+ "grad_norm": 0.27734375,
426
+ "learning_rate": 6.891191187907455e-05,
427
+ "loss": 0.49821176528930666,
428
+ "mean_token_accuracy": 0.8958980679512024,
429
+ "num_tokens": 817238.0,
430
+ "step": 420
431
+ },
432
+ {
433
+ "entropy": 0.5404686734080315,
434
+ "epoch": 1.911111111111111,
435
+ "grad_norm": 0.296875,
436
+ "learning_rate": 6.429288022172068e-05,
437
+ "loss": 0.5022043704986572,
438
+ "mean_token_accuracy": 0.8958747044205666,
439
+ "num_tokens": 836646.0,
440
+ "step": 430
441
+ },
442
+ {
443
+ "entropy": 0.5426375091075897,
444
+ "epoch": 1.9555555555555557,
445
+ "grad_norm": 0.279296875,
446
+ "learning_rate": 5.9759602008468996e-05,
447
+ "loss": 0.5035938262939453,
448
+ "mean_token_accuracy": 0.8944751426577568,
449
+ "num_tokens": 856282.0,
450
+ "step": 440
451
+ },
452
+ {
453
+ "entropy": 0.5425564736127854,
454
+ "epoch": 2.0,
455
+ "grad_norm": 0.263671875,
456
+ "learning_rate": 5.532296426191539e-05,
457
+ "loss": 0.5004886627197266,
458
+ "mean_token_accuracy": 0.8947419315576554,
459
+ "num_tokens": 875656.0,
460
+ "step": 450
461
+ },
462
+ {
463
+ "entropy": 0.526330380141735,
464
+ "epoch": 2.0444444444444443,
465
+ "grad_norm": 0.279296875,
466
+ "learning_rate": 5.0993621915007785e-05,
467
+ "loss": 0.44887552261352537,
468
+ "mean_token_accuracy": 0.9024429768323898,
469
+ "num_tokens": 894973.0,
470
+ "step": 460
471
+ },
472
+ {
473
+ "entropy": 0.4774711772799492,
474
+ "epoch": 2.088888888888889,
475
+ "grad_norm": 0.365234375,
476
+ "learning_rate": 4.678197222239035e-05,
477
+ "loss": 0.44640169143676756,
478
+ "mean_token_accuracy": 0.9045588001608849,
479
+ "num_tokens": 914579.0,
480
+ "step": 470
481
+ },
482
+ {
483
+ "entropy": 0.4789727419614792,
484
+ "epoch": 2.1333333333333333,
485
+ "grad_norm": 0.3828125,
486
+ "learning_rate": 4.269812979058235e-05,
487
+ "loss": 0.44276885986328124,
488
+ "mean_token_accuracy": 0.9066318318247795,
489
+ "num_tokens": 933807.0,
490
+ "step": 480
491
+ },
492
+ {
493
+ "entropy": 0.47537712305784224,
494
+ "epoch": 2.1777777777777776,
495
+ "grad_norm": 0.359375,
496
+ "learning_rate": 3.875190228695862e-05,
497
+ "loss": 0.4367781639099121,
498
+ "mean_token_accuracy": 0.9070042923092843,
499
+ "num_tokens": 953241.0,
500
+ "step": 490
501
+ },
502
+ {
503
+ "entropy": 0.4877164676785469,
504
+ "epoch": 2.2222222222222223,
505
+ "grad_norm": 0.328125,
506
+ "learning_rate": 3.4952766885868346e-05,
507
+ "loss": 0.4422135353088379,
508
+ "mean_token_accuracy": 0.905805604159832,
509
+ "num_tokens": 972758.0,
510
+ "step": 500
511
+ },
512
+ {
513
+ "entropy": 0.4915403999388218,
514
+ "epoch": 2.2666666666666666,
515
+ "grad_norm": 0.376953125,
516
+ "learning_rate": 3.130984750845885e-05,
517
+ "loss": 0.4357293128967285,
518
+ "mean_token_accuracy": 0.9067289993166924,
519
+ "num_tokens": 992534.0,
520
+ "step": 510
521
+ },
522
+ {
523
+ "entropy": 0.4957615494728088,
524
+ "epoch": 2.311111111111111,
525
+ "grad_norm": 0.33984375,
526
+ "learning_rate": 2.7831892910864434e-05,
527
+ "loss": 0.44167218208312986,
528
+ "mean_token_accuracy": 0.9066730305552483,
529
+ "num_tokens": 1012003.0,
530
+ "step": 520
531
+ },
532
+ {
533
+ "entropy": 0.48269262462854384,
534
+ "epoch": 2.3555555555555556,
535
+ "grad_norm": 0.375,
536
+ "learning_rate": 2.4527255673383565e-05,
537
+ "loss": 0.437894344329834,
538
+ "mean_token_accuracy": 0.904815036058426,
539
+ "num_tokens": 1031505.0,
540
+ "step": 530
541
+ },
542
+ {
543
+ "entropy": 0.47741171419620515,
544
+ "epoch": 2.4,
545
+ "grad_norm": 0.4140625,
546
+ "learning_rate": 2.140387214110322e-05,
547
+ "loss": 0.4400351047515869,
548
+ "mean_token_accuracy": 0.9068154886364936,
549
+ "num_tokens": 1050914.0,
550
+ "step": 540
551
+ },
552
+ {
553
+ "entropy": 0.4727069653570652,
554
+ "epoch": 2.4444444444444446,
555
+ "grad_norm": 0.35546875,
556
+ "learning_rate": 1.846924336414474e-05,
557
+ "loss": 0.43111190795898435,
558
+ "mean_token_accuracy": 0.9075597256422043,
559
+ "num_tokens": 1070218.0,
560
+ "step": 550
561
+ },
562
+ {
563
+ "entropy": 0.4883471392095089,
564
+ "epoch": 2.488888888888889,
565
+ "grad_norm": 0.314453125,
566
+ "learning_rate": 1.5730417083304573e-05,
567
+ "loss": 0.4493571758270264,
568
+ "mean_token_accuracy": 0.9073263511061669,
569
+ "num_tokens": 1089667.0,
570
+ "step": 560
571
+ },
572
+ {
573
+ "entropy": 0.4878625735640526,
574
+ "epoch": 2.533333333333333,
575
+ "grad_norm": 0.375,
576
+ "learning_rate": 1.3193970804352952e-05,
577
+ "loss": 0.44793548583984377,
578
+ "mean_token_accuracy": 0.9050837978720665,
579
+ "num_tokens": 1109201.0,
580
+ "step": 570
581
+ },
582
+ {
583
+ "entropy": 0.48537297546863556,
584
+ "epoch": 2.5777777777777775,
585
+ "grad_norm": 0.337890625,
586
+ "learning_rate": 1.08659960016387e-05,
587
+ "loss": 0.4368610382080078,
588
+ "mean_token_accuracy": 0.9097044110298157,
589
+ "num_tokens": 1128649.0,
590
+ "step": 580
591
+ },
592
+ {
593
+ "entropy": 0.4859867602586746,
594
+ "epoch": 2.6222222222222222,
595
+ "grad_norm": 0.328125,
596
+ "learning_rate": 8.75208348893667e-06,
597
+ "loss": 0.4249687194824219,
598
+ "mean_token_accuracy": 0.9087634190917016,
599
+ "num_tokens": 1148353.0,
600
+ "step": 590
601
+ },
602
+ {
603
+ "entropy": 0.4833585321903229,
604
+ "epoch": 2.6666666666666665,
605
+ "grad_norm": 0.3828125,
606
+ "learning_rate": 6.857309992670624e-06,
607
+ "loss": 0.4455591678619385,
608
+ "mean_token_accuracy": 0.9075253725051879,
609
+ "num_tokens": 1167638.0,
610
+ "step": 600
611
+ },
612
+ {
613
+ "entropy": 0.4794360339641571,
614
+ "epoch": 2.7111111111111112,
615
+ "grad_norm": 0.3203125,
616
+ "learning_rate": 5.186225959757207e-06,
617
+ "loss": 0.43627562522888186,
618
+ "mean_token_accuracy": 0.907138803601265,
619
+ "num_tokens": 1186894.0,
620
+ "step": 610
621
+ },
622
+ {
623
+ "entropy": 0.4800324112176895,
624
+ "epoch": 2.7555555555555555,
625
+ "grad_norm": 0.365234375,
626
+ "learning_rate": 3.7428446293514386e-06,
627
+ "loss": 0.42800016403198243,
628
+ "mean_token_accuracy": 0.9089159920811654,
629
+ "num_tokens": 1206224.0,
630
+ "step": 620
631
+ },
632
+ {
633
+ "entropy": 0.48999542444944383,
634
+ "epoch": 2.8,
635
+ "grad_norm": 0.31640625,
636
+ "learning_rate": 2.5306323947385746e-06,
637
+ "loss": 0.44663453102111816,
638
+ "mean_token_accuracy": 0.9060632780194282,
639
+ "num_tokens": 1225569.0,
640
+ "step": 630
641
+ },
642
+ {
643
+ "entropy": 0.4874999448657036,
644
+ "epoch": 2.8444444444444446,
645
+ "grad_norm": 0.326171875,
646
+ "learning_rate": 1.5525004785192143e-06,
647
+ "loss": 0.44841961860656737,
648
+ "mean_token_accuracy": 0.9045136958360672,
649
+ "num_tokens": 1244959.0,
650
+ "step": 640
651
+ },
652
+ {
653
+ "entropy": 0.49936808943748473,
654
+ "epoch": 2.888888888888889,
655
+ "grad_norm": 0.4296875,
656
+ "learning_rate": 8.107979410802769e-07,
657
+ "loss": 0.4453989028930664,
658
+ "mean_token_accuracy": 0.9050952970981598,
659
+ "num_tokens": 1264904.0,
660
+ "step": 650
661
+ },
662
+ {
663
+ "entropy": 0.49325661584734914,
664
+ "epoch": 2.9333333333333336,
665
+ "grad_norm": 0.373046875,
666
+ "learning_rate": 3.0730603914255193e-07,
667
+ "loss": 0.44756379127502444,
668
+ "mean_token_accuracy": 0.9043820068240166,
669
+ "num_tokens": 1284311.0,
670
+ "step": 660
671
+ },
672
+ {
673
+ "entropy": 0.5068465434014797,
674
+ "epoch": 2.977777777777778,
675
+ "grad_norm": 0.373046875,
676
+ "learning_rate": 4.323394793315228e-08,
677
+ "loss": 0.4593667030334473,
678
+ "mean_token_accuracy": 0.9011617928743363,
679
+ "num_tokens": 1303914.0,
680
+ "step": 670
681
+ },
682
+ {
683
+ "epoch": 3.0,
684
+ "eval_entropy": 0.5207519015669823,
685
+ "eval_loss": 0.5975516438484192,
686
+ "eval_mean_token_accuracy": 0.8822531878948212,
687
+ "eval_num_tokens": 1313484.0,
688
+ "eval_runtime": 57.0793,
689
+ "eval_samples_per_second": 7.008,
690
+ "eval_steps_per_second": 1.752,
691
+ "step": 675
692
+ },
693
+ {
694
+ "epoch": 3.0,
695
+ "step": 675,
696
+ "total_flos": 2.3441226939026637e+17,
697
+ "train_loss": 0.5646523337894016,
698
+ "train_runtime": 6392.6844,
699
+ "train_samples_per_second": 1.689,
700
+ "train_steps_per_second": 0.106
701
+ }
702
+ ],
703
+ "logging_steps": 10,
704
+ "max_steps": 675,
705
+ "num_input_tokens_seen": 0,
706
+ "num_train_epochs": 3,
707
+ "save_steps": 9999,
708
+ "stateful_callbacks": {
709
+ "TrainerControl": {
710
+ "args": {
711
+ "should_epoch_stop": false,
712
+ "should_evaluate": false,
713
+ "should_log": false,
714
+ "should_save": true,
715
+ "should_training_stop": true
716
+ },
717
+ "attributes": {}
718
+ }
719
+ },
720
+ "total_flos": 2.3441226939026637e+17,
721
+ "train_batch_size": 1,
722
+ "trial_name": null,
723
+ "trial_params": null
724
+ }