Text Generation
PEFT
Safetensors
English
lora
qlora
sft
trl
text-to-sql
sql
conversational
Eval Results (legacy)
Instructions to use SASVAAI/GLM-4.7-Flash-sql-create-context with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use SASVAAI/GLM-4.7-Flash-sql-create-context with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("zai-org/GLM-4.7-Flash") model = PeftModel.from_pretrained(base_model, "SASVAAI/GLM-4.7-Flash-sql-create-context") - Notebooks
- Google Colab
- Kaggle
Add LoRA adapter, model card, and eval artifacts
#1
by Poojanbuselvan - opened
- .gitattributes +1 -0
- LICENSE +33 -0
- README.md +410 -1
- adapter_config.json +47 -0
- adapter_model.safetensors +3 -0
- all_results.json +7 -0
- chat_template.jinja +86 -0
- eval_results.json +8 -0
- predictions.jsonl +0 -0
- tokenizer.json +3 -0
- tokenizer_config.json +13 -0
- train_results.json +7 -0
- trainer_state.json +724 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
LICENSE
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
These LoRA adapter weights are released under the MIT License, matching the
|
| 2 |
+
license declared by the base model, zai-org/GLM-4.7-Flash. Note that the base
|
| 3 |
+
model's repository declares `mit` in its metadata but does not itself ship a
|
| 4 |
+
LICENSE file; the standard MIT text is reproduced below and applies to this
|
| 5 |
+
adapter.
|
| 6 |
+
|
| 7 |
+
The training data is derived from b-mc2/sql-create-context, which is licensed
|
| 8 |
+
CC-BY-4.0. That license is not superseded by this one — downstream use of these
|
| 9 |
+
weights should carry its attribution requirement.
|
| 10 |
+
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
MIT License
|
| 14 |
+
|
| 15 |
+
Copyright (c) 2026 SASVA AI Model Cognition Labs (MCL)
|
| 16 |
+
|
| 17 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 18 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 19 |
+
in the Software without restriction, including without limitation the rights
|
| 20 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 21 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 22 |
+
furnished to do so, subject to the following conditions:
|
| 23 |
+
|
| 24 |
+
The above copyright notice and this permission notice shall be included in all
|
| 25 |
+
copies or substantial portions of the Software.
|
| 26 |
+
|
| 27 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 28 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 29 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 30 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 31 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 32 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 33 |
+
SOFTWARE.
|
README.md
CHANGED
|
@@ -1,3 +1,412 @@
|
|
| 1 |
---
|
| 2 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
# ---- Identity -------------------------------------------------------------
|
| 3 |
+
base_model: zai-org/GLM-4.7-Flash
|
| 4 |
+
base_model_relation: adapter
|
| 5 |
+
library_name: peft
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
language:
|
| 8 |
+
- en
|
| 9 |
+
license: mit # verified via the Hub API: zai-org/GLM-4.7-Flash reports `mit`.
|
| 10 |
+
# The training data is CC-BY-4.0 — see Licence below.
|
| 11 |
+
|
| 12 |
+
# ---- Discovery ------------------------------------------------------------
|
| 13 |
+
tags:
|
| 14 |
+
- lora
|
| 15 |
+
- qlora
|
| 16 |
+
- peft
|
| 17 |
+
- sft
|
| 18 |
+
- trl
|
| 19 |
+
- text-to-sql
|
| 20 |
+
- sql
|
| 21 |
+
|
| 22 |
+
datasets:
|
| 23 |
+
- b-mc2/sql-create-context
|
| 24 |
+
|
| 25 |
+
metrics:
|
| 26 |
+
- exact_match
|
| 27 |
+
- bleu
|
| 28 |
+
- rouge
|
| 29 |
+
|
| 30 |
+
# ---- Structured evaluation ------------------------------------------------
|
| 31 |
+
model-index:
|
| 32 |
+
- name: glm-4.7-flash-sql-create-context-lora
|
| 33 |
+
results:
|
| 34 |
+
- task:
|
| 35 |
+
type: text-generation
|
| 36 |
+
name: Text-to-SQL (natural language + CREATE TABLE schema -> SQL)
|
| 37 |
+
dataset:
|
| 38 |
+
type: sql-create-context-val
|
| 39 |
+
name: sql-create-context derived validation split (400 pairs)
|
| 40 |
+
split: validation
|
| 41 |
+
metrics:
|
| 42 |
+
- type: exact_match
|
| 43 |
+
name: Exact match
|
| 44 |
+
value: 0.805
|
| 45 |
+
- type: bleu
|
| 46 |
+
name: BLEU
|
| 47 |
+
value: 0.940082
|
| 48 |
+
- type: rouge
|
| 49 |
+
name: ROUGE-L
|
| 50 |
+
value: 0.986055
|
| 51 |
+
args:
|
| 52 |
+
rouge_type: rougeL
|
| 53 |
---
|
| 54 |
+
|
| 55 |
+
# GLM-4.7-Flash Text-to-SQL (LoRA)
|
| 56 |
+
|
| 57 |
+
Given a natural-language question and a `CREATE TABLE` schema, emits exactly one
|
| 58 |
+
SQL query answering that question against that schema. For natural-language
|
| 59 |
+
query interfaces over a known relational schema.
|
| 60 |
+
|
| 61 |
+
This is a **LoRA adapter for**
|
| 62 |
+
[zai-org/GLM-4.7-Flash](https://huggingface.co/zai-org/GLM-4.7-Flash), trained
|
| 63 |
+
with **QLoRA (4-bit NF4 base, bf16 compute)** via
|
| 64 |
+
[TRL](https://github.com/huggingface/trl) SFT.
|
| 65 |
+
|
| 66 |
+
## Model details
|
| 67 |
+
|
| 68 |
+
| | |
|
| 69 |
+
|---|---|
|
| 70 |
+
| Developed by | SASVA AI Model Cognition Labs (MCL) Team |
|
| 71 |
+
| Base model | [`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash) |
|
| 72 |
+
| Base parameters | 30B total / 3B active (MoE) — 29,943,396,864 in the merged bf16 build |
|
| 73 |
+
| Architecture family | `glm4_moe_lite` (`Glm4MoeLiteForCausalLM`), 47 layers, hidden size 2048, vocab 154,880 |
|
| 74 |
+
| Adaptation | LoRA (`r=64`, `alpha=128`, `dropout=0.05`) |
|
| 75 |
+
| Trainable modules | `q_a_proj`, `q_b_proj`, `kv_a_proj_with_mqa`, `kv_b_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` |
|
| 76 |
+
| Training method | `qlora` (4-bit NF4, double quant, bf16 compute) |
|
| 77 |
+
| Refinement | none |
|
| 78 |
+
| Language | English (questions) / SQL (outputs) |
|
| 79 |
+
| License | MIT (inherited from the base model) |
|
| 80 |
+
|
| 81 |
+
Trainable parameters: **118,140,928** — **0.3945%** of the base. The adapter file
|
| 82 |
+
is 472,671,000 bytes (752 fp32 tensors: a `lora_A` + `lora_B` pair for each of
|
| 83 |
+
the 8 target modules across all 47 layers).
|
| 84 |
+
|
| 85 |
+
GLM-4.7-Flash uses Multi-head Latent Attention, so the attention target modules
|
| 86 |
+
are the MLA projections (`q_a_proj`/`q_b_proj`/`kv_a_proj_with_mqa`/`kv_b_proj`),
|
| 87 |
+
not `q_proj`/`k_proj`/`v_proj`. Targeting the conventional names would silently
|
| 88 |
+
adapt nothing.
|
| 89 |
+
|
| 90 |
+
## Intended use
|
| 91 |
+
|
| 92 |
+
**Direct use.** Translate one English question plus one `CREATE TABLE` schema
|
| 93 |
+
into one SQL query. The model was trained on a specific prompt shape and that
|
| 94 |
+
shape is part of the contract:
|
| 95 |
+
|
| 96 |
+
- System prompt (verbatim): *"You are a text-to-SQL engine. Given a
|
| 97 |
+
natural-language question and a CREATE TABLE schema, output exactly one SQL
|
| 98 |
+
query that answers the question against the provided schema. Output only the
|
| 99 |
+
raw SQL query on a single line with no explanation, no markdown formatting,
|
| 100 |
+
and no additional text."*
|
| 101 |
+
- User turn: the question, a blank line, then the schema inside a fenced code
|
| 102 |
+
block.
|
| 103 |
+
- Applied through the tokenizer's chat template (`chat_template.jinja`, shipped
|
| 104 |
+
in this repo) with `enable_thinking=False`. Do not concatenate strings by hand.
|
| 105 |
+
- The query is the **first line** of the generation; discard anything after it.
|
| 106 |
+
|
| 107 |
+
**Out of scope.**
|
| 108 |
+
- **Not validated against a live database.** The model is scored on string
|
| 109 |
+
similarity to a reference query, never on execution. A syntactically perfect
|
| 110 |
+
query can still be semantically wrong. Parse and, where you can, dry-run
|
| 111 |
+
against the real schema before trusting output.
|
| 112 |
+
- **Never interpolate output into a privileged connection.** Treat generated SQL
|
| 113 |
+
as untrusted input: run it read-only, with least privilege, on a connection
|
| 114 |
+
that cannot write or drop.
|
| 115 |
+
- Multi-table joins, CTEs, window functions, subqueries, and DDL/DML are largely
|
| 116 |
+
out of distribution — the training data is dominated by single-table
|
| 117 |
+
`SELECT`s. Measured join accuracy is poor (see Limitations).
|
| 118 |
+
- Dialect is not controllable. The model reproduces the source corpus's
|
| 119 |
+
conventions (double-quoted string literals, lower-cased comparison values),
|
| 120 |
+
which are not portable to every engine.
|
| 121 |
+
- Not a general-purpose assistant. It emits a bare query, never prose.
|
| 122 |
+
|
| 123 |
+
## How to get started
|
| 124 |
+
|
| 125 |
+
```python
|
| 126 |
+
import torch
|
| 127 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
|
| 128 |
+
from peft import PeftModel
|
| 129 |
+
|
| 130 |
+
BASE = "zai-org/GLM-4.7-Flash"
|
| 131 |
+
ADAPTER = "SASVAAI/GLM-4.7-Flash-sql-create-context"
|
| 132 |
+
|
| 133 |
+
# 4-bit NF4 matches the numerics the adapter was trained against. A bf16 base
|
| 134 |
+
# also works and scores the same (see Merged-weights equivalence) but needs
|
| 135 |
+
# ~60 GB rather than ~22 GB.
|
| 136 |
+
bnb = BitsAndBytesConfig(
|
| 137 |
+
load_in_4bit=True,
|
| 138 |
+
bnb_4bit_quant_type="nf4",
|
| 139 |
+
bnb_4bit_compute_dtype=torch.bfloat16,
|
| 140 |
+
bnb_4bit_use_double_quant=True,
|
| 141 |
+
)
|
| 142 |
+
|
| 143 |
+
tokenizer = AutoTokenizer.from_pretrained(ADAPTER)
|
| 144 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 145 |
+
BASE, quantization_config=bnb, dtype=torch.bfloat16, device_map="auto"
|
| 146 |
+
)
|
| 147 |
+
model = PeftModel.from_pretrained(model, ADAPTER)
|
| 148 |
+
model.eval()
|
| 149 |
+
|
| 150 |
+
SYSTEM = (
|
| 151 |
+
"You are a text-to-SQL engine. Given a natural-language question and a "
|
| 152 |
+
"CREATE TABLE schema, output exactly one SQL query that answers the question "
|
| 153 |
+
"against the provided schema. Output only the raw SQL query on a single line "
|
| 154 |
+
"with no explanation, no markdown formatting, and no additional text."
|
| 155 |
+
)
|
| 156 |
+
|
| 157 |
+
question = "Which kingdom has Suin as its capital?"
|
| 158 |
+
schema = "CREATE TABLE table_name_65 (name_of_kingdom VARCHAR, capital VARCHAR)"
|
| 159 |
+
|
| 160 |
+
messages = [
|
| 161 |
+
{"role": "system", "content": SYSTEM},
|
| 162 |
+
{"role": "user", "content": f"{question}\n\n```\n{schema}\n```"},
|
| 163 |
+
]
|
| 164 |
+
prompt = tokenizer.apply_chat_template(
|
| 165 |
+
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
|
| 166 |
+
)
|
| 167 |
+
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
| 168 |
+
|
| 169 |
+
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
|
| 170 |
+
text = tokenizer.decode(out[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
|
| 171 |
+
print(text.strip().splitlines()[0])
|
| 172 |
+
# -> SELECT name_of_kingdom FROM table_name_65 WHERE capital = "suin"
|
| 173 |
+
```
|
| 174 |
+
|
| 175 |
+
The base model is ~59 GB in bfloat16, or ~22 GB per GPU under 4-bit NF4.
|
| 176 |
+
|
| 177 |
+
> Decoding matters. This model was evaluated with greedy decoding
|
| 178 |
+
> (`do_sample=False`, `max_new_tokens=128`). Sampling will not reproduce the
|
| 179 |
+
> reported numbers.
|
| 180 |
+
|
| 181 |
+
**Serving note.** These adapter weights do **not** load as a vLLM LoRA on this
|
| 182 |
+
architecture — vLLM's MLA path asserts inside
|
| 183 |
+
`DeepSeekV2FusedQkvAProjLinear` because `q_a_proj` and `kv_a_proj_with_mqa` are
|
| 184 |
+
fused into one module that a LoRA cannot be attached to. A merged build of these
|
| 185 |
+
weights serves under vLLM without complaint. For vLLM deployment, merge first
|
| 186 |
+
(`peft.merge_and_unload()`).
|
| 187 |
+
|
| 188 |
+
## Training details
|
| 189 |
+
|
| 190 |
+
**Data.** A 4,000-pair subset of
|
| 191 |
+
[`b-mc2/sql-create-context`](https://huggingface.co/datasets/b-mc2/sql-create-context)
|
| 192 |
+
(78,577 pairs, itself derived from WikiSQL and Spider), split 90/10 by this
|
| 193 |
+
project's data-generation stage. Each record is
|
| 194 |
+
`{"instruction": <question>, "input": <CREATE TABLE ...>, "output": <SQL>}`.
|
| 195 |
+
The subset selection is a generated artifact, not a published split — the
|
| 196 |
+
`id`/`instruction`/`input`/`gold` tuples in `predictions.jsonl` are the
|
| 197 |
+
authoritative record of what was evaluated.
|
| 198 |
+
|
| 199 |
+
| | |
|
| 200 |
+
|---|---|
|
| 201 |
+
| Train samples | 3,600 |
|
| 202 |
+
| Validation samples | 400 |
|
| 203 |
+
| Prompt format | chat template + system prompt (see Intended use) |
|
| 204 |
+
|
| 205 |
+
### Method
|
| 206 |
+
|
| 207 |
+
| | |
|
| 208 |
+
|---|---|
|
| 209 |
+
| SFT method | `qlora` |
|
| 210 |
+
| Base quantisation during training | 4-bit NF4, double quant, bf16 compute |
|
| 211 |
+
| Refinement stage | none |
|
| 212 |
+
| Hardware | 4x NVIDIA H200 (141 GB), `torchrun --nproc_per_node=4` |
|
| 213 |
+
|
| 214 |
+
`qlora` is one of three methods considered for this model, alongside 8-bit LoRA
|
| 215 |
+
and attention-only 4-bit LoRA. All three were tried; `qlora` scored highest
|
| 216 |
+
every time it was compared against the other two on the same data.
|
| 217 |
+
|
| 218 |
+
No refinement stage ran; the published weights are the SFT adapter.
|
| 219 |
+
|
| 220 |
+
### Final hyperparameters
|
| 221 |
+
|
| 222 |
+
| Hyperparameter | Value | Source |
|
| 223 |
+
|---|---|---|
|
| 224 |
+
| `learning_rate` | 0.0002 | `training_args.bin` |
|
| 225 |
+
| `lr_scheduler_type` | cosine | `training_args.bin` |
|
| 226 |
+
| `num_train_epochs` | 3 | `training_args.bin` |
|
| 227 |
+
| `per_device_train_batch_size` | 1 | `training_args.bin` |
|
| 228 |
+
| `gradient_accumulation_steps` | 4 | `training_args.bin` |
|
| 229 |
+
| `max_length` | 2048 | `training_args.bin` |
|
| 230 |
+
| `warmup_ratio` | 0.05 | `training_args.bin` |
|
| 231 |
+
| `weight_decay` | 0.01 | `training_args.bin` |
|
| 232 |
+
| `optim` / `max_grad_norm` / `seed` | `adamw_torch` / 1.0 / 42 | `training_args.bin` |
|
| 233 |
+
| `bf16` / `gradient_checkpointing` | true / true (`use_reentrant=False`) | `training_args.bin` |
|
| 234 |
+
| `neftune_noise_alpha` / `packing` | None / false | `training_args.bin` |
|
| 235 |
+
| `lora_r` / `lora_alpha` / `lora_dropout` | 64 / 128 / 0.05 | `adapter_config.json` |
|
| 236 |
+
| `target_modules` | the 8 listed in Model details | `adapter_config.json` |
|
| 237 |
+
|
| 238 |
+
**Effective batch size: 16** (`1 x 4 x 4`). Optimizer steps: 675.
|
| 239 |
+
|
| 240 |
+
KD parameters are omitted deliberately — this is a `qlora` run, not
|
| 241 |
+
`bf16_lora_kd`, so `KD_ALPHA`/`KD_BETA`/`KD_TEMPERATURE` carry inert defaults
|
| 242 |
+
that would imply distillation that did not happen.
|
| 243 |
+
|
| 244 |
+
This configuration is **not a unique optimum**. Other configurations reached the
|
| 245 |
+
same `exact_match`; this one was published for being the simplest and cheapest
|
| 246 |
+
of them — fewest epochs, shortest training time — and because it scored
|
| 247 |
+
marginally higher on BLEU and ROUGE-L. Treat the values as a good working point,
|
| 248 |
+
not a tuned maximum.
|
| 249 |
+
|
| 250 |
+
**Observed training metrics.**
|
| 251 |
+
|
| 252 |
+
| | |
|
| 253 |
+
|---|---|
|
| 254 |
+
| Final train loss | 0.4594 |
|
| 255 |
+
| Mean train loss | 0.5646523337894016 |
|
| 256 |
+
| Train runtime | 6392.6844s |
|
| 257 |
+
| Total FLOPs | 2.3441226939026637e+17 |
|
| 258 |
+
| Throughput | 1.689 samples/s, 0.106 steps/s |
|
| 259 |
+
|
| 260 |
+
No eval loss was computed during training; the loop scores on generation, not
|
| 261 |
+
perplexity. Loss falls from 2.3882 at step 10 to 1.1077 at step 20 and 0.5940 by
|
| 262 |
+
step 170, then improves slowly to ~0.45 by step 660. The task is essentially
|
| 263 |
+
learned within the first quarter of epoch 1; epochs 2 and 3 together buy roughly
|
| 264 |
+
0.14 of training loss.
|
| 265 |
+
|
| 266 |
+
## Evaluation
|
| 267 |
+
|
| 268 |
+
**Protocol.** All 400 validation pairs, no sampling. Predictions generated
|
| 269 |
+
greedily (`do_sample=False`, `max_new_tokens=128`) through the same chat template
|
| 270 |
+
used in training, against a 4-bit NF4 base to match training numerics. The
|
| 271 |
+
predicted query is the **first line** of the generation, stripped. No constrained
|
| 272 |
+
decoding and no SQL grammar were applied. Exact match is byte equality against
|
| 273 |
+
the reference query; BLEU and ROUGE-L are computed over the same strings.
|
| 274 |
+
|
| 275 |
+
| Metric | Value |
|
| 276 |
+
|---|---|
|
| 277 |
+
| Exact match | 0.805000 (322 / 400) |
|
| 278 |
+
| BLEU | 0.940082 |
|
| 279 |
+
| ROUGE-L | 0.986055 |
|
| 280 |
+
| Samples | 400 |
|
| 281 |
+
|
| 282 |
+
**Baseline for comparison.** **Not measured.** The untuned
|
| 283 |
+
`zai-org/GLM-4.7-Flash` was never scored on this split, so these numbers
|
| 284 |
+
quantify the fine-tuned model's performance but do not establish how much of it
|
| 285 |
+
the fine-tuning is responsible for.
|
| 286 |
+
|
| 287 |
+
**This is a validation split, not a held-out test set.** Hyperparameters were
|
| 288 |
+
selected against it, so expect optimistic bias. A clean estimate needs a third
|
| 289 |
+
split that was never used for selection.
|
| 290 |
+
|
| 291 |
+
## Limitations and bias
|
| 292 |
+
|
| 293 |
+
**Exact match understates the model; BLEU and ROUGE-L overstate it.** The
|
| 294 |
+
0.805 / 0.940 / 0.986 spread is the story of this model. Exact match is byte
|
| 295 |
+
equality, so a semantically identical query loses the point on quoting or a
|
| 296 |
+
missing `DISTINCT`:
|
| 297 |
+
|
| 298 |
+
```
|
| 299 |
+
question: Find the states where have some college students in tryout and their decisions are yes.
|
| 300 |
+
schema: CREATE TABLE tryout (cName VARCHAR, decision VARCHAR);
|
| 301 |
+
CREATE TABLE college (state VARCHAR, cName VARCHAR)
|
| 302 |
+
gold: SELECT DISTINCT T1.state FROM college AS T1 JOIN tryout AS T2 ON T1.cName = T2.cName WHERE T2.decision = 'yes'
|
| 303 |
+
pred: SELECT T1.state FROM college AS T1 JOIN tryout AS T2 ON T1.cName = T2.cName WHERE T2.decision = "yes"
|
| 304 |
+
```
|
| 305 |
+
|
| 306 |
+
Relaxing to case- and whitespace-insensitive comparison moves exact match from
|
| 307 |
+
0.805 to **0.8125** (325/400) — so only 3 of the 78 misses are pure formatting.
|
| 308 |
+
The other 75 are real semantic or structural errors. ROUGE-L at 0.986 mostly
|
| 309 |
+
measures that both strings are short SQL over the same table names; it is not
|
| 310 |
+
evidence of correctness.
|
| 311 |
+
|
| 312 |
+
**Output is never degenerate.** All 400 golds and all 400 predictions begin with
|
| 313 |
+
`SELECT`; the model never emitted prose, markdown, or an empty string. Format
|
| 314 |
+
compliance is not the failure mode.
|
| 315 |
+
|
| 316 |
+
**Joins are the failure mode.** Miss rate by gold-query feature:
|
| 317 |
+
|
| 318 |
+
| Gold query contains | n | misses | miss rate |
|
| 319 |
+
|---|---|---|---|
|
| 320 |
+
| `JOIN` | 13 | 10 | **76.9%** |
|
| 321 |
+
| `ORDER BY` | 3 | 1 | 33.3% |
|
| 322 |
+
| aggregate (`COUNT`/`SUM`/`AVG`/`MIN`/`MAX`) | 138 | 37 | 26.8% |
|
| 323 |
+
| `GROUP BY` | 8 | 2 | 25.0% |
|
| 324 |
+
| multi-predicate `WHERE` (`AND`/`OR`) | 123 | 27 | 22.0% |
|
| 325 |
+
| single-predicate `SELECT..WHERE` | 181 | 22 | 12.2% |
|
| 326 |
+
|
| 327 |
+
The model is reliable on the shape it saw constantly (one table, one predicate)
|
| 328 |
+
and unreliable on the shape it barely saw. **Do not deploy this on a
|
| 329 |
+
multi-table schema.** Note the join, `ORDER BY`, and `GROUP BY` rows rest on
|
| 330 |
+
13, 3, and 8 examples respectively — read them as a strong warning, not a precise
|
| 331 |
+
rate.
|
| 332 |
+
|
| 333 |
+
Complexity tracks length: correct predictions have a mean gold length of 10.6
|
| 334 |
+
tokens, misses 12.8.
|
| 335 |
+
|
| 336 |
+
**Dialect is baked in.** The model emits the source corpus's double-quoted
|
| 337 |
+
string literals and lower-cases comparison values. On engines where `"x"` is an
|
| 338 |
+
identifier rather than a string (PostgreSQL, ANSI mode), output will not run
|
| 339 |
+
unmodified.
|
| 340 |
+
|
| 341 |
+
**No execution or injection safety.** Correctness is measured only as string
|
| 342 |
+
similarity to a reference. Nothing here prevents a generated query from being
|
| 343 |
+
expensive, wrong, or destructive against a real database.
|
| 344 |
+
|
| 345 |
+
**Inherits all biases and limitations of the base model.** This adapter changes
|
| 346 |
+
0.3945% of the parameters and was not evaluated for social bias, safety, or
|
| 347 |
+
fairness.
|
| 348 |
+
|
| 349 |
+
## Merged-weights equivalence
|
| 350 |
+
|
| 351 |
+
A merged build of these weights (base + adapter folded into one standalone bf16
|
| 352 |
+
model, `W + (alpha/r) * B @ A`) was evaluated on the identical split:
|
| 353 |
+
|
| 354 |
+
| Metric | Adapter (4-bit base) | Merged (bf16) | Delta |
|
| 355 |
+
|---|---|---|---|
|
| 356 |
+
| Exact match | 0.805000 | 0.805000 | 0.000000 |
|
| 357 |
+
| BLEU | 0.940082 | 0.942079 | +0.001997 |
|
| 358 |
+
| ROUGE-L | 0.986055 | 0.987663 | +0.001608 |
|
| 359 |
+
|
| 360 |
+
**39 of 400 predictions differ** (9.75%) — far more churn than a bf16-trained
|
| 361 |
+
adapter would show, because merging into an unquantised base genuinely changes
|
| 362 |
+
the numerics the adapter was fitted against. The changes cancel exactly:
|
| 363 |
+
**10 predictions flip correct→incorrect and 10 flip incorrect→correct**, leaving
|
| 364 |
+
exact match identical and BLEU/ROUGE-L marginally higher.
|
| 365 |
+
|
| 366 |
+
So a merged distribution is behaviourally equivalent in aggregate but not
|
| 367 |
+
prediction-for-prediction. Merging is also the practical route to vLLM serving
|
| 368 |
+
(see Serving note). MIT permits distributing derivative works, so publishing a
|
| 369 |
+
merged build is allowed.
|
| 370 |
+
|
| 371 |
+
## Environmental impact
|
| 372 |
+
|
| 373 |
+
| | |
|
| 374 |
+
|---|---|
|
| 375 |
+
| Hardware | 4x NVIDIA H200 (141 GB) |
|
| 376 |
+
| Training time | 106.5 minutes (6392.6844s) |
|
| 377 |
+
| Cloud provider / region | on-premise |
|
| 378 |
+
|
| 379 |
+
Covers the training of these published weights only. It excludes the wider
|
| 380 |
+
hyperparameter search that selected them, which cost substantially more.
|
| 381 |
+
|
| 382 |
+
## Framework versions
|
| 383 |
+
|
| 384 |
+
- PEFT 0.18.1
|
| 385 |
+
- TRL: 1.0.0
|
| 386 |
+
- Transformers: 5.7.0.dev0
|
| 387 |
+
- Pytorch: 2.5.1+cu121
|
| 388 |
+
- Datasets: 4.8.4
|
| 389 |
+
- Tokenizers: 0.22.2
|
| 390 |
+
- bitsandbytes: 0.49.2
|
| 391 |
+
|
| 392 |
+
`transformers` is a git-main build: GLM-4.7-Flash's `Glm4MoeLite` architecture is
|
| 393 |
+
not in the stable PyPI release.
|
| 394 |
+
|
| 395 |
+
## Licence
|
| 396 |
+
|
| 397 |
+
Adapter weights: **MIT**, inherited from
|
| 398 |
+
[`zai-org/GLM-4.7-Flash`](https://huggingface.co/zai-org/GLM-4.7-Flash)
|
| 399 |
+
(verified via the Hub API). Training data:
|
| 400 |
+
[`b-mc2/sql-create-context`](https://huggingface.co/datasets/b-mc2/sql-create-context),
|
| 401 |
+
licensed **CC-BY-4.0** — downstream use should carry that attribution.
|
| 402 |
+
|
| 403 |
+
## Citation
|
| 404 |
+
|
| 405 |
+
```bibtex
|
| 406 |
+
@misc{glm47flash_sql_create_context_lora_2026,
|
| 407 |
+
title = {GLM-4.7-Flash Text-to-SQL (LoRA)},
|
| 408 |
+
author = {{SASVA AI Model Cognition Labs (MCL) Team}},
|
| 409 |
+
year = {2026},
|
| 410 |
+
url = {https://huggingface.co/SASVAAI/GLM-4.7-Flash-sql-create-context}
|
| 411 |
+
}
|
| 412 |
+
```
|
adapter_config.json
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "zai-org/GLM-4.7-Flash",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 128,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0.05,
|
| 22 |
+
"megatron_config": null,
|
| 23 |
+
"megatron_core": "megatron.core",
|
| 24 |
+
"modules_to_save": null,
|
| 25 |
+
"peft_type": "LORA",
|
| 26 |
+
"peft_version": "0.18.1",
|
| 27 |
+
"qalora_group_size": 16,
|
| 28 |
+
"r": 64,
|
| 29 |
+
"rank_pattern": {},
|
| 30 |
+
"revision": null,
|
| 31 |
+
"target_modules": [
|
| 32 |
+
"up_proj",
|
| 33 |
+
"down_proj",
|
| 34 |
+
"gate_proj",
|
| 35 |
+
"q_a_proj",
|
| 36 |
+
"kv_a_proj_with_mqa",
|
| 37 |
+
"o_proj",
|
| 38 |
+
"kv_b_proj",
|
| 39 |
+
"q_b_proj"
|
| 40 |
+
],
|
| 41 |
+
"target_parameters": null,
|
| 42 |
+
"task_type": "CAUSAL_LM",
|
| 43 |
+
"trainable_token_indices": null,
|
| 44 |
+
"use_dora": false,
|
| 45 |
+
"use_qalora": false,
|
| 46 |
+
"use_rslora": false
|
| 47 |
+
}
|
adapter_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f7c4b6367c7ac51b2700e675946c6e0dffb7d037d46dadf8096fecd3757b9d41
|
| 3 |
+
size 472671000
|
all_results.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"total_flos": 2.3441226939026637e+17,
|
| 3 |
+
"train_loss": 0.5646523337894016,
|
| 4 |
+
"train_runtime": 6392.6844,
|
| 5 |
+
"train_samples_per_second": 1.689,
|
| 6 |
+
"train_steps_per_second": 0.106
|
| 7 |
+
}
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,86 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
[gMASK]<sop>
|
| 2 |
+
{%- if tools -%}
|
| 3 |
+
<|system|>
|
| 4 |
+
# Tools
|
| 5 |
+
|
| 6 |
+
You may call one or more functions to assist with the user query.
|
| 7 |
+
|
| 8 |
+
You are provided with function signatures within <tools></tools> XML tags:
|
| 9 |
+
<tools>
|
| 10 |
+
{% for tool in tools %}
|
| 11 |
+
{{ tool | tojson(ensure_ascii=False) }}
|
| 12 |
+
{% endfor %}
|
| 13 |
+
</tools>
|
| 14 |
+
|
| 15 |
+
For each function call, output the function name and arguments within the following XML format:
|
| 16 |
+
<tool_call>{function-name}<arg_key>{arg-key-1}</arg_key><arg_value>{arg-value-1}</arg_value><arg_key>{arg-key-2}</arg_key><arg_value>{arg-value-2}</arg_value>...</tool_call>{%- endif -%}
|
| 17 |
+
{%- macro visible_text(content) -%}
|
| 18 |
+
{%- if content is string -%}
|
| 19 |
+
{{- content }}
|
| 20 |
+
{%- elif content is iterable and content is not mapping -%}
|
| 21 |
+
{%- for item in content -%}
|
| 22 |
+
{%- if item is mapping and item.type == 'text' -%}
|
| 23 |
+
{{- item.text }}
|
| 24 |
+
{%- elif item is string -%}
|
| 25 |
+
{{- item }}
|
| 26 |
+
{%- endif -%}
|
| 27 |
+
{%- endfor -%}
|
| 28 |
+
{%- else -%}
|
| 29 |
+
{{- content }}
|
| 30 |
+
{%- endif -%}
|
| 31 |
+
{%- endmacro -%}
|
| 32 |
+
{%- set ns = namespace(last_user_index=-1) %}
|
| 33 |
+
{%- for m in messages %}
|
| 34 |
+
{%- if m.role == 'user' %}
|
| 35 |
+
{% set ns.last_user_index = loop.index0 -%}
|
| 36 |
+
{%- endif %}
|
| 37 |
+
{%- endfor %}
|
| 38 |
+
{% for m in messages %}
|
| 39 |
+
{%- if m.role == 'user' -%}<|user|>{{ visible_text(m.content) }}
|
| 40 |
+
{%- elif m.role == 'assistant' -%}
|
| 41 |
+
<|assistant|>
|
| 42 |
+
{%- set reasoning_content = '' %}
|
| 43 |
+
{%- set content = visible_text(m.content) %}
|
| 44 |
+
{%- if m.reasoning_content is string %}
|
| 45 |
+
{%- set reasoning_content = m.reasoning_content %}
|
| 46 |
+
{%- else %}
|
| 47 |
+
{%- if '</think>' in content %}
|
| 48 |
+
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 49 |
+
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
| 50 |
+
{%- endif %}
|
| 51 |
+
{%- endif %}
|
| 52 |
+
{%- if ((clear_thinking is defined and not clear_thinking) or loop.index0 > ns.last_user_index) and reasoning_content -%}
|
| 53 |
+
{{ '<think>' + reasoning_content.strip() + '</think>'}}
|
| 54 |
+
{%- else -%}
|
| 55 |
+
{{ '</think>' }}
|
| 56 |
+
{%- endif -%}
|
| 57 |
+
{%- if content.strip() -%}
|
| 58 |
+
{{ content.strip() }}
|
| 59 |
+
{%- endif -%}
|
| 60 |
+
{% if m.tool_calls %}
|
| 61 |
+
{% for tc in m.tool_calls %}
|
| 62 |
+
{%- if tc.function %}
|
| 63 |
+
{%- set tc = tc.function %}
|
| 64 |
+
{%- endif %}
|
| 65 |
+
{{- '<tool_call>' + tc.name -}}
|
| 66 |
+
{% set _args = tc.arguments %}{% for k, v in _args.items() %}<arg_key>{{ k }}</arg_key><arg_value>{{ v | tojson(ensure_ascii=False) if v is not string else v }}</arg_value>{% endfor %}</tool_call>{% endfor %}
|
| 67 |
+
{% endif %}
|
| 68 |
+
{%- elif m.role == 'tool' -%}
|
| 69 |
+
{%- if m.content is string -%}
|
| 70 |
+
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
| 71 |
+
{{- '<|observation|>' }}
|
| 72 |
+
{%- endif %}
|
| 73 |
+
{{- '<tool_response>' }}
|
| 74 |
+
{{- m.content }}
|
| 75 |
+
{{- '</tool_response>' }}
|
| 76 |
+
{%- else -%}
|
| 77 |
+
<|observation|>{% for tr in m.content %}
|
| 78 |
+
<tool_response>{{ tr.output if tr.output is defined else tr }}</tool_response>{% endfor -%}
|
| 79 |
+
{% endif -%}
|
| 80 |
+
{%- elif m.role == 'system' -%}
|
| 81 |
+
<|system|>{{ visible_text(m.content) }}
|
| 82 |
+
{%- endif -%}
|
| 83 |
+
{%- endfor -%}
|
| 84 |
+
{%- if add_generation_prompt -%}
|
| 85 |
+
<|assistant|>{{- '</think>' if (enable_thinking is defined and not enable_thinking) else '<think>' -}}
|
| 86 |
+
{%- endif -%}
|
eval_results.json
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"metrics": {
|
| 3 |
+
"exact_match": 0.805,
|
| 4 |
+
"bleu": 0.9400823275452048,
|
| 5 |
+
"rouge_l": 0.9860549005631489
|
| 6 |
+
},
|
| 7 |
+
"num_samples": 400
|
| 8 |
+
}
|
predictions.jsonl
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:19e773648cb4e65de8660ea6365e10acca112d42a854923df93db4a6f333a82d
|
| 3 |
+
size 20217442
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"backend": "tokenizers",
|
| 3 |
+
"clean_up_tokenization_spaces": false,
|
| 4 |
+
"do_lower_case": false,
|
| 5 |
+
"eos_token": "<|endoftext|>",
|
| 6 |
+
"is_local": false,
|
| 7 |
+
"local_files_only": false,
|
| 8 |
+
"model_max_length": 128000,
|
| 9 |
+
"pad_token": "<|endoftext|>",
|
| 10 |
+
"padding_side": "right",
|
| 11 |
+
"remove_space": false,
|
| 12 |
+
"tokenizer_class": "TokenizersBackend"
|
| 13 |
+
}
|
train_results.json
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"total_flos": 2.3441226939026637e+17,
|
| 3 |
+
"train_loss": 0.5646523337894016,
|
| 4 |
+
"train_runtime": 6392.6844,
|
| 5 |
+
"train_samples_per_second": 1.689,
|
| 6 |
+
"train_steps_per_second": 0.106
|
| 7 |
+
}
|
trainer_state.json
ADDED
|
@@ -0,0 +1,724 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": 675,
|
| 3 |
+
"best_metric": 0.5975516438484192,
|
| 4 |
+
"best_model_checkpoint": "/home/svadouser1/pooja/llm_autotuning/checkpoints/exp_20260827_221617/checkpoint-675",
|
| 5 |
+
"epoch": 3.0,
|
| 6 |
+
"eval_steps": 9999,
|
| 7 |
+
"global_step": 675,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"entropy": 1.461071628332138,
|
| 14 |
+
"epoch": 0.044444444444444446,
|
| 15 |
+
"grad_norm": 1.6171875,
|
| 16 |
+
"learning_rate": 5.294117647058824e-05,
|
| 17 |
+
"loss": 2.388203811645508,
|
| 18 |
+
"mean_token_accuracy": 0.6123090386390686,
|
| 19 |
+
"num_tokens": 19486.0,
|
| 20 |
+
"step": 10
|
| 21 |
+
},
|
| 22 |
+
{
|
| 23 |
+
"entropy": 0.9562640145421029,
|
| 24 |
+
"epoch": 0.08888888888888889,
|
| 25 |
+
"grad_norm": 0.6875,
|
| 26 |
+
"learning_rate": 0.00011176470588235294,
|
| 27 |
+
"loss": 1.1077078819274901,
|
| 28 |
+
"mean_token_accuracy": 0.8430786401033401,
|
| 29 |
+
"num_tokens": 38826.0,
|
| 30 |
+
"step": 20
|
| 31 |
+
},
|
| 32 |
+
{
|
| 33 |
+
"entropy": 0.77465368360281,
|
| 34 |
+
"epoch": 0.13333333333333333,
|
| 35 |
+
"grad_norm": 0.455078125,
|
| 36 |
+
"learning_rate": 0.00017058823529411766,
|
| 37 |
+
"loss": 0.8623821258544921,
|
| 38 |
+
"mean_token_accuracy": 0.864073084294796,
|
| 39 |
+
"num_tokens": 58244.0,
|
| 40 |
+
"step": 30
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"entropy": 0.7483784720301628,
|
| 44 |
+
"epoch": 0.17777777777777778,
|
| 45 |
+
"grad_norm": 0.423828125,
|
| 46 |
+
"learning_rate": 0.00019996997576394573,
|
| 47 |
+
"loss": 0.7713714599609375,
|
| 48 |
+
"mean_token_accuracy": 0.8697837173938752,
|
| 49 |
+
"num_tokens": 77711.0,
|
| 50 |
+
"step": 40
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"entropy": 0.7391687169671058,
|
| 54 |
+
"epoch": 0.2222222222222222,
|
| 55 |
+
"grad_norm": 0.376953125,
|
| 56 |
+
"learning_rate": 0.00019972989003925543,
|
| 57 |
+
"loss": 0.7115508556365967,
|
| 58 |
+
"mean_token_accuracy": 0.8751530349254608,
|
| 59 |
+
"num_tokens": 97104.0,
|
| 60 |
+
"step": 50
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"entropy": 0.7042164742946625,
|
| 64 |
+
"epoch": 0.26666666666666666,
|
| 65 |
+
"grad_norm": 0.330078125,
|
| 66 |
+
"learning_rate": 0.00019925029517454195,
|
| 67 |
+
"loss": 0.6672511577606202,
|
| 68 |
+
"mean_token_accuracy": 0.8725973218679428,
|
| 69 |
+
"num_tokens": 116498.0,
|
| 70 |
+
"step": 60
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"entropy": 0.6745826601982117,
|
| 74 |
+
"epoch": 0.3111111111111111,
|
| 75 |
+
"grad_norm": 0.361328125,
|
| 76 |
+
"learning_rate": 0.00019853234295442638,
|
| 77 |
+
"loss": 0.6365277290344238,
|
| 78 |
+
"mean_token_accuracy": 0.872233435511589,
|
| 79 |
+
"num_tokens": 136089.0,
|
| 80 |
+
"step": 70
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"entropy": 0.6878040120005607,
|
| 84 |
+
"epoch": 0.35555555555555557,
|
| 85 |
+
"grad_norm": 0.37109375,
|
| 86 |
+
"learning_rate": 0.00019757775759738272,
|
| 87 |
+
"loss": 0.6395976066589355,
|
| 88 |
+
"mean_token_accuracy": 0.8741081401705741,
|
| 89 |
+
"num_tokens": 155486.0,
|
| 90 |
+
"step": 80
|
| 91 |
+
},
|
| 92 |
+
{
|
| 93 |
+
"entropy": 0.6898319900035859,
|
| 94 |
+
"epoch": 0.4,
|
| 95 |
+
"grad_norm": 0.330078125,
|
| 96 |
+
"learning_rate": 0.00019638883161489224,
|
| 97 |
+
"loss": 0.6368544578552247,
|
| 98 |
+
"mean_token_accuracy": 0.8765213042497635,
|
| 99 |
+
"num_tokens": 175116.0,
|
| 100 |
+
"step": 90
|
| 101 |
+
},
|
| 102 |
+
{
|
| 103 |
+
"entropy": 0.656438185274601,
|
| 104 |
+
"epoch": 0.4444444444444444,
|
| 105 |
+
"grad_norm": 0.33203125,
|
| 106 |
+
"learning_rate": 0.0001949684203057978,
|
| 107 |
+
"loss": 0.6183670043945313,
|
| 108 |
+
"mean_token_accuracy": 0.8779006212949753,
|
| 109 |
+
"num_tokens": 194495.0,
|
| 110 |
+
"step": 100
|
| 111 |
+
},
|
| 112 |
+
{
|
| 113 |
+
"entropy": 0.6711793914437294,
|
| 114 |
+
"epoch": 0.4888888888888889,
|
| 115 |
+
"grad_norm": 0.388671875,
|
| 116 |
+
"learning_rate": 0.00019331993489907977,
|
| 117 |
+
"loss": 0.633416748046875,
|
| 118 |
+
"mean_token_accuracy": 0.8727847129106522,
|
| 119 |
+
"num_tokens": 214311.0,
|
| 120 |
+
"step": 110
|
| 121 |
+
},
|
| 122 |
+
{
|
| 123 |
+
"entropy": 0.6290019288659096,
|
| 124 |
+
"epoch": 0.5333333333333333,
|
| 125 |
+
"grad_norm": 0.33203125,
|
| 126 |
+
"learning_rate": 0.00019144733436152284,
|
| 127 |
+
"loss": 0.5889303684234619,
|
| 128 |
+
"mean_token_accuracy": 0.8814716726541519,
|
| 129 |
+
"num_tokens": 233676.0,
|
| 130 |
+
"step": 120
|
| 131 |
+
},
|
| 132 |
+
{
|
| 133 |
+
"entropy": 0.6458568304777146,
|
| 134 |
+
"epoch": 0.5777777777777777,
|
| 135 |
+
"grad_norm": 0.318359375,
|
| 136 |
+
"learning_rate": 0.00018935511588994715,
|
| 137 |
+
"loss": 0.6043062686920166,
|
| 138 |
+
"mean_token_accuracy": 0.8794952735304833,
|
| 139 |
+
"num_tokens": 253188.0,
|
| 140 |
+
"step": 130
|
| 141 |
+
},
|
| 142 |
+
{
|
| 143 |
+
"entropy": 0.6591948792338371,
|
| 144 |
+
"epoch": 0.6222222222222222,
|
| 145 |
+
"grad_norm": 0.32421875,
|
| 146 |
+
"learning_rate": 0.000187048304110838,
|
| 147 |
+
"loss": 0.6233312129974365,
|
| 148 |
+
"mean_token_accuracy": 0.873698017001152,
|
| 149 |
+
"num_tokens": 273185.0,
|
| 150 |
+
"step": 140
|
| 151 |
+
},
|
| 152 |
+
{
|
| 153 |
+
"entropy": 0.6109217047691345,
|
| 154 |
+
"epoch": 0.6666666666666666,
|
| 155 |
+
"grad_norm": 0.291015625,
|
| 156 |
+
"learning_rate": 0.00018453243901331195,
|
| 157 |
+
"loss": 0.5925732135772706,
|
| 158 |
+
"mean_token_accuracy": 0.884858050942421,
|
| 159 |
+
"num_tokens": 292431.0,
|
| 160 |
+
"step": 150
|
| 161 |
+
},
|
| 162 |
+
{
|
| 163 |
+
"entropy": 0.6377103522419929,
|
| 164 |
+
"epoch": 0.7111111111111111,
|
| 165 |
+
"grad_norm": 0.3671875,
|
| 166 |
+
"learning_rate": 0.00018181356264439905,
|
| 167 |
+
"loss": 0.5884737968444824,
|
| 168 |
+
"mean_token_accuracy": 0.8802034169435501,
|
| 169 |
+
"num_tokens": 311687.0,
|
| 170 |
+
"step": 160
|
| 171 |
+
},
|
| 172 |
+
{
|
| 173 |
+
"entropy": 0.6447647094726563,
|
| 174 |
+
"epoch": 0.7555555555555555,
|
| 175 |
+
"grad_norm": 0.3046875,
|
| 176 |
+
"learning_rate": 0.0001788982045985939,
|
| 177 |
+
"loss": 0.5939778327941895,
|
| 178 |
+
"mean_token_accuracy": 0.8817022785544395,
|
| 179 |
+
"num_tokens": 330813.0,
|
| 180 |
+
"step": 170
|
| 181 |
+
},
|
| 182 |
+
{
|
| 183 |
+
"entropy": 0.6029744669795036,
|
| 184 |
+
"epoch": 0.8,
|
| 185 |
+
"grad_norm": 0.3125,
|
| 186 |
+
"learning_rate": 0.00017579336633652317,
|
| 187 |
+
"loss": 0.5752908706665039,
|
| 188 |
+
"mean_token_accuracy": 0.8855120778083801,
|
| 189 |
+
"num_tokens": 350308.0,
|
| 190 |
+
"step": 180
|
| 191 |
+
},
|
| 192 |
+
{
|
| 193 |
+
"entropy": 0.6441277623176574,
|
| 194 |
+
"epoch": 0.8444444444444444,
|
| 195 |
+
"grad_norm": 0.259765625,
|
| 196 |
+
"learning_rate": 0.00017250650437038964,
|
| 197 |
+
"loss": 0.5958236217498779,
|
| 198 |
+
"mean_token_accuracy": 0.8808144956827164,
|
| 199 |
+
"num_tokens": 370083.0,
|
| 200 |
+
"step": 190
|
| 201 |
+
},
|
| 202 |
+
{
|
| 203 |
+
"entropy": 0.5766214028000831,
|
| 204 |
+
"epoch": 0.8888888888888888,
|
| 205 |
+
"grad_norm": 0.2734375,
|
| 206 |
+
"learning_rate": 0.0001690455123565743,
|
| 207 |
+
"loss": 0.5667627811431885,
|
| 208 |
+
"mean_token_accuracy": 0.8866458177566529,
|
| 209 |
+
"num_tokens": 389516.0,
|
| 210 |
+
"step": 200
|
| 211 |
+
},
|
| 212 |
+
{
|
| 213 |
+
"entropy": 0.6425503984093666,
|
| 214 |
+
"epoch": 0.9333333333333333,
|
| 215 |
+
"grad_norm": 0.3046875,
|
| 216 |
+
"learning_rate": 0.00016541870213840243,
|
| 217 |
+
"loss": 0.6137699127197266,
|
| 218 |
+
"mean_token_accuracy": 0.8819929376244545,
|
| 219 |
+
"num_tokens": 408940.0,
|
| 220 |
+
"step": 210
|
| 221 |
+
},
|
| 222 |
+
{
|
| 223 |
+
"entropy": 0.6159977808594703,
|
| 224 |
+
"epoch": 0.9777777777777777,
|
| 225 |
+
"grad_norm": 0.265625,
|
| 226 |
+
"learning_rate": 0.0001616347837846011,
|
| 227 |
+
"loss": 0.5788507461547852,
|
| 228 |
+
"mean_token_accuracy": 0.8852896198630333,
|
| 229 |
+
"num_tokens": 428166.0,
|
| 230 |
+
"step": 220
|
| 231 |
+
},
|
| 232 |
+
{
|
| 233 |
+
"entropy": 0.6020576372742653,
|
| 234 |
+
"epoch": 1.0222222222222221,
|
| 235 |
+
"grad_norm": 0.2353515625,
|
| 236 |
+
"learning_rate": 0.000157702844671387,
|
| 237 |
+
"loss": 0.5500371932983399,
|
| 238 |
+
"mean_token_accuracy": 0.8901642486453056,
|
| 239 |
+
"num_tokens": 447438.0,
|
| 240 |
+
"step": 230
|
| 241 |
+
},
|
| 242 |
+
{
|
| 243 |
+
"entropy": 0.5630205765366554,
|
| 244 |
+
"epoch": 1.0666666666666667,
|
| 245 |
+
"grad_norm": 0.275390625,
|
| 246 |
+
"learning_rate": 0.00015363232765842089,
|
| 247 |
+
"loss": 0.5382141590118408,
|
| 248 |
+
"mean_token_accuracy": 0.8885036528110504,
|
| 249 |
+
"num_tokens": 467009.0,
|
| 250 |
+
"step": 240
|
| 251 |
+
},
|
| 252 |
+
{
|
| 253 |
+
"entropy": 0.5631943918764591,
|
| 254 |
+
"epoch": 1.1111111111111112,
|
| 255 |
+
"grad_norm": 0.306640625,
|
| 256 |
+
"learning_rate": 0.00014943300841104094,
|
| 257 |
+
"loss": 0.5194426536560058,
|
| 258 |
+
"mean_token_accuracy": 0.8929360061883926,
|
| 259 |
+
"num_tokens": 486591.0,
|
| 260 |
+
"step": 250
|
| 261 |
+
},
|
| 262 |
+
{
|
| 263 |
+
"entropy": 0.5655099518597126,
|
| 264 |
+
"epoch": 1.1555555555555554,
|
| 265 |
+
"grad_norm": 0.259765625,
|
| 266 |
+
"learning_rate": 0.0001451149719232366,
|
| 267 |
+
"loss": 0.5290531158447266,
|
| 268 |
+
"mean_token_accuracy": 0.8912677451968193,
|
| 269 |
+
"num_tokens": 506180.0,
|
| 270 |
+
"step": 260
|
| 271 |
+
},
|
| 272 |
+
{
|
| 273 |
+
"entropy": 0.5660845704376698,
|
| 274 |
+
"epoch": 1.2,
|
| 275 |
+
"grad_norm": 0.326171875,
|
| 276 |
+
"learning_rate": 0.00014068858829774608,
|
| 277 |
+
"loss": 0.5366004943847656,
|
| 278 |
+
"mean_token_accuracy": 0.8909024119377136,
|
| 279 |
+
"num_tokens": 525696.0,
|
| 280 |
+
"step": 270
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"entropy": 0.5818345256149768,
|
| 284 |
+
"epoch": 1.2444444444444445,
|
| 285 |
+
"grad_norm": 0.26171875,
|
| 286 |
+
"learning_rate": 0.0001361644878414428,
|
| 287 |
+
"loss": 0.5424720764160156,
|
| 288 |
+
"mean_token_accuracy": 0.8896975666284561,
|
| 289 |
+
"num_tokens": 545392.0,
|
| 290 |
+
"step": 280
|
| 291 |
+
},
|
| 292 |
+
{
|
| 293 |
+
"entropy": 0.5546154774725437,
|
| 294 |
+
"epoch": 1.2888888888888888,
|
| 295 |
+
"grad_norm": 0.3359375,
|
| 296 |
+
"learning_rate": 0.00013155353553582057,
|
| 297 |
+
"loss": 0.5176132202148438,
|
| 298 |
+
"mean_token_accuracy": 0.8928953528404235,
|
| 299 |
+
"num_tokens": 565036.0,
|
| 300 |
+
"step": 290
|
| 301 |
+
},
|
| 302 |
+
{
|
| 303 |
+
"entropy": 0.5676225990056991,
|
| 304 |
+
"epoch": 1.3333333333333333,
|
| 305 |
+
"grad_norm": 0.275390625,
|
| 306 |
+
"learning_rate": 0.0001268668049438902,
|
| 307 |
+
"loss": 0.5130613803863525,
|
| 308 |
+
"mean_token_accuracy": 0.8928169339895249,
|
| 309 |
+
"num_tokens": 584507.0,
|
| 310 |
+
"step": 300
|
| 311 |
+
},
|
| 312 |
+
{
|
| 313 |
+
"entropy": 0.5529272347688675,
|
| 314 |
+
"epoch": 1.3777777777777778,
|
| 315 |
+
"grad_norm": 0.326171875,
|
| 316 |
+
"learning_rate": 0.0001221155516161506,
|
| 317 |
+
"loss": 0.5317890644073486,
|
| 318 |
+
"mean_token_accuracy": 0.8913554430007935,
|
| 319 |
+
"num_tokens": 604008.0,
|
| 320 |
+
"step": 310
|
| 321 |
+
},
|
| 322 |
+
{
|
| 323 |
+
"entropy": 0.5658974438905716,
|
| 324 |
+
"epoch": 1.4222222222222223,
|
| 325 |
+
"grad_norm": 0.26953125,
|
| 326 |
+
"learning_rate": 0.0001173111860595032,
|
| 327 |
+
"loss": 0.5273444652557373,
|
| 328 |
+
"mean_token_accuracy": 0.8927861481904984,
|
| 329 |
+
"num_tokens": 623480.0,
|
| 330 |
+
"step": 320
|
| 331 |
+
},
|
| 332 |
+
{
|
| 333 |
+
"entropy": 0.5697685681283474,
|
| 334 |
+
"epoch": 1.4666666666666668,
|
| 335 |
+
"grad_norm": 0.322265625,
|
| 336 |
+
"learning_rate": 0.00011246524633402573,
|
| 337 |
+
"loss": 0.5287697792053223,
|
| 338 |
+
"mean_token_accuracy": 0.8912984907627106,
|
| 339 |
+
"num_tokens": 642979.0,
|
| 340 |
+
"step": 330
|
| 341 |
+
},
|
| 342 |
+
{
|
| 343 |
+
"entropy": 0.5488456651568413,
|
| 344 |
+
"epoch": 1.511111111111111,
|
| 345 |
+
"grad_norm": 0.326171875,
|
| 346 |
+
"learning_rate": 0.00010758937034341787,
|
| 347 |
+
"loss": 0.513739824295044,
|
| 348 |
+
"mean_token_accuracy": 0.8952319726347924,
|
| 349 |
+
"num_tokens": 662396.0,
|
| 350 |
+
"step": 340
|
| 351 |
+
},
|
| 352 |
+
{
|
| 353 |
+
"entropy": 0.55315260887146,
|
| 354 |
+
"epoch": 1.5555555555555556,
|
| 355 |
+
"grad_norm": 0.26953125,
|
| 356 |
+
"learning_rate": 0.00010269526788566408,
|
| 357 |
+
"loss": 0.5225533485412598,
|
| 358 |
+
"mean_token_accuracy": 0.8937178790569306,
|
| 359 |
+
"num_tokens": 681644.0,
|
| 360 |
+
"step": 350
|
| 361 |
+
},
|
| 362 |
+
{
|
| 363 |
+
"entropy": 0.5738558314740658,
|
| 364 |
+
"epoch": 1.6,
|
| 365 |
+
"grad_norm": 0.29296875,
|
| 366 |
+
"learning_rate": 9.779469253103684e-05,
|
| 367 |
+
"loss": 0.5239317417144775,
|
| 368 |
+
"mean_token_accuracy": 0.8908430561423302,
|
| 369 |
+
"num_tokens": 701030.0,
|
| 370 |
+
"step": 360
|
| 371 |
+
},
|
| 372 |
+
{
|
| 373 |
+
"entropy": 0.5420261971652508,
|
| 374 |
+
"epoch": 1.6444444444444444,
|
| 375 |
+
"grad_norm": 0.259765625,
|
| 376 |
+
"learning_rate": 9.289941339497719e-05,
|
| 377 |
+
"loss": 0.5274893283843994,
|
| 378 |
+
"mean_token_accuracy": 0.8941561266779899,
|
| 379 |
+
"num_tokens": 720294.0,
|
| 380 |
+
"step": 370
|
| 381 |
+
},
|
| 382 |
+
{
|
| 383 |
+
"entropy": 0.577882957458496,
|
| 384 |
+
"epoch": 1.6888888888888889,
|
| 385 |
+
"grad_norm": 0.33203125,
|
| 386 |
+
"learning_rate": 8.802118687364284e-05,
|
| 387 |
+
"loss": 0.5202207088470459,
|
| 388 |
+
"mean_token_accuracy": 0.8931182771921158,
|
| 389 |
+
"num_tokens": 739607.0,
|
| 390 |
+
"step": 380
|
| 391 |
+
},
|
| 392 |
+
{
|
| 393 |
+
"entropy": 0.5461296208202839,
|
| 394 |
+
"epoch": 1.7333333333333334,
|
| 395 |
+
"grad_norm": 0.30859375,
|
| 396 |
+
"learning_rate": 8.317172841000173e-05,
|
| 397 |
+
"loss": 0.5105932235717774,
|
| 398 |
+
"mean_token_accuracy": 0.8917569547891617,
|
| 399 |
+
"num_tokens": 759370.0,
|
| 400 |
+
"step": 390
|
| 401 |
+
},
|
| 402 |
+
{
|
| 403 |
+
"entropy": 0.5631573006510735,
|
| 404 |
+
"epoch": 1.7777777777777777,
|
| 405 |
+
"grad_norm": 0.298828125,
|
| 406 |
+
"learning_rate": 7.836268435827875e-05,
|
| 407 |
+
"loss": 0.5263056755065918,
|
| 408 |
+
"mean_token_accuracy": 0.8906426534056664,
|
| 409 |
+
"num_tokens": 778969.0,
|
| 410 |
+
"step": 400
|
| 411 |
+
},
|
| 412 |
+
{
|
| 413 |
+
"entropy": 0.5476421102881431,
|
| 414 |
+
"epoch": 1.8222222222222222,
|
| 415 |
+
"grad_norm": 0.255859375,
|
| 416 |
+
"learning_rate": 7.360560401432401e-05,
|
| 417 |
+
"loss": 0.5025547504425049,
|
| 418 |
+
"mean_token_accuracy": 0.8946620702743531,
|
| 419 |
+
"num_tokens": 798105.0,
|
| 420 |
+
"step": 410
|
| 421 |
+
},
|
| 422 |
+
{
|
| 423 |
+
"entropy": 0.5343898519873619,
|
| 424 |
+
"epoch": 1.8666666666666667,
|
| 425 |
+
"grad_norm": 0.27734375,
|
| 426 |
+
"learning_rate": 6.891191187907455e-05,
|
| 427 |
+
"loss": 0.49821176528930666,
|
| 428 |
+
"mean_token_accuracy": 0.8958980679512024,
|
| 429 |
+
"num_tokens": 817238.0,
|
| 430 |
+
"step": 420
|
| 431 |
+
},
|
| 432 |
+
{
|
| 433 |
+
"entropy": 0.5404686734080315,
|
| 434 |
+
"epoch": 1.911111111111111,
|
| 435 |
+
"grad_norm": 0.296875,
|
| 436 |
+
"learning_rate": 6.429288022172068e-05,
|
| 437 |
+
"loss": 0.5022043704986572,
|
| 438 |
+
"mean_token_accuracy": 0.8958747044205666,
|
| 439 |
+
"num_tokens": 836646.0,
|
| 440 |
+
"step": 430
|
| 441 |
+
},
|
| 442 |
+
{
|
| 443 |
+
"entropy": 0.5426375091075897,
|
| 444 |
+
"epoch": 1.9555555555555557,
|
| 445 |
+
"grad_norm": 0.279296875,
|
| 446 |
+
"learning_rate": 5.9759602008468996e-05,
|
| 447 |
+
"loss": 0.5035938262939453,
|
| 448 |
+
"mean_token_accuracy": 0.8944751426577568,
|
| 449 |
+
"num_tokens": 856282.0,
|
| 450 |
+
"step": 440
|
| 451 |
+
},
|
| 452 |
+
{
|
| 453 |
+
"entropy": 0.5425564736127854,
|
| 454 |
+
"epoch": 2.0,
|
| 455 |
+
"grad_norm": 0.263671875,
|
| 456 |
+
"learning_rate": 5.532296426191539e-05,
|
| 457 |
+
"loss": 0.5004886627197266,
|
| 458 |
+
"mean_token_accuracy": 0.8947419315576554,
|
| 459 |
+
"num_tokens": 875656.0,
|
| 460 |
+
"step": 450
|
| 461 |
+
},
|
| 462 |
+
{
|
| 463 |
+
"entropy": 0.526330380141735,
|
| 464 |
+
"epoch": 2.0444444444444443,
|
| 465 |
+
"grad_norm": 0.279296875,
|
| 466 |
+
"learning_rate": 5.0993621915007785e-05,
|
| 467 |
+
"loss": 0.44887552261352537,
|
| 468 |
+
"mean_token_accuracy": 0.9024429768323898,
|
| 469 |
+
"num_tokens": 894973.0,
|
| 470 |
+
"step": 460
|
| 471 |
+
},
|
| 472 |
+
{
|
| 473 |
+
"entropy": 0.4774711772799492,
|
| 474 |
+
"epoch": 2.088888888888889,
|
| 475 |
+
"grad_norm": 0.365234375,
|
| 476 |
+
"learning_rate": 4.678197222239035e-05,
|
| 477 |
+
"loss": 0.44640169143676756,
|
| 478 |
+
"mean_token_accuracy": 0.9045588001608849,
|
| 479 |
+
"num_tokens": 914579.0,
|
| 480 |
+
"step": 470
|
| 481 |
+
},
|
| 482 |
+
{
|
| 483 |
+
"entropy": 0.4789727419614792,
|
| 484 |
+
"epoch": 2.1333333333333333,
|
| 485 |
+
"grad_norm": 0.3828125,
|
| 486 |
+
"learning_rate": 4.269812979058235e-05,
|
| 487 |
+
"loss": 0.44276885986328124,
|
| 488 |
+
"mean_token_accuracy": 0.9066318318247795,
|
| 489 |
+
"num_tokens": 933807.0,
|
| 490 |
+
"step": 480
|
| 491 |
+
},
|
| 492 |
+
{
|
| 493 |
+
"entropy": 0.47537712305784224,
|
| 494 |
+
"epoch": 2.1777777777777776,
|
| 495 |
+
"grad_norm": 0.359375,
|
| 496 |
+
"learning_rate": 3.875190228695862e-05,
|
| 497 |
+
"loss": 0.4367781639099121,
|
| 498 |
+
"mean_token_accuracy": 0.9070042923092843,
|
| 499 |
+
"num_tokens": 953241.0,
|
| 500 |
+
"step": 490
|
| 501 |
+
},
|
| 502 |
+
{
|
| 503 |
+
"entropy": 0.4877164676785469,
|
| 504 |
+
"epoch": 2.2222222222222223,
|
| 505 |
+
"grad_norm": 0.328125,
|
| 506 |
+
"learning_rate": 3.4952766885868346e-05,
|
| 507 |
+
"loss": 0.4422135353088379,
|
| 508 |
+
"mean_token_accuracy": 0.905805604159832,
|
| 509 |
+
"num_tokens": 972758.0,
|
| 510 |
+
"step": 500
|
| 511 |
+
},
|
| 512 |
+
{
|
| 513 |
+
"entropy": 0.4915403999388218,
|
| 514 |
+
"epoch": 2.2666666666666666,
|
| 515 |
+
"grad_norm": 0.376953125,
|
| 516 |
+
"learning_rate": 3.130984750845885e-05,
|
| 517 |
+
"loss": 0.4357293128967285,
|
| 518 |
+
"mean_token_accuracy": 0.9067289993166924,
|
| 519 |
+
"num_tokens": 992534.0,
|
| 520 |
+
"step": 510
|
| 521 |
+
},
|
| 522 |
+
{
|
| 523 |
+
"entropy": 0.4957615494728088,
|
| 524 |
+
"epoch": 2.311111111111111,
|
| 525 |
+
"grad_norm": 0.33984375,
|
| 526 |
+
"learning_rate": 2.7831892910864434e-05,
|
| 527 |
+
"loss": 0.44167218208312986,
|
| 528 |
+
"mean_token_accuracy": 0.9066730305552483,
|
| 529 |
+
"num_tokens": 1012003.0,
|
| 530 |
+
"step": 520
|
| 531 |
+
},
|
| 532 |
+
{
|
| 533 |
+
"entropy": 0.48269262462854384,
|
| 534 |
+
"epoch": 2.3555555555555556,
|
| 535 |
+
"grad_norm": 0.375,
|
| 536 |
+
"learning_rate": 2.4527255673383565e-05,
|
| 537 |
+
"loss": 0.437894344329834,
|
| 538 |
+
"mean_token_accuracy": 0.904815036058426,
|
| 539 |
+
"num_tokens": 1031505.0,
|
| 540 |
+
"step": 530
|
| 541 |
+
},
|
| 542 |
+
{
|
| 543 |
+
"entropy": 0.47741171419620515,
|
| 544 |
+
"epoch": 2.4,
|
| 545 |
+
"grad_norm": 0.4140625,
|
| 546 |
+
"learning_rate": 2.140387214110322e-05,
|
| 547 |
+
"loss": 0.4400351047515869,
|
| 548 |
+
"mean_token_accuracy": 0.9068154886364936,
|
| 549 |
+
"num_tokens": 1050914.0,
|
| 550 |
+
"step": 540
|
| 551 |
+
},
|
| 552 |
+
{
|
| 553 |
+
"entropy": 0.4727069653570652,
|
| 554 |
+
"epoch": 2.4444444444444446,
|
| 555 |
+
"grad_norm": 0.35546875,
|
| 556 |
+
"learning_rate": 1.846924336414474e-05,
|
| 557 |
+
"loss": 0.43111190795898435,
|
| 558 |
+
"mean_token_accuracy": 0.9075597256422043,
|
| 559 |
+
"num_tokens": 1070218.0,
|
| 560 |
+
"step": 550
|
| 561 |
+
},
|
| 562 |
+
{
|
| 563 |
+
"entropy": 0.4883471392095089,
|
| 564 |
+
"epoch": 2.488888888888889,
|
| 565 |
+
"grad_norm": 0.314453125,
|
| 566 |
+
"learning_rate": 1.5730417083304573e-05,
|
| 567 |
+
"loss": 0.4493571758270264,
|
| 568 |
+
"mean_token_accuracy": 0.9073263511061669,
|
| 569 |
+
"num_tokens": 1089667.0,
|
| 570 |
+
"step": 560
|
| 571 |
+
},
|
| 572 |
+
{
|
| 573 |
+
"entropy": 0.4878625735640526,
|
| 574 |
+
"epoch": 2.533333333333333,
|
| 575 |
+
"grad_norm": 0.375,
|
| 576 |
+
"learning_rate": 1.3193970804352952e-05,
|
| 577 |
+
"loss": 0.44793548583984377,
|
| 578 |
+
"mean_token_accuracy": 0.9050837978720665,
|
| 579 |
+
"num_tokens": 1109201.0,
|
| 580 |
+
"step": 570
|
| 581 |
+
},
|
| 582 |
+
{
|
| 583 |
+
"entropy": 0.48537297546863556,
|
| 584 |
+
"epoch": 2.5777777777777775,
|
| 585 |
+
"grad_norm": 0.337890625,
|
| 586 |
+
"learning_rate": 1.08659960016387e-05,
|
| 587 |
+
"loss": 0.4368610382080078,
|
| 588 |
+
"mean_token_accuracy": 0.9097044110298157,
|
| 589 |
+
"num_tokens": 1128649.0,
|
| 590 |
+
"step": 580
|
| 591 |
+
},
|
| 592 |
+
{
|
| 593 |
+
"entropy": 0.4859867602586746,
|
| 594 |
+
"epoch": 2.6222222222222222,
|
| 595 |
+
"grad_norm": 0.328125,
|
| 596 |
+
"learning_rate": 8.75208348893667e-06,
|
| 597 |
+
"loss": 0.4249687194824219,
|
| 598 |
+
"mean_token_accuracy": 0.9087634190917016,
|
| 599 |
+
"num_tokens": 1148353.0,
|
| 600 |
+
"step": 590
|
| 601 |
+
},
|
| 602 |
+
{
|
| 603 |
+
"entropy": 0.4833585321903229,
|
| 604 |
+
"epoch": 2.6666666666666665,
|
| 605 |
+
"grad_norm": 0.3828125,
|
| 606 |
+
"learning_rate": 6.857309992670624e-06,
|
| 607 |
+
"loss": 0.4455591678619385,
|
| 608 |
+
"mean_token_accuracy": 0.9075253725051879,
|
| 609 |
+
"num_tokens": 1167638.0,
|
| 610 |
+
"step": 600
|
| 611 |
+
},
|
| 612 |
+
{
|
| 613 |
+
"entropy": 0.4794360339641571,
|
| 614 |
+
"epoch": 2.7111111111111112,
|
| 615 |
+
"grad_norm": 0.3203125,
|
| 616 |
+
"learning_rate": 5.186225959757207e-06,
|
| 617 |
+
"loss": 0.43627562522888186,
|
| 618 |
+
"mean_token_accuracy": 0.907138803601265,
|
| 619 |
+
"num_tokens": 1186894.0,
|
| 620 |
+
"step": 610
|
| 621 |
+
},
|
| 622 |
+
{
|
| 623 |
+
"entropy": 0.4800324112176895,
|
| 624 |
+
"epoch": 2.7555555555555555,
|
| 625 |
+
"grad_norm": 0.365234375,
|
| 626 |
+
"learning_rate": 3.7428446293514386e-06,
|
| 627 |
+
"loss": 0.42800016403198243,
|
| 628 |
+
"mean_token_accuracy": 0.9089159920811654,
|
| 629 |
+
"num_tokens": 1206224.0,
|
| 630 |
+
"step": 620
|
| 631 |
+
},
|
| 632 |
+
{
|
| 633 |
+
"entropy": 0.48999542444944383,
|
| 634 |
+
"epoch": 2.8,
|
| 635 |
+
"grad_norm": 0.31640625,
|
| 636 |
+
"learning_rate": 2.5306323947385746e-06,
|
| 637 |
+
"loss": 0.44663453102111816,
|
| 638 |
+
"mean_token_accuracy": 0.9060632780194282,
|
| 639 |
+
"num_tokens": 1225569.0,
|
| 640 |
+
"step": 630
|
| 641 |
+
},
|
| 642 |
+
{
|
| 643 |
+
"entropy": 0.4874999448657036,
|
| 644 |
+
"epoch": 2.8444444444444446,
|
| 645 |
+
"grad_norm": 0.326171875,
|
| 646 |
+
"learning_rate": 1.5525004785192143e-06,
|
| 647 |
+
"loss": 0.44841961860656737,
|
| 648 |
+
"mean_token_accuracy": 0.9045136958360672,
|
| 649 |
+
"num_tokens": 1244959.0,
|
| 650 |
+
"step": 640
|
| 651 |
+
},
|
| 652 |
+
{
|
| 653 |
+
"entropy": 0.49936808943748473,
|
| 654 |
+
"epoch": 2.888888888888889,
|
| 655 |
+
"grad_norm": 0.4296875,
|
| 656 |
+
"learning_rate": 8.107979410802769e-07,
|
| 657 |
+
"loss": 0.4453989028930664,
|
| 658 |
+
"mean_token_accuracy": 0.9050952970981598,
|
| 659 |
+
"num_tokens": 1264904.0,
|
| 660 |
+
"step": 650
|
| 661 |
+
},
|
| 662 |
+
{
|
| 663 |
+
"entropy": 0.49325661584734914,
|
| 664 |
+
"epoch": 2.9333333333333336,
|
| 665 |
+
"grad_norm": 0.373046875,
|
| 666 |
+
"learning_rate": 3.0730603914255193e-07,
|
| 667 |
+
"loss": 0.44756379127502444,
|
| 668 |
+
"mean_token_accuracy": 0.9043820068240166,
|
| 669 |
+
"num_tokens": 1284311.0,
|
| 670 |
+
"step": 660
|
| 671 |
+
},
|
| 672 |
+
{
|
| 673 |
+
"entropy": 0.5068465434014797,
|
| 674 |
+
"epoch": 2.977777777777778,
|
| 675 |
+
"grad_norm": 0.373046875,
|
| 676 |
+
"learning_rate": 4.323394793315228e-08,
|
| 677 |
+
"loss": 0.4593667030334473,
|
| 678 |
+
"mean_token_accuracy": 0.9011617928743363,
|
| 679 |
+
"num_tokens": 1303914.0,
|
| 680 |
+
"step": 670
|
| 681 |
+
},
|
| 682 |
+
{
|
| 683 |
+
"epoch": 3.0,
|
| 684 |
+
"eval_entropy": 0.5207519015669823,
|
| 685 |
+
"eval_loss": 0.5975516438484192,
|
| 686 |
+
"eval_mean_token_accuracy": 0.8822531878948212,
|
| 687 |
+
"eval_num_tokens": 1313484.0,
|
| 688 |
+
"eval_runtime": 57.0793,
|
| 689 |
+
"eval_samples_per_second": 7.008,
|
| 690 |
+
"eval_steps_per_second": 1.752,
|
| 691 |
+
"step": 675
|
| 692 |
+
},
|
| 693 |
+
{
|
| 694 |
+
"epoch": 3.0,
|
| 695 |
+
"step": 675,
|
| 696 |
+
"total_flos": 2.3441226939026637e+17,
|
| 697 |
+
"train_loss": 0.5646523337894016,
|
| 698 |
+
"train_runtime": 6392.6844,
|
| 699 |
+
"train_samples_per_second": 1.689,
|
| 700 |
+
"train_steps_per_second": 0.106
|
| 701 |
+
}
|
| 702 |
+
],
|
| 703 |
+
"logging_steps": 10,
|
| 704 |
+
"max_steps": 675,
|
| 705 |
+
"num_input_tokens_seen": 0,
|
| 706 |
+
"num_train_epochs": 3,
|
| 707 |
+
"save_steps": 9999,
|
| 708 |
+
"stateful_callbacks": {
|
| 709 |
+
"TrainerControl": {
|
| 710 |
+
"args": {
|
| 711 |
+
"should_epoch_stop": false,
|
| 712 |
+
"should_evaluate": false,
|
| 713 |
+
"should_log": false,
|
| 714 |
+
"should_save": true,
|
| 715 |
+
"should_training_stop": true
|
| 716 |
+
},
|
| 717 |
+
"attributes": {}
|
| 718 |
+
}
|
| 719 |
+
},
|
| 720 |
+
"total_flos": 2.3441226939026637e+17,
|
| 721 |
+
"train_batch_size": 1,
|
| 722 |
+
"trial_name": null,
|
| 723 |
+
"trial_params": null
|
| 724 |
+
}
|