Nihilux commited on
Commit
015b050
·
verified ·
1 Parent(s): cdec759

checkpoint step 1800

Browse files
.gitattributes CHANGED
@@ -48,3 +48,4 @@ checkpoint-1400/tokenizer.json filter=lfs diff=lfs merge=lfs -text
48
  checkpoint-1500/tokenizer.json filter=lfs diff=lfs merge=lfs -text
49
  checkpoint-1600/tokenizer.json filter=lfs diff=lfs merge=lfs -text
50
  checkpoint-1700/tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
48
  checkpoint-1500/tokenizer.json filter=lfs diff=lfs merge=lfs -text
49
  checkpoint-1600/tokenizer.json filter=lfs diff=lfs merge=lfs -text
50
  checkpoint-1700/tokenizer.json filter=lfs diff=lfs merge=lfs -text
51
+ checkpoint-1800/tokenizer.json filter=lfs diff=lfs merge=lfs -text
checkpoint-1800/README.md ADDED
@@ -0,0 +1,210 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: unsloth/qwen3-30b-a3b-base
3
+ library_name: peft
4
+ pipeline_tag: text-generation
5
+ tags:
6
+ - base_model:adapter:unsloth/qwen3-30b-a3b-base
7
+ - lora
8
+ - sft
9
+ - transformers
10
+ - trl
11
+ - unsloth
12
+ ---
13
+
14
+ # Model Card for Model ID
15
+
16
+ <!-- Provide a quick summary of what the model is/does. -->
17
+
18
+
19
+
20
+ ## Model Details
21
+
22
+ ### Model Description
23
+
24
+ <!-- Provide a longer summary of what this model is. -->
25
+
26
+
27
+
28
+ - **Developed by:** [More Information Needed]
29
+ - **Funded by [optional]:** [More Information Needed]
30
+ - **Shared by [optional]:** [More Information Needed]
31
+ - **Model type:** [More Information Needed]
32
+ - **Language(s) (NLP):** [More Information Needed]
33
+ - **License:** [More Information Needed]
34
+ - **Finetuned from model [optional]:** [More Information Needed]
35
+
36
+ ### Model Sources [optional]
37
+
38
+ <!-- Provide the basic links for the model. -->
39
+
40
+ - **Repository:** [More Information Needed]
41
+ - **Paper [optional]:** [More Information Needed]
42
+ - **Demo [optional]:** [More Information Needed]
43
+
44
+ ## Uses
45
+
46
+ <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
47
+
48
+ ### Direct Use
49
+
50
+ <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
51
+
52
+ [More Information Needed]
53
+
54
+ ### Downstream Use [optional]
55
+
56
+ <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
57
+
58
+ [More Information Needed]
59
+
60
+ ### Out-of-Scope Use
61
+
62
+ <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
63
+
64
+ [More Information Needed]
65
+
66
+ ## Bias, Risks, and Limitations
67
+
68
+ <!-- This section is meant to convey both technical and sociotechnical limitations. -->
69
+
70
+ [More Information Needed]
71
+
72
+ ### Recommendations
73
+
74
+ <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
75
+
76
+ Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
77
+
78
+ ## How to Get Started with the Model
79
+
80
+ Use the code below to get started with the model.
81
+
82
+ [More Information Needed]
83
+
84
+ ## Training Details
85
+
86
+ ### Training Data
87
+
88
+ <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
89
+
90
+ [More Information Needed]
91
+
92
+ ### Training Procedure
93
+
94
+ <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
95
+
96
+ #### Preprocessing [optional]
97
+
98
+ [More Information Needed]
99
+
100
+
101
+ #### Training Hyperparameters
102
+
103
+ - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
104
+
105
+ #### Speeds, Sizes, Times [optional]
106
+
107
+ <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
108
+
109
+ [More Information Needed]
110
+
111
+ ## Evaluation
112
+
113
+ <!-- This section describes the evaluation protocols and provides the results. -->
114
+
115
+ ### Testing Data, Factors & Metrics
116
+
117
+ #### Testing Data
118
+
119
+ <!-- This should link to a Dataset Card if possible. -->
120
+
121
+ [More Information Needed]
122
+
123
+ #### Factors
124
+
125
+ <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
126
+
127
+ [More Information Needed]
128
+
129
+ #### Metrics
130
+
131
+ <!-- These are the evaluation metrics being used, ideally with a description of why. -->
132
+
133
+ [More Information Needed]
134
+
135
+ ### Results
136
+
137
+ [More Information Needed]
138
+
139
+ #### Summary
140
+
141
+
142
+
143
+ ## Model Examination [optional]
144
+
145
+ <!-- Relevant interpretability work for the model goes here -->
146
+
147
+ [More Information Needed]
148
+
149
+ ## Environmental Impact
150
+
151
+ <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
152
+
153
+ Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
154
+
155
+ - **Hardware Type:** [More Information Needed]
156
+ - **Hours used:** [More Information Needed]
157
+ - **Cloud Provider:** [More Information Needed]
158
+ - **Compute Region:** [More Information Needed]
159
+ - **Carbon Emitted:** [More Information Needed]
160
+
161
+ ## Technical Specifications [optional]
162
+
163
+ ### Model Architecture and Objective
164
+
165
+ [More Information Needed]
166
+
167
+ ### Compute Infrastructure
168
+
169
+ [More Information Needed]
170
+
171
+ #### Hardware
172
+
173
+ [More Information Needed]
174
+
175
+ #### Software
176
+
177
+ [More Information Needed]
178
+
179
+ ## Citation [optional]
180
+
181
+ <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
182
+
183
+ **BibTeX:**
184
+
185
+ [More Information Needed]
186
+
187
+ **APA:**
188
+
189
+ [More Information Needed]
190
+
191
+ ## Glossary [optional]
192
+
193
+ <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
194
+
195
+ [More Information Needed]
196
+
197
+ ## More Information [optional]
198
+
199
+ [More Information Needed]
200
+
201
+ ## Model Card Authors [optional]
202
+
203
+ [More Information Needed]
204
+
205
+ ## Model Card Contact
206
+
207
+ [More Information Needed]
208
+ ### Framework versions
209
+
210
+ - PEFT 0.20.0
checkpoint-1800/adapter_config.json ADDED
@@ -0,0 +1,57 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": {
6
+ "base_model_class": "Qwen3MoeForCausalLM",
7
+ "parent_library": "transformers.models.qwen3_moe.modeling_qwen3_moe",
8
+ "unsloth_fixed": true
9
+ },
10
+ "base_model_name_or_path": "unsloth/qwen3-30b-a3b-base",
11
+ "bias": "none",
12
+ "corda_config": null,
13
+ "ensure_weight_tying": false,
14
+ "eva_config": null,
15
+ "exclude_modules": null,
16
+ "fan_in_fan_out": false,
17
+ "inference_mode": true,
18
+ "init_lora_weights": true,
19
+ "layer_replication": null,
20
+ "layers_pattern": null,
21
+ "layers_to_transform": null,
22
+ "loftq_config": {},
23
+ "lora_alpha": 64,
24
+ "lora_bias": false,
25
+ "lora_dropout": 0,
26
+ "lora_ga_config": null,
27
+ "megatron_config": null,
28
+ "megatron_core": "megatron.core",
29
+ "modules_to_save": null,
30
+ "monteclora_config": null,
31
+ "peft_type": "LORA",
32
+ "peft_version": "0.20.0",
33
+ "qalora_group_size": 16,
34
+ "r": 32,
35
+ "rank_pattern": {},
36
+ "revision": null,
37
+ "target_modules": [
38
+ "o_proj",
39
+ "gate_proj",
40
+ "q_proj",
41
+ "up_proj",
42
+ "down_proj",
43
+ "v_proj",
44
+ "k_proj"
45
+ ],
46
+ "target_parameters": [
47
+ "mlp.experts.gate_up_proj",
48
+ "mlp.experts.down_proj"
49
+ ],
50
+ "task_type": "CAUSAL_LM",
51
+ "trainable_token_indices": null,
52
+ "use_bdlora": null,
53
+ "use_dora": false,
54
+ "use_qalora": false,
55
+ "use_rslora": true,
56
+ "velora_config": null
57
+ }
checkpoint-1800/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a1d2ae9fa0797c5f6e9f3f1fd5ff53c9f12c200dc1a05cd82c60729edd18c74d
3
+ size 5140199640
checkpoint-1800/optimizer.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:56c492f1b7e7aadb4be5765df5743236b8b8957887ba6b7cb017c429207f523d
3
+ size 2612400547
checkpoint-1800/scheduler.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d873a4f9239c98d6624b353e6bee56153a0bc656a3d9a559f175392e2d9f793b
3
+ size 1465
checkpoint-1800/tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506
3
+ size 11422650
checkpoint-1800/tokenizer_config.json ADDED
@@ -0,0 +1,225 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": null,
5
+ "clean_up_tokenization_spaces": false,
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "is_local": false,
9
+ "model_max_length": 32768,
10
+ "pad_token": "<|vision_pad|>",
11
+ "padding_side": "right",
12
+ "split_special_tokens": false,
13
+ "tokenizer_class": "Qwen2Tokenizer",
14
+ "unk_token": null,
15
+ "added_tokens_decoder": {
16
+ "151643": {
17
+ "content": "<|endoftext|>",
18
+ "single_word": false,
19
+ "lstrip": false,
20
+ "rstrip": false,
21
+ "normalized": false,
22
+ "special": true
23
+ },
24
+ "151644": {
25
+ "content": "<|im_start|>",
26
+ "single_word": false,
27
+ "lstrip": false,
28
+ "rstrip": false,
29
+ "normalized": false,
30
+ "special": true
31
+ },
32
+ "151645": {
33
+ "content": "<|im_end|>",
34
+ "single_word": false,
35
+ "lstrip": false,
36
+ "rstrip": false,
37
+ "normalized": false,
38
+ "special": true
39
+ },
40
+ "151646": {
41
+ "content": "<|object_ref_start|>",
42
+ "single_word": false,
43
+ "lstrip": false,
44
+ "rstrip": false,
45
+ "normalized": false,
46
+ "special": true
47
+ },
48
+ "151647": {
49
+ "content": "<|object_ref_end|>",
50
+ "single_word": false,
51
+ "lstrip": false,
52
+ "rstrip": false,
53
+ "normalized": false,
54
+ "special": true
55
+ },
56
+ "151648": {
57
+ "content": "<|box_start|>",
58
+ "single_word": false,
59
+ "lstrip": false,
60
+ "rstrip": false,
61
+ "normalized": false,
62
+ "special": true
63
+ },
64
+ "151649": {
65
+ "content": "<|box_end|>",
66
+ "single_word": false,
67
+ "lstrip": false,
68
+ "rstrip": false,
69
+ "normalized": false,
70
+ "special": true
71
+ },
72
+ "151650": {
73
+ "content": "<|quad_start|>",
74
+ "single_word": false,
75
+ "lstrip": false,
76
+ "rstrip": false,
77
+ "normalized": false,
78
+ "special": true
79
+ },
80
+ "151651": {
81
+ "content": "<|quad_end|>",
82
+ "single_word": false,
83
+ "lstrip": false,
84
+ "rstrip": false,
85
+ "normalized": false,
86
+ "special": true
87
+ },
88
+ "151652": {
89
+ "content": "<|vision_start|>",
90
+ "single_word": false,
91
+ "lstrip": false,
92
+ "rstrip": false,
93
+ "normalized": false,
94
+ "special": true
95
+ },
96
+ "151653": {
97
+ "content": "<|vision_end|>",
98
+ "single_word": false,
99
+ "lstrip": false,
100
+ "rstrip": false,
101
+ "normalized": false,
102
+ "special": true
103
+ },
104
+ "151654": {
105
+ "content": "<|vision_pad|>",
106
+ "single_word": false,
107
+ "lstrip": false,
108
+ "rstrip": false,
109
+ "normalized": false,
110
+ "special": true
111
+ },
112
+ "151655": {
113
+ "content": "<|image_pad|>",
114
+ "single_word": false,
115
+ "lstrip": false,
116
+ "rstrip": false,
117
+ "normalized": false,
118
+ "special": true
119
+ },
120
+ "151656": {
121
+ "content": "<|video_pad|>",
122
+ "single_word": false,
123
+ "lstrip": false,
124
+ "rstrip": false,
125
+ "normalized": false,
126
+ "special": true
127
+ },
128
+ "151657": {
129
+ "content": "<tool_call>",
130
+ "single_word": false,
131
+ "lstrip": false,
132
+ "rstrip": false,
133
+ "normalized": false,
134
+ "special": false
135
+ },
136
+ "151658": {
137
+ "content": "</tool_call>",
138
+ "single_word": false,
139
+ "lstrip": false,
140
+ "rstrip": false,
141
+ "normalized": false,
142
+ "special": false
143
+ },
144
+ "151659": {
145
+ "content": "<|fim_prefix|>",
146
+ "single_word": false,
147
+ "lstrip": false,
148
+ "rstrip": false,
149
+ "normalized": false,
150
+ "special": false
151
+ },
152
+ "151660": {
153
+ "content": "<|fim_middle|>",
154
+ "single_word": false,
155
+ "lstrip": false,
156
+ "rstrip": false,
157
+ "normalized": false,
158
+ "special": false
159
+ },
160
+ "151661": {
161
+ "content": "<|fim_suffix|>",
162
+ "single_word": false,
163
+ "lstrip": false,
164
+ "rstrip": false,
165
+ "normalized": false,
166
+ "special": false
167
+ },
168
+ "151662": {
169
+ "content": "<|fim_pad|>",
170
+ "single_word": false,
171
+ "lstrip": false,
172
+ "rstrip": false,
173
+ "normalized": false,
174
+ "special": false
175
+ },
176
+ "151663": {
177
+ "content": "<|repo_name|>",
178
+ "single_word": false,
179
+ "lstrip": false,
180
+ "rstrip": false,
181
+ "normalized": false,
182
+ "special": false
183
+ },
184
+ "151664": {
185
+ "content": "<|file_sep|>",
186
+ "single_word": false,
187
+ "lstrip": false,
188
+ "rstrip": false,
189
+ "normalized": false,
190
+ "special": false
191
+ },
192
+ "151665": {
193
+ "content": "<tool_response>",
194
+ "single_word": false,
195
+ "lstrip": false,
196
+ "rstrip": false,
197
+ "normalized": false,
198
+ "special": false
199
+ },
200
+ "151666": {
201
+ "content": "</tool_response>",
202
+ "single_word": false,
203
+ "lstrip": false,
204
+ "rstrip": false,
205
+ "normalized": false,
206
+ "special": false
207
+ },
208
+ "151667": {
209
+ "content": "<think>",
210
+ "single_word": false,
211
+ "lstrip": false,
212
+ "rstrip": false,
213
+ "normalized": false,
214
+ "special": false
215
+ },
216
+ "151668": {
217
+ "content": "</think>",
218
+ "single_word": false,
219
+ "lstrip": false,
220
+ "rstrip": false,
221
+ "normalized": false,
222
+ "special": false
223
+ }
224
+ }
225
+ }
checkpoint-1800/trainer_state.json ADDED
@@ -0,0 +1,1323 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "best_global_step": null,
3
+ "best_metric": null,
4
+ "best_model_checkpoint": null,
5
+ "epoch": 0.6945533033485669,
6
+ "eval_steps": 150,
7
+ "global_step": 1800,
8
+ "is_hyper_param_search": false,
9
+ "is_local_process_zero": true,
10
+ "is_world_process_zero": true,
11
+ "log_history": [
12
+ {
13
+ "epoch": 0.003858629463047594,
14
+ "grad_norm": 0.9612510204315186,
15
+ "learning_rate": 1.8e-05,
16
+ "loss": 1.0067429542541504,
17
+ "step": 10
18
+ },
19
+ {
20
+ "epoch": 0.007717258926095188,
21
+ "grad_norm": 0.6660895347595215,
22
+ "learning_rate": 3.8e-05,
23
+ "loss": 0.8466403007507324,
24
+ "step": 20
25
+ },
26
+ {
27
+ "epoch": 0.011575888389142782,
28
+ "grad_norm": 0.5681023001670837,
29
+ "learning_rate": 5.8e-05,
30
+ "loss": 0.7557645797729492,
31
+ "step": 30
32
+ },
33
+ {
34
+ "epoch": 0.015434517852190376,
35
+ "grad_norm": 0.7179134488105774,
36
+ "learning_rate": 7.800000000000001e-05,
37
+ "loss": 0.7409078121185303,
38
+ "step": 40
39
+ },
40
+ {
41
+ "epoch": 0.019293147315237968,
42
+ "grad_norm": 0.5565946102142334,
43
+ "learning_rate": 9.8e-05,
44
+ "loss": 0.7274651050567627,
45
+ "step": 50
46
+ },
47
+ {
48
+ "epoch": 0.023151776778285563,
49
+ "grad_norm": 0.5450749397277832,
50
+ "learning_rate": 0.000118,
51
+ "loss": 0.741389274597168,
52
+ "step": 60
53
+ },
54
+ {
55
+ "epoch": 0.027010406241333156,
56
+ "grad_norm": 0.4871410131454468,
57
+ "learning_rate": 0.000138,
58
+ "loss": 0.725303316116333,
59
+ "step": 70
60
+ },
61
+ {
62
+ "epoch": 0.03086903570438075,
63
+ "grad_norm": 0.5452415347099304,
64
+ "learning_rate": 0.00015800000000000002,
65
+ "loss": 0.724397611618042,
66
+ "step": 80
67
+ },
68
+ {
69
+ "epoch": 0.03472766516742835,
70
+ "grad_norm": 0.438076376914978,
71
+ "learning_rate": 0.00017800000000000002,
72
+ "loss": 0.7216084480285645,
73
+ "step": 90
74
+ },
75
+ {
76
+ "epoch": 0.038586294630475935,
77
+ "grad_norm": 0.3792989253997803,
78
+ "learning_rate": 0.00019800000000000002,
79
+ "loss": 0.7241204738616943,
80
+ "step": 100
81
+ },
82
+ {
83
+ "epoch": 0.04244492409352353,
84
+ "grad_norm": 0.31958940625190735,
85
+ "learning_rate": 0.0001999935634368633,
86
+ "loss": 0.7156620502471924,
87
+ "step": 110
88
+ },
89
+ {
90
+ "epoch": 0.04630355355657113,
91
+ "grad_norm": 0.3532995283603668,
92
+ "learning_rate": 0.00019997131465276176,
93
+ "loss": 0.716402530670166,
94
+ "step": 120
95
+ },
96
+ {
97
+ "epoch": 0.05016218301961872,
98
+ "grad_norm": 0.3720763921737671,
99
+ "learning_rate": 0.0001999331777190454,
100
+ "loss": 0.725121021270752,
101
+ "step": 130
102
+ },
103
+ {
104
+ "epoch": 0.05402081248266631,
105
+ "grad_norm": 0.327168345451355,
106
+ "learning_rate": 0.0001998791586967059,
107
+ "loss": 0.711461877822876,
108
+ "step": 140
109
+ },
110
+ {
111
+ "epoch": 0.05787944194571391,
112
+ "grad_norm": 0.31293419003486633,
113
+ "learning_rate": 0.00019980926617082901,
114
+ "loss": 0.7175386905670166,
115
+ "step": 150
116
+ },
117
+ {
118
+ "epoch": 0.05787944194571391,
119
+ "eval_loss": 0.7297696471214294,
120
+ "eval_runtime": 244.1933,
121
+ "eval_samples_per_second": 3.419,
122
+ "eval_steps_per_second": 1.712,
123
+ "step": 150
124
+ },
125
+ {
126
+ "epoch": 0.0617380714087615,
127
+ "grad_norm": 0.3199326694011688,
128
+ "learning_rate": 0.0001997235112492302,
129
+ "loss": 0.7084239006042481,
130
+ "step": 160
131
+ },
132
+ {
133
+ "epoch": 0.06559670087180909,
134
+ "grad_norm": 0.3334254026412964,
135
+ "learning_rate": 0.00019962190756068907,
136
+ "loss": 0.7120396614074707,
137
+ "step": 170
138
+ },
139
+ {
140
+ "epoch": 0.0694553303348567,
141
+ "grad_norm": 0.32970502972602844,
142
+ "learning_rate": 0.00019950447125278376,
143
+ "loss": 0.7141448974609375,
144
+ "step": 180
145
+ },
146
+ {
147
+ "epoch": 0.07331395979790428,
148
+ "grad_norm": 0.3339720368385315,
149
+ "learning_rate": 0.00019937122098932428,
150
+ "loss": 0.7231975555419922,
151
+ "step": 190
152
+ },
153
+ {
154
+ "epoch": 0.07717258926095187,
155
+ "grad_norm": 0.302263081073761,
156
+ "learning_rate": 0.00019922217794738657,
157
+ "loss": 0.7377682685852051,
158
+ "step": 200
159
+ },
160
+ {
161
+ "epoch": 0.08103121872399947,
162
+ "grad_norm": 0.3148849904537201,
163
+ "learning_rate": 0.00019905736581394689,
164
+ "loss": 0.7064791202545166,
165
+ "step": 210
166
+ },
167
+ {
168
+ "epoch": 0.08488984818704706,
169
+ "grad_norm": 0.3036375641822815,
170
+ "learning_rate": 0.00019887681078211707,
171
+ "loss": 0.7082621097564697,
172
+ "step": 220
173
+ },
174
+ {
175
+ "epoch": 0.08874847765009465,
176
+ "grad_norm": 0.353236585855484,
177
+ "learning_rate": 0.00019868054154698202,
178
+ "loss": 0.7202902793884277,
179
+ "step": 230
180
+ },
181
+ {
182
+ "epoch": 0.09260710711314225,
183
+ "grad_norm": 0.2983817458152771,
184
+ "learning_rate": 0.0001984685893010392,
185
+ "loss": 0.6949288845062256,
186
+ "step": 240
187
+ },
188
+ {
189
+ "epoch": 0.09646573657618984,
190
+ "grad_norm": 0.3325659930706024,
191
+ "learning_rate": 0.00019824098772924114,
192
+ "loss": 0.7241635799407959,
193
+ "step": 250
194
+ },
195
+ {
196
+ "epoch": 0.10032436603923744,
197
+ "grad_norm": 0.29093560576438904,
198
+ "learning_rate": 0.00019799777300364218,
199
+ "loss": 0.7079925060272216,
200
+ "step": 260
201
+ },
202
+ {
203
+ "epoch": 0.10418299550228503,
204
+ "grad_norm": 0.3151165246963501,
205
+ "learning_rate": 0.0001977389837776497,
206
+ "loss": 0.7299640655517579,
207
+ "step": 270
208
+ },
209
+ {
210
+ "epoch": 0.10804162496533262,
211
+ "grad_norm": 0.30055317282676697,
212
+ "learning_rate": 0.000197464661179881,
213
+ "loss": 0.7072536468505859,
214
+ "step": 280
215
+ },
216
+ {
217
+ "epoch": 0.11190025442838022,
218
+ "grad_norm": 0.28629517555236816,
219
+ "learning_rate": 0.00019717484880762685,
220
+ "loss": 0.7206380844116211,
221
+ "step": 290
222
+ },
223
+ {
224
+ "epoch": 0.1196175133544754,
225
+ "grad_norm": 0.31153738498687744,
226
+ "learning_rate": 0.0001965489414302289,
227
+ "loss": 0.722527265548706,
228
+ "step": 310
229
+ },
230
+ {
231
+ "epoch": 0.123476142817523,
232
+ "grad_norm": 0.2931964695453644,
233
+ "learning_rate": 0.00019621294589872003,
234
+ "loss": 0.7037133693695068,
235
+ "step": 320
236
+ },
237
+ {
238
+ "epoch": 0.1273347722805706,
239
+ "grad_norm": 0.33466067910194397,
240
+ "learning_rate": 0.0001958616595241865,
241
+ "loss": 0.727421760559082,
242
+ "step": 330
243
+ },
244
+ {
245
+ "epoch": 0.13119340174361818,
246
+ "grad_norm": 0.28575319051742554,
247
+ "learning_rate": 0.00019549513813554785,
248
+ "loss": 0.6994576930999756,
249
+ "step": 340
250
+ },
251
+ {
252
+ "epoch": 0.13505203120666578,
253
+ "grad_norm": 0.31457364559173584,
254
+ "learning_rate": 0.0001951134399829799,
255
+ "loss": 0.734303331375122,
256
+ "step": 350
257
+ },
258
+ {
259
+ "epoch": 0.1389106606697134,
260
+ "grad_norm": 0.30264604091644287,
261
+ "learning_rate": 0.00019471662572865736,
262
+ "loss": 0.7409284591674805,
263
+ "step": 360
264
+ },
265
+ {
266
+ "epoch": 0.14276929013276096,
267
+ "grad_norm": 0.28342851996421814,
268
+ "learning_rate": 0.00019430475843711293,
269
+ "loss": 0.7335753440856934,
270
+ "step": 370
271
+ },
272
+ {
273
+ "epoch": 0.14662791959580856,
274
+ "grad_norm": 0.30577903985977173,
275
+ "learning_rate": 0.00019387790356521463,
276
+ "loss": 0.7155701637268066,
277
+ "step": 380
278
+ },
279
+ {
280
+ "epoch": 0.15048654905885617,
281
+ "grad_norm": 0.28058379888534546,
282
+ "learning_rate": 0.000193436128951763,
283
+ "loss": 0.6958878993988037,
284
+ "step": 390
285
+ },
286
+ {
287
+ "epoch": 0.15434517852190374,
288
+ "grad_norm": 0.30901074409484863,
289
+ "learning_rate": 0.0001929795048067095,
290
+ "loss": 0.7224118709564209,
291
+ "step": 400
292
+ },
293
+ {
294
+ "epoch": 0.15820380798495134,
295
+ "grad_norm": 0.28316858410835266,
296
+ "learning_rate": 0.0001925081036999984,
297
+ "loss": 0.7156408786773681,
298
+ "step": 410
299
+ },
300
+ {
301
+ "epoch": 0.16206243744799895,
302
+ "grad_norm": 0.29417017102241516,
303
+ "learning_rate": 0.00019202200055003346,
304
+ "loss": 0.7048601150512696,
305
+ "step": 420
306
+ },
307
+ {
308
+ "epoch": 0.16592106691104652,
309
+ "grad_norm": 0.29546213150024414,
310
+ "learning_rate": 0.00019152127261177126,
311
+ "loss": 0.7047354698181152,
312
+ "step": 430
313
+ },
314
+ {
315
+ "epoch": 0.16977969637409412,
316
+ "grad_norm": 0.2986455261707306,
317
+ "learning_rate": 0.0001910059994644434,
318
+ "loss": 0.6936575889587402,
319
+ "step": 440
320
+ },
321
+ {
322
+ "epoch": 0.17363832583714173,
323
+ "grad_norm": 0.35280054807662964,
324
+ "learning_rate": 0.0001904762629989091,
325
+ "loss": 0.6981217384338378,
326
+ "step": 450
327
+ },
328
+ {
329
+ "epoch": 0.17363832583714173,
330
+ "eval_loss": 0.7733834385871887,
331
+ "eval_runtime": 241.1126,
332
+ "eval_samples_per_second": 3.463,
333
+ "eval_steps_per_second": 1.734,
334
+ "step": 450
335
+ },
336
+ {
337
+ "epoch": 0.1774969553001893,
338
+ "grad_norm": 0.28504979610443115,
339
+ "learning_rate": 0.00018993214740464063,
340
+ "loss": 0.7187217235565185,
341
+ "step": 460
342
+ },
343
+ {
344
+ "epoch": 0.1813555847632369,
345
+ "grad_norm": 0.294179230928421,
346
+ "learning_rate": 0.00018937373915634323,
347
+ "loss": 0.723856258392334,
348
+ "step": 470
349
+ },
350
+ {
351
+ "epoch": 0.1852142142262845,
352
+ "grad_norm": 0.3039489984512329,
353
+ "learning_rate": 0.00018880112700021205,
354
+ "loss": 0.7158075332641601,
355
+ "step": 480
356
+ },
357
+ {
358
+ "epoch": 0.18907284368933208,
359
+ "grad_norm": 0.3352907598018646,
360
+ "learning_rate": 0.0001882144019398278,
361
+ "loss": 0.7123313903808594,
362
+ "step": 490
363
+ },
364
+ {
365
+ "epoch": 0.19293147315237968,
366
+ "grad_norm": 0.31155890226364136,
367
+ "learning_rate": 0.00018761365722169403,
368
+ "loss": 0.7325988292694092,
369
+ "step": 500
370
+ },
371
+ {
372
+ "epoch": 0.1967901026154273,
373
+ "grad_norm": 0.3036295175552368,
374
+ "learning_rate": 0.00018699898832041757,
375
+ "loss": 0.7070968627929688,
376
+ "step": 510
377
+ },
378
+ {
379
+ "epoch": 0.2006487320784749,
380
+ "grad_norm": 0.2792280316352844,
381
+ "learning_rate": 0.00018637049292353513,
382
+ "loss": 0.7023006916046143,
383
+ "step": 520
384
+ },
385
+ {
386
+ "epoch": 0.20450736154152246,
387
+ "grad_norm": 0.30598390102386475,
388
+ "learning_rate": 0.00018572827091598793,
389
+ "loss": 0.708428144454956,
390
+ "step": 530
391
+ },
392
+ {
393
+ "epoch": 0.20836599100457007,
394
+ "grad_norm": 0.28768792748451233,
395
+ "learning_rate": 0.00018507242436424765,
396
+ "loss": 0.708616304397583,
397
+ "step": 540
398
+ },
399
+ {
400
+ "epoch": 0.21222462046761767,
401
+ "grad_norm": 0.32237133383750916,
402
+ "learning_rate": 0.00018440305750009483,
403
+ "loss": 0.7210727214813233,
404
+ "step": 550
405
+ },
406
+ {
407
+ "epoch": 0.21608324993066524,
408
+ "grad_norm": 0.2892608940601349,
409
+ "learning_rate": 0.000183720276704054,
410
+ "loss": 0.6931821823120117,
411
+ "step": 560
412
+ },
413
+ {
414
+ "epoch": 0.21994187939371285,
415
+ "grad_norm": 0.3152562081813812,
416
+ "learning_rate": 0.00018302419048848667,
417
+ "loss": 0.7213119506835938,
418
+ "step": 570
419
+ },
420
+ {
421
+ "epoch": 0.22380050885676045,
422
+ "grad_norm": 0.31499791145324707,
423
+ "learning_rate": 0.00018231490948034592,
424
+ "loss": 0.6939045906066894,
425
+ "step": 580
426
+ },
427
+ {
428
+ "epoch": 0.22765913831980802,
429
+ "grad_norm": 0.2861924469470978,
430
+ "learning_rate": 0.00018159254640359487,
431
+ "loss": 0.7252971649169921,
432
+ "step": 590
433
+ },
434
+ {
435
+ "epoch": 0.23537639724590323,
436
+ "grad_norm": 0.2836938500404358,
437
+ "learning_rate": 0.00018010903531734363,
438
+ "loss": 0.6902801513671875,
439
+ "step": 610
440
+ },
441
+ {
442
+ "epoch": 0.2392350267089508,
443
+ "grad_norm": 0.3197932243347168,
444
+ "learning_rate": 0.00017934812307793583,
445
+ "loss": 0.6933451175689698,
446
+ "step": 620
447
+ },
448
+ {
449
+ "epoch": 0.2430936561719984,
450
+ "grad_norm": 0.29024383425712585,
451
+ "learning_rate": 0.00017857460027263225,
452
+ "loss": 0.6954158306121826,
453
+ "step": 630
454
+ },
455
+ {
456
+ "epoch": 0.246952285635046,
457
+ "grad_norm": 0.29053837060928345,
458
+ "learning_rate": 0.00017778858983515743,
459
+ "loss": 0.7062996864318848,
460
+ "step": 640
461
+ },
462
+ {
463
+ "epoch": 0.2508109150980936,
464
+ "grad_norm": 0.2862823009490967,
465
+ "learning_rate": 0.00017699021668385895,
466
+ "loss": 0.726755428314209,
467
+ "step": 650
468
+ },
469
+ {
470
+ "epoch": 0.2546695445611412,
471
+ "grad_norm": 0.31322675943374634,
472
+ "learning_rate": 0.00017617960770185444,
473
+ "loss": 0.7403687953948974,
474
+ "step": 660
475
+ },
476
+ {
477
+ "epoch": 0.25852817402418876,
478
+ "grad_norm": 0.2916722595691681,
479
+ "learning_rate": 0.00017535689171686644,
480
+ "loss": 0.6931378364562988,
481
+ "step": 670
482
+ },
483
+ {
484
+ "epoch": 0.26238680348723636,
485
+ "grad_norm": 0.3099428117275238,
486
+ "learning_rate": 0.00017452219948074814,
487
+ "loss": 0.7085556983947754,
488
+ "step": 680
489
+ },
490
+ {
491
+ "epoch": 0.26624543295028397,
492
+ "grad_norm": 0.30417969822883606,
493
+ "learning_rate": 0.0001736756636487035,
494
+ "loss": 0.7230380058288575,
495
+ "step": 690
496
+ },
497
+ {
498
+ "epoch": 0.27010406241333157,
499
+ "grad_norm": 0.28197744488716125,
500
+ "learning_rate": 0.0001728174187582045,
501
+ "loss": 0.722746467590332,
502
+ "step": 700
503
+ },
504
+ {
505
+ "epoch": 0.27396269187637917,
506
+ "grad_norm": 0.30067917704582214,
507
+ "learning_rate": 0.00017194760120760986,
508
+ "loss": 0.7116784572601318,
509
+ "step": 710
510
+ },
511
+ {
512
+ "epoch": 0.2778213213394268,
513
+ "grad_norm": 0.2828962802886963,
514
+ "learning_rate": 0.00017106634923448724,
515
+ "loss": 0.6991414546966552,
516
+ "step": 720
517
+ },
518
+ {
519
+ "epoch": 0.2816799508024743,
520
+ "grad_norm": 0.303771436214447,
521
+ "learning_rate": 0.00017017380289364388,
522
+ "loss": 0.700380277633667,
523
+ "step": 730
524
+ },
525
+ {
526
+ "epoch": 0.2855385802655219,
527
+ "grad_norm": 0.3401305377483368,
528
+ "learning_rate": 0.00016927010403486786,
529
+ "loss": 0.717545461654663,
530
+ "step": 740
531
+ },
532
+ {
533
+ "epoch": 0.2893972097285695,
534
+ "grad_norm": 0.29057246446609497,
535
+ "learning_rate": 0.00016835539628038445,
536
+ "loss": 0.7112515449523926,
537
+ "step": 750
538
+ },
539
+ {
540
+ "epoch": 0.2893972097285695,
541
+ "eval_loss": 0.7159802317619324,
542
+ "eval_runtime": 247.114,
543
+ "eval_samples_per_second": 3.379,
544
+ "eval_steps_per_second": 1.692,
545
+ "step": 750
546
+ },
547
+ {
548
+ "epoch": 0.29325583919161713,
549
+ "grad_norm": 0.2816619277000427,
550
+ "learning_rate": 0.0001674298250020307,
551
+ "loss": 0.7210293769836426,
552
+ "step": 760
553
+ },
554
+ {
555
+ "epoch": 0.29711446865466473,
556
+ "grad_norm": 0.35117030143737793,
557
+ "learning_rate": 0.00016649353729815172,
558
+ "loss": 0.7004368305206299,
559
+ "step": 770
560
+ },
561
+ {
562
+ "epoch": 0.30097309811771233,
563
+ "grad_norm": 0.3049272298812866,
564
+ "learning_rate": 0.00016554668197022295,
565
+ "loss": 0.693110990524292,
566
+ "step": 780
567
+ },
568
+ {
569
+ "epoch": 0.3048317275807599,
570
+ "grad_norm": 0.33454808592796326,
571
+ "learning_rate": 0.0001645894094992015,
572
+ "loss": 0.7131325244903565,
573
+ "step": 790
574
+ },
575
+ {
576
+ "epoch": 0.3086903570438075,
577
+ "grad_norm": 0.2931094765663147,
578
+ "learning_rate": 0.00016362187202161076,
579
+ "loss": 0.7059287071228028,
580
+ "step": 800
581
+ },
582
+ {
583
+ "epoch": 0.3125489865068551,
584
+ "grad_norm": 0.31233930587768555,
585
+ "learning_rate": 0.00016264422330536154,
586
+ "loss": 0.7155674934387207,
587
+ "step": 810
588
+ },
589
+ {
590
+ "epoch": 0.3164076159699027,
591
+ "grad_norm": 0.30221912264823914,
592
+ "learning_rate": 0.00016165661872531443,
593
+ "loss": 0.7006759166717529,
594
+ "step": 820
595
+ },
596
+ {
597
+ "epoch": 0.3202662454329503,
598
+ "grad_norm": 0.28081852197647095,
599
+ "learning_rate": 0.00016065921523858635,
600
+ "loss": 0.7008425712585449,
601
+ "step": 830
602
+ },
603
+ {
604
+ "epoch": 0.3241248748959979,
605
+ "grad_norm": 0.31421926617622375,
606
+ "learning_rate": 0.00015965217135960607,
607
+ "loss": 0.7249965190887451,
608
+ "step": 840
609
+ },
610
+ {
611
+ "epoch": 0.3279835043590455,
612
+ "grad_norm": 0.2977774143218994,
613
+ "learning_rate": 0.0001586356471349215,
614
+ "loss": 0.6932929039001465,
615
+ "step": 850
616
+ },
617
+ {
618
+ "epoch": 0.33184213382209304,
619
+ "grad_norm": 0.28536227345466614,
620
+ "learning_rate": 0.0001576098041177646,
621
+ "loss": 0.7123911857604981,
622
+ "step": 860
623
+ },
624
+ {
625
+ "epoch": 0.33570076328514065,
626
+ "grad_norm": 0.2825419008731842,
627
+ "learning_rate": 0.00015657480534237561,
628
+ "loss": 0.6955167293548584,
629
+ "step": 870
630
+ },
631
+ {
632
+ "epoch": 0.33955939274818825,
633
+ "grad_norm": 0.2923476994037628,
634
+ "learning_rate": 0.00015553081529809281,
635
+ "loss": 0.6951467990875244,
636
+ "step": 880
637
+ },
638
+ {
639
+ "epoch": 0.34341802221123585,
640
+ "grad_norm": 0.290323406457901,
641
+ "learning_rate": 0.00015447799990321066,
642
+ "loss": 0.6957611083984375,
643
+ "step": 890
644
+ },
645
+ {
646
+ "epoch": 0.34727665167428345,
647
+ "grad_norm": 0.29604029655456543,
648
+ "learning_rate": 0.00015341652647861084,
649
+ "loss": 0.7186079025268555,
650
+ "step": 900
651
+ },
652
+ {
653
+ "epoch": 0.34727665167428345,
654
+ "eval_loss": 0.7107371687889099,
655
+ "eval_runtime": 261.1674,
656
+ "eval_samples_per_second": 3.197,
657
+ "eval_steps_per_second": 1.601,
658
+ "step": 900
659
+ },
660
+ {
661
+ "epoch": 0.35113528113733106,
662
+ "grad_norm": 0.27018797397613525,
663
+ "learning_rate": 0.00015234656372117042,
664
+ "loss": 0.7154855728149414,
665
+ "step": 910
666
+ },
667
+ {
668
+ "epoch": 0.3549939106003786,
669
+ "grad_norm": 0.26977527141571045,
670
+ "learning_rate": 0.00015126828167695146,
671
+ "loss": 0.699330997467041,
672
+ "step": 920
673
+ },
674
+ {
675
+ "epoch": 0.3588525400634262,
676
+ "grad_norm": 0.30887559056282043,
677
+ "learning_rate": 0.00015018185171417604,
678
+ "loss": 0.6892439842224121,
679
+ "step": 930
680
+ },
681
+ {
682
+ "epoch": 0.3627111695264738,
683
+ "grad_norm": 0.31451812386512756,
684
+ "learning_rate": 0.00014908744649599104,
685
+ "loss": 0.7145347595214844,
686
+ "step": 940
687
+ },
688
+ {
689
+ "epoch": 0.3665697989895214,
690
+ "grad_norm": 0.2898563742637634,
691
+ "learning_rate": 0.00014798523995302758,
692
+ "loss": 0.7055789947509765,
693
+ "step": 950
694
+ },
695
+ {
696
+ "epoch": 0.370428428452569,
697
+ "grad_norm": 0.26642370223999023,
698
+ "learning_rate": 0.00014687540725575845,
699
+ "loss": 0.6984900951385498,
700
+ "step": 960
701
+ },
702
+ {
703
+ "epoch": 0.3742870579156166,
704
+ "grad_norm": 0.32259368896484375,
705
+ "learning_rate": 0.0001457581247866591,
706
+ "loss": 0.6736190795898438,
707
+ "step": 970
708
+ },
709
+ {
710
+ "epoch": 0.37814568737866416,
711
+ "grad_norm": 0.2674982249736786,
712
+ "learning_rate": 0.00014463357011217532,
713
+ "loss": 0.6823287010192871,
714
+ "step": 980
715
+ },
716
+ {
717
+ "epoch": 0.38200431684171177,
718
+ "grad_norm": 0.295978307723999,
719
+ "learning_rate": 0.0001435019219545034,
720
+ "loss": 0.7077735900878906,
721
+ "step": 990
722
+ },
723
+ {
724
+ "epoch": 0.38586294630475937,
725
+ "grad_norm": 0.3013753294944763,
726
+ "learning_rate": 0.0001423633601631862,
727
+ "loss": 0.7043970108032227,
728
+ "step": 1000
729
+ },
730
+ {
731
+ "epoch": 0.38972157576780697,
732
+ "grad_norm": 0.26629018783569336,
733
+ "learning_rate": 0.00014121806568653025,
734
+ "loss": 0.6968683719635009,
735
+ "step": 1010
736
+ },
737
+ {
738
+ "epoch": 0.3935802052308546,
739
+ "grad_norm": 0.2660974860191345,
740
+ "learning_rate": 0.00014006622054284806,
741
+ "loss": 0.6901365756988526,
742
+ "step": 1020
743
+ },
744
+ {
745
+ "epoch": 0.3974388346939022,
746
+ "grad_norm": 0.2733144760131836,
747
+ "learning_rate": 0.0001389080077915307,
748
+ "loss": 0.6891340255737305,
749
+ "step": 1030
750
+ },
751
+ {
752
+ "epoch": 0.4012974641569498,
753
+ "grad_norm": 0.30119362473487854,
754
+ "learning_rate": 0.00013774361150395442,
755
+ "loss": 0.687501335144043,
756
+ "step": 1040
757
+ },
758
+ {
759
+ "epoch": 0.4051560936199973,
760
+ "grad_norm": 0.29165130853652954,
761
+ "learning_rate": 0.00013657321673422691,
762
+ "loss": 0.7198416233062744,
763
+ "step": 1050
764
+ },
765
+ {
766
+ "epoch": 0.4051560936199973,
767
+ "eval_loss": 0.7047804594039917,
768
+ "eval_runtime": 261.4782,
769
+ "eval_samples_per_second": 3.193,
770
+ "eval_steps_per_second": 1.599,
771
+ "step": 1050
772
+ },
773
+ {
774
+ "epoch": 0.40901472308304493,
775
+ "grad_norm": 0.2916712760925293,
776
+ "learning_rate": 0.00013539700948977717,
777
+ "loss": 0.6978848934173584,
778
+ "step": 1060
779
+ },
780
+ {
781
+ "epoch": 0.41287335254609253,
782
+ "grad_norm": 0.29535216093063354,
783
+ "learning_rate": 0.0001342151767017938,
784
+ "loss": 0.6890017986297607,
785
+ "step": 1070
786
+ },
787
+ {
788
+ "epoch": 0.41673198200914013,
789
+ "grad_norm": 0.28752076625823975,
790
+ "learning_rate": 0.00013302790619551674,
791
+ "loss": 0.717583703994751,
792
+ "step": 1080
793
+ },
794
+ {
795
+ "epoch": 0.42059061147218774,
796
+ "grad_norm": 0.32082512974739075,
797
+ "learning_rate": 0.00013183538666038648,
798
+ "loss": 0.7101277351379395,
799
+ "step": 1090
800
+ },
801
+ {
802
+ "epoch": 0.42444924093523534,
803
+ "grad_norm": 0.27900826930999756,
804
+ "learning_rate": 0.00013063780762005654,
805
+ "loss": 0.6807409286499023,
806
+ "step": 1100
807
+ },
808
+ {
809
+ "epoch": 0.4283078703982829,
810
+ "grad_norm": 0.3178718090057373,
811
+ "learning_rate": 0.0001294353594022728,
812
+ "loss": 0.686630392074585,
813
+ "step": 1110
814
+ },
815
+ {
816
+ "epoch": 0.4321664998613305,
817
+ "grad_norm": 0.2647388279438019,
818
+ "learning_rate": 0.00012822823310862524,
819
+ "loss": 0.670403003692627,
820
+ "step": 1120
821
+ },
822
+ {
823
+ "epoch": 0.4360251293243781,
824
+ "grad_norm": 0.28969231247901917,
825
+ "learning_rate": 0.00012701662058417688,
826
+ "loss": 0.6713655948638916,
827
+ "step": 1130
828
+ },
829
+ {
830
+ "epoch": 0.4398837587874257,
831
+ "grad_norm": 0.28609490394592285,
832
+ "learning_rate": 0.00012580071438697427,
833
+ "loss": 0.6948023319244385,
834
+ "step": 1140
835
+ },
836
+ {
837
+ "epoch": 0.4437423882504733,
838
+ "grad_norm": 0.27543318271636963,
839
+ "learning_rate": 0.0001245807077574449,
840
+ "loss": 0.7000330924987793,
841
+ "step": 1150
842
+ },
843
+ {
844
+ "epoch": 0.4476010177135209,
845
+ "grad_norm": 0.28665491938591003,
846
+ "learning_rate": 0.00012335679458768607,
847
+ "loss": 0.6907845497131347,
848
+ "step": 1160
849
+ },
850
+ {
851
+ "epoch": 0.45145964717656845,
852
+ "grad_norm": 0.26809993386268616,
853
+ "learning_rate": 0.00012212916939064998,
854
+ "loss": 0.6751095294952393,
855
+ "step": 1170
856
+ },
857
+ {
858
+ "epoch": 0.45531827663961605,
859
+ "grad_norm": 0.28608420491218567,
860
+ "learning_rate": 0.00012089802726923061,
861
+ "loss": 0.6977941513061523,
862
+ "step": 1180
863
+ },
864
+ {
865
+ "epoch": 0.45917690610266365,
866
+ "grad_norm": 0.2964702844619751,
867
+ "learning_rate": 0.00011966356388525646,
868
+ "loss": 0.685225248336792,
869
+ "step": 1190
870
+ },
871
+ {
872
+ "epoch": 0.46689416502875886,
873
+ "grad_norm": 0.3035467863082886,
874
+ "learning_rate": 0.00011718545858497087,
875
+ "loss": 0.6892534732818604,
876
+ "step": 1210
877
+ },
878
+ {
879
+ "epoch": 0.47075279449180646,
880
+ "grad_norm": 0.2953267991542816,
881
+ "learning_rate": 0.00011594221050671095,
882
+ "loss": 0.6970641136169433,
883
+ "step": 1220
884
+ },
885
+ {
886
+ "epoch": 0.47461142395485406,
887
+ "grad_norm": 0.2962872087955475,
888
+ "learning_rate": 0.00011469642877940779,
889
+ "loss": 0.6811253547668457,
890
+ "step": 1230
891
+ },
892
+ {
893
+ "epoch": 0.4784700534179016,
894
+ "grad_norm": 0.2830030024051666,
895
+ "learning_rate": 0.00011344831139151972,
896
+ "loss": 0.6688879489898681,
897
+ "step": 1240
898
+ },
899
+ {
900
+ "epoch": 0.4823286828809492,
901
+ "grad_norm": 0.29539620876312256,
902
+ "learning_rate": 0.00011219805670270496,
903
+ "loss": 0.700579833984375,
904
+ "step": 1250
905
+ },
906
+ {
907
+ "epoch": 0.4861873123439968,
908
+ "grad_norm": 0.29297176003456116,
909
+ "learning_rate": 0.00011094586341229656,
910
+ "loss": 0.6867257118225097,
911
+ "step": 1260
912
+ },
913
+ {
914
+ "epoch": 0.4900459418070444,
915
+ "grad_norm": 0.29863327741622925,
916
+ "learning_rate": 0.00010969193052772396,
917
+ "loss": 0.71397385597229,
918
+ "step": 1270
919
+ },
920
+ {
921
+ "epoch": 0.493904571270092,
922
+ "grad_norm": 0.2799651622772217,
923
+ "learning_rate": 0.00010843645733288519,
924
+ "loss": 0.7007513046264648,
925
+ "step": 1280
926
+ },
927
+ {
928
+ "epoch": 0.4977632007331396,
929
+ "grad_norm": 0.2826097905635834,
930
+ "learning_rate": 0.00010717964335647535,
931
+ "loss": 0.6815872669219971,
932
+ "step": 1290
933
+ },
934
+ {
935
+ "epoch": 0.5016218301961872,
936
+ "grad_norm": 0.27504852414131165,
937
+ "learning_rate": 0.00010592168834027598,
938
+ "loss": 0.6807018280029297,
939
+ "step": 1300
940
+ },
941
+ {
942
+ "epoch": 0.5054804596592348,
943
+ "grad_norm": 0.28094950318336487,
944
+ "learning_rate": 0.00010466279220741078,
945
+ "loss": 0.6823868751525879,
946
+ "step": 1310
947
+ },
948
+ {
949
+ "epoch": 0.5093390891222824,
950
+ "grad_norm": 0.2797650694847107,
951
+ "learning_rate": 0.00010340315503057243,
952
+ "loss": 0.674619722366333,
953
+ "step": 1320
954
+ },
955
+ {
956
+ "epoch": 0.5131977185853299,
957
+ "grad_norm": 0.28894102573394775,
958
+ "learning_rate": 0.00010214297700022549,
959
+ "loss": 0.7022358417510987,
960
+ "step": 1330
961
+ },
962
+ {
963
+ "epoch": 0.5170563480483775,
964
+ "grad_norm": 0.29386165738105774,
965
+ "learning_rate": 0.00010088245839279082,
966
+ "loss": 0.67178316116333,
967
+ "step": 1340
968
+ },
969
+ {
970
+ "epoch": 0.5209149775114251,
971
+ "grad_norm": 0.28537702560424805,
972
+ "learning_rate": 9.96217995388162e-05,
973
+ "loss": 0.6890518188476562,
974
+ "step": 1350
975
+ },
976
+ {
977
+ "epoch": 0.5209149775114251,
978
+ "eval_loss": 0.6932345032691956,
979
+ "eval_runtime": 269.3818,
980
+ "eval_samples_per_second": 3.1,
981
+ "eval_steps_per_second": 1.552,
982
+ "step": 1350
983
+ },
984
+ {
985
+ "epoch": 0.5247736069744727,
986
+ "grad_norm": 0.3049728274345398,
987
+ "learning_rate": 9.83612007911384e-05,
988
+ "loss": 0.6744989871978759,
989
+ "step": 1360
990
+ },
991
+ {
992
+ "epoch": 0.5286322364375203,
993
+ "grad_norm": 0.33253172039985657,
994
+ "learning_rate": 9.710086249304164e-05,
995
+ "loss": 0.6773325443267822,
996
+ "step": 1370
997
+ },
998
+ {
999
+ "epoch": 0.5324908659005679,
1000
+ "grad_norm": 0.27811330556869507,
1001
+ "learning_rate": 9.584098494641772e-05,
1002
+ "loss": 0.6876926898956299,
1003
+ "step": 1380
1004
+ },
1005
+ {
1006
+ "epoch": 0.5363494953636155,
1007
+ "grad_norm": 0.29328352212905884,
1008
+ "learning_rate": 9.458176837993246e-05,
1009
+ "loss": 0.689525842666626,
1010
+ "step": 1390
1011
+ },
1012
+ {
1013
+ "epoch": 0.5402081248266631,
1014
+ "grad_norm": 0.27268341183662415,
1015
+ "learning_rate": 9.332341291720408e-05,
1016
+ "loss": 0.6747959136962891,
1017
+ "step": 1400
1018
+ },
1019
+ {
1020
+ "epoch": 0.5440667542897107,
1021
+ "grad_norm": 0.28394851088523865,
1022
+ "learning_rate": 9.206611854499805e-05,
1023
+ "loss": 0.6711764335632324,
1024
+ "step": 1410
1025
+ },
1026
+ {
1027
+ "epoch": 0.5479253837527583,
1028
+ "grad_norm": 0.28641200065612793,
1029
+ "learning_rate": 9.081008508144388e-05,
1030
+ "loss": 0.6956794261932373,
1031
+ "step": 1420
1032
+ },
1033
+ {
1034
+ "epoch": 0.551784013215806,
1035
+ "grad_norm": 0.2705199718475342,
1036
+ "learning_rate": 8.955551214427856e-05,
1037
+ "loss": 0.7046959400177002,
1038
+ "step": 1430
1039
+ },
1040
+ {
1041
+ "epoch": 0.5556426426788535,
1042
+ "grad_norm": 0.26797372102737427,
1043
+ "learning_rate": 8.830259911912173e-05,
1044
+ "loss": 0.6727779865264892,
1045
+ "step": 1440
1046
+ },
1047
+ {
1048
+ "epoch": 0.5595012721419012,
1049
+ "grad_norm": 0.2805560529232025,
1050
+ "learning_rate": 8.705154512778821e-05,
1051
+ "loss": 0.6727550506591797,
1052
+ "step": 1450
1053
+ },
1054
+ {
1055
+ "epoch": 0.5633599016049486,
1056
+ "grad_norm": 0.28046268224716187,
1057
+ "learning_rate": 8.580254899664195e-05,
1058
+ "loss": 0.6817981243133545,
1059
+ "step": 1460
1060
+ },
1061
+ {
1062
+ "epoch": 0.5672185310679962,
1063
+ "grad_norm": 0.27203667163848877,
1064
+ "learning_rate": 8.455580922499716e-05,
1065
+ "loss": 0.6813424110412598,
1066
+ "step": 1470
1067
+ },
1068
+ {
1069
+ "epoch": 0.5710771605310438,
1070
+ "grad_norm": 0.29054710268974304,
1071
+ "learning_rate": 8.331152395357141e-05,
1072
+ "loss": 0.6807193279266357,
1073
+ "step": 1480
1074
+ },
1075
+ {
1076
+ "epoch": 0.5749357899940915,
1077
+ "grad_norm": 0.32792478799819946,
1078
+ "learning_rate": 8.206989093299572e-05,
1079
+ "loss": 0.6816314220428467,
1080
+ "step": 1490
1081
+ },
1082
+ {
1083
+ "epoch": 0.578794419457139,
1084
+ "grad_norm": 0.2703556418418884,
1085
+ "learning_rate": 8.083110749238659e-05,
1086
+ "loss": 0.6654800415039063,
1087
+ "step": 1500
1088
+ },
1089
+ {
1090
+ "epoch": 0.578794419457139,
1091
+ "eval_loss": 0.6862806081771851,
1092
+ "eval_runtime": 245.6426,
1093
+ "eval_samples_per_second": 3.399,
1094
+ "eval_steps_per_second": 1.702,
1095
+ "step": 1500
1096
+ },
1097
+ {
1098
+ "epoch": 0.5826530489201867,
1099
+ "grad_norm": 0.27336329221725464,
1100
+ "learning_rate": 7.959537050798512e-05,
1101
+ "loss": 0.6732516288757324,
1102
+ "step": 1510
1103
+ },
1104
+ {
1105
+ "epoch": 0.5865116783832343,
1106
+ "grad_norm": 0.27847999334335327,
1107
+ "learning_rate": 7.836287637186801e-05,
1108
+ "loss": 0.68833327293396,
1109
+ "step": 1520
1110
+ },
1111
+ {
1112
+ "epoch": 0.5903703078462819,
1113
+ "grad_norm": 0.25630271434783936,
1114
+ "learning_rate": 7.713382096073545e-05,
1115
+ "loss": 0.6711580276489257,
1116
+ "step": 1530
1117
+ },
1118
+ {
1119
+ "epoch": 0.5942289373093295,
1120
+ "grad_norm": 0.2752869427204132,
1121
+ "learning_rate": 7.59083996047812e-05,
1122
+ "loss": 0.6594626426696777,
1123
+ "step": 1540
1124
+ },
1125
+ {
1126
+ "epoch": 0.5980875667723771,
1127
+ "grad_norm": 0.27679699659347534,
1128
+ "learning_rate": 7.468680705664914e-05,
1129
+ "loss": 0.6733936309814453,
1130
+ "step": 1550
1131
+ },
1132
+ {
1133
+ "epoch": 0.6019461962354247,
1134
+ "grad_norm": 0.2734837830066681,
1135
+ "learning_rate": 7.346923746048202e-05,
1136
+ "loss": 0.6738894939422607,
1137
+ "step": 1560
1138
+ },
1139
+ {
1140
+ "epoch": 0.6058048256984723,
1141
+ "grad_norm": 0.27832263708114624,
1142
+ "learning_rate": 7.225588432106633e-05,
1143
+ "loss": 0.6458727359771729,
1144
+ "step": 1570
1145
+ },
1146
+ {
1147
+ "epoch": 0.6096634551615198,
1148
+ "grad_norm": 0.28797298669815063,
1149
+ "learning_rate": 7.104694047307963e-05,
1150
+ "loss": 0.6783483982086181,
1151
+ "step": 1580
1152
+ },
1153
+ {
1154
+ "epoch": 0.6135220846245674,
1155
+ "grad_norm": 0.27791404724121094,
1156
+ "learning_rate": 6.984259805044342e-05,
1157
+ "loss": 0.6734447479248047,
1158
+ "step": 1590
1159
+ },
1160
+ {
1161
+ "epoch": 0.617380714087615,
1162
+ "grad_norm": 0.26526719331741333,
1163
+ "learning_rate": 6.864304845578826e-05,
1164
+ "loss": 0.671996021270752,
1165
+ "step": 1600
1166
+ },
1167
+ {
1168
+ "epoch": 0.6212393435506626,
1169
+ "grad_norm": 0.28982335329055786,
1170
+ "learning_rate": 6.74484823300345e-05,
1171
+ "loss": 0.6816313743591309,
1172
+ "step": 1610
1173
+ },
1174
+ {
1175
+ "epoch": 0.6250979730137102,
1176
+ "grad_norm": 0.262482613325119,
1177
+ "learning_rate": 6.625908952209418e-05,
1178
+ "loss": 0.6751953601837158,
1179
+ "step": 1620
1180
+ },
1181
+ {
1182
+ "epoch": 0.6289566024767578,
1183
+ "grad_norm": 0.39887505769729614,
1184
+ "learning_rate": 6.507505905869914e-05,
1185
+ "loss": 0.6822636127471924,
1186
+ "step": 1630
1187
+ },
1188
+ {
1189
+ "epoch": 0.6328152319398054,
1190
+ "grad_norm": 0.2669890224933624,
1191
+ "learning_rate": 6.389657911435943e-05,
1192
+ "loss": 0.6915277957916259,
1193
+ "step": 1640
1194
+ },
1195
+ {
1196
+ "epoch": 0.636673861402853,
1197
+ "grad_norm": 0.2670292556285858,
1198
+ "learning_rate": 6.272383698145723e-05,
1199
+ "loss": 0.652358102798462,
1200
+ "step": 1650
1201
+ },
1202
+ {
1203
+ "epoch": 0.636673861402853,
1204
+ "eval_loss": 0.6799824237823486,
1205
+ "eval_runtime": 238.2058,
1206
+ "eval_samples_per_second": 3.505,
1207
+ "eval_steps_per_second": 1.755,
1208
+ "step": 1650
1209
+ },
1210
+ {
1211
+ "epoch": 0.6405324908659006,
1212
+ "grad_norm": 0.27272072434425354,
1213
+ "learning_rate": 6.15570190404811e-05,
1214
+ "loss": 0.6684075832366944,
1215
+ "step": 1660
1216
+ },
1217
+ {
1218
+ "epoch": 0.6443911203289482,
1219
+ "grad_norm": 0.25919416546821594,
1220
+ "learning_rate": 6.039631073040507e-05,
1221
+ "loss": 0.6651289463043213,
1222
+ "step": 1670
1223
+ },
1224
+ {
1225
+ "epoch": 0.6482497497919958,
1226
+ "grad_norm": 0.27560484409332275,
1227
+ "learning_rate": 5.924189651921728e-05,
1228
+ "loss": 0.6770682811737061,
1229
+ "step": 1680
1230
+ },
1231
+ {
1232
+ "epoch": 0.6521083792550434,
1233
+ "grad_norm": 0.29968157410621643,
1234
+ "learning_rate": 5.8093959874603176e-05,
1235
+ "loss": 0.6616378784179687,
1236
+ "step": 1690
1237
+ },
1238
+ {
1239
+ "epoch": 0.6598256381811385,
1240
+ "grad_norm": 0.2539561092853546,
1241
+ "learning_rate": 5.581824797953925e-05,
1242
+ "loss": 0.6652106761932373,
1243
+ "step": 1710
1244
+ },
1245
+ {
1246
+ "epoch": 0.6636842676441861,
1247
+ "grad_norm": 0.2645578980445862,
1248
+ "learning_rate": 5.4690834401347034e-05,
1249
+ "loss": 0.6816202163696289,
1250
+ "step": 1720
1251
+ },
1252
+ {
1253
+ "epoch": 0.6675428971072337,
1254
+ "grad_norm": 0.2860375642776489,
1255
+ "learning_rate": 5.357062167676426e-05,
1256
+ "loss": 0.6432219982147217,
1257
+ "step": 1730
1258
+ },
1259
+ {
1260
+ "epoch": 0.6714015265702813,
1261
+ "grad_norm": 0.29771652817726135,
1262
+ "learning_rate": 5.2457787837933715e-05,
1263
+ "loss": 0.6574324131011963,
1264
+ "step": 1740
1265
+ },
1266
+ {
1267
+ "epoch": 0.6752601560333289,
1268
+ "grad_norm": 0.26455527544021606,
1269
+ "learning_rate": 5.135250974429342e-05,
1270
+ "loss": 0.6606158256530762,
1271
+ "step": 1750
1272
+ },
1273
+ {
1274
+ "epoch": 0.6791187854963765,
1275
+ "grad_norm": 0.2809986174106598,
1276
+ "learning_rate": 5.02549630544688e-05,
1277
+ "loss": 0.6675248622894288,
1278
+ "step": 1760
1279
+ },
1280
+ {
1281
+ "epoch": 0.6829774149594241,
1282
+ "grad_norm": 0.2710455060005188,
1283
+ "learning_rate": 4.916532219835592e-05,
1284
+ "loss": 0.6551553249359131,
1285
+ "step": 1770
1286
+ },
1287
+ {
1288
+ "epoch": 0.6868360444224717,
1289
+ "grad_norm": 0.3084876239299774,
1290
+ "learning_rate": 4.808376034939965e-05,
1291
+ "loss": 0.6641845703125,
1292
+ "step": 1780
1293
+ },
1294
+ {
1295
+ "epoch": 0.6906946738855193,
1296
+ "grad_norm": 0.2739737033843994,
1297
+ "learning_rate": 4.701044939707181e-05,
1298
+ "loss": 0.640526533126831,
1299
+ "step": 1790
1300
+ }
1301
+ ],
1302
+ "logging_steps": 10,
1303
+ "max_steps": 2592,
1304
+ "num_input_tokens_seen": 0,
1305
+ "num_train_epochs": 1,
1306
+ "save_steps": 500,
1307
+ "stateful_callbacks": {
1308
+ "TrainerControl": {
1309
+ "args": {
1310
+ "should_epoch_stop": false,
1311
+ "should_evaluate": false,
1312
+ "should_log": false,
1313
+ "should_save": false,
1314
+ "should_training_stop": false
1315
+ },
1316
+ "attributes": {}
1317
+ }
1318
+ },
1319
+ "total_flos": 2.157629021565309e+19,
1320
+ "train_batch_size": 2,
1321
+ "trial_name": null,
1322
+ "trial_params": null
1323
+ }
checkpoint-1800/training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:36e6279f3948829ea711876c2172486c9470f648c9ce969627b2e58872c6b5f7
3
+ size 5713