Image-Text-to-Text
PEFT
Safetensors
vision-language
multimodal
llava
lora
siglip2
n-atlas
nigerian-languages
UncleanCode commited on
Commit
2a2540a
·
verified ·
1 Parent(s): afb7034

Upload via givemeanode export_data

Browse files
README.md CHANGED
@@ -1,3 +1,277 @@
1
  ---
2
- license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ license: other
3
+ license_name: awarri-open-source-research-and-innovation-license
4
+ license_link: https://huggingface.co/NCAIR1/N-ATLaS
5
+ base_model:
6
+ - NCAIR1/N-ATLaS
7
+ - google/siglip2-base-patch16-224
8
+ library_name: peft
9
+ pipeline_tag: image-text-to-text
10
+ language:
11
+ - en
12
+ - ha
13
+ - ig
14
+ - yo
15
+ datasets:
16
+ - liuhaotian/LLaVA-Pretrain
17
+ - liuhaotian/LLaVA-Instruct-150K
18
+ - lmms-lab/POPE
19
+ tags:
20
+ - vision-language
21
+ - multimodal
22
+ - llava
23
+ - lora
24
+ - siglip2
25
+ - n-atlas
26
+ - nigerian-languages
27
  ---
28
+
29
+ # AtlasVision
30
+
31
+ **AtlasVision** gives [N-ATLaS](https://huggingface.co/NCAIR1/N-ATLaS) — Nigeria's multilingual Llama-3-8B model for
32
+ English, Hausa, Igbo and Yoruba — the ability to see. It connects a frozen
33
+ [SigLIP2](https://huggingface.co/google/siglip2-base-patch16-224) image encoder to N-ATLaS through a trained MLP projector,
34
+ following the two-stage LLaVA recipe:
35
+
36
+ 1. **Stage 1 — alignment:** only the projector is trained, on 300k image–caption pairs, so the LLM can "read" image features.
37
+ 2. **Stage 2 — visual instruction tuning:** the projector keeps training and LoRA adapters are added to N-ATLaS, on
38
+ LLaVA-Instruct-150K conversations, so the model answers questions and holds conversations about images.
39
+
40
+ This repository contains **only the trained parts** (projectors and LoRA adapter, ~0.9 GB).
41
+ The base models are downloaded from their own repositories at load time, so their licenses and access conditions
42
+ (N-ATLaS is gated) continue to apply.
43
+
44
+ ## Model summary
45
+
46
+ | | |
47
+ |---|---|
48
+ | Language model | `NCAIR1/N-ATLaS` (Llama-3 8B, frozen; LoRA in stage 2) |
49
+ | Vision encoder | `google/siglip2-base-patch16-224` (ViT-B/16, 224 px, 196 patch tokens, frozen) |
50
+ | Projector | 2-layer MLP 768 → 4096 → 4096 with GELU (19.9M params) |
51
+ | LoRA (stage 2) | r = 64, alpha = 128, dropout 0.05 on q/k/v/o/gate/up/down projections (167.8M params) |
52
+ | Image tokens | 196 per image, inserted after `<|begin_of_text|>` and the user header |
53
+ | Precision | bf16 weights and compute; projector and LoRA trained in fp32 |
54
+ | Languages | Training data is English; N-ATLaS contributes Hausa, Igbo and Yoruba |
55
+
56
+ ## Repository contents
57
+
58
+ ```
59
+ stage1/projector.safetensors # stage-1 projector (image captioning / alignment)
60
+ stage2/projector.safetensors # stage-2 projector (use together with the LoRA adapter)
61
+ stage2/lora_adapter/ # PEFT LoRA adapter for NCAIR1/N-ATLaS
62
+ chat.py # standalone inference script (CLI + Python API)
63
+ code/ # exact training, evaluation and launch scripts used
64
+ eval/ # raw evaluation results (JSON) and report
65
+ logs/ # per-step training logs (loss, grad norm, LR, speed)
66
+ ```
67
+
68
+ ## Quick start
69
+
70
+ Accept the conditions on the [N-ATLaS model page](https://huggingface.co/NCAIR1/N-ATLaS) first, then:
71
+
72
+ ```bash
73
+ pip install -U torch transformers peft safetensors pillow accelerate huggingface_hub
74
+ export HF_TOKEN=hf_... # a token from the account that was granted N-ATLaS access
75
+ wget https://huggingface.co/FUTO-NIGERIA/AtlasVision/resolve/main/chat.py
76
+
77
+ python chat.py --image photo.jpg --question "What is happening in this picture?"
78
+ python chat.py --image photo.jpg --question "Kedu ihe dị na foto a?" # Igbo
79
+ python chat.py --stage 1 --image photo.jpg # stage-1 captioner
80
+ python chat.py --image photo.jpg --interactive # follow-up questions
81
+ python chat.py --image photo.jpg --load-in-4bit # ~8 GB GPU (pip install bitsandbytes)
82
+ ```
83
+
84
+ bf16 needs about 18 GB of GPU memory (A100, L4, RTX 4090…). `--load-in-4bit` runs on ~8 GB GPUs such as a free Colab T4.
85
+
86
+ From Python:
87
+
88
+ ```python
89
+ from chat import AtlasVision, load_image
90
+
91
+ model = AtlasVision(stage=2) # downloads N-ATLaS, SigLIP2 and this repo's weights
92
+ image = load_image("https://example.com/street.jpg")
93
+ print(model.ask(image, "Describe this image in detail."))
94
+ print(model.ask(image, "How many people are there?")) # follow-ups keep the conversation
95
+ ```
96
+
97
+ ### Prompt format
98
+
99
+ The image embeddings are spliced into the Llama-3 chat format:
100
+
101
+ ```
102
+ <|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n [196 image embeddings] {question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n
103
+ ```
104
+
105
+ Follow-up turns append `{answer}<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n{question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n`.
106
+ `chat.py` builds this for you.
107
+
108
+ ## Training
109
+
110
+ Both stages ran on a single NVIDIA H100 80GB.
111
+
112
+ | | Stage 1 — alignment | Stage 2 — instruction tuning |
113
+ |---|---|---|
114
+ | Data | LLaVA-Pretrain (BLIP captions of LAION/CC/SBU), random 300k of 558k | LLaVA-Instruct-150K on COCO train2017 (156,712 train / 1,000 held out) |
115
+ | Trained | projector | projector + LoRA |
116
+ | Effective batch | 32 (16 × 2 accumulation) | 32 (8 × 4 accumulation) |
117
+ | Learning rate | 2e-4, 3% warmup, cosine | LoRA 2e-4, projector 2e-5, 3% warmup, cosine |
118
+ | Max text length | 128 tokens | 1,024 tokens |
119
+ | Optimizer steps | 9375 | 4897 |
120
+ | Loss (first log → last log) | 8.551 → 2.454 | 2.102 → 1.15 |
121
+ | Wall-clock time | ~91 min | ~206 min |
122
+ | Peak GPU memory | 45.2 GB | 35.9 GB (gradient checkpointing) |
123
+
124
+ Other details: AdamW (no weight decay), gradient clipping at 1.0, bf16 autocast, loss only on assistant tokens,
125
+ length-bucketed batches in stage 2. Per-step logs are in `logs/`.
126
+
127
+ ## Evaluation
128
+
129
+ ### Stage 1 — does the model actually use the image?
130
+
131
+ Measured on 2,000 LLaVA-Pretrain captions that were **not** in the training subset (caption loss, lower is better):
132
+
133
+ | Setting | Held-out loss |
134
+ |---|---|
135
+ | Untrained (random) projector | 8.7249 |
136
+ | **Trained stage-1 projector** | **2.4523** |
137
+ | Trained projector, captions paired with the wrong images | 5.4332 |
138
+
139
+ Held-out loss matches the final training loss (no over-fitting), and pairing captions with the wrong images
140
+ roughly doubles the loss — the language model is relying on the visual content, not guessing generic captions.
141
+
142
+ ### Stage 1 vs stage 2
143
+
144
+ | Metric | Stage 1 | Stage 2 |
145
+ |---|---|---|
146
+ | Held-out LLaVA-Instruct loss (500 unseen conversations, lower is better) | 2.2521 | **1.1271** |
147
+
148
+ **POPE** (object hallucination; yes/no questions about COCO val2014 images, scored from the Yes/No token probabilities;
149
+ balanced 50% yes / 50% no, so a yes-ratio near 0.5 is ideal):
150
+
151
+ | POPE split | Stage 1: accuracy / F1 / yes-ratio | Stage 2: accuracy / F1 / yes-ratio | Stage 2: precision / recall |
152
+ |---|---|---|---|
153
+ | random | 0.500 / 0.200 / 0.13 | **0.543** / **0.686** / 0.95 | 0.522 / 0.997 |
154
+ | popular | 0.503 / 0.201 / 0.12 | **0.522** / **0.676** / 0.98 | 0.511 / 0.997 |
155
+ | adversarial | 0.502 / 0.201 / 0.12 | **0.514** / **0.672** / 0.98 | 0.507 / 0.997 |
156
+
157
+ **How to read this: stage 2 has a strong "yes" bias.** It answers "yes" to 95%–98% of POPE
158
+ questions, so it almost never misses an object that is present (recall ≈ 1.0) but also confirms most objects that are
159
+ *absent*, leaving accuracy close to chance. Stage 1 shows the opposite bias (it was never trained to answer questions).
160
+ This is a known effect of training on LLaVA-Instruct-150K alone: its conversations almost never contain questions whose
161
+ answer is "no". The descriptive examples below are far stronger than these numbers suggest, but **do not use AtlasVision
162
+ for yes/no verification of image content**. The usual fix is to add short-answer VQA data (VQAv2, GQA) and negative
163
+ examples (e.g. LRV-Instruction) to stage 2, as LLaVA-1.5 does — it reaches POPE F1 ≈ 0.86 with 665k mixed samples at 336 px.
164
+
165
+ ### Examples — stage 1 captions on unseen images
166
+
167
+ | Reference caption | AtlasVision stage 1 |
168
+ |---|---|
169
+ | cute homecoming prom dress with lace top and satin skirt | a - line sweetheart sweetheart lace tulle prom dress |
170
+ | a pair of silicone bracelets with a four codes message | a pair of silicone bracelets with the word code on them |
171
+ | a set of fingerprint icons in white on a black background | the iphone 6s and iphone 6s plus are shown in a black and white image |
172
+ | a business woman drawing modern concept of a website creation | a man writing the word modern web design on a whiteboard |
173
+ | the beach in el nido national park, puerto puerto | the limestone cliffs and limestone islands in the background |
174
+ | three pieces of paper with the words, democratic decentified dp controlled centralized ccp | a diagram showing the different types of democracy |
175
+
176
+ ### Examples — "Describe this image in detail." (COCO val2014, unseen)
177
+
178
+ **COCO_val2014_000000310196**
179
+
180
+ - *Stage 1:* a skier in the snow on a mountain
181
+ - *Stage 2:* The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.
182
+
183
+ There are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together.
184
+
185
+ **COCO_val2014_000000210789**
186
+
187
+ - *Stage 1:* a woman and her child in the rain
188
+ - *Stage 2:* The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.
189
+
190
+ In the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene.
191
+
192
+ **COCO_val2014_000000429109**
193
+
194
+ - *Stage 1:* a bus and several other vehicles parked in front of a building
195
+ - *Stage 2:* The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.
196
+
197
+ Numerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.
198
+
199
+ In addition to the vehicles, there are traffic lights at various points along the street
200
+
201
+ **COCO_val2014_000000211674**
202
+
203
+ - *Stage 1:* a bus with a red and white logo on it, carrying passengers
204
+ - *Stage 2:* The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.
205
+
206
+ In addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by.
207
+
208
+ ### Questions in Nigerian languages (stage 2)
209
+
210
+ The stage-2 instruction data is English only. The model understands the questions below (its answers match the image) but replies in English; adding translated instruction data would be needed for answers in these languages.
211
+
212
+ | Language | Question | Answer |
213
+ |---|---|---|
214
+ | Igbo | Kedu ihe dị na foto a? | The image features a person skiing down a snow-covered slope, with the skier wearing red pants. |
215
+ | Yoruba | Kí ni ó wà nínú àwòrán yìí? | The image features a person skiing down a snow-covered slope, with the skier wearing red pants. |
216
+ | Hausa | Me ke cikin wannan hoton? | In the image, a person is skiing down a snow-covered slope. |
217
+
218
+ ### Text-only check (no image)
219
+
220
+ Does the stage-2 LoRA damage N-ATLaS's original text abilities? Same Igbo question, greedy decoding:
221
+
222
+ **Q:** Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.
223
+
224
+ - *Base N-ATLaS:* Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
225
+ - *With stage-2 LoRA:* Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu "positron" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej
226
+
227
+ Both answer fluently in Igbo, so the adapter keeps N-ATLaS's text abilities intact (responses are cut at 120 tokens).
228
+
229
+ ## Limitations
230
+
231
+ - **Hallucination and yes-bias.** Long descriptions are fluent but often add plausible details that are not in the
232
+ image (exact counts, extra cars, a handbag), and stage 2 answers "yes" to most yes/no questions (see POPE above).
233
+ Treat counts, small objects and yes/no answers as unreliable.
234
+ - **English-only visual training.** All image–text training data is English. Answers to Hausa, Igbo and Yoruba questions
235
+ come from N-ATLaS's own multilingual ability and are noticeably less reliable; they may switch to English.
236
+ - **Low resolution.** Images are resized to 224×224, so small text, fine details and dense documents are hard.
237
+ - **One image per conversation**, no video, no bounding boxes or grounding.
238
+ - **Data biases.** Web captions (LAION/CC/SBU) and COCO carry Western-centric content; performance on Nigerian scenes,
239
+ people and text has not been measured and is likely weaker.
240
+ - **Not for high-stakes use** (medical, legal, identity, surveillance or safety-critical decisions).
241
+
242
+ ## Licenses and terms
243
+
244
+ This repository contains weights trained on top of other people's models and data. Using it means complying with all of:
245
+
246
+ - **N-ATLaS**: Awarri Open-Source Research and Innovation License (attribution required; caps deployments at 1,000
247
+ active end-users per 30 days) and the Llama 3 Community License it inherits. Access is gated — accept the
248
+ conditions on the [model page](https://huggingface.co/NCAIR1/N-ATLaS).
249
+ - **SigLIP2**: Apache-2.0.
250
+ - **LLaVA-Instruct-150K** (stage 2 data): CC BY-NC 4.0, generated with GPT-4 and subject to OpenAI's terms —
251
+ **stage-2 weights are for non-commercial research use.**
252
+ - **LLaVA-Pretrain** (stage 1 data): captions/images from LAION, Conceptual Captions and SBU under their respective terms.
253
+ - **COCO images**: Flickr images under their individual Creative Commons licenses.
254
+
255
+ ## Acknowledgements
256
+
257
+ - N-ATLaS by Awarri Technologies and the Federal Ministry of Communications, Innovation and Digital Economy (Nigeria),
258
+ as part of the Nigerian Languages AI Initiative.
259
+ - SigLIP2 by Google; LLaVA data and recipe by Liu et al.; POPE by Li et al.
260
+ - Trained on an H100 from San Francisco Compute (Autoresearch).
261
+
262
+ ## Citation
263
+
264
+ ```bibtex
265
+ @misc{atlasvision2026,
266
+ title = {AtlasVision: a vision-language extension of N-ATLaS},
267
+ author = {FUTO-NIGERIA},
268
+ year = {2026},
269
+ url = {https://huggingface.co/FUTO-NIGERIA/AtlasVision}
270
+ }
271
+ @inproceedings{liu2023llava,
272
+ title = {Visual Instruction Tuning},
273
+ author = {Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae},
274
+ booktitle = {NeurIPS},
275
+ year = {2023}
276
+ }
277
+ ```
chat.py ADDED
@@ -0,0 +1,156 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """AtlasVision inference: SigLIP2 vision encoder + N-ATLaS (Llama-3 8B) via a trained MLP projector.
3
+
4
+ pip install -U torch transformers peft safetensors pillow accelerate huggingface_hub
5
+ # optional for 4-bit on small GPUs (e.g. Colab T4): pip install bitsandbytes
6
+ export HF_TOKEN=hf_... # needs accepted access to the gated NCAIR1/N-ATLaS
7
+
8
+ python chat.py --image photo.jpg --question "What is happening in this picture?" # stage 2
9
+ python chat.py --stage 1 --image photo.jpg # stage 1 captioner
10
+ python chat.py --image https://example.com/cat.jpg --question "Kedu ihe dị na foto a?"
11
+ python chat.py --image photo.jpg --load-in-4bit # ~8 GB GPU
12
+ python chat.py --image photo.jpg --interactive # several questions
13
+
14
+ Needs ~18 GB of GPU memory in bf16 (A100, L4, RTX 4090...); --load-in-4bit fits ~8 GB GPUs such as a Colab T4.
15
+ """
16
+ import argparse
17
+ import io
18
+ import os
19
+ import sys
20
+
21
+ import torch
22
+ import torch.nn as nn
23
+ from PIL import Image
24
+
25
+ REPO = "FUTO-NIGERIA/AtlasVision"
26
+ LLM = "NCAIR1/N-ATLaS"
27
+ VISION = "google/siglip2-base-patch16-224"
28
+ USER_HEADER = "<|start_header_id|>user<|end_header_id|>\n\n"
29
+ ASSIST_HEADER = "<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"
30
+ EOT = "<|eot_id|>"
31
+
32
+
33
+ class ProjectionMLP(nn.Module):
34
+ def __init__(self, vision_dim, text_dim):
35
+ super().__init__()
36
+ self.net = nn.Sequential(nn.Linear(vision_dim, text_dim), nn.GELU(), nn.Linear(text_dim, text_dim))
37
+
38
+ def forward(self, x):
39
+ return self.net(x)
40
+
41
+
42
+ def fetch(repo, filename):
43
+ if os.path.isdir(repo):
44
+ return os.path.join(repo, filename)
45
+ from huggingface_hub import hf_hub_download
46
+ return hf_hub_download(repo, filename)
47
+
48
+
49
+ def load_image(src):
50
+ if src.startswith(("http://", "https://")):
51
+ import urllib.request
52
+ with urllib.request.urlopen(src) as r:
53
+ return Image.open(io.BytesIO(r.read())).convert("RGB")
54
+ return Image.open(src).convert("RGB")
55
+
56
+
57
+ class AtlasVision:
58
+ def __init__(self, stage=2, repo=REPO, llm=LLM, vision=VISION, load_in_4bit=False, device=None):
59
+ from safetensors.torch import load_file
60
+ from transformers import AutoImageProcessor, AutoModel, AutoModelForCausalLM, AutoTokenizer
61
+
62
+ self.device = torch.device(device or ("cuda" if torch.cuda.is_available() else "cpu"))
63
+ cuda = self.device.type == "cuda"
64
+ self.dtype = torch.bfloat16 if (not cuda or torch.cuda.is_bf16_supported()) else torch.float16
65
+
66
+ self.tok = AutoTokenizer.from_pretrained(llm)
67
+ self.pad_id = self.tok.pad_token_id if self.tok.pad_token_id is not None else self.tok.eos_token_id
68
+ self.eot_id = self.tok.convert_tokens_to_ids(EOT)
69
+ self.processor = AutoImageProcessor.from_pretrained(vision)
70
+
71
+ full_vision = AutoModel.from_pretrained(vision, dtype=self.dtype)
72
+ self.vision = full_vision.vision_model.to(self.device).eval()
73
+ vision_dim = full_vision.config.vision_config.hidden_size
74
+ del full_vision
75
+
76
+ kw = {"dtype": self.dtype}
77
+ if load_in_4bit:
78
+ from transformers import BitsAndBytesConfig
79
+ kw["quantization_config"] = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
80
+ bnb_4bit_compute_dtype=self.dtype)
81
+ kw["device_map"] = {"": self.device.index or 0}
82
+ self.llm = AutoModelForCausalLM.from_pretrained(llm, **kw)
83
+ if not load_in_4bit:
84
+ self.llm.to(self.device)
85
+ text_dim = self.llm.config.hidden_size
86
+
87
+ if stage == 2:
88
+ from peft import PeftModel
89
+ if os.path.isdir(repo):
90
+ self.llm = PeftModel.from_pretrained(self.llm, os.path.join(repo, "stage2/lora_adapter"))
91
+ else:
92
+ self.llm = PeftModel.from_pretrained(self.llm, repo, subfolder="stage2/lora_adapter")
93
+ self.llm.eval()
94
+
95
+ self.projector = ProjectionMLP(vision_dim, text_dim)
96
+ self.projector.load_state_dict(load_file(fetch(repo, f"stage{stage}/projector.safetensors")))
97
+ self.projector.to(self.device, dtype=torch.float32).eval()
98
+
99
+ self.prefix = torch.tensor([self.tok(USER_HEADER, add_special_tokens=True).input_ids], device=self.device)
100
+ self.history = [] # (question, answer) turns about the current image
101
+
102
+ @torch.no_grad()
103
+ def ask(self, image, question, max_new_tokens=256, temperature=0.0):
104
+ pv = self.processor(images=image, return_tensors="pt").pixel_values.to(self.device, self.dtype)
105
+ img = self.projector(self.vision(pixel_values=pv).last_hidden_state.float()).to(self.dtype)
106
+
107
+ text = ""
108
+ for i, (q, a) in enumerate(self.history):
109
+ text += (q if i == 0 else USER_HEADER + q) + ASSIST_HEADER + a + EOT
110
+ text += (question if not self.history else USER_HEADER + question) + ASSIST_HEADER
111
+ ids = torch.tensor([self.tok(text, add_special_tokens=False).input_ids], device=self.device)
112
+
113
+ emb = self.llm.get_input_embeddings()
114
+ embeds = torch.cat([emb(self.prefix).to(self.dtype), img, emb(ids).to(self.dtype)], dim=1)
115
+ mask = torch.ones(embeds.shape[:2], dtype=torch.long, device=self.device)
116
+ gen = dict(max_new_tokens=max_new_tokens, repetition_penalty=1.1, eos_token_id=self.eot_id, pad_token_id=self.pad_id)
117
+ if temperature > 0:
118
+ gen.update(do_sample=True, temperature=temperature, top_p=0.9)
119
+ else:
120
+ gen.update(do_sample=False)
121
+ out = self.llm.generate(inputs_embeds=embeds, attention_mask=mask, **gen)
122
+ answer = self.tok.decode(out[0], skip_special_tokens=True).strip()
123
+ self.history.append((question, answer))
124
+ return answer
125
+
126
+
127
+ def main():
128
+ ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
129
+ ap.add_argument("--image", required=True, help="path or URL")
130
+ ap.add_argument("--question", default=None)
131
+ ap.add_argument("--stage", type=int, default=2, choices=[1, 2])
132
+ ap.add_argument("--repo", default=REPO, help="HF repo id, or a local folder containing stage1/ and stage2/")
133
+ ap.add_argument("--llm", default=LLM)
134
+ ap.add_argument("--vision", default=VISION)
135
+ ap.add_argument("--load-in-4bit", action="store_true")
136
+ ap.add_argument("--max-new-tokens", type=int, default=256)
137
+ ap.add_argument("--temperature", type=float, default=0.0)
138
+ ap.add_argument("--interactive", action="store_true")
139
+ a = ap.parse_args()
140
+
141
+ question = a.question or ("Describe this image briefly." if a.stage == 1 else "Describe this image in detail.")
142
+ model = AtlasVision(a.stage, a.repo, a.llm, a.vision, a.load_in_4bit)
143
+ image = load_image(a.image)
144
+ print(f"\nQ: {question}\nA: {model.ask(image, question, a.max_new_tokens, a.temperature)}", flush=True)
145
+ while a.interactive:
146
+ try:
147
+ q = input("\nQ (empty to quit): ").strip()
148
+ except EOFError:
149
+ break
150
+ if not q:
151
+ break
152
+ print(f"A: {model.ask(image, q, a.max_new_tokens, a.temperature)}", flush=True)
153
+
154
+
155
+ if __name__ == "__main__":
156
+ sys.exit(main())
code/eval_stage1.py ADDED
@@ -0,0 +1,85 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Stage-1 verification on samples NOT seen in training.
3
+ 1) checkpoint integrity 2) held-out loss: random projector vs trained vs trained-with-shuffled-images
4
+ 3) captions on held-out images 4) COCO zip readiness for stage 2"""
5
+ import json, os, time, zipfile
6
+ os.environ.setdefault("HF_HUB_OFFLINE", "1")
7
+ import torch
8
+ from torch.utils.data import DataLoader
9
+ import train as T
10
+
11
+ OUT = os.path.expanduser("~/eval"); os.makedirs(OUT, exist_ok=True)
12
+ res = {}
13
+ t0 = time.time()
14
+
15
+ # 1. checkpoint integrity
16
+ sd = torch.load(os.path.join(T.CKPT_DIR, "projector_final.pt"), map_location="cpu")
17
+ res["checkpoint"] = {k: list(v.shape) for k, v in sd.items()}
18
+ res["checkpoint_all_finite"] = all(torch.isfinite(v).all().item() for v in sd.values())
19
+ print("checkpoint:", res["checkpoint"], "finite:", res["checkpoint_all_finite"], flush=True)
20
+
21
+ model, tok, pad_id, proc = T.build_model()
22
+ ann = json.load(open(T.JSON_PATH))
23
+ perm = torch.randperm(len(ann), generator=torch.Generator().manual_seed(T.SEED)).tolist()
24
+ held = [ann[i] for i in perm[T.NUM_SAMPLES:T.NUM_SAMPLES + 2000]] # never trained on
25
+ prefix = T.find_zip_prefix(held, T.ZIP_PATH)
26
+ ds = T.LLaVAPretrainDataset(held, T.ZIP_PATH, prefix, proc, tok, T.MAX_TEXT_LEN)
27
+ dl = DataLoader(ds, batch_size=32, num_workers=8, collate_fn=T.make_collate(pad_id))
28
+
29
+
30
+ @torch.no_grad()
31
+ def val_loss(shuffle_images=False):
32
+ model.eval(); tot, n = 0.0, 0
33
+ for b in dl:
34
+ pv = b["pixel_values"].cuda()
35
+ if shuffle_images:
36
+ pv = pv.roll(1, dims=0) # each caption paired with a different image
37
+ with torch.autocast("cuda", dtype=T.DTYPE):
38
+ l = model(pv, b["input_ids"].cuda(), b["attention_mask"].cuda(), b["labels"].cuda()).float()
39
+ k = int((b["labels"] != -100).sum()); tot += l.item() * k; n += k
40
+ return round(tot / n, 4)
41
+
42
+
43
+ # 2. held-out loss
44
+ torch.manual_seed(0)
45
+ model.projector = T.ProjectionMLP(model.projector.net[0].in_features, model.projector.net[2].out_features).cuda()
46
+ res["heldout_loss_random_projector"] = val_loss()
47
+ model.projector.load_state_dict(sd)
48
+ res["heldout_loss_trained"] = val_loss()
49
+ res["heldout_loss_trained_shuffled_images"] = val_loss(shuffle_images=True)
50
+ print("held-out loss:", {k: v for k, v in res.items() if k.startswith("heldout")}, flush=True)
51
+
52
+ # 3. captions on held-out images
53
+ eot = tok.convert_tokens_to_ids(T.EOT)
54
+ caps = []
55
+ for item in held[:8]:
56
+ img, ok = None, True
57
+ with zipfile.ZipFile(T.ZIP_PATH) as zf, zf.open(prefix + item["image"]) as f:
58
+ from PIL import Image
59
+ img = Image.open(f).convert("RGB")
60
+ pv = proc(images=img, return_tensors="pt").pixel_values.cuda()
61
+ ids = torch.tensor([tok("Describe this image briefly." + T.ASSIST_HEADER, add_special_tokens=False).input_ids], device="cuda")
62
+ with torch.no_grad(), torch.autocast("cuda", dtype=T.DTYPE):
63
+ e, m, _ = model.build_inputs(pv, ids, torch.ones_like(ids))
64
+ out = model.llm.generate(inputs_embeds=e, attention_mask=m, max_new_tokens=50, do_sample=False,
65
+ eos_token_id=eot, pad_token_id=pad_id)
66
+ caps.append({"image": item["image"], "reference": item["conversations"][1]["value"],
67
+ "model": tok.decode(out[0], skip_special_tokens=True).strip()})
68
+ res["captions"] = caps
69
+ for c in caps:
70
+ print(f"\n[{c['image']}]\n ref: {c['reference']}\n model: {c['model']}", flush=True)
71
+
72
+ # 4. stage-2 data readiness
73
+ with zipfile.ZipFile(os.path.expanduser("~/data/coco/train2017.zip")) as zf:
74
+ names = set(zf.namelist())
75
+ inst = json.load(open(os.path.expanduser("~/data/llava_instruct/llava_instruct_150k.json")))
76
+ hits = sum(("train2017/" + a["image"]) in names for a in inst[:2000])
77
+ res["coco_zip_files"] = len(names)
78
+ res["instruct_conversations"] = len(inst)
79
+ res["instruct_images_found_of_2000"] = hits
80
+ print("\nstage-2 data:", res["coco_zip_files"], "files in COCO zip;", len(inst), "conversations;",
81
+ hits, "/2000 images found", flush=True)
82
+
83
+ res["eval_minutes"] = round((time.time() - t0) / 60, 1)
84
+ json.dump(res, open(os.path.join(OUT, "stage1_eval.json"), "w"), indent=1)
85
+ print("EVAL_DONE", flush=True)
code/eval_stage2.py ADDED
@@ -0,0 +1,231 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Evaluate stage 1 vs stage 2 and write ~/eval/stage2_eval.json + ~/eval/report.md
3
+ 1) held-out LLaVA-Instruct loss 2) POPE (random / popular / adversarial): acc, precision, recall, F1, yes-ratio
4
+ 3) detailed descriptions on POPE images 4) image questions in Igbo / Yoruba / Hausa 5) text-only check of N-ATLaS
5
+ """
6
+ import glob
7
+ import io
8
+ import json
9
+ import os
10
+ import time
11
+
12
+ os.environ.setdefault("HF_HUB_OFFLINE", "1")
13
+ import torch
14
+ from PIL import Image
15
+ from torch.utils.data import DataLoader, Subset
16
+
17
+ import train as T
18
+ import train_stage2 as S
19
+
20
+ HOME = os.path.expanduser("~")
21
+ OUT = T.env("EVAL_DIR", f"{HOME}/eval")
22
+ POPE_DIR = T.env("POPE_DIR", f"{HOME}/data/pope")
23
+ POPE_LIMIT = T.env("POPE_LIMIT", 0, int) # 0 = all questions per split
24
+ HELD_N = T.env("HELD_N", 500, int)
25
+ os.makedirs(OUT, exist_ok=True)
26
+ DEV, DT = T.DEVICE, T.DTYPE
27
+ t0 = time.time()
28
+
29
+ proj1 = S._load_proj(S.STAGE1_PROJECTOR)
30
+ proj2 = torch.load(os.path.join(S.CKPT_DIR, "projector_stage2.pt"), map_location="cpu")
31
+ from peft import PeftModel
32
+
33
+ model, tok, pad_id, processor = T.build_model()
34
+ model.llm = PeftModel.from_pretrained(model.llm, os.path.join(S.CKPT_DIR, "lora_adapter")).to(DEV)
35
+ model.llm.eval()
36
+ model.eval()
37
+ EOT = tok.convert_tokens_to_ids(T.EOT)
38
+
39
+
40
+ class Stage:
41
+ """Context manager: stage 1 = stage-1 projector, adapters off; stage 2 = stage-2 projector, adapters on."""
42
+ def __init__(self, n):
43
+ self.n = n
44
+
45
+ def __enter__(self):
46
+ model.projector.load_state_dict(proj1 if self.n == 1 else proj2)
47
+ model.projector.to(DEV, dtype=torch.float32)
48
+ self.ctx = model.llm.disable_adapter() if self.n == 1 else None
49
+ if self.ctx:
50
+ self.ctx.__enter__()
51
+
52
+ def __exit__(self, *a):
53
+ if self.ctx:
54
+ self.ctx.__exit__(*a)
55
+
56
+
57
+ res = {"stage1": {}, "stage2": {}}
58
+
59
+ # ---------------- 1. held-out instruct loss ----------------
60
+ ann = json.load(open(S.INSTRUCT_JSON))
61
+ _, held = S.split(ann)
62
+ prefix = T.find_zip_prefix([ann[i] for i in held[:200]], S.COCO_ZIP)
63
+ ds = S.InstructDataset(ann, S.COCO_ZIP, prefix, processor, tok, S.MAX_TEXT_LEN)
64
+ dl = DataLoader(Subset(ds, held[:HELD_N]), batch_size=8, num_workers=T.NUM_WORKERS,
65
+ collate_fn=T.make_collate(pad_id))
66
+
67
+
68
+ @torch.no_grad()
69
+ def heldout_loss():
70
+ tot, n = 0.0, 0
71
+ for b in dl:
72
+ with torch.autocast(device_type=DEV.type, dtype=DT):
73
+ l = model(b["pixel_values"].to(DEV), b["input_ids"].to(DEV), b["attention_mask"].to(DEV),
74
+ b["labels"].to(DEV)).float()
75
+ k = int((b["labels"] != -100).sum())
76
+ tot, n = tot + l.item() * k, n + k
77
+ return round(tot / n, 4)
78
+
79
+
80
+ for s in (1, 2):
81
+ with Stage(s):
82
+ res[f"stage{s}"]["heldout_instruct_loss"] = heldout_loss()
83
+ T.log(f"held-out instruct loss: stage1 {res['stage1']['heldout_instruct_loss']} | "
84
+ f"stage2 {res['stage2']['heldout_instruct_loss']}")
85
+
86
+ # ---------------- 2. POPE ----------------
87
+ import pyarrow.parquet as pq
88
+
89
+ SUFFIX = " Answer the question using a single word or phrase."
90
+ yes_ids = sorted({tok(w, add_special_tokens=False).input_ids[0] for w in ("Yes", "yes", " Yes", " yes")})
91
+ no_ids = sorted({tok(w, add_special_tokens=False).input_ids[0] for w in ("No", "no", " No", " no")})
92
+
93
+
94
+ def to_image(cell):
95
+ if isinstance(cell, dict):
96
+ cell = cell.get("bytes") or open(cell["path"], "rb").read()
97
+ return Image.open(io.BytesIO(cell)).convert("RGB")
98
+
99
+
100
+ def pope_rows(split):
101
+ files = sorted(glob.glob(f"{POPE_DIR}/**/{split}-*.parquet", recursive=True))
102
+ if files:
103
+ rows = pq.read_table(files[0]).to_pylist()
104
+ else: # some versions ship one 'test' table with a 'category' column
105
+ rows = [r for f in sorted(glob.glob(f"{POPE_DIR}/**/test-*.parquet", recursive=True))
106
+ for r in pq.read_table(f).to_pylist() if r.get("category") == split]
107
+ return rows[:POPE_LIMIT] if POPE_LIMIT else rows
108
+
109
+
110
+ @torch.no_grad()
111
+ def pope_eval(rows, bs=32):
112
+ tp = fp = tn = fn = 0
113
+ for s in range(0, len(rows), bs):
114
+ chunk = rows[s:s + bs]
115
+ pv = torch.stack([processor(images=to_image(r["image"]), return_tensors="pt").pixel_values[0] for r in chunk]).to(DEV)
116
+ seqs = [tok(r["question"].strip() + SUFFIX + T.ASSIST_HEADER, add_special_tokens=False).input_ids for r in chunk]
117
+ L = max(map(len, seqs))
118
+ ids = torch.full((len(seqs), L), pad_id, dtype=torch.long)
119
+ mask = torch.zeros_like(ids)
120
+ for i, q in enumerate(seqs):
121
+ ids[i, :len(q)] = torch.tensor(q)
122
+ mask[i, :len(q)] = 1
123
+ ids, mask = ids.to(DEV), mask.to(DEV)
124
+ with torch.autocast(device_type=DEV.type, dtype=DT):
125
+ e, m, _ = model.build_inputs(pv, ids, mask)
126
+ logits = model.llm(inputs_embeds=e, attention_mask=m).logits
127
+ n_fixed = e.shape[1] - L
128
+ for i, r in enumerate(chunk):
129
+ last = logits[i, n_fixed + len(seqs[i]) - 1].float()
130
+ pred_yes = last[yes_ids].max() > last[no_ids].max()
131
+ gold_yes = str(r["answer"]).strip().lower().startswith("yes")
132
+ tp += pred_yes and gold_yes
133
+ fp += pred_yes and not gold_yes
134
+ tn += (not pred_yes) and (not gold_yes)
135
+ fn += (not pred_yes) and gold_yes
136
+ tp, fp, tn, fn = map(int, (tp, fp, tn, fn))
137
+ n = tp + fp + tn + fn
138
+ prec = tp / max(1, tp + fp)
139
+ rec = tp / max(1, tp + fn)
140
+ return {"n": n, "accuracy": round((tp + tn) / max(1, n), 4), "precision": round(prec, 4),
141
+ "recall": round(rec, 4), "f1": round(2 * prec * rec / max(1e-9, prec + rec), 4),
142
+ "yes_ratio": round((tp + fp) / max(1, n), 4)}
143
+
144
+
145
+ pope_samples = []
146
+ for split in ("random", "popular", "adversarial"):
147
+ rows = pope_rows(split)
148
+ if not rows:
149
+ T.log(f"POPE {split}: no data found")
150
+ continue
151
+ pope_samples = pope_samples or rows
152
+ for s in (1, 2):
153
+ with Stage(s):
154
+ res[f"stage{s}"][f"pope_{split}"] = pope_eval(rows)
155
+ T.log(f"POPE {split}: stage1 {res['stage1'][f'pope_{split}']} | stage2 {res['stage2'][f'pope_{split}']}")
156
+
157
+
158
+ # ---------------- 3/4. generations ----------------
159
+ @torch.no_grad()
160
+ def answer(image, question, max_new_tokens=120):
161
+ pv = processor(images=image, return_tensors="pt").pixel_values.to(DEV)
162
+ ids = torch.tensor([tok(question + T.ASSIST_HEADER, add_special_tokens=False).input_ids], device=DEV)
163
+ with torch.autocast(device_type=DEV.type, dtype=DT):
164
+ e, m, _ = model.build_inputs(pv, ids, torch.ones_like(ids))
165
+ out = model.llm.generate(inputs_embeds=e, attention_mask=m, max_new_tokens=max_new_tokens,
166
+ do_sample=False, eos_token_id=EOT, pad_token_id=pad_id, repetition_penalty=1.1)
167
+ return tok.decode(out[0], skip_special_tokens=True).strip()
168
+
169
+
170
+ seen, gen_images = set(), []
171
+ for r in pope_samples:
172
+ key = r.get("image_source") or r.get("question_id")
173
+ if key not in seen:
174
+ seen.add(key)
175
+ gen_images.append((str(key), to_image(r["image"])))
176
+ if len(gen_images) == 4:
177
+ break
178
+
179
+ res["descriptions"] = []
180
+ for key, img in gen_images:
181
+ row = {"image": key}
182
+ for s in (1, 2):
183
+ with Stage(s):
184
+ row[f"stage{s}"] = answer(img, "Describe this image in detail.")
185
+ res["descriptions"].append(row)
186
+
187
+ MULTI = {"igbo": "Kedu ihe dị na foto a?", "yoruba": "Kí ni ó wà nínú àwòrán yìí?", "hausa": "Me ke cikin wannan hoton?"}
188
+ res["multilingual"] = []
189
+ if gen_images:
190
+ with Stage(2):
191
+ for lang, q in MULTI.items():
192
+ res["multilingual"].append({"image": gen_images[0][0], "language": lang, "question": q,
193
+ "stage2": answer(gen_images[0][1], q)})
194
+
195
+ # ---------------- 5. text-only check ----------------
196
+ @torch.no_grad()
197
+ def text_only(q, max_new_tokens=120):
198
+ ids = tok(T.USER_HEADER + q + T.ASSIST_HEADER, add_special_tokens=True, return_tensors="pt").input_ids.to(DEV)
199
+ with torch.autocast(device_type=DEV.type, dtype=DT):
200
+ out = model.llm.generate(input_ids=ids, attention_mask=torch.ones_like(ids), max_new_tokens=max_new_tokens,
201
+ do_sample=False, eos_token_id=EOT, pad_token_id=pad_id, repetition_penalty=1.1)
202
+ return tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip()
203
+
204
+
205
+ q_text = "Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo."
206
+ with Stage(1):
207
+ base_answer = text_only(q_text)
208
+ with Stage(2):
209
+ lora_answer = text_only(q_text)
210
+ res["text_only"] = {"question": q_text, "base_n_atlas": base_answer, "with_stage2_lora": lora_answer}
211
+ res["eval_minutes"] = round((time.time() - t0) / 60, 1)
212
+ json.dump(res, open(os.path.join(OUT, "stage2_eval.json"), "w"), indent=1, ensure_ascii=False)
213
+
214
+ # ---------------- report ----------------
215
+ L = ["# Atlas-Vision evaluation\n", "## Scores (stage 1 → stage 2)\n", "| Metric | Stage 1 | Stage 2 |", "|---|---|---|",
216
+ f"| Held-out instruct loss (lower is better) | {res['stage1']['heldout_instruct_loss']} | {res['stage2']['heldout_instruct_loss']} |"]
217
+ for split in ("random", "popular", "adversarial"):
218
+ if f"pope_{split}" in res["stage1"]:
219
+ a, b = res["stage1"][f"pope_{split}"], res["stage2"][f"pope_{split}"]
220
+ L.append(f"| POPE {split}: accuracy / F1 / yes-ratio | {a['accuracy']} / {a['f1']} / {a['yes_ratio']} | "
221
+ f"{b['accuracy']} / {b['f1']} / {b['yes_ratio']} |")
222
+ L.append("\n## Detailed descriptions\n")
223
+ for d in res["descriptions"]:
224
+ L += [f"**{d['image']}**\n", f"- Stage 1: {d['stage1']}", f"- Stage 2: {d['stage2']}\n"]
225
+ L.append("## Questions in Nigerian languages (stage 2)\n")
226
+ for m in res["multilingual"]:
227
+ L.append(f"- **{m['language']}** — {m['question']}\n → {m['stage2']}")
228
+ L += ["\n## Text-only check (no image)\n", f"Question: {q_text}\n",
229
+ f"- Base N-ATLaS: {base_answer}", f"- With stage-2 LoRA: {lora_answer}"]
230
+ open(os.path.join(OUT, "report.md"), "w").write("\n".join(L) + "\n")
231
+ T.log(f"EVAL_DONE in {res['eval_minutes']} min -> {OUT}/stage2_eval.json, {OUT}/report.md")
code/infer.py ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Caption images with the trained projector.
3
+ Usage: python3 infer.py [projector.pt] [image paths...] (default: 8 held-out images from the zip)"""
4
+ import json
5
+ import os
6
+ import sys
7
+ import zipfile
8
+
9
+ import torch
10
+ from PIL import Image
11
+
12
+ os.environ.setdefault("HF_HUB_OFFLINE", "1")
13
+ import train as T # reuses the exact model layout and prompt format used in training
14
+
15
+ ckpt = sys.argv[1] if len(sys.argv) > 1 else os.path.join(T.CKPT_DIR, "projector_final.pt")
16
+ model, tok, pad_id, processor = T.build_model()
17
+ state = torch.load(ckpt, map_location="cpu")
18
+ state = state.get("projector_state_dict", state)
19
+ model.projector.load_state_dict(state)
20
+ model.eval()
21
+
22
+ if len(sys.argv) > 2:
23
+ images = [(p, Image.open(p).convert("RGB")) for p in sys.argv[2:]]
24
+ else: # images NOT in the training subset
25
+ ann = json.load(open(T.JSON_PATH))
26
+ g = torch.Generator().manual_seed(T.SEED)
27
+ held_out = torch.randperm(len(ann), generator=g)[T.NUM_SAMPLES:T.NUM_SAMPLES + 8].tolist()
28
+ zf = zipfile.ZipFile(T.ZIP_PATH)
29
+ prefix = T.find_zip_prefix([ann[i] for i in held_out], T.ZIP_PATH, n_check=8)
30
+ images = []
31
+ for i in held_out:
32
+ with zf.open(prefix + ann[i]["image"]) as f:
33
+ images.append((f"{ann[i]['image']} | reference: {ann[i]['conversations'][1]['value']}",
34
+ Image.open(f).convert("RGB")))
35
+
36
+ prompt = "Describe this image briefly."
37
+ for name, img in images:
38
+ pv = processor(images=img, return_tensors="pt").pixel_values.to(T.DEVICE)
39
+ ids = torch.tensor([tok(prompt + T.ASSIST_HEADER, add_special_tokens=False).input_ids], device=T.DEVICE)
40
+ mask = torch.ones_like(ids)
41
+ with torch.no_grad(), torch.autocast(device_type=T.DEVICE.type, dtype=T.DTYPE):
42
+ embeds, full_mask, _ = model.build_inputs(pv, ids, mask)
43
+ out = model.llm.generate(inputs_embeds=embeds, attention_mask=full_mask, max_new_tokens=60,
44
+ do_sample=False, eos_token_id=tok.convert_tokens_to_ids(T.EOT), pad_token_id=pad_id)
45
+ print(f"\n{name}\n -> {tok.decode(out[0], skip_special_tokens=True).strip()}", flush=True)
code/prepare.py ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Download everything onto the node's persistent disk. Safe to re-run (skips finished files).
3
+ Needs HF_TOKEN in the environment for the gated N-ATLaS repo."""
4
+ import os
5
+ import sys
6
+ import time
7
+
8
+ from huggingface_hub import hf_hub_download, snapshot_download
9
+
10
+ HOME = os.path.expanduser("~")
11
+ DATA_DIR = f"{HOME}/data/llava_pretrain"
12
+ LLM_NAME = os.environ.get("LLM_NAME", "NCAIR1/N-ATLaS")
13
+ VISION_NAME = os.environ.get("VISION_NAME", "google/siglip2-base-patch16-224")
14
+
15
+
16
+ def log(msg):
17
+ print(f"[{time.strftime('%H:%M:%S')}] {msg}", flush=True)
18
+
19
+
20
+ if not os.environ.get("HF_TOKEN"):
21
+ sys.exit("HF_TOKEN is not set - needed for the gated N-ATLaS repo")
22
+
23
+ # 1. Fail fast if the token can't see the gated model
24
+ try:
25
+ hf_hub_download(LLM_NAME, "config.json")
26
+ log(f"gated access OK for {LLM_NAME}")
27
+ except Exception as e:
28
+ sys.exit(f"Cannot access {LLM_NAME}: {type(e).__name__}: {e}\n"
29
+ "Check that access was granted on the model page and the token can read gated repos.")
30
+
31
+ t = time.time()
32
+ snapshot_download(LLM_NAME, allow_patterns=["*.json", "*.safetensors", "tokenizer*", "*.txt", "*.model", "*.jinja"])
33
+ log(f"LLM downloaded ({time.time() - t:.0f}s)")
34
+
35
+ t = time.time()
36
+ snapshot_download(VISION_NAME)
37
+ log(f"vision encoder downloaded ({time.time() - t:.0f}s)")
38
+
39
+ t = time.time()
40
+ os.makedirs(DATA_DIR, exist_ok=True)
41
+ hf_hub_download("liuhaotian/LLaVA-Pretrain", "blip_laion_cc_sbu_558k.json", repo_type="dataset", local_dir=DATA_DIR)
42
+ hf_hub_download("liuhaotian/LLaVA-Pretrain", "images.zip", repo_type="dataset", local_dir=DATA_DIR)
43
+ size_gb = os.path.getsize(f"{DATA_DIR}/images.zip") / 1e9
44
+ log(f"dataset downloaded: images.zip {size_gb:.1f} GB ({time.time() - t:.0f}s)")
code/run.sh ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # Idempotent: safe to re-run after a stop/wake. Resumes training from ~/checkpoints/stage1/latest.pt.
3
+ set -uo pipefail
4
+ cd "$HOME/atlas"
5
+ export PYTHONUNBUFFERED=1
6
+ export HF_HOME="$HOME/.cache/huggingface"
7
+ export HF_HUB_ENABLE_HF_TRANSFER=1
8
+
9
+ echo "=== [1/4] python packages ==="
10
+ PIPFLAGS="--user --break-system-packages"
11
+ python3 -c "import sys; sys.exit(0 if sys.prefix != sys.base_prefix else 1)" && PIPFLAGS=""
12
+ python3 -m pip install -q $PIPFLAGS -U "transformers>=4.56" accelerate "huggingface_hub>=0.34" hf_transfer pillow sentencepiece protobuf || exit 10
13
+ export PATH="$HOME/.local/bin:$PATH"
14
+ python3 -c "import torch, transformers; print('torch', torch.__version__, '| transformers', transformers.__version__, '| cuda', torch.cuda.is_available())"
15
+ nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
16
+
17
+ echo "=== [2/4] downloads ==="
18
+ python3 prepare.py || exit 11
19
+ df -h "$HOME" | tail -1
20
+
21
+ export HF_HUB_OFFLINE=1 # everything is cached now; training never needs the token
22
+ GC="${GRAD_CKPT:-0}"
23
+ if [ "${SKIP_SMOKE:-0}" != "1" ]; then
24
+ echo "=== [3/4] smoke test (20 steps) ==="
25
+ rm -rf "$HOME/checkpoints/smoke"
26
+ MAX_STEPS=20 LOG_EVERY=5 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/smoke" GRAD_CKPT=$GC python3 train.py
27
+ rc=$?
28
+ if [ $rc -eq 3 ] && [ "$GC" = "0" ]; then
29
+ echo "OOM without gradient checkpointing -> retrying with GRAD_CKPT=1"
30
+ GC=1
31
+ rm -rf "$HOME/checkpoints/smoke"
32
+ MAX_STEPS=20 LOG_EVERY=5 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/smoke" GRAD_CKPT=1 python3 train.py
33
+ rc=$?
34
+ fi
35
+ if [ $rc -ne 0 ]; then echo "SMOKE TEST FAILED (exit $rc)"; exit 12; fi
36
+ fi
37
+
38
+ echo "=== [4/4] full stage-1 run (GRAD_CKPT=$GC) ==="
39
+ GRAD_CKPT=$GC python3 train.py
40
+ rc=$?
41
+ if [ $rc -eq 3 ] && [ "$GC" = "0" ]; then
42
+ echo "OOM in full run -> resuming with GRAD_CKPT=1"
43
+ GRAD_CKPT=1 python3 train.py
44
+ rc=$?
45
+ fi
46
+ echo "=== run.sh finished (exit $rc) ==="
47
+ exit $rc
code/run_stage2.sh ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # Stage 2: smoke test -> full LoRA instruction tuning -> evaluation. Safe to re-run (resumes).
3
+ set -uo pipefail
4
+ cd "$HOME/atlas"
5
+ export PYTHONUNBUFFERED=1 HF_HOME="$HOME/.cache/huggingface" HF_HUB_OFFLINE=1 PATH="$HOME/.local/bin:$PATH"
6
+ python3 -m pip install -q --user --break-system-packages peft pyarrow || exit 10
7
+ python3 -c "import peft, pyarrow; print('peft', peft.__version__, '| pyarrow', pyarrow.__version__)"
8
+ BS="${BATCH_SIZE:-8}"; GA="${GRAD_ACCUM:-4}"
9
+
10
+ if [ "${SKIP_SMOKE:-0}" != "1" ]; then
11
+ echo "=== [1/3] smoke test (10 steps) ==="
12
+ rm -rf "$HOME/checkpoints/stage2_smoke"
13
+ MAX_STEPS=10 LOG_EVERY=2 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/stage2_smoke" BATCH_SIZE=$BS GRAD_ACCUM=$GA python3 train_stage2.py
14
+ rc=$?
15
+ if [ $rc -eq 3 ]; then
16
+ echo "OOM -> retrying with BATCH_SIZE=4 GRAD_ACCUM=8"; BS=4; GA=8
17
+ rm -rf "$HOME/checkpoints/stage2_smoke"
18
+ MAX_STEPS=10 LOG_EVERY=2 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/stage2_smoke" BATCH_SIZE=$BS GRAD_ACCUM=$GA python3 train_stage2.py
19
+ rc=$?
20
+ fi
21
+ if [ $rc -ne 0 ]; then echo "SMOKE TEST FAILED (exit $rc)"; exit 12; fi
22
+ fi
23
+
24
+ echo "=== [2/3] full stage-2 run (BATCH_SIZE=$BS GRAD_ACCUM=$GA) ==="
25
+ BATCH_SIZE=$BS GRAD_ACCUM=$GA python3 train_stage2.py
26
+ rc=$?
27
+ if [ $rc -eq 3 ] && [ "$BS" -gt 4 ]; then
28
+ echo "OOM in full run -> continuing with BATCH_SIZE=4 GRAD_ACCUM=8"
29
+ BATCH_SIZE=4 GRAD_ACCUM=8 python3 train_stage2.py
30
+ rc=$?
31
+ fi
32
+ if [ $rc -ne 0 ]; then echo "TRAINING FAILED (exit $rc)"; exit $rc; fi
33
+
34
+ echo "=== [3/3] evaluation ==="
35
+ python3 eval_stage2.py
36
+ rc=$?
37
+ echo "=== run_stage2.sh finished (exit $rc) ==="
38
+ exit $rc
code/train.py ADDED
@@ -0,0 +1,384 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Stage 1 (projector alignment): SigLIP2 vision tower + N-ATLaS LLM, only the MLP projector trains.
3
+
4
+ Fixes vs. the Colab notebook:
5
+ * LLM loaded in bf16 (notebook loaded fp16), bf16 autocast, fp32 projector, no GradScaler
6
+ * dynamic padding + attention mask (notebook padded every sample to 512 tokens; captions are ~10-25)
7
+ * image tokens placed after <BOS><user header>, not before BOS
8
+ * fixed random subset, cosine LR schedule, resume by samples seen (instant, batch-size independent)
9
+ * zip path check so missing images fail loudly instead of silently training on white images
10
+ All settings come from environment variables (see CONFIG below).
11
+ """
12
+ import json
13
+ import math
14
+ import os
15
+ import sys
16
+ import time
17
+ import zipfile
18
+
19
+ import torch
20
+ import torch.nn as nn
21
+ from PIL import Image
22
+ from torch.utils.data import DataLoader, Dataset, Subset
23
+ from transformers import AutoImageProcessor, AutoModel, AutoModelForCausalLM, AutoTokenizer
24
+
25
+
26
+ def env(name, default, cast=str):
27
+ v = os.environ.get(name)
28
+ return cast(v) if v not in (None, "") else default
29
+
30
+
31
+ # ----------------------------- CONFIG -----------------------------
32
+ HOME = os.path.expanduser("~")
33
+ LLM_NAME = env("LLM_NAME", "NCAIR1/N-ATLaS")
34
+ VISION_NAME = env("VISION_NAME", "google/siglip2-base-patch16-224")
35
+ DATA_DIR = env("DATA_DIR", f"{HOME}/data/llava_pretrain")
36
+ JSON_PATH = env("JSON_PATH", f"{DATA_DIR}/blip_laion_cc_sbu_558k.json")
37
+ ZIP_PATH = env("ZIP_PATH", f"{DATA_DIR}/images.zip")
38
+ CKPT_DIR = env("CKPT_DIR", f"{HOME}/checkpoints/stage1")
39
+ NUM_SAMPLES = env("NUM_SAMPLES", 150_000, int)
40
+ BATCH_SIZE = env("BATCH_SIZE", 16, int)
41
+ GRAD_ACCUM = env("GRAD_ACCUM", 1, int)
42
+ LR = env("LR", 2e-4, float)
43
+ WARMUP_RATIO = env("WARMUP_RATIO", 0.03, float)
44
+ MAX_TEXT_LEN = env("MAX_TEXT_LEN", 128, int)
45
+ MAX_STEPS = env("MAX_STEPS", 0, int) # >0 caps optimizer steps (smoke tests)
46
+ GRAD_CKPT = env("GRAD_CKPT", 0, int)
47
+ LOG_EVERY = env("LOG_EVERY", 25, int)
48
+ SAVE_EVERY = env("SAVE_EVERY", 500, int)
49
+ SEED = env("SEED", 42, int)
50
+ NUM_WORKERS = env("NUM_WORKERS", max(1, int(float(os.environ.get("GMN_CPU_LIMIT", "8"))) - 2), int)
51
+ DEVICE = torch.device(env("DEVICE", "cuda" if torch.cuda.is_available() else "cpu"))
52
+ DTYPE = torch.bfloat16
53
+
54
+ USER_HEADER = "<|start_header_id|>user<|end_header_id|>\n\n"
55
+ ASSIST_HEADER = "<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"
56
+ EOT = "<|eot_id|>"
57
+
58
+
59
+ def log(msg):
60
+ print(f"[{time.strftime('%H:%M:%S')}] {msg}", flush=True)
61
+
62
+
63
+ # ----------------------------- DATA -----------------------------
64
+ def find_zip_prefix(annotations, zip_path, n_check=500):
65
+ """Return the prefix to put before item['image'] inside the zip, verifying images exist."""
66
+ with zipfile.ZipFile(zip_path) as zf:
67
+ names = set(zf.namelist())
68
+ sample = [a["image"] for a in annotations[:n_check]]
69
+ prefix = ""
70
+ if sum(s in names for s in sample) < 0.95 * len(sample):
71
+ first = sample[0]
72
+ hits = [n for n in names if n.endswith("/" + first)]
73
+ if hits:
74
+ prefix = hits[0][: -len(first)]
75
+ found = sum((prefix + s) in names for s in sample)
76
+ log(f"zip check: {found}/{len(sample)} images found (prefix={prefix!r}, {len(names):,} files in zip)")
77
+ if found < 0.95 * len(sample):
78
+ raise SystemExit(f"Too many images missing from {zip_path}; example: {sample[0]!r}")
79
+ return prefix
80
+
81
+
82
+ class LLaVAPretrainDataset(Dataset):
83
+ def __init__(self, annotations, zip_path, zip_prefix, processor, tokenizer, max_text_len):
84
+ self.ann = annotations
85
+ self.zip_path = zip_path
86
+ self.zip_prefix = zip_prefix
87
+ self.processor = processor
88
+ self.tok = tokenizer
89
+ self.max_len = max_text_len
90
+ self.zip = None # opened lazily, once per dataloader worker
91
+
92
+ def __len__(self):
93
+ return len(self.ann)
94
+
95
+ def __getitem__(self, idx):
96
+ if self.zip is None:
97
+ self.zip = zipfile.ZipFile(self.zip_path)
98
+ item = self.ann[idx]
99
+ user, answer = "", ""
100
+ for turn in item["conversations"]:
101
+ if turn["from"] == "human":
102
+ user = turn["value"].replace("<image>", "").strip()
103
+ elif turn["from"] == "gpt":
104
+ answer = turn["value"].strip()
105
+ ok = True
106
+ try:
107
+ with self.zip.open(self.zip_prefix + item["image"]) as f:
108
+ image = Image.open(f).convert("RGB")
109
+ except Exception:
110
+ image = Image.new("RGB", (224, 224), "white")
111
+ ok = False
112
+ pixel_values = self.processor(images=image, return_tensors="pt").pixel_values[0]
113
+
114
+ q = self.tok(user + ASSIST_HEADER, add_special_tokens=False).input_ids
115
+ a = self.tok(answer + EOT, add_special_tokens=False).input_ids
116
+ ids = (q + a)[: self.max_len]
117
+ labels = ([-100] * len(q) + a)[: self.max_len]
118
+ return {"pixel_values": pixel_values, "input_ids": ids, "labels": labels, "ok": ok}
119
+
120
+
121
+ def make_collate(pad_id):
122
+ def collate(batch):
123
+ L = max(len(b["input_ids"]) for b in batch)
124
+ B = len(batch)
125
+ ids = torch.full((B, L), pad_id, dtype=torch.long)
126
+ labels = torch.full((B, L), -100, dtype=torch.long)
127
+ mask = torch.zeros((B, L), dtype=torch.long)
128
+ for i, b in enumerate(batch):
129
+ n = len(b["input_ids"])
130
+ ids[i, :n] = torch.tensor(b["input_ids"])
131
+ labels[i, :n] = torch.tensor(b["labels"])
132
+ mask[i, :n] = 1
133
+ return {
134
+ "pixel_values": torch.stack([b["pixel_values"] for b in batch]),
135
+ "input_ids": ids,
136
+ "labels": labels,
137
+ "attention_mask": mask,
138
+ "n_bad": sum(not b["ok"] for b in batch),
139
+ }
140
+ return collate
141
+
142
+
143
+ # ----------------------------- MODEL -----------------------------
144
+ class ProjectionMLP(nn.Module):
145
+ """Same layout as the notebook (net.0 / net.2), so Colab checkpoints load."""
146
+ def __init__(self, vision_dim, text_dim):
147
+ super().__init__()
148
+ self.net = nn.Sequential(nn.Linear(vision_dim, text_dim), nn.GELU(), nn.Linear(text_dim, text_dim))
149
+
150
+ def forward(self, x):
151
+ return self.net(x)
152
+
153
+
154
+ class AtlasVision(nn.Module):
155
+ """[<BOS> user-header] [image tokens] [user text, assistant header, answer]"""
156
+ def __init__(self, vision, llm, prefix_ids, vision_dim, text_dim):
157
+ super().__init__()
158
+ self.vision = vision
159
+ self.llm = llm
160
+ self.projector = ProjectionMLP(vision_dim, text_dim)
161
+ self.register_buffer("prefix_ids", torch.tensor(prefix_ids, dtype=torch.long)[None], persistent=False)
162
+
163
+ def build_inputs(self, pixel_values, input_ids, attention_mask, labels=None):
164
+ B = input_ids.size(0)
165
+ with torch.no_grad():
166
+ feats = self.vision(pixel_values=pixel_values.to(DTYPE)).last_hidden_state
167
+ img = self.projector(feats.float()).to(DTYPE)
168
+ emb = self.llm.get_input_embeddings()
169
+ pre = emb(self.prefix_ids.expand(B, -1)).to(DTYPE)
170
+ txt = emb(input_ids).to(DTYPE)
171
+ n_fixed = pre.size(1) + img.size(1)
172
+ inputs_embeds = torch.cat([pre, img, txt], dim=1)
173
+ mask = torch.cat([attention_mask.new_ones(B, n_fixed), attention_mask], dim=1)
174
+ full_labels = None
175
+ if labels is not None:
176
+ full_labels = torch.cat([labels.new_full((B, n_fixed), -100), labels], dim=1)
177
+ return inputs_embeds, mask, full_labels
178
+
179
+ def forward(self, pixel_values, input_ids, attention_mask, labels):
180
+ inputs_embeds, mask, full_labels = self.build_inputs(pixel_values, input_ids, attention_mask, labels)
181
+ return self.llm(inputs_embeds=inputs_embeds, attention_mask=mask, labels=full_labels, use_cache=False).loss
182
+
183
+
184
+ def load_pretrained(cls, name, **kw):
185
+ try:
186
+ return cls.from_pretrained(name, dtype=DTYPE, **kw)
187
+ except TypeError:
188
+ return cls.from_pretrained(name, torch_dtype=DTYPE, **kw)
189
+
190
+
191
+ def build_model():
192
+ tok = AutoTokenizer.from_pretrained(LLM_NAME)
193
+ pad_id = tok.pad_token_id if tok.pad_token_id is not None else tok.eos_token_id
194
+ processor = AutoImageProcessor.from_pretrained(VISION_NAME)
195
+
196
+ full_vision = load_pretrained(AutoModel, VISION_NAME)
197
+ vision = full_vision.vision_model
198
+ vision_dim = full_vision.config.vision_config.hidden_size
199
+ del full_vision # drops the unused SigLIP text tower
200
+
201
+ llm = load_pretrained(AutoModelForCausalLM, LLM_NAME)
202
+ text_dim = llm.config.hidden_size
203
+ llm.config.use_cache = False
204
+
205
+ vision.to(DEVICE).eval()
206
+ llm.to(DEVICE)
207
+ for p in vision.parameters():
208
+ p.requires_grad_(False)
209
+ for p in llm.parameters():
210
+ p.requires_grad_(False)
211
+ if GRAD_CKPT:
212
+ llm.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
213
+ llm.train() # HF only checkpoints in train mode; Llama has no dropout
214
+ else:
215
+ llm.eval()
216
+
217
+ prefix_ids = tok(USER_HEADER, add_special_tokens=True).input_ids
218
+ model = AtlasVision(vision, llm, prefix_ids, vision_dim, text_dim)
219
+ model.projector.to(DEVICE, dtype=torch.float32)
220
+ model.prefix_ids = model.prefix_ids.to(DEVICE)
221
+ return model, tok, pad_id, processor
222
+
223
+
224
+ # ----------------------------- TRAIN -----------------------------
225
+ def lr_at(step, total, warmup):
226
+ if step < warmup:
227
+ return LR * (step + 1) / warmup
228
+ progress = (step - warmup) / max(1, total - warmup)
229
+ return LR * 0.5 * (1 + math.cos(math.pi * min(1.0, progress)))
230
+
231
+
232
+ def save_checkpoint(path, model, opt, samples_seen, step, extra=None):
233
+ tmp = path + ".tmp"
234
+ torch.save({
235
+ "projector_state_dict": model.projector.state_dict(),
236
+ "optimizer_state_dict": opt.state_dict(),
237
+ "samples_seen": samples_seen,
238
+ "opt_step": step,
239
+ "config": {"LR": LR, "BATCH_SIZE": BATCH_SIZE, "GRAD_ACCUM": GRAD_ACCUM, "NUM_SAMPLES": NUM_SAMPLES,
240
+ "SEED": SEED, "LLM_NAME": LLM_NAME, "VISION_NAME": VISION_NAME, **(extra or {})},
241
+ }, tmp)
242
+ os.replace(tmp, path) # atomic: a stop mid-save never corrupts the last good checkpoint
243
+
244
+
245
+ def write_result(d):
246
+ path = os.environ.get("GMN_RESULT_PATH")
247
+ if path:
248
+ with open(path, "w") as f:
249
+ json.dump(d, f)
250
+
251
+
252
+ def main():
253
+ torch.manual_seed(SEED)
254
+ os.makedirs(CKPT_DIR, exist_ok=True)
255
+ ckpt_path = os.path.join(CKPT_DIR, "latest.pt")
256
+ log(f"device={DEVICE} batch={BATCH_SIZE}x{GRAD_ACCUM} lr={LR} samples={NUM_SAMPLES} "
257
+ f"grad_ckpt={GRAD_CKPT} workers={NUM_WORKERS} max_steps={MAX_STEPS or 'full'}")
258
+
259
+ model, tok, pad_id, processor = build_model()
260
+ n_train = sum(p.numel() for p in model.parameters() if p.requires_grad)
261
+ log(f"trainable params: {n_train:,} (projector only)")
262
+
263
+ with open(JSON_PATH) as f:
264
+ annotations = json.load(f)
265
+ g = torch.Generator().manual_seed(SEED)
266
+ keep = torch.randperm(len(annotations), generator=g)[:NUM_SAMPLES].tolist()
267
+ annotations = [annotations[i] for i in keep]
268
+ zip_prefix = find_zip_prefix(annotations, ZIP_PATH)
269
+ dataset = LLaVAPretrainDataset(annotations, ZIP_PATH, zip_prefix, processor, tok, MAX_TEXT_LEN)
270
+
271
+ eff_bs = BATCH_SIZE * GRAD_ACCUM
272
+ total_steps = math.ceil(len(dataset) / eff_bs)
273
+ if MAX_STEPS:
274
+ total_steps = min(total_steps, MAX_STEPS)
275
+ warmup = max(1, int(total_steps * WARMUP_RATIO))
276
+
277
+ params = [p for p in model.projector.parameters()]
278
+ opt = torch.optim.AdamW(params, lr=LR, weight_decay=0.0)
279
+
280
+ samples_seen, step = 0, 0
281
+ if os.path.exists(ckpt_path):
282
+ ck = torch.load(ckpt_path, map_location="cpu")
283
+ model.projector.load_state_dict(ck["projector_state_dict"])
284
+ opt.load_state_dict(ck["optimizer_state_dict"])
285
+ samples_seen = ck["samples_seen"]
286
+ step = samples_seen // eff_bs # works even if batch size changed
287
+ log(f"resumed: {samples_seen:,} samples seen -> optimizer step {step}/{total_steps}")
288
+ del ck
289
+ if step >= total_steps:
290
+ log("already complete")
291
+ return
292
+
293
+ order = torch.randperm(len(dataset), generator=torch.Generator().manual_seed(SEED + 1))[samples_seen:].tolist()
294
+ loader = DataLoader(
295
+ Subset(dataset, order), batch_size=BATCH_SIZE, shuffle=False, num_workers=NUM_WORKERS,
296
+ pin_memory=DEVICE.type == "cuda", collate_fn=make_collate(pad_id), drop_last=True,
297
+ persistent_workers=NUM_WORKERS > 0, prefetch_factor=4 if NUM_WORKERS > 0 else None,
298
+ )
299
+ log(f"schedule: {total_steps} optimizer steps, {warmup} warmup, starting at {step}")
300
+
301
+ model.projector.train()
302
+ if DEVICE.type == "cuda":
303
+ torch.cuda.reset_peak_memory_stats()
304
+ opt.zero_grad(set_to_none=True)
305
+ micro, loss_sum, loss_n, bad_imgs, skipped = 0, 0.0, 0, 0, 0
306
+ last_loss, t_window, steps_window = None, time.time(), 0
307
+ t_start = time.time()
308
+ log_f = open(os.path.join(CKPT_DIR, "train_log.jsonl"), "a")
309
+
310
+ for batch in loader:
311
+ if step >= total_steps:
312
+ break
313
+ try:
314
+ pv = batch["pixel_values"].to(DEVICE, non_blocking=True)
315
+ ids = batch["input_ids"].to(DEVICE, non_blocking=True)
316
+ mask = batch["attention_mask"].to(DEVICE, non_blocking=True)
317
+ labels = batch["labels"].to(DEVICE, non_blocking=True)
318
+ bad_imgs += batch["n_bad"]
319
+ with torch.autocast(device_type=DEVICE.type, dtype=DTYPE):
320
+ loss = model(pv, ids, mask, labels).float()
321
+ if not torch.isfinite(loss):
322
+ skipped += 1
323
+ log(f"non-finite loss at step {step}; dropping this accumulation window (skipped={skipped})")
324
+ opt.zero_grad(set_to_none=True)
325
+ micro = 0
326
+ continue
327
+ (loss / GRAD_ACCUM).backward()
328
+ except torch.cuda.OutOfMemoryError:
329
+ log("CUDA OOM - rerun with GRAD_CKPT=1 or a smaller BATCH_SIZE")
330
+ sys.exit(3)
331
+
332
+ loss_sum += loss.item()
333
+ loss_n += 1
334
+ micro += 1
335
+ if micro < GRAD_ACCUM:
336
+ continue
337
+ micro = 0
338
+
339
+ lr = lr_at(step, total_steps, warmup)
340
+ for grp in opt.param_groups:
341
+ grp["lr"] = lr
342
+ grad_norm = torch.nn.utils.clip_grad_norm_(params, max_norm=1.0).item()
343
+ opt.step()
344
+ opt.zero_grad(set_to_none=True)
345
+ step += 1
346
+ steps_window += 1
347
+ samples_seen += eff_bs
348
+
349
+ if step % LOG_EVERY == 0 or step == 1 or step == total_steps:
350
+ dt = time.time() - t_window
351
+ sps = dt / max(1, steps_window)
352
+ avg = loss_sum / max(1, loss_n)
353
+ last_loss = avg
354
+ peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
355
+ eta_h = (total_steps - step) * sps / 3600
356
+ log(f"step {step}/{total_steps} | loss {avg:.4f} | grad {grad_norm:.2f} | lr {lr:.2e} | "
357
+ f"{sps:.3f} s/step | {eff_bs / sps:.1f} img/s | ETA {eta_h:.2f} h | peak {peak:.1f} GB | "
358
+ f"bad_imgs {bad_imgs}")
359
+ log_f.write(json.dumps({"step": step, "loss": avg, "grad_norm": grad_norm, "lr": lr,
360
+ "s_per_step": sps, "peak_gb": peak, "samples_seen": samples_seen,
361
+ "time": time.time()}) + "\n")
362
+ log_f.flush()
363
+ loss_sum, loss_n, t_window, steps_window = 0.0, 0, time.time(), 0
364
+
365
+ if step % SAVE_EVERY == 0:
366
+ save_checkpoint(ckpt_path, model, opt, samples_seen, step)
367
+ log(f"checkpoint saved at step {step} ({samples_seen:,} samples)")
368
+
369
+ save_checkpoint(ckpt_path, model, opt, samples_seen, step)
370
+ done = step >= total_steps
371
+ if done and not MAX_STEPS:
372
+ torch.save(model.projector.state_dict(), os.path.join(CKPT_DIR, "projector_final.pt"))
373
+ log("saved projector_final.pt")
374
+ peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
375
+ elapsed = time.time() - t_start
376
+ log(f"finished: step {step}/{total_steps}, last avg loss {last_loss}, {elapsed / 60:.1f} min, "
377
+ f"peak {peak:.1f} GB, skipped {skipped}, bad_imgs {bad_imgs}")
378
+ write_result({"complete": done, "smoke": bool(MAX_STEPS), "step": step, "total_steps": total_steps,
379
+ "last_loss": last_loss, "peak_gb": round(peak, 1), "skipped": skipped, "bad_imgs": bad_imgs,
380
+ "grad_ckpt": GRAD_CKPT, "batch_size": BATCH_SIZE})
381
+
382
+
383
+ if __name__ == "__main__":
384
+ main()
code/train_stage2.py ADDED
@@ -0,0 +1,294 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Stage 2 (visual instruction tuning): LLaVA-Instruct-150K on COCO images.
3
+
4
+ Starts from the stage-1 projector. Vision tower frozen; the LLM gets LoRA adapters;
5
+ the projector keeps training at a lower LR. Same prompt layout as stage 1:
6
+ [<BOS> user-header] [image tokens] [q1 <eot> assistant-header] a1 <eot> [user-header q2 <eot> assistant-header] a2 <eot> ...
7
+ Loss only on assistant answers (+ their <eot>). Gradient checkpointing on, length-bucketed batches,
8
+ resume by samples seen. All settings via environment variables.
9
+ """
10
+ import json
11
+ import math
12
+ import os
13
+ import sys
14
+ import time
15
+ import zipfile
16
+
17
+ import torch
18
+ from PIL import Image
19
+ from torch.utils.data import DataLoader, Dataset, Subset
20
+
21
+ import train as T # stage-1 module: model layout, prompt strings, collate, zip check
22
+
23
+ env = T.env
24
+ HOME = os.path.expanduser("~")
25
+ INSTRUCT_JSON = env("INSTRUCT_JSON", f"{HOME}/data/llava_instruct/llava_instruct_150k.json")
26
+ COCO_ZIP = env("COCO_ZIP", f"{HOME}/data/coco/train2017.zip")
27
+ STAGE1_PROJECTOR = env("STAGE1_PROJECTOR", f"{HOME}/checkpoints/stage1/projector_final.pt")
28
+ CKPT_DIR = env("CKPT_DIR", f"{HOME}/checkpoints/stage2")
29
+ NUM_SAMPLES = env("NUM_SAMPLES", 0, int) # 0 = all except the held-out set
30
+ HELDOUT = env("HELDOUT", 1000, int)
31
+ BATCH_SIZE = env("BATCH_SIZE", 8, int)
32
+ GRAD_ACCUM = env("GRAD_ACCUM", 4, int)
33
+ LORA_LR = env("LORA_LR", 2e-4, float)
34
+ PROJ_LR = env("PROJ_LR", 2e-5, float)
35
+ LORA_R = env("LORA_R", 64, int)
36
+ LORA_ALPHA = env("LORA_ALPHA", 128, int)
37
+ WARMUP_RATIO = env("WARMUP_RATIO", 0.03, float)
38
+ MAX_TEXT_LEN = env("MAX_TEXT_LEN", 1024, int)
39
+ MAX_STEPS = env("MAX_STEPS", 0, int)
40
+ LOG_EVERY = env("LOG_EVERY", 25, int)
41
+ SAVE_EVERY = env("SAVE_EVERY", 250, int)
42
+ SEED = env("SEED", 42, int)
43
+ NUM_WORKERS = T.NUM_WORKERS
44
+ DEVICE, DTYPE, log = T.DEVICE, T.DTYPE, T.log
45
+ LORA_TARGETS = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
46
+
47
+
48
+ # ----------------------------- DATA -----------------------------
49
+ def split(ann):
50
+ perm = torch.randperm(len(ann), generator=torch.Generator().manual_seed(SEED)).tolist()
51
+ held, train = perm[-HELDOUT:], perm[:-HELDOUT]
52
+ if NUM_SAMPLES:
53
+ train = train[:NUM_SAMPLES]
54
+ return train, held
55
+
56
+
57
+ def build_ids(tok, conversations, max_len):
58
+ ids, labels = [], []
59
+ first = True
60
+ for turn in conversations:
61
+ v = turn["value"].replace("<image>", "").strip()
62
+ if turn["from"] == "human":
63
+ text = (v if first else T.USER_HEADER + v) + T.ASSIST_HEADER
64
+ t = tok(text, add_special_tokens=False).input_ids
65
+ ids += t
66
+ labels += [-100] * len(t)
67
+ first = False
68
+ else:
69
+ t = tok(v + T.EOT, add_special_tokens=False).input_ids
70
+ ids += t
71
+ labels += t
72
+ return ids[:max_len], labels[:max_len]
73
+
74
+
75
+ class InstructDataset(Dataset):
76
+ def __init__(self, ann, zip_path, zip_prefix, processor, tok, max_len):
77
+ self.ann, self.zip_path, self.prefix = ann, zip_path, zip_prefix
78
+ self.processor, self.tok, self.max_len = processor, tok, max_len
79
+ self.zip = None
80
+
81
+ def __len__(self):
82
+ return len(self.ann)
83
+
84
+ def load_image(self, item):
85
+ if self.zip is None:
86
+ self.zip = zipfile.ZipFile(self.zip_path)
87
+ try:
88
+ with self.zip.open(self.prefix + item["image"]) as f:
89
+ return Image.open(f).convert("RGB"), True
90
+ except Exception:
91
+ return Image.new("RGB", (224, 224), "white"), False
92
+
93
+ def __getitem__(self, idx):
94
+ item = self.ann[idx]
95
+ image, ok = self.load_image(item)
96
+ pv = self.processor(images=image, return_tensors="pt").pixel_values[0]
97
+ ids, labels = build_ids(self.tok, item["conversations"], self.max_len)
98
+ return {"pixel_values": pv, "input_ids": ids, "labels": labels, "ok": ok}
99
+
100
+
101
+ def bucketed_order(indices, lengths, mb, seed):
102
+ """Shuffle, sort by length inside chunks of 64 batches, keep only full batches, shuffle batches."""
103
+ g = torch.Generator().manual_seed(seed)
104
+ perm = [indices[i] for i in torch.randperm(len(indices), generator=g).tolist()]
105
+ chunk = mb * 64
106
+ batches = []
107
+ for s in range(0, len(perm), chunk):
108
+ c = sorted(perm[s:s + chunk], key=lambda i: lengths[i])
109
+ batches += [c[j:j + mb] for j in range(0, len(c), mb) if len(c[j:j + mb]) == mb]
110
+ order = torch.randperm(len(batches), generator=g).tolist()
111
+ return [i for b in order for i in batches[b]]
112
+
113
+
114
+ # ----------------------------- MODEL -----------------------------
115
+ def build_stage2_model(lora_state=None, projector_state=None, train_mode=True):
116
+ from peft import LoraConfig, get_peft_model, set_peft_model_state_dict
117
+
118
+ model, tok, pad_id, processor = T.build_model() # frozen vision + frozen LLM (bf16), fresh projector
119
+ proj = projector_state if projector_state is not None else _load_proj(STAGE1_PROJECTOR)
120
+ model.projector.load_state_dict(proj)
121
+ model.projector.to(DEVICE, dtype=torch.float32)
122
+
123
+ cfg = LoraConfig(r=LORA_R, lora_alpha=LORA_ALPHA, lora_dropout=0.05, target_modules=LORA_TARGETS,
124
+ bias="none", task_type="CAUSAL_LM")
125
+ model.llm = get_peft_model(model.llm, cfg)
126
+ if lora_state is not None:
127
+ set_peft_model_state_dict(model.llm, lora_state)
128
+ for n, p in model.llm.named_parameters(): # LoRA weights in fp32 for stable AdamW updates
129
+ if p.requires_grad:
130
+ p.data = p.data.float()
131
+ if train_mode:
132
+ model.llm.base_model.model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
133
+ model.llm.train()
134
+ else:
135
+ model.llm.eval()
136
+ model.vision.eval()
137
+ return model, tok, pad_id, processor
138
+
139
+
140
+ def _load_proj(path):
141
+ sd = torch.load(path, map_location="cpu")
142
+ return sd.get("projector_state_dict", sd)
143
+
144
+
145
+ def lr_scale(step, total, warmup):
146
+ if step < warmup:
147
+ return (step + 1) / warmup
148
+ progress = (step - warmup) / max(1, total - warmup)
149
+ return 0.5 * (1 + math.cos(math.pi * min(1.0, progress)))
150
+
151
+
152
+ def save_checkpoint(path, model, opt, samples_seen, step):
153
+ from peft import get_peft_model_state_dict
154
+ tmp = path + ".tmp"
155
+ torch.save({
156
+ "lora_state_dict": get_peft_model_state_dict(model.llm),
157
+ "projector_state_dict": model.projector.state_dict(),
158
+ "optimizer_state_dict": opt.state_dict(),
159
+ "samples_seen": samples_seen, "opt_step": step,
160
+ "config": {"LORA_R": LORA_R, "LORA_ALPHA": LORA_ALPHA, "LORA_LR": LORA_LR, "PROJ_LR": PROJ_LR,
161
+ "BATCH_SIZE": BATCH_SIZE, "GRAD_ACCUM": GRAD_ACCUM, "SEED": SEED},
162
+ }, tmp)
163
+ os.replace(tmp, path)
164
+
165
+
166
+ # ----------------------------- TRAIN -----------------------------
167
+ def main():
168
+ torch.manual_seed(SEED)
169
+ os.makedirs(CKPT_DIR, exist_ok=True)
170
+ ckpt_path = os.path.join(CKPT_DIR, "latest.pt")
171
+ log(f"stage2: batch={BATCH_SIZE}x{GRAD_ACCUM} lora_r={LORA_R} lora_lr={LORA_LR} proj_lr={PROJ_LR} "
172
+ f"max_text_len={MAX_TEXT_LEN} workers={NUM_WORKERS} max_steps={MAX_STEPS or 'full'}")
173
+
174
+ ck = torch.load(ckpt_path, map_location="cpu") if os.path.exists(ckpt_path) else None
175
+ model, tok, pad_id, processor = build_stage2_model(
176
+ lora_state=ck["lora_state_dict"] if ck else None,
177
+ projector_state=ck["projector_state_dict"] if ck else None)
178
+ lora_params = [p for n, p in model.llm.named_parameters() if p.requires_grad]
179
+ proj_params = list(model.projector.parameters())
180
+ log(f"trainable: LoRA {sum(p.numel() for p in lora_params):,} + projector {sum(p.numel() for p in proj_params):,}")
181
+
182
+ ann = json.load(open(INSTRUCT_JSON))
183
+ train_idx, _ = split(ann)
184
+ prefix = T.find_zip_prefix([ann[i] for i in train_idx[:2000]], COCO_ZIP)
185
+ dataset = InstructDataset(ann, COCO_ZIP, prefix, processor, tok, MAX_TEXT_LEN)
186
+ lengths = [sum(len(t["value"]) for t in a["conversations"]) for a in ann]
187
+
188
+ eff_bs = BATCH_SIZE * GRAD_ACCUM
189
+ order = bucketed_order(train_idx, lengths, BATCH_SIZE, SEED + 1)
190
+ total_steps = len(order) // eff_bs
191
+ if MAX_STEPS:
192
+ total_steps = min(total_steps, MAX_STEPS)
193
+ warmup = max(1, int(total_steps * WARMUP_RATIO))
194
+
195
+ opt = torch.optim.AdamW([{"params": lora_params, "lr": LORA_LR, "base_lr": LORA_LR},
196
+ {"params": proj_params, "lr": PROJ_LR, "base_lr": PROJ_LR}], weight_decay=0.0)
197
+ samples_seen, step = 0, 0
198
+ if ck:
199
+ opt.load_state_dict(ck["optimizer_state_dict"])
200
+ samples_seen = ck["samples_seen"]
201
+ step = samples_seen // eff_bs
202
+ log(f"resumed: {samples_seen:,} samples -> step {step}/{total_steps}")
203
+ del ck
204
+ if step >= total_steps:
205
+ log("already complete")
206
+ return
207
+
208
+ loader = DataLoader(
209
+ Subset(dataset, order[samples_seen:]), batch_size=BATCH_SIZE, shuffle=False, num_workers=NUM_WORKERS,
210
+ pin_memory=True, collate_fn=T.make_collate(pad_id), drop_last=True,
211
+ persistent_workers=NUM_WORKERS > 0, prefetch_factor=4 if NUM_WORKERS > 0 else None)
212
+ log(f"schedule: {total_steps} optimizer steps ({len(order):,} samples), {warmup} warmup, starting at {step}")
213
+
214
+ params = lora_params + proj_params
215
+ if DEVICE.type == "cuda":
216
+ torch.cuda.reset_peak_memory_stats()
217
+ opt.zero_grad(set_to_none=True)
218
+ micro, loss_sum, loss_n, bad, skipped = 0, 0.0, 0, 0, 0
219
+ last_loss, t_window, steps_window, t_start = None, time.time(), 0, time.time()
220
+ log_f = open(os.path.join(CKPT_DIR, "train_log.jsonl"), "a")
221
+
222
+ for batch in loader:
223
+ if step >= total_steps:
224
+ break
225
+ try:
226
+ bad += batch["n_bad"]
227
+ with torch.autocast(device_type=DEVICE.type, dtype=DTYPE):
228
+ loss = model(batch["pixel_values"].to(DEVICE, non_blocking=True),
229
+ batch["input_ids"].to(DEVICE, non_blocking=True),
230
+ batch["attention_mask"].to(DEVICE, non_blocking=True),
231
+ batch["labels"].to(DEVICE, non_blocking=True)).float()
232
+ if not torch.isfinite(loss):
233
+ skipped += 1
234
+ log(f"non-finite loss at step {step}; dropping accumulation window (skipped={skipped})")
235
+ opt.zero_grad(set_to_none=True)
236
+ micro = 0
237
+ continue
238
+ (loss / GRAD_ACCUM).backward()
239
+ except torch.cuda.OutOfMemoryError:
240
+ log("CUDA OOM - rerun with a smaller BATCH_SIZE (and larger GRAD_ACCUM)")
241
+ sys.exit(3)
242
+
243
+ loss_sum += loss.item()
244
+ loss_n += 1
245
+ micro += 1
246
+ if micro < GRAD_ACCUM:
247
+ continue
248
+ micro = 0
249
+
250
+ s = lr_scale(step, total_steps, warmup)
251
+ for grp in opt.param_groups:
252
+ grp["lr"] = grp["base_lr"] * s
253
+ grad_norm = torch.nn.utils.clip_grad_norm_(params, max_norm=1.0).item()
254
+ opt.step()
255
+ opt.zero_grad(set_to_none=True)
256
+ step += 1
257
+ steps_window += 1
258
+ samples_seen += eff_bs
259
+
260
+ if step % LOG_EVERY == 0 or step == 1 or step == total_steps:
261
+ dt = time.time() - t_window
262
+ sps = dt / max(1, steps_window)
263
+ avg = loss_sum / max(1, loss_n)
264
+ last_loss = avg
265
+ peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
266
+ log(f"step {step}/{total_steps} | loss {avg:.4f} | grad {grad_norm:.2f} | lr {LORA_LR * s:.2e} | "
267
+ f"{sps:.3f} s/step | {eff_bs / sps:.1f} conv/s | ETA {(total_steps - step) * sps / 3600:.2f} h | "
268
+ f"peak {peak:.1f} GB | bad_imgs {bad}")
269
+ log_f.write(json.dumps({"step": step, "loss": avg, "grad_norm": grad_norm, "lr": LORA_LR * s,
270
+ "s_per_step": sps, "peak_gb": peak, "samples_seen": samples_seen,
271
+ "time": time.time()}) + "\n")
272
+ log_f.flush()
273
+ loss_sum, loss_n, t_window, steps_window = 0.0, 0, time.time(), 0
274
+
275
+ if step % SAVE_EVERY == 0:
276
+ save_checkpoint(ckpt_path, model, opt, samples_seen, step)
277
+ log(f"checkpoint saved at step {step} ({samples_seen:,} samples)")
278
+
279
+ save_checkpoint(ckpt_path, model, opt, samples_seen, step)
280
+ done = step >= total_steps
281
+ if done and not MAX_STEPS:
282
+ model.llm.save_pretrained(os.path.join(CKPT_DIR, "lora_adapter"))
283
+ torch.save(model.projector.state_dict(), os.path.join(CKPT_DIR, "projector_stage2.pt"))
284
+ log("saved lora_adapter/ and projector_stage2.pt")
285
+ peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
286
+ log(f"finished: step {step}/{total_steps}, last avg loss {last_loss}, {(time.time() - t_start) / 60:.1f} min, "
287
+ f"peak {peak:.1f} GB, skipped {skipped}, bad_imgs {bad}")
288
+ T.write_result({"complete": done, "smoke": bool(MAX_STEPS), "step": step, "total_steps": total_steps,
289
+ "last_loss": last_loss, "peak_gb": round(peak, 1), "skipped": skipped, "bad_imgs": bad,
290
+ "batch_size": BATCH_SIZE})
291
+
292
+
293
+ if __name__ == "__main__":
294
+ main()
eval/report.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Atlas-Vision evaluation
2
+
3
+ ## Scores (stage 1 → stage 2)
4
+
5
+ | Metric | Stage 1 | Stage 2 |
6
+ |---|---|---|
7
+ | Held-out instruct loss (lower is better) | 2.2521 | 1.1271 |
8
+ | POPE random: accuracy / F1 / yes-ratio | 0.4997 / 0.2003 / 0.1257 | 0.5427 / 0.6856 / 0.9547 |
9
+ | POPE popular: accuracy / F1 / yes-ratio | 0.5027 / 0.2013 / 0.1227 | 0.522 / 0.676 / 0.9753 |
10
+ | POPE adversarial: accuracy / F1 / yes-ratio | 0.5023 / 0.2012 / 0.123 | 0.5137 / 0.6722 / 0.9837 |
11
+
12
+ ## Detailed descriptions
13
+
14
+ **COCO_val2014_000000310196**
15
+
16
+ - Stage 1: a skier in the snow on a mountain
17
+ - Stage 2: The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.
18
+
19
+ There are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together.
20
+
21
+ **COCO_val2014_000000210789**
22
+
23
+ - Stage 1: a woman and her child in the rain
24
+ - Stage 2: The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.
25
+
26
+ In the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene.
27
+
28
+ **COCO_val2014_000000429109**
29
+
30
+ - Stage 1: a bus and several other vehicles parked in front of a building
31
+ - Stage 2: The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.
32
+
33
+ Numerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.
34
+
35
+ In addition to the vehicles, there are traffic lights at various points along the street
36
+
37
+ **COCO_val2014_000000211674**
38
+
39
+ - Stage 1: a bus with a red and white logo on it, carrying passengers
40
+ - Stage 2: The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.
41
+
42
+ In addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by.
43
+
44
+ ## Questions in Nigerian languages (stage 2)
45
+
46
+ - **igbo** — Kedu ihe dị na foto a?
47
+ → The image features a person skiing down a snow-covered slope, with the skier wearing red pants.
48
+ - **yoruba** — Kí ni ó wà nínú àwòrán yìí?
49
+ → The image features a person skiing down a snow-covered slope, with the skier wearing red pants.
50
+ - **hausa** — Me ke cikin wannan hoton?
51
+ → In the image, a person is skiing down a snow-covered slope.
52
+
53
+ ## Text-only check (no image)
54
+
55
+ Question: Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.
56
+
57
+ - Base N-ATLaS: Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
58
+ - With stage-2 LoRA: Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu "positron" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej
eval/stage1_eval.json ADDED
@@ -0,0 +1,68 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "checkpoint": {
3
+ "net.0.weight": [
4
+ 4096,
5
+ 768
6
+ ],
7
+ "net.0.bias": [
8
+ 4096
9
+ ],
10
+ "net.2.weight": [
11
+ 4096,
12
+ 4096
13
+ ],
14
+ "net.2.bias": [
15
+ 4096
16
+ ]
17
+ },
18
+ "checkpoint_all_finite": true,
19
+ "heldout_loss_random_projector": 8.7249,
20
+ "heldout_loss_trained": 2.4523,
21
+ "heldout_loss_trained_shuffled_images": 5.4332,
22
+ "captions": [
23
+ {
24
+ "image": "00390/003907752.jpg",
25
+ "reference": "cute homecoming prom dress with lace top and satin skirt",
26
+ "model": "a - line sweetheart sweetheart lace tulle prom dress"
27
+ },
28
+ {
29
+ "image": "00193/001937181.jpg",
30
+ "reference": "a pair of silicone bracelets with a four codes message",
31
+ "model": "a pair of silicone bracelets with the word code on them"
32
+ },
33
+ {
34
+ "image": "00530/005308076.jpg",
35
+ "reference": "a set of fingerprint icons in white on a black background",
36
+ "model": "the iphone 6s and iphone 6s plus are shown in a black and white image"
37
+ },
38
+ {
39
+ "image": "00399/003998740.jpg",
40
+ "reference": "a business woman drawing modern concept of a website creation",
41
+ "model": "a man writing the word modern web design on a whiteboard"
42
+ },
43
+ {
44
+ "image": "00391/003917454.jpg",
45
+ "reference": "the beach in el nido national park, puerto puerto",
46
+ "model": "the limestone cliffs and limestone islands in the background"
47
+ },
48
+ {
49
+ "image": "00077/000776787.jpg",
50
+ "reference": "three pieces of paper with the words, democratic decentified dp controlled centralized ccp",
51
+ "model": "a diagram showing the different types of democracy"
52
+ },
53
+ {
54
+ "image": "00099/000999958.jpg",
55
+ "reference": "the beginer's guide to social security book",
56
+ "model": "a beginner's guide to social security"
57
+ },
58
+ {
59
+ "image": "00446/004469100.jpg",
60
+ "reference": "a flag of bolivia, with a smiley face water bottle",
61
+ "model": "a red and yellow flag with a red and yellow flag on it water bottle"
62
+ }
63
+ ],
64
+ "coco_zip_files": 118288,
65
+ "instruct_conversations": 157712,
66
+ "instruct_images_found_of_2000": 2000,
67
+ "eval_minutes": 1.9
68
+ }
eval/stage2_eval.json ADDED
@@ -0,0 +1,104 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage1": {
3
+ "heldout_instruct_loss": 2.2521,
4
+ "pope_random": {
5
+ "n": 3000,
6
+ "accuracy": 0.4997,
7
+ "precision": 0.4987,
8
+ "recall": 0.1253,
9
+ "f1": 0.2003,
10
+ "yes_ratio": 0.1257
11
+ },
12
+ "pope_popular": {
13
+ "n": 3000,
14
+ "accuracy": 0.5027,
15
+ "precision": 0.5109,
16
+ "recall": 0.1253,
17
+ "f1": 0.2013,
18
+ "yes_ratio": 0.1227
19
+ },
20
+ "pope_adversarial": {
21
+ "n": 3000,
22
+ "accuracy": 0.5023,
23
+ "precision": 0.5095,
24
+ "recall": 0.1253,
25
+ "f1": 0.2012,
26
+ "yes_ratio": 0.123
27
+ }
28
+ },
29
+ "stage2": {
30
+ "heldout_instruct_loss": 1.1271,
31
+ "pope_random": {
32
+ "n": 3000,
33
+ "accuracy": 0.5427,
34
+ "precision": 0.5223,
35
+ "recall": 0.9973,
36
+ "f1": 0.6856,
37
+ "yes_ratio": 0.9547
38
+ },
39
+ "pope_popular": {
40
+ "n": 3000,
41
+ "accuracy": 0.522,
42
+ "precision": 0.5113,
43
+ "recall": 0.9973,
44
+ "f1": 0.676,
45
+ "yes_ratio": 0.9753
46
+ },
47
+ "pope_adversarial": {
48
+ "n": 3000,
49
+ "accuracy": 0.5137,
50
+ "precision": 0.5069,
51
+ "recall": 0.9973,
52
+ "f1": 0.6722,
53
+ "yes_ratio": 0.9837
54
+ }
55
+ },
56
+ "descriptions": [
57
+ {
58
+ "image": "COCO_val2014_000000310196",
59
+ "stage1": "a skier in the snow on a mountain",
60
+ "stage2": "The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.\n\nThere are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together."
61
+ },
62
+ {
63
+ "image": "COCO_val2014_000000210789",
64
+ "stage1": "a woman and her child in the rain",
65
+ "stage2": "The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.\n\nIn the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene."
66
+ },
67
+ {
68
+ "image": "COCO_val2014_000000429109",
69
+ "stage1": "a bus and several other vehicles parked in front of a building",
70
+ "stage2": "The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.\n\nNumerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.\n\nIn addition to the vehicles, there are traffic lights at various points along the street"
71
+ },
72
+ {
73
+ "image": "COCO_val2014_000000211674",
74
+ "stage1": "a bus with a red and white logo on it, carrying passengers",
75
+ "stage2": "The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.\n\nIn addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by."
76
+ }
77
+ ],
78
+ "multilingual": [
79
+ {
80
+ "image": "COCO_val2014_000000310196",
81
+ "language": "igbo",
82
+ "question": "Kedu ihe dị na foto a?",
83
+ "stage2": "The image features a person skiing down a snow-covered slope, with the skier wearing red pants."
84
+ },
85
+ {
86
+ "image": "COCO_val2014_000000310196",
87
+ "language": "yoruba",
88
+ "question": "Kí ni ó wà nínú àwòrán yìí?",
89
+ "stage2": "The image features a person skiing down a snow-covered slope, with the skier wearing red pants."
90
+ },
91
+ {
92
+ "image": "COCO_val2014_000000310196",
93
+ "language": "hausa",
94
+ "question": "Me ke cikin wannan hoton?",
95
+ "stage2": "In the image, a person is skiing down a snow-covered slope."
96
+ }
97
+ ],
98
+ "text_only": {
99
+ "question": "Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.",
100
+ "base_n_atlas": "Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us",
101
+ "with_stage2_lora": "Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu \"positron\" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej"
102
+ },
103
+ "eval_minutes": 5.2
104
+ }
logs/stage1_train_log.jsonl ADDED
@@ -0,0 +1,376 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"step": 1, "loss": 8.55087947845459, "grad_norm": 87.72406768798828, "lr": 7.117437722419929e-07, "s_per_step": 8.396198749542236, "peak_gb": 41.247548928, "samples_seen": 32, "time": 1791130833.9630904}
2
+ {"step": 25, "loss": 6.397367646296819, "grad_norm": 15.911425590515137, "lr": 1.779359430604982e-05, "s_per_step": 0.6239869793256124, "peak_gb": 42.138478592, "samples_seen": 800, "time": 1791130848.9393055}
3
+ {"step": 50, "loss": 4.635109643936158, "grad_norm": 7.2413105964660645, "lr": 3.558718861209964e-05, "s_per_step": 0.5889950180053711, "peak_gb": 42.138478592, "samples_seen": 1600, "time": 1791130863.664765}
4
+ {"step": 75, "loss": 4.093637528419495, "grad_norm": 4.90704870223999, "lr": 5.338078291814947e-05, "s_per_step": 0.5850510597229004, "peak_gb": 42.138478592, "samples_seen": 2400, "time": 1791130878.2915204}
5
+ {"step": 100, "loss": 3.8673372650146485, "grad_norm": 4.5811944007873535, "lr": 7.117437722419929e-05, "s_per_step": 0.623746509552002, "peak_gb": 44.713082368, "samples_seen": 3200, "time": 1791130893.8860688}
6
+ {"step": 125, "loss": 3.740880389213562, "grad_norm": 3.184147834777832, "lr": 8.896797153024912e-05, "s_per_step": 0.5818255233764649, "peak_gb": 44.713082368, "samples_seen": 4000, "time": 1791130908.432502}
7
+ {"step": 150, "loss": 3.7184673833847044, "grad_norm": 2.9738147258758545, "lr": 0.00010676156583629893, "s_per_step": 0.576975965499878, "peak_gb": 44.713082368, "samples_seen": 4800, "time": 1791130922.857381}
8
+ {"step": 175, "loss": 3.5475265645980834, "grad_norm": 2.332338333129883, "lr": 0.00012455516014234878, "s_per_step": 0.5761029434204101, "peak_gb": 44.713082368, "samples_seen": 5600, "time": 1791130937.2603934}
9
+ {"step": 200, "loss": 3.5738428688049315, "grad_norm": 1.9974111318588257, "lr": 0.00014234875444839857, "s_per_step": 0.580996150970459, "peak_gb": 44.713082368, "samples_seen": 6400, "time": 1791130951.7857678}
10
+ {"step": 225, "loss": 3.5230879592895508, "grad_norm": 1.8189729452133179, "lr": 0.00016014234875444842, "s_per_step": 0.5787319278717041, "peak_gb": 44.713082368, "samples_seen": 7200, "time": 1791130966.2545114}
11
+ {"step": 250, "loss": 3.4794790887832643, "grad_norm": 2.1197030544281006, "lr": 0.00017793594306049823, "s_per_step": 0.5812134552001953, "peak_gb": 44.713082368, "samples_seen": 8000, "time": 1791130980.7852197}
12
+ {"step": 275, "loss": 3.5135861301422118, "grad_norm": 2.010641098022461, "lr": 0.00019572953736654805, "s_per_step": 0.5902993679046631, "peak_gb": 44.713082368, "samples_seen": 8800, "time": 1791130995.5431538}
13
+ {"step": 300, "loss": 3.332580523490906, "grad_norm": 1.7257816791534424, "lr": 0.00019999806668125934, "s_per_step": 0.5805612754821777, "peak_gb": 44.713082368, "samples_seen": 9600, "time": 1791131010.0576127}
14
+ {"step": 325, "loss": 3.331582636833191, "grad_norm": 1.817513346672058, "lr": 0.0001999889671230344, "s_per_step": 0.5767327976226807, "peak_gb": 44.713082368, "samples_seen": 10400, "time": 1791131024.4763641}
15
+ {"step": 350, "loss": 3.3034389925003054, "grad_norm": 1.4937862157821655, "lr": 0.00019997240961861618, "s_per_step": 0.5783243465423584, "peak_gb": 44.713082368, "samples_seen": 11200, "time": 1791131038.9348876}
16
+ {"step": 375, "loss": 3.310972599983215, "grad_norm": 1.4323818683624268, "lr": 0.00019994839540299076, "s_per_step": 0.5792804718017578, "peak_gb": 44.713082368, "samples_seen": 12000, "time": 1791131053.417312}
17
+ {"step": 400, "loss": 3.287726469039917, "grad_norm": 1.5991160869598389, "lr": 0.000199916926267323, "s_per_step": 0.578808069229126, "peak_gb": 44.713082368, "samples_seen": 12800, "time": 1791131067.8879645}
18
+ {"step": 425, "loss": 3.2084068727493285, "grad_norm": 1.271116018295288, "lr": 0.00019987800455882306, "s_per_step": 0.5803878498077393, "peak_gb": 44.713082368, "samples_seen": 13600, "time": 1791131082.3980415}
19
+ {"step": 450, "loss": 3.2198647975921633, "grad_norm": 1.286621332168579, "lr": 0.00019983163318057137, "s_per_step": 0.5836947727203369, "peak_gb": 44.713082368, "samples_seen": 14400, "time": 1791131096.991321}
20
+ {"step": 475, "loss": 3.1718136644363404, "grad_norm": 1.3426055908203125, "lr": 0.00019977781559130187, "s_per_step": 0.5814045429229736, "peak_gb": 44.713082368, "samples_seen": 15200, "time": 1791131111.52718}
21
+ {"step": 500, "loss": 3.176059513092041, "grad_norm": 1.3492554426193237, "lr": 0.00019971655580514438, "s_per_step": 0.5872831630706787, "peak_gb": 44.713082368, "samples_seen": 16000, "time": 1791131126.2099152}
22
+ {"step": 525, "loss": 3.1693886518478394, "grad_norm": 1.2558445930480957, "lr": 0.00019964785839132488, "s_per_step": 0.6033973217010498, "peak_gb": 44.713082368, "samples_seen": 16800, "time": 1791131141.295349}
23
+ {"step": 550, "loss": 3.14515998840332, "grad_norm": 1.2257124185562134, "lr": 0.0001995717284738248, "s_per_step": 0.5826546669006347, "peak_gb": 44.713082368, "samples_seen": 17600, "time": 1791131155.8622203}
24
+ {"step": 575, "loss": 3.12627272605896, "grad_norm": 1.2098281383514404, "lr": 0.000199488171730999, "s_per_step": 0.582458963394165, "peak_gb": 44.713082368, "samples_seen": 18400, "time": 1791131170.424219}
25
+ {"step": 600, "loss": 3.135997772216797, "grad_norm": 1.3937653303146362, "lr": 0.0001993971943951519, "s_per_step": 0.5858418941497803, "peak_gb": 44.713082368, "samples_seen": 19200, "time": 1791131185.0707946}
26
+ {"step": 625, "loss": 3.0999442672729494, "grad_norm": 1.0516718626022339, "lr": 0.000199298803252073, "s_per_step": 0.5833398532867432, "peak_gb": 44.713082368, "samples_seen": 20000, "time": 1791131199.654984}
27
+ {"step": 650, "loss": 3.080396270751953, "grad_norm": 1.1961040496826172, "lr": 0.00019919300564053046, "s_per_step": 0.6152436542510986, "peak_gb": 44.713082368, "samples_seen": 20800, "time": 1791131215.037142}
28
+ {"step": 675, "loss": 3.0702759647369384, "grad_norm": 1.0449596643447876, "lr": 0.00019907980945172384, "s_per_step": 0.5828781223297119, "peak_gb": 44.713082368, "samples_seen": 21600, "time": 1791131229.609633}
29
+ {"step": 700, "loss": 3.0974017333984376, "grad_norm": 1.0658111572265625, "lr": 0.00019895922312869555, "s_per_step": 0.5841869640350342, "peak_gb": 44.713082368, "samples_seen": 22400, "time": 1791131244.2148454}
30
+ {"step": 725, "loss": 3.074553599357605, "grad_norm": 0.9470886588096619, "lr": 0.00019883125566570094, "s_per_step": 0.5814198112487793, "peak_gb": 44.713082368, "samples_seen": 23200, "time": 1791131258.7509072}
31
+ {"step": 750, "loss": 3.1141643714904785, "grad_norm": 1.1800023317337036, "lr": 0.00019869591660753763, "s_per_step": 0.5784324645996094, "peak_gb": 44.713082368, "samples_seen": 24000, "time": 1791131273.212172}
32
+ {"step": 775, "loss": 3.1024985694885254, "grad_norm": 1.0973631143569946, "lr": 0.00019855321604883352, "s_per_step": 0.5959490013122558, "peak_gb": 44.713082368, "samples_seen": 24800, "time": 1791131288.1113417}
33
+ {"step": 800, "loss": 3.095311689376831, "grad_norm": 1.2262718677520752, "lr": 0.00019840316463329378, "s_per_step": 0.5826352882385254, "peak_gb": 44.713082368, "samples_seen": 25600, "time": 1791131302.6778867}
34
+ {"step": 825, "loss": 3.0253290700912476, "grad_norm": 1.2020397186279297, "lr": 0.000198245773552907, "s_per_step": 0.5797882652282715, "peak_gb": 44.713082368, "samples_seen": 26400, "time": 1791131317.173062}
35
+ {"step": 850, "loss": 3.0069773197174072, "grad_norm": 1.1738593578338623, "lr": 0.00019808105454711055, "s_per_step": 0.5793154048919678, "peak_gb": 44.713082368, "samples_seen": 27200, "time": 1791131331.6564252}
36
+ {"step": 875, "loss": 2.975392165184021, "grad_norm": 0.9737287163734436, "lr": 0.0001979090199019147, "s_per_step": 0.5790924072265625, "peak_gb": 44.713082368, "samples_seen": 28000, "time": 1791131346.1341693}
37
+ {"step": 900, "loss": 2.9498213052749636, "grad_norm": 1.2708204984664917, "lr": 0.0001977296824489864, "s_per_step": 0.5816666603088378, "peak_gb": 44.713082368, "samples_seen": 28800, "time": 1791131360.676415}
38
+ {"step": 925, "loss": 2.9142259979248046, "grad_norm": 1.089045763015747, "lr": 0.00019754305556469223, "s_per_step": 0.5844484996795655, "peak_gb": 44.713082368, "samples_seen": 29600, "time": 1791131375.2880468}
39
+ {"step": 950, "loss": 3.0302752017974854, "grad_norm": 1.205527663230896, "lr": 0.0001973491531691006, "s_per_step": 0.5822625255584717, "peak_gb": 44.713082368, "samples_seen": 30400, "time": 1791131389.8449948}
40
+ {"step": 975, "loss": 2.9161105060577395, "grad_norm": 1.1548945903778076, "lr": 0.00019714798972494347, "s_per_step": 0.5808561515808105, "peak_gb": 44.713082368, "samples_seen": 31200, "time": 1791131404.3668756}
41
+ {"step": 1000, "loss": 2.919966125488281, "grad_norm": 1.1152652502059937, "lr": 0.00019693958023653767, "s_per_step": 0.5795290184020996, "peak_gb": 44.713082368, "samples_seen": 32000, "time": 1791131418.8555546}
42
+ {"step": 1025, "loss": 2.952225284576416, "grad_norm": 1.1843825578689575, "lr": 0.00019672394024866576, "s_per_step": 0.5974857711791992, "peak_gb": 44.713082368, "samples_seen": 32800, "time": 1791131433.7931862}
43
+ {"step": 1050, "loss": 2.9690034532547, "grad_norm": 1.1854422092437744, "lr": 0.00019650108584541654, "s_per_step": 0.5811002635955811, "peak_gb": 44.713082368, "samples_seen": 33600, "time": 1791131448.3211079}
44
+ {"step": 1075, "loss": 2.9967500257492063, "grad_norm": 1.2190947532653809, "lr": 0.00019627103364898538, "s_per_step": 0.580099172592163, "peak_gb": 44.713082368, "samples_seen": 34400, "time": 1791131462.8241384}
45
+ {"step": 1100, "loss": 2.990977711677551, "grad_norm": 1.0495576858520508, "lr": 0.00019603380081843449, "s_per_step": 0.5808154773712159, "peak_gb": 44.713082368, "samples_seen": 35200, "time": 1791131477.3449311}
46
+ {"step": 1125, "loss": 2.8920970916748048, "grad_norm": 0.9397656321525574, "lr": 0.0001957894050484129, "s_per_step": 0.5838562393188477, "peak_gb": 44.713082368, "samples_seen": 36000, "time": 1791131491.941773}
47
+ {"step": 1150, "loss": 2.914958839416504, "grad_norm": 1.3053033351898193, "lr": 0.00019553786456783686, "s_per_step": 0.582167854309082, "peak_gb": 44.713082368, "samples_seen": 36800, "time": 1791131506.4963646}
48
+ {"step": 1175, "loss": 2.9034262990951536, "grad_norm": 1.0504006147384644, "lr": 0.00019527919813853, "s_per_step": 0.5793894004821777, "peak_gb": 44.713082368, "samples_seen": 37600, "time": 1791131520.981513}
49
+ {"step": 1200, "loss": 2.9084231233596802, "grad_norm": 1.1124707460403442, "lr": 0.00019501342505382405, "s_per_step": 0.5810192966461182, "peak_gb": 44.713082368, "samples_seen": 38400, "time": 1791131535.5074165}
50
+ {"step": 1225, "loss": 2.827880344390869, "grad_norm": 1.4236536026000977, "lr": 0.00019474056513711977, "s_per_step": 0.5862373161315918, "peak_gb": 44.713082368, "samples_seen": 39200, "time": 1791131550.1637754}
51
+ {"step": 1250, "loss": 2.8731605625152588, "grad_norm": 1.1672555208206177, "lr": 0.00019446063874040834, "s_per_step": 0.5826034450531006, "peak_gb": 44.713082368, "samples_seen": 40000, "time": 1791131564.7292533}
52
+ {"step": 1275, "loss": 2.9100488948822023, "grad_norm": 1.0153155326843262, "lr": 0.00019417366674275333, "s_per_step": 0.6047325611114502, "peak_gb": 44.713082368, "samples_seen": 40800, "time": 1791131579.8479767}
53
+ {"step": 1300, "loss": 2.8539438247680664, "grad_norm": 1.0282107591629028, "lr": 0.00019387967054873356, "s_per_step": 0.5824565982818604, "peak_gb": 44.713082368, "samples_seen": 41600, "time": 1791131594.4098606}
54
+ {"step": 1325, "loss": 2.8969113874435424, "grad_norm": 1.3478723764419556, "lr": 0.00019357867208684621, "s_per_step": 0.582642650604248, "peak_gb": 44.713082368, "samples_seen": 42400, "time": 1791131608.9763298}
55
+ {"step": 1350, "loss": 2.8323376035690306, "grad_norm": 1.1149699687957764, "lr": 0.00019327069380787164, "s_per_step": 0.5824921321868897, "peak_gb": 44.713082368, "samples_seen": 43200, "time": 1791131623.5390694}
56
+ {"step": 1375, "loss": 2.8894362831115723, "grad_norm": 1.1063823699951172, "lr": 0.00019295575868319857, "s_per_step": 0.5840862846374512, "peak_gb": 44.713082368, "samples_seen": 44000, "time": 1791131638.1416295}
57
+ {"step": 1400, "loss": 2.834498815536499, "grad_norm": 1.1182633638381958, "lr": 0.00019263389020311082, "s_per_step": 0.5812054634094238, "peak_gb": 44.713082368, "samples_seen": 44800, "time": 1791131652.6722355}
58
+ {"step": 1425, "loss": 2.8067142009735107, "grad_norm": 1.203196406364441, "lr": 0.00019230511237503515, "s_per_step": 0.5813498687744141, "peak_gb": 44.713082368, "samples_seen": 45600, "time": 1791131667.2064037}
59
+ {"step": 1450, "loss": 2.800572829246521, "grad_norm": 1.1607813835144043, "lr": 0.00019196944972175062, "s_per_step": 0.5796218490600586, "peak_gb": 44.713082368, "samples_seen": 46400, "time": 1791131681.697381}
60
+ {"step": 1475, "loss": 2.824986810684204, "grad_norm": 1.1562610864639282, "lr": 0.00019162692727955953, "s_per_step": 0.5842508506774903, "peak_gb": 44.713082368, "samples_seen": 47200, "time": 1791131696.3040268}
61
+ {"step": 1500, "loss": 2.8580026912689207, "grad_norm": 1.2550817728042603, "lr": 0.00019127757059642, "s_per_step": 0.5836690044403077, "peak_gb": 44.713082368, "samples_seen": 48000, "time": 1791131710.8980353}
62
+ {"step": 1525, "loss": 2.7956640100479127, "grad_norm": 1.0501290559768677, "lr": 0.00019092140573004042, "s_per_step": 0.5987179946899414, "peak_gb": 44.713082368, "samples_seen": 48800, "time": 1791131725.866389}
63
+ {"step": 1550, "loss": 2.904049711227417, "grad_norm": 1.2289342880249023, "lr": 0.0001905584592459358, "s_per_step": 0.579942455291748, "peak_gb": 44.713082368, "samples_seen": 49600, "time": 1791131740.3653781}
64
+ {"step": 1575, "loss": 2.801865735054016, "grad_norm": 1.4003002643585205, "lr": 0.0001901887582154464, "s_per_step": 0.5801137828826904, "peak_gb": 44.713082368, "samples_seen": 50400, "time": 1791131754.8686228}
65
+ {"step": 1600, "loss": 2.81393018245697, "grad_norm": 1.0660154819488525, "lr": 0.00018981233021371843, "s_per_step": 0.5823390769958496, "peak_gb": 44.713082368, "samples_seen": 51200, "time": 1791131769.4274774}
66
+ {"step": 1625, "loss": 2.8056074666976927, "grad_norm": 1.1312650442123413, "lr": 0.00018942920331764746, "s_per_step": 0.5844465160369873, "peak_gb": 44.713082368, "samples_seen": 52000, "time": 1791131784.0390692}
67
+ {"step": 1650, "loss": 2.77773108959198, "grad_norm": 1.0975704193115234, "lr": 0.00018903940610378407, "s_per_step": 0.582863941192627, "peak_gb": 44.713082368, "samples_seen": 52800, "time": 1791131798.6111288}
68
+ {"step": 1675, "loss": 2.7603001546859742, "grad_norm": 1.241600513458252, "lr": 0.00018864296764620238, "s_per_step": 0.5836851119995117, "peak_gb": 44.713082368, "samples_seen": 53600, "time": 1791131813.2036917}
69
+ {"step": 1700, "loss": 2.7555297422409057, "grad_norm": 1.1494439840316772, "lr": 0.00018823991751433167, "s_per_step": 0.5850475883483887, "peak_gb": 44.713082368, "samples_seen": 54400, "time": 1791131827.8303254}
70
+ {"step": 1725, "loss": 2.8411933469772337, "grad_norm": 0.991641104221344, "lr": 0.00018783028577075065, "s_per_step": 0.5855899620056152, "peak_gb": 44.713082368, "samples_seen": 55200, "time": 1791131842.4704814}
71
+ {"step": 1750, "loss": 2.8297329354286194, "grad_norm": 1.3848960399627686, "lr": 0.00018741410296894528, "s_per_step": 0.5831822967529297, "peak_gb": 44.713082368, "samples_seen": 56000, "time": 1791131857.0504653}
72
+ {"step": 1775, "loss": 2.7489702224731447, "grad_norm": 1.1250214576721191, "lr": 0.0001869914001510298, "s_per_step": 0.6029328918457031, "peak_gb": 44.713082368, "samples_seen": 56800, "time": 1791131872.1245182}
73
+ {"step": 1800, "loss": 2.840885500907898, "grad_norm": 1.192891240119934, "lr": 0.00018656220884543143, "s_per_step": 0.5853533840179443, "peak_gb": 44.713082368, "samples_seen": 57600, "time": 1791131886.7587817}
74
+ {"step": 1825, "loss": 2.7824712562561036, "grad_norm": 1.121734857559204, "lr": 0.00018612656106453871, "s_per_step": 0.5848813915252685, "peak_gb": 44.713082368, "samples_seen": 58400, "time": 1791131901.3812835}
75
+ {"step": 1850, "loss": 2.7347464036941527, "grad_norm": 1.163681149482727, "lr": 0.00018568448930231373, "s_per_step": 0.5818957328796387, "peak_gb": 44.713082368, "samples_seen": 59200, "time": 1791131915.9291348}
76
+ {"step": 1875, "loss": 2.7876648902893066, "grad_norm": 1.1300899982452393, "lr": 0.00018523602653186858, "s_per_step": 0.5813259029388428, "peak_gb": 44.713082368, "samples_seen": 60000, "time": 1791131930.462728}
77
+ {"step": 1900, "loss": 2.7792696571350097, "grad_norm": 1.1007814407348633, "lr": 0.00018478120620300578, "s_per_step": 0.5799147510528564, "peak_gb": 44.713082368, "samples_seen": 60800, "time": 1791131944.9613276}
78
+ {"step": 1925, "loss": 2.760815134048462, "grad_norm": 1.2840055227279663, "lr": 0.00018432006223972357, "s_per_step": 0.581923418045044, "peak_gb": 44.713082368, "samples_seen": 61600, "time": 1791131959.510147}
79
+ {"step": 1950, "loss": 2.72098069190979, "grad_norm": 1.1493003368377686, "lr": 0.00018385262903768547, "s_per_step": 0.5796582794189453, "peak_gb": 44.713082368, "samples_seen": 62400, "time": 1791131974.0020506}
80
+ {"step": 1975, "loss": 2.7809674453735354, "grad_norm": 1.0490471124649048, "lr": 0.00018337894146165467, "s_per_step": 0.5820777893066407, "peak_gb": 44.713082368, "samples_seen": 63200, "time": 1791131988.5544076}
81
+ {"step": 2000, "loss": 2.7531254482269287, "grad_norm": 1.0774011611938477, "lr": 0.00018289903484289383, "s_per_step": 0.5824406623840332, "peak_gb": 44.713082368, "samples_seen": 64000, "time": 1791132003.1158311}
82
+ {"step": 2025, "loss": 2.724808783531189, "grad_norm": 1.0720311403274536, "lr": 0.00018241294497652958, "s_per_step": 0.5992901039123535, "peak_gb": 44.713082368, "samples_seen": 64800, "time": 1791132018.0985038}
83
+ {"step": 2050, "loss": 2.790845193862915, "grad_norm": 1.2732620239257812, "lr": 0.00018192070811888272, "s_per_step": 0.5811270427703857, "peak_gb": 44.713082368, "samples_seen": 65600, "time": 1791132032.6270998}
84
+ {"step": 2075, "loss": 2.7737058067321776, "grad_norm": 1.0213253498077393, "lr": 0.00018142236098476398, "s_per_step": 0.5809852790832519, "peak_gb": 44.713082368, "samples_seen": 66400, "time": 1791132047.1521604}
85
+ {"step": 2100, "loss": 2.739048023223877, "grad_norm": 1.082443118095398, "lr": 0.00018091794074473537, "s_per_step": 0.5808903026580811, "peak_gb": 44.713082368, "samples_seen": 67200, "time": 1791132061.6748774}
86
+ {"step": 2125, "loss": 2.7185260105133056, "grad_norm": 1.1437773704528809, "lr": 0.00018040748502233802, "s_per_step": 0.5793789958953858, "peak_gb": 44.713082368, "samples_seen": 68000, "time": 1791132076.1597369}
87
+ {"step": 2150, "loss": 2.773498544692993, "grad_norm": 1.1243386268615723, "lr": 0.0001798910318912856, "s_per_step": 0.579986400604248, "peak_gb": 44.713082368, "samples_seen": 68800, "time": 1791132090.6597726}
88
+ {"step": 2175, "loss": 2.7247875928878784, "grad_norm": 1.1453449726104736, "lr": 0.00017936861987262475, "s_per_step": 0.6377030372619629, "peak_gb": 44.713082368, "samples_seen": 69600, "time": 1791132106.602996}
89
+ {"step": 2200, "loss": 2.758620800971985, "grad_norm": 1.0647848844528198, "lr": 0.0001788402879318618, "s_per_step": 0.583631067276001, "peak_gb": 44.713082368, "samples_seen": 70400, "time": 1791132121.194404}
90
+ {"step": 2225, "loss": 2.737944917678833, "grad_norm": 0.9973115921020508, "lr": 0.00017830607547605624, "s_per_step": 0.5805112552642823, "peak_gb": 44.713082368, "samples_seen": 71200, "time": 1791132135.7078857}
91
+ {"step": 2250, "loss": 2.7589354658126832, "grad_norm": 1.3257142305374146, "lr": 0.00017776602235088182, "s_per_step": 0.5818637657165527, "peak_gb": 44.713082368, "samples_seen": 72000, "time": 1791132150.255094}
92
+ {"step": 2275, "loss": 2.691265125274658, "grad_norm": 1.2526953220367432, "lr": 0.0001772201688376541, "s_per_step": 0.5970073795318603, "peak_gb": 44.713082368, "samples_seen": 72800, "time": 1791132165.180914}
93
+ {"step": 2300, "loss": 2.6809875345230103, "grad_norm": 1.1973172426223755, "lr": 0.00017666855565032642, "s_per_step": 0.581067304611206, "peak_gb": 44.713082368, "samples_seen": 73600, "time": 1791132179.7082095}
94
+ {"step": 2325, "loss": 2.7636598443984983, "grad_norm": 1.240551471710205, "lr": 0.00017611122393245275, "s_per_step": 0.5833901596069336, "peak_gb": 44.713082368, "samples_seen": 74400, "time": 1791132194.29364}
95
+ {"step": 2350, "loss": 2.6926284646987915, "grad_norm": 1.4566211700439453, "lr": 0.00017554821525411906, "s_per_step": 0.5827942657470703, "peak_gb": 44.713082368, "samples_seen": 75200, "time": 1791132208.8641186}
96
+ {"step": 2375, "loss": 2.7419660758972166, "grad_norm": 1.1245578527450562, "lr": 0.00017497957160884285, "s_per_step": 0.582745018005371, "peak_gb": 44.713082368, "samples_seen": 76000, "time": 1791132223.4334397}
97
+ {"step": 2400, "loss": 2.663253164291382, "grad_norm": 1.3107283115386963, "lr": 0.00017440533541044056, "s_per_step": 0.5869213008880615, "peak_gb": 44.713082368, "samples_seen": 76800, "time": 1791132238.1070802}
98
+ {"step": 2425, "loss": 2.719187321662903, "grad_norm": 1.1205676794052124, "lr": 0.00017382554948986445, "s_per_step": 0.5820108604431152, "peak_gb": 44.713082368, "samples_seen": 77600, "time": 1791132252.6579833}
99
+ {"step": 2450, "loss": 2.7780220317840576, "grad_norm": 1.3026337623596191, "lr": 0.00017324025709200766, "s_per_step": 0.5817010879516602, "peak_gb": 44.713082368, "samples_seen": 78400, "time": 1791132267.2012224}
100
+ {"step": 2475, "loss": 2.670451397895813, "grad_norm": 1.2014914751052856, "lr": 0.00017264950187247875, "s_per_step": 0.5823606967926025, "peak_gb": 44.713082368, "samples_seen": 79200, "time": 1791132281.7608585}
101
+ {"step": 2500, "loss": 2.716837215423584, "grad_norm": 1.10873281955719, "lr": 0.00017205332789434557, "s_per_step": 0.5853345394134521, "peak_gb": 44.713082368, "samples_seen": 80000, "time": 1791132296.3948615}
102
+ {"step": 2525, "loss": 2.7432607793807984, "grad_norm": 1.098906397819519, "lr": 0.00017145177962484857, "s_per_step": 0.5997240447998047, "peak_gb": 44.713082368, "samples_seen": 80800, "time": 1791132311.388538}
103
+ {"step": 2550, "loss": 2.72246383190155, "grad_norm": 1.1802773475646973, "lr": 0.00017084490193208435, "s_per_step": 0.5805338096618652, "peak_gb": 44.713082368, "samples_seen": 81600, "time": 1791132325.902455}
104
+ {"step": 2575, "loss": 2.7470813512802126, "grad_norm": 1.0650112628936768, "lr": 0.00017023274008165878, "s_per_step": 0.5804518604278565, "peak_gb": 44.713082368, "samples_seen": 82400, "time": 1791132340.4143882}
105
+ {"step": 2600, "loss": 2.733012390136719, "grad_norm": 1.1363471746444702, "lr": 0.00016961533973331081, "s_per_step": 0.5814594936370849, "peak_gb": 44.713082368, "samples_seen": 83200, "time": 1791132354.9514756}
106
+ {"step": 2625, "loss": 2.6893992137908938, "grad_norm": 1.1025614738464355, "lr": 0.00016899274693750696, "s_per_step": 0.5805016994476319, "peak_gb": 44.713082368, "samples_seen": 84000, "time": 1791132369.4646068}
107
+ {"step": 2650, "loss": 2.6821892070770263, "grad_norm": 1.081494688987732, "lr": 0.00016836500813200632, "s_per_step": 0.5812886905670166, "peak_gb": 44.713082368, "samples_seen": 84800, "time": 1791132383.9974716}
108
+ {"step": 2675, "loss": 2.6333597564697264, "grad_norm": 1.1299649477005005, "lr": 0.00016773217013839705, "s_per_step": 0.5820088291168213, "peak_gb": 44.713082368, "samples_seen": 85600, "time": 1791132398.5483217}
109
+ {"step": 2700, "loss": 2.6820573711395266, "grad_norm": 1.2370376586914062, "lr": 0.0001670942801586039, "s_per_step": 0.5822090339660645, "peak_gb": 44.713082368, "samples_seen": 86400, "time": 1791132413.1041813}
110
+ {"step": 2725, "loss": 2.695576796531677, "grad_norm": 1.1963305473327637, "lr": 0.0001664513857713677, "s_per_step": 0.5848300170898437, "peak_gb": 44.713082368, "samples_seen": 87200, "time": 1791132427.7255838}
111
+ {"step": 2750, "loss": 2.668781318664551, "grad_norm": 1.170694351196289, "lr": 0.00016580353492869635, "s_per_step": 0.5830868148803711, "peak_gb": 44.713082368, "samples_seen": 88000, "time": 1791132442.3036287}
112
+ {"step": 2775, "loss": 2.6674189043045042, "grad_norm": 1.4151184558868408, "lr": 0.00016515077595228837, "s_per_step": 0.6010197257995605, "peak_gb": 44.713082368, "samples_seen": 88800, "time": 1791132457.330008}
113
+ {"step": 2800, "loss": 2.716653389930725, "grad_norm": 1.140033483505249, "lr": 0.0001644931575299287, "s_per_step": 0.5838634395599365, "peak_gb": 44.713082368, "samples_seen": 89600, "time": 1791132471.9272482}
114
+ {"step": 2825, "loss": 2.7051201915740966, "grad_norm": 1.2430063486099243, "lr": 0.0001638307287118571, "s_per_step": 0.5820985507965087, "peak_gb": 44.713082368, "samples_seen": 90400, "time": 1791132486.4804387}
115
+ {"step": 2850, "loss": 2.6607730865478514, "grad_norm": 1.1695795059204102, "lr": 0.00016316353890710958, "s_per_step": 0.5837668800354003, "peak_gb": 44.713082368, "samples_seen": 91200, "time": 1791132501.0753098}
116
+ {"step": 2875, "loss": 2.6640982151031496, "grad_norm": 1.0638611316680908, "lr": 0.00016249163787983325, "s_per_step": 0.5824350929260254, "peak_gb": 44.713082368, "samples_seen": 92000, "time": 1791132515.6368423}
117
+ {"step": 2900, "loss": 2.6686540985107423, "grad_norm": 1.0542538166046143, "lr": 0.00016181507574557436, "s_per_step": 0.5836594104766846, "peak_gb": 44.713082368, "samples_seen": 92800, "time": 1791132530.229168}
118
+ {"step": 2925, "loss": 2.6495557260513305, "grad_norm": 1.2142537832260132, "lr": 0.00016113390296754036, "s_per_step": 0.5849769687652588, "peak_gb": 44.713082368, "samples_seen": 93600, "time": 1791132544.8542972}
119
+ {"step": 2950, "loss": 2.6701879692077637, "grad_norm": 1.2481223344802856, "lr": 0.00016044817035283605, "s_per_step": 0.5815197849273681, "peak_gb": 44.713082368, "samples_seen": 94400, "time": 1791132559.3929307}
120
+ {"step": 2975, "loss": 2.729030981063843, "grad_norm": 0.9954540729522705, "lr": 0.00015975792904867388, "s_per_step": 0.5838561344146729, "peak_gb": 44.713082368, "samples_seen": 95200, "time": 1791132573.989938}
121
+ {"step": 3000, "loss": 2.649655590057373, "grad_norm": 1.0716608762741089, "lr": 0.00015906323053855898, "s_per_step": 0.5867070293426514, "peak_gb": 44.713082368, "samples_seen": 96000, "time": 1791132588.6582751}
122
+ {"step": 3025, "loss": 2.6586125326156616, "grad_norm": 1.4092375040054321, "lr": 0.0001583641266384493, "s_per_step": 0.604687147140503, "peak_gb": 44.713082368, "samples_seen": 96800, "time": 1791132603.776329}
123
+ {"step": 3050, "loss": 2.6345360708236694, "grad_norm": 1.3088021278381348, "lr": 0.00015766066949289052, "s_per_step": 0.5844075584411621, "peak_gb": 44.713082368, "samples_seen": 97600, "time": 1791132618.3874304}
124
+ {"step": 3075, "loss": 2.697037835121155, "grad_norm": 1.2829176187515259, "lr": 0.00015695291157112693, "s_per_step": 0.5823062896728516, "peak_gb": 44.713082368, "samples_seen": 98400, "time": 1791132632.9458525}
125
+ {"step": 3100, "loss": 2.6461993885040282, "grad_norm": 1.1458219289779663, "lr": 0.0001562409056631878, "s_per_step": 0.5842306709289551, "peak_gb": 44.713082368, "samples_seen": 99200, "time": 1791132647.552255}
126
+ {"step": 3125, "loss": 2.6994865417480467, "grad_norm": 1.0527013540267944, "lr": 0.00015552470487594983, "s_per_step": 0.5832872867584229, "peak_gb": 44.713082368, "samples_seen": 100000, "time": 1791132662.1353507}
127
+ {"step": 3150, "loss": 2.6455965662002563, "grad_norm": 1.210248351097107, "lr": 0.00015480436262917615, "s_per_step": 0.5897402572631836, "peak_gb": 44.713082368, "samples_seen": 100800, "time": 1791132676.879246}
128
+ {"step": 3175, "loss": 2.7154378509521484, "grad_norm": 1.226425051689148, "lr": 0.0001540799326515317, "s_per_step": 0.5857205295562744, "peak_gb": 44.713082368, "samples_seen": 101600, "time": 1791132691.5226414}
129
+ {"step": 3200, "loss": 2.60671639919281, "grad_norm": 1.0465337038040161, "lr": 0.00015335146897657587, "s_per_step": 0.5847133922576905, "peak_gb": 44.713082368, "samples_seen": 102400, "time": 1791132706.1408818}
130
+ {"step": 3225, "loss": 2.5945321559906005, "grad_norm": 1.1408718824386597, "lr": 0.00015261902593873222, "s_per_step": 0.5843933010101319, "peak_gb": 44.713082368, "samples_seen": 103200, "time": 1791132720.7511504}
131
+ {"step": 3250, "loss": 2.627700114250183, "grad_norm": 1.0017207860946655, "lr": 0.00015188265816923585, "s_per_step": 0.5855321598052978, "peak_gb": 44.713082368, "samples_seen": 104000, "time": 1791132735.3898623}
132
+ {"step": 3275, "loss": 2.6447770547866822, "grad_norm": 1.066293478012085, "lr": 0.0001511424205920585, "s_per_step": 0.6015197467803955, "peak_gb": 44.713082368, "samples_seen": 104800, "time": 1791132750.4282053}
133
+ {"step": 3300, "loss": 2.6553041982650756, "grad_norm": 1.275646686553955, "lr": 0.00015039836841981185, "s_per_step": 0.5847999572753906, "peak_gb": 44.713082368, "samples_seen": 105600, "time": 1791132765.0488286}
134
+ {"step": 3325, "loss": 2.5941198587417604, "grad_norm": 1.0782554149627686, "lr": 0.00014965055714962952, "s_per_step": 0.5845184135437012, "peak_gb": 44.713082368, "samples_seen": 106400, "time": 1791132779.6623993}
135
+ {"step": 3350, "loss": 2.628866515159607, "grad_norm": 1.201255440711975, "lr": 0.0001488990425590276, "s_per_step": 0.5825023555755615, "peak_gb": 44.713082368, "samples_seen": 107200, "time": 1791132794.225617}
136
+ {"step": 3375, "loss": 2.6890606594085695, "grad_norm": 1.0845966339111328, "lr": 0.00014814388070174408, "s_per_step": 0.5846240043640136, "peak_gb": 44.713082368, "samples_seen": 108000, "time": 1791132808.841806}
137
+ {"step": 3400, "loss": 2.590783829689026, "grad_norm": 1.1155602931976318, "lr": 0.00014738512790355842, "s_per_step": 0.5839724159240722, "peak_gb": 44.713082368, "samples_seen": 108800, "time": 1791132823.4414928}
138
+ {"step": 3425, "loss": 2.6391691303253175, "grad_norm": 1.0816454887390137, "lr": 0.00014662284075808995, "s_per_step": 0.5850570201873779, "peak_gb": 44.713082368, "samples_seen": 109600, "time": 1791132838.0683677}
139
+ {"step": 3450, "loss": 2.6323579359054565, "grad_norm": 1.131683588027954, "lr": 0.00014585707612257672, "s_per_step": 0.5823802852630615, "peak_gb": 44.713082368, "samples_seen": 110400, "time": 1791132852.6284838}
140
+ {"step": 3475, "loss": 2.637928156852722, "grad_norm": 0.9959607720375061, "lr": 0.00014508789111363494, "s_per_step": 0.5853830051422119, "peak_gb": 44.713082368, "samples_seen": 111200, "time": 1791132867.2637053}
141
+ {"step": 3500, "loss": 2.65216420173645, "grad_norm": 1.3305087089538574, "lr": 0.00014431534310299838, "s_per_step": 0.5855358028411866, "peak_gb": 44.713082368, "samples_seen": 112000, "time": 1791132881.9027245}
142
+ {"step": 3525, "loss": 2.6313996505737305, "grad_norm": 1.2155920267105103, "lr": 0.00014353948971323942, "s_per_step": 0.6034837055206299, "peak_gb": 44.713082368, "samples_seen": 112800, "time": 1791132896.9904048}
143
+ {"step": 3550, "loss": 2.6454056882858277, "grad_norm": 1.2063369750976562, "lr": 0.00014276038881347105, "s_per_step": 0.5832828998565673, "peak_gb": 44.713082368, "samples_seen": 113600, "time": 1791132911.5728936}
144
+ {"step": 3575, "loss": 2.6533055901527405, "grad_norm": 1.2403618097305298, "lr": 0.00014197809851503049, "s_per_step": 0.5824703311920166, "peak_gb": 44.713082368, "samples_seen": 114400, "time": 1791132926.135257}
145
+ {"step": 3600, "loss": 2.645680537223816, "grad_norm": 1.168968915939331, "lr": 0.0001411926771671449, "s_per_step": 0.5845646953582764, "peak_gb": 44.713082368, "samples_seen": 115200, "time": 1791132940.7499871}
146
+ {"step": 3625, "loss": 2.5960699319839478, "grad_norm": 1.1931697130203247, "lr": 0.0001404041833525792, "s_per_step": 0.5853945255279541, "peak_gb": 44.713082368, "samples_seen": 116000, "time": 1791132955.3854115}
147
+ {"step": 3650, "loss": 2.6213397121429445, "grad_norm": 0.9856603741645813, "lr": 0.00013961267588326635, "s_per_step": 0.5852436065673828, "peak_gb": 44.713082368, "samples_seen": 116800, "time": 1791132970.017149}
148
+ {"step": 3675, "loss": 2.5998027467727662, "grad_norm": 1.1824589967727661, "lr": 0.0001388182137959211, "s_per_step": 0.5843393993377686, "peak_gb": 44.713082368, "samples_seen": 117600, "time": 1791132984.6262264}
149
+ {"step": 3700, "loss": 2.629755802154541, "grad_norm": 1.2840546369552612, "lr": 0.00013802085634763612, "s_per_step": 0.5829662704467773, "peak_gb": 44.713082368, "samples_seen": 118400, "time": 1791132999.2009673}
150
+ {"step": 3725, "loss": 2.6009298038482664, "grad_norm": 1.2750321626663208, "lr": 0.00013722066301146254, "s_per_step": 0.5825348472595215, "peak_gb": 44.713082368, "samples_seen": 119200, "time": 1791133013.764688}
151
+ {"step": 3750, "loss": 2.612462890148163, "grad_norm": 1.1066721677780151, "lr": 0.00013641769347197367, "s_per_step": 0.5850222778320312, "peak_gb": 44.713082368, "samples_seen": 120000, "time": 1791133028.3908353}
152
+ {"step": 3775, "loss": 2.591866898536682, "grad_norm": 1.233768343925476, "lr": 0.00013561200762081353, "s_per_step": 0.6027588081359864, "peak_gb": 44.713082368, "samples_seen": 120800, "time": 1791133043.4603734}
153
+ {"step": 3800, "loss": 2.578413248062134, "grad_norm": 1.108993649482727, "lr": 0.00013480366555222948, "s_per_step": 0.5819605350494385, "peak_gb": 44.713082368, "samples_seen": 121600, "time": 1791133058.0100117}
154
+ {"step": 3825, "loss": 2.547907509803772, "grad_norm": 0.9940484166145325, "lr": 0.00013399272755859002, "s_per_step": 0.5811294841766358, "peak_gb": 44.713082368, "samples_seen": 122400, "time": 1791133072.5388749}
155
+ {"step": 3850, "loss": 2.564617967605591, "grad_norm": 1.148879885673523, "lr": 0.00013317925412588777, "s_per_step": 0.5868311309814453, "peak_gb": 44.713082368, "samples_seen": 123200, "time": 1791133087.210223}
156
+ {"step": 3875, "loss": 2.587318305969238, "grad_norm": 1.0706918239593506, "lr": 0.00013236330592922784, "s_per_step": 0.5894640159606933, "peak_gb": 44.713082368, "samples_seen": 124000, "time": 1791133101.9471831}
157
+ {"step": 3900, "loss": 2.599490542411804, "grad_norm": 1.0891342163085938, "lr": 0.00013154494382830226, "s_per_step": 0.5840770053863525, "peak_gb": 44.713082368, "samples_seen": 124800, "time": 1791133116.5494475}
158
+ {"step": 3925, "loss": 2.628100709915161, "grad_norm": 1.1885673999786377, "lr": 0.00013072422886285063, "s_per_step": 0.5845713424682617, "peak_gb": 44.713082368, "samples_seen": 125600, "time": 1791133131.1641}
159
+ {"step": 3950, "loss": 2.6188903713226317, "grad_norm": 1.0750778913497925, "lr": 0.00012990122224810726, "s_per_step": 0.5842515659332276, "peak_gb": 44.713082368, "samples_seen": 126400, "time": 1791133145.7707798}
160
+ {"step": 3975, "loss": 2.577172713279724, "grad_norm": 1.1962721347808838, "lr": 0.00012907598537023528, "s_per_step": 0.5835914134979248, "peak_gb": 44.713082368, "samples_seen": 127200, "time": 1791133160.360931}
161
+ {"step": 4000, "loss": 2.5662565279006957, "grad_norm": 1.1467726230621338, "lr": 0.00012824857978174808, "s_per_step": 0.5806757926940918, "peak_gb": 44.713082368, "samples_seen": 128000, "time": 1791133174.8781514}
162
+ {"step": 4025, "loss": 2.6355027961730957, "grad_norm": 1.1590839624404907, "lr": 0.00012741906719691812, "s_per_step": 0.595490140914917, "peak_gb": 44.713082368, "samples_seen": 128800, "time": 1791133189.7657502}
163
+ {"step": 4050, "loss": 2.5592614030838012, "grad_norm": 0.9599524140357971, "lr": 0.0001265875094871738, "s_per_step": 0.5824095058441162, "peak_gb": 44.713082368, "samples_seen": 129600, "time": 1791133204.3263824}
164
+ {"step": 4075, "loss": 2.5762628889083863, "grad_norm": 0.9201562404632568, "lr": 0.0001257539686764847, "s_per_step": 0.5829432201385498, "peak_gb": 44.713082368, "samples_seen": 130400, "time": 1791133218.9003162}
165
+ {"step": 4100, "loss": 2.6393421840667726, "grad_norm": 1.352101445198059, "lr": 0.0001249185069367353, "s_per_step": 0.5822771739959717, "peak_gb": 44.713082368, "samples_seen": 131200, "time": 1791133233.4575992}
166
+ {"step": 4125, "loss": 2.545742371082306, "grad_norm": 1.2281466722488403, "lr": 0.00012408118658308783, "s_per_step": 0.5824446392059326, "peak_gb": 44.713082368, "samples_seen": 132000, "time": 1791133248.0190716}
167
+ {"step": 4150, "loss": 2.588506646156311, "grad_norm": 1.1411926746368408, "lr": 0.00012324207006933417, "s_per_step": 0.5804715347290039, "peak_gb": 44.713082368, "samples_seen": 132800, "time": 1791133262.5312164}
168
+ {"step": 4175, "loss": 2.5443468284606934, "grad_norm": 1.0073899030685425, "lr": 0.00012240121998323766, "s_per_step": 0.58156418800354, "peak_gb": 44.713082368, "samples_seen": 133600, "time": 1791133277.0706964}
169
+ {"step": 4200, "loss": 2.5501110649108885, "grad_norm": 1.205407977104187, "lr": 0.00012155869904186475, "s_per_step": 0.5811350154876709, "peak_gb": 44.713082368, "samples_seen": 134400, "time": 1791133291.5995052}
170
+ {"step": 4225, "loss": 2.602772789001465, "grad_norm": 1.2279781103134155, "lr": 0.00012071457008690715, "s_per_step": 0.5812279796600341, "peak_gb": 44.713082368, "samples_seen": 135200, "time": 1791133306.1305652}
171
+ {"step": 4250, "loss": 2.6113069915771483, "grad_norm": 1.1278338432312012, "lr": 0.00011986889607999463, "s_per_step": 0.580757417678833, "peak_gb": 44.713082368, "samples_seen": 136000, "time": 1791133320.6498582}
172
+ {"step": 4275, "loss": 2.566324009895325, "grad_norm": 1.2942112684249878, "lr": 0.00011902174009799881, "s_per_step": 0.5944104099273682, "peak_gb": 44.713082368, "samples_seen": 136800, "time": 1791133335.5104773}
173
+ {"step": 4300, "loss": 2.5078661108016966, "grad_norm": 1.2482675313949585, "lr": 0.00011817316532832833, "s_per_step": 0.5815795421600342, "peak_gb": 44.713082368, "samples_seen": 137600, "time": 1791133350.050364}
174
+ {"step": 4325, "loss": 2.6592298412322997, "grad_norm": 1.0717958211898804, "lr": 0.00011732323506421603, "s_per_step": 0.583767204284668, "peak_gb": 44.713082368, "samples_seen": 138400, "time": 1791133364.644909}
175
+ {"step": 4350, "loss": 2.6119984006881714, "grad_norm": 1.00667142868042, "lr": 0.00011647201269999792, "s_per_step": 0.5827497482299805, "peak_gb": 44.713082368, "samples_seen": 139200, "time": 1791133379.2140245}
176
+ {"step": 4375, "loss": 2.5818211317062376, "grad_norm": 1.1398861408233643, "lr": 0.0001156195617263847, "s_per_step": 0.6331087684631348, "peak_gb": 44.713082368, "samples_seen": 140000, "time": 1791133395.0420854}
177
+ {"step": 4400, "loss": 2.5258120822906496, "grad_norm": 1.2284903526306152, "lr": 0.00011476594572572632, "s_per_step": 0.5820397758483886, "peak_gb": 44.713082368, "samples_seen": 140800, "time": 1791133409.5934033}
178
+ {"step": 4425, "loss": 2.5755055570602416, "grad_norm": 1.141483187675476, "lr": 0.00011391122836726932, "s_per_step": 0.580828332901001, "peak_gb": 44.713082368, "samples_seen": 141600, "time": 1791133424.1144583}
179
+ {"step": 4450, "loss": 2.566441774368286, "grad_norm": 1.1366175413131714, "lr": 0.00011305547340240805, "s_per_step": 0.5789041805267334, "peak_gb": 44.713082368, "samples_seen": 142400, "time": 1791133438.5874844}
180
+ {"step": 4475, "loss": 2.561001148223877, "grad_norm": 1.0844303369522095, "lr": 0.00011219874465992948, "s_per_step": 0.5818092918395996, "peak_gb": 44.713082368, "samples_seen": 143200, "time": 1791133453.13306}
181
+ {"step": 4500, "loss": 2.5567925786972046, "grad_norm": 1.1574947834014893, "lr": 0.00011134110604125239, "s_per_step": 0.5787328243255615, "peak_gb": 44.713082368, "samples_seen": 144000, "time": 1791133467.6017637}
182
+ {"step": 4525, "loss": 2.590312509536743, "grad_norm": 1.257983684539795, "lr": 0.00011048262151566114, "s_per_step": 0.6001885604858398, "peak_gb": 44.713082368, "samples_seen": 144800, "time": 1791133482.606845}
183
+ {"step": 4550, "loss": 2.629988832473755, "grad_norm": 1.1604526042938232, "lr": 0.00010962335511553434, "s_per_step": 0.5829675579071045, "peak_gb": 44.713082368, "samples_seen": 145600, "time": 1791133497.1814847}
184
+ {"step": 4575, "loss": 2.5439812088012697, "grad_norm": 1.1712712049484253, "lr": 0.00010876337093156886, "s_per_step": 0.5811667346954346, "peak_gb": 44.713082368, "samples_seen": 146400, "time": 1791133511.711039}
185
+ {"step": 4600, "loss": 2.5177254247665406, "grad_norm": 1.1479181051254272, "lr": 0.00010790273310799933, "s_per_step": 0.5817756175994873, "peak_gb": 44.713082368, "samples_seen": 147200, "time": 1791133526.2557957}
186
+ {"step": 4625, "loss": 2.5491437196731566, "grad_norm": 1.2170847654342651, "lr": 0.00010704150583781388, "s_per_step": 0.5829072570800782, "peak_gb": 44.713082368, "samples_seen": 148000, "time": 1791133540.8288167}
187
+ {"step": 4650, "loss": 2.6045084619522094, "grad_norm": 1.189692497253418, "lr": 0.00010617975335796613, "s_per_step": 0.5828746032714843, "peak_gb": 44.713082368, "samples_seen": 148800, "time": 1791133555.4010704}
188
+ {"step": 4675, "loss": 2.576749534606934, "grad_norm": 1.1746689081192017, "lr": 0.00010531753994458382, "s_per_step": 0.5798631191253663, "peak_gb": 44.713082368, "samples_seen": 149600, "time": 1791133569.898002}
189
+ {"step": 4700, "loss": 2.5879501485824585, "grad_norm": 1.182418704032898, "lr": 0.00010445492990817471, "s_per_step": 0.58068115234375, "peak_gb": 44.713082368, "samples_seen": 150400, "time": 1791133584.415375}
190
+ {"step": 4725, "loss": 2.5357809329032897, "grad_norm": 1.4740160703659058, "lr": 0.00010359198758882973, "s_per_step": 0.5839321517944336, "peak_gb": 44.713082368, "samples_seen": 151200, "time": 1791133599.0140464}
191
+ {"step": 4750, "loss": 2.55805965423584, "grad_norm": 1.172712802886963, "lr": 0.00010272877735142407, "s_per_step": 0.5821543216705323, "peak_gb": 44.713082368, "samples_seen": 152000, "time": 1791133613.568319}
192
+ {"step": 4775, "loss": 2.5926496958732606, "grad_norm": 1.194968581199646, "lr": 0.00010186536358081623, "s_per_step": 0.5943687915802002, "peak_gb": 44.713082368, "samples_seen": 152800, "time": 1791133628.427965}
193
+ {"step": 4800, "loss": 2.5890002751350405, "grad_norm": 1.1471878290176392, "lr": 0.00010100181067704579, "s_per_step": 0.5851948833465577, "peak_gb": 44.713082368, "samples_seen": 153600, "time": 1791133643.0582457}
194
+ {"step": 4825, "loss": 2.5580673170089723, "grad_norm": 1.1443222761154175, "lr": 0.00010013818305053007, "s_per_step": 0.5806027603149414, "peak_gb": 44.713082368, "samples_seen": 154400, "time": 1791133657.573739}
195
+ {"step": 4850, "loss": 2.5510888147354125, "grad_norm": 1.1550148725509644, "lr": 9.927454511725963e-05, "s_per_step": 0.5874297904968262, "peak_gb": 44.947974656, "samples_seen": 155200, "time": 1791133672.2599056}
196
+ {"step": 4875, "loss": 2.5219507360458375, "grad_norm": 1.0392953157424927, "lr": 9.841096129399391e-05, "s_per_step": 0.5823034572601319, "peak_gb": 44.947974656, "samples_seen": 156000, "time": 1791133686.8179238}
197
+ {"step": 4900, "loss": 2.5420716381073, "grad_norm": 1.2059009075164795, "lr": 9.754749599345636e-05, "s_per_step": 0.585919017791748, "peak_gb": 44.947974656, "samples_seen": 156800, "time": 1791133701.4663377}
198
+ {"step": 4925, "loss": 2.5494388055801394, "grad_norm": 1.132745385169983, "lr": 9.668421361953004e-05, "s_per_step": 0.5843277454376221, "peak_gb": 44.947974656, "samples_seen": 157600, "time": 1791133716.0755186}
199
+ {"step": 4950, "loss": 2.534179072380066, "grad_norm": 1.1565256118774414, "lr": 9.582117856245404e-05, "s_per_step": 0.5862970352172852, "peak_gb": 44.947974656, "samples_seen": 158400, "time": 1791133730.7334623}
200
+ {"step": 4975, "loss": 2.575369110107422, "grad_norm": 1.273497223854065, "lr": 9.495845519402057e-05, "s_per_step": 0.582844295501709, "peak_gb": 44.947974656, "samples_seen": 159200, "time": 1791133745.3050113}
201
+ {"step": 5000, "loss": 2.5750187397003175, "grad_norm": 1.1561633348464966, "lr": 9.409610786277377e-05, "s_per_step": 0.5810141372680664, "peak_gb": 44.947974656, "samples_seen": 160000, "time": 1791133759.830802}
202
+ {"step": 5025, "loss": 2.560807399749756, "grad_norm": 1.3111895322799683, "lr": 9.323420088921002e-05, "s_per_step": 0.595534782409668, "peak_gb": 44.947974656, "samples_seen": 160800, "time": 1791133774.7197492}
203
+ {"step": 5050, "loss": 2.5090624332427978, "grad_norm": 1.128811240196228, "lr": 9.237279856098037e-05, "s_per_step": 0.5866300392150879, "peak_gb": 44.947974656, "samples_seen": 161600, "time": 1791133789.385984}
204
+ {"step": 5075, "loss": 2.5157794332504273, "grad_norm": 1.1408394575119019, "lr": 9.15119651280956e-05, "s_per_step": 0.5873813438415527, "peak_gb": 44.947974656, "samples_seen": 162400, "time": 1791133804.070975}
205
+ {"step": 5100, "loss": 2.4853475761413573, "grad_norm": 1.1142104864120483, "lr": 9.065176479813393e-05, "s_per_step": 0.5891704750061035, "peak_gb": 44.947974656, "samples_seen": 163200, "time": 1791133818.8006737}
206
+ {"step": 5125, "loss": 2.535648889541626, "grad_norm": 1.1073273420333862, "lr": 8.979226173145183e-05, "s_per_step": 0.5848058700561524, "peak_gb": 44.947974656, "samples_seen": 164000, "time": 1791133833.421258}
207
+ {"step": 5150, "loss": 2.582467722892761, "grad_norm": 1.2277833223342896, "lr": 8.893352003639856e-05, "s_per_step": 0.5909334659576416, "peak_gb": 44.947974656, "samples_seen": 164800, "time": 1791133848.1950076}
208
+ {"step": 5175, "loss": 2.624933614730835, "grad_norm": 1.0772781372070312, "lr": 8.807560376453437e-05, "s_per_step": 0.580198278427124, "peak_gb": 44.947974656, "samples_seen": 165600, "time": 1791133862.700406}
209
+ {"step": 5200, "loss": 2.5316876649856566, "grad_norm": 1.251028299331665, "lr": 8.721857690585314e-05, "s_per_step": 0.5918057537078858, "peak_gb": 44.947974656, "samples_seen": 166400, "time": 1791133877.495987}
210
+ {"step": 5225, "loss": 2.5712487649917604, "grad_norm": 1.1420977115631104, "lr": 8.636250338400954e-05, "s_per_step": 0.5819726657867431, "peak_gb": 44.947974656, "samples_seen": 167200, "time": 1791133892.0457838}
211
+ {"step": 5250, "loss": 2.5827745389938355, "grad_norm": 1.37429678440094, "lr": 8.550744705155087e-05, "s_per_step": 0.5890950107574463, "peak_gb": 44.947974656, "samples_seen": 168000, "time": 1791133906.7736533}
212
+ {"step": 5275, "loss": 2.493775963783264, "grad_norm": 1.1473252773284912, "lr": 8.465347168515478e-05, "s_per_step": 0.5988393306732178, "peak_gb": 44.947974656, "samples_seen": 168800, "time": 1791133921.745074}
213
+ {"step": 5300, "loss": 2.5703449940681455, "grad_norm": 1.315991997718811, "lr": 8.380064098087212e-05, "s_per_step": 0.5884280204772949, "peak_gb": 44.947974656, "samples_seen": 169600, "time": 1791133936.4563115}
214
+ {"step": 5325, "loss": 2.5584110116958616, "grad_norm": 1.3650532960891724, "lr": 8.294901854937599e-05, "s_per_step": 0.5825414657592773, "peak_gb": 44.947974656, "samples_seen": 170400, "time": 1791133951.0203662}
215
+ {"step": 5350, "loss": 2.5070259499549867, "grad_norm": 1.0850861072540283, "lr": 8.20986679112173e-05, "s_per_step": 0.5887926959991455, "peak_gb": 44.947974656, "samples_seen": 171200, "time": 1791133965.7406182}
216
+ {"step": 5375, "loss": 2.5562100505828855, "grad_norm": 1.1863794326782227, "lr": 8.124965249208671e-05, "s_per_step": 0.5837687969207763, "peak_gb": 44.947974656, "samples_seen": 172000, "time": 1791133980.3353226}
217
+ {"step": 5400, "loss": 2.5509732866287234, "grad_norm": 1.3163355588912964, "lr": 8.040203561808406e-05, "s_per_step": 0.5839274787902832, "peak_gb": 44.947974656, "samples_seen": 172800, "time": 1791133994.934007}
218
+ {"step": 5425, "loss": 2.518836889266968, "grad_norm": 1.0753092765808105, "lr": 7.955588051099495e-05, "s_per_step": 0.5834451007843018, "peak_gb": 44.947974656, "samples_seen": 173600, "time": 1791134009.5206103}
219
+ {"step": 5450, "loss": 2.486795573234558, "grad_norm": 1.2447965145111084, "lr": 7.87112502835751e-05, "s_per_step": 0.590032205581665, "peak_gb": 45.225538048, "samples_seen": 174400, "time": 1791134024.2720983}
220
+ {"step": 5475, "loss": 2.5397833967208863, "grad_norm": 1.2422091960906982, "lr": 7.786820793484306e-05, "s_per_step": 0.5841742324829101, "peak_gb": 45.225538048, "samples_seen": 175200, "time": 1791134038.8771572}
221
+ {"step": 5500, "loss": 2.582821946144104, "grad_norm": 1.1261637210845947, "lr": 7.702681634538104e-05, "s_per_step": 0.584408836364746, "peak_gb": 45.225538048, "samples_seen": 176000, "time": 1791134053.4881816}
222
+ {"step": 5525, "loss": 2.5332029485702514, "grad_norm": 1.3875887393951416, "lr": 7.618713827264505e-05, "s_per_step": 0.6076767444610596, "peak_gb": 45.225538048, "samples_seen": 176800, "time": 1791134068.680784}
223
+ {"step": 5550, "loss": 2.577248592376709, "grad_norm": 1.1470293998718262, "lr": 7.534923634628381e-05, "s_per_step": 0.582500295639038, "peak_gb": 45.225538048, "samples_seen": 177600, "time": 1791134083.2440093}
224
+ {"step": 5575, "loss": 2.521993110179901, "grad_norm": 1.2665389776229858, "lr": 7.451317306346738e-05, "s_per_step": 0.5894964313507081, "peak_gb": 45.225538048, "samples_seen": 178400, "time": 1791134097.9821842}
225
+ {"step": 5600, "loss": 2.538360843658447, "grad_norm": 1.314635992050171, "lr": 7.367901078422563e-05, "s_per_step": 0.5822025871276856, "peak_gb": 45.225538048, "samples_seen": 179200, "time": 1791134112.5379786}
226
+ {"step": 5625, "loss": 2.519887990951538, "grad_norm": 1.2479692697525024, "lr": 7.284681172679697e-05, "s_per_step": 0.5847748279571533, "peak_gb": 45.225538048, "samples_seen": 180000, "time": 1791134127.1580513}
227
+ {"step": 5650, "loss": 2.55150728225708, "grad_norm": 1.3254190683364868, "lr": 7.201663796298764e-05, "s_per_step": 0.5816287231445313, "peak_gb": 45.225538048, "samples_seen": 180800, "time": 1791134141.6998255}
228
+ {"step": 5675, "loss": 2.480019929409027, "grad_norm": 1.3282145261764526, "lr": 7.118855141354186e-05, "s_per_step": 0.5884755229949952, "peak_gb": 45.225538048, "samples_seen": 181600, "time": 1791134156.4120905}
229
+ {"step": 5700, "loss": 2.495427899360657, "grad_norm": 1.325265884399414, "lr": 7.036261384352343e-05, "s_per_step": 0.5823914813995361, "peak_gb": 45.225538048, "samples_seen": 182400, "time": 1791134170.972266}
230
+ {"step": 5725, "loss": 2.478832368850708, "grad_norm": 1.0613092184066772, "lr": 6.953888685770869e-05, "s_per_step": 0.5816488647460938, "peak_gb": 45.225538048, "samples_seen": 183200, "time": 1791134185.5139039}
231
+ {"step": 5750, "loss": 2.52920090675354, "grad_norm": 1.207517385482788, "lr": 6.871743189599156e-05, "s_per_step": 0.5833349800109864, "peak_gb": 45.225538048, "samples_seen": 184000, "time": 1791134200.0977366}
232
+ {"step": 5775, "loss": 2.5796267938613893, "grad_norm": 1.1419939994812012, "lr": 6.789831022880101e-05, "s_per_step": 0.5960377788543701, "peak_gb": 45.225538048, "samples_seen": 184800, "time": 1791134214.9990556}
233
+ {"step": 5800, "loss": 2.52916264295578, "grad_norm": 1.3384249210357666, "lr": 6.708158295253092e-05, "s_per_step": 0.5907851696014405, "peak_gb": 45.225538048, "samples_seen": 185600, "time": 1791134229.7691712}
234
+ {"step": 5825, "loss": 2.5286863899230956, "grad_norm": 1.351222276687622, "lr": 6.626731098498311e-05, "s_per_step": 0.5822148704528809, "peak_gb": 45.225538048, "samples_seen": 186400, "time": 1791134244.3249931}
235
+ {"step": 5850, "loss": 2.5227443981170654, "grad_norm": 1.1627928018569946, "lr": 6.545555506082355e-05, "s_per_step": 0.5858915615081787, "peak_gb": 45.225538048, "samples_seen": 187200, "time": 1791134258.9727561}
236
+ {"step": 5875, "loss": 2.534986820220947, "grad_norm": 1.2816951274871826, "lr": 6.464637572705237e-05, "s_per_step": 0.5801717948913574, "peak_gb": 45.225538048, "samples_seen": 188000, "time": 1791134273.4775524}
237
+ {"step": 5900, "loss": 2.489676251411438, "grad_norm": 1.4309545755386353, "lr": 6.383983333848773e-05, "s_per_step": 0.5923868274688721, "peak_gb": 45.225538048, "samples_seen": 188800, "time": 1791134288.2877378}
238
+ {"step": 5925, "loss": 2.4901873302459716, "grad_norm": 1.322837471961975, "lr": 6.303598805326419e-05, "s_per_step": 0.5796713924407959, "peak_gb": 45.225538048, "samples_seen": 189600, "time": 1791134302.7800236}
239
+ {"step": 5950, "loss": 2.53445716381073, "grad_norm": 1.3041024208068848, "lr": 6.223489982834558e-05, "s_per_step": 0.5908190441131592, "peak_gb": 45.225538048, "samples_seen": 190400, "time": 1791134317.550958}
240
+ {"step": 5975, "loss": 2.523839635848999, "grad_norm": 1.1598639488220215, "lr": 6.1436628415053e-05, "s_per_step": 0.5810464763641358, "peak_gb": 45.225538048, "samples_seen": 191200, "time": 1791134332.0775537}
241
+ {"step": 6000, "loss": 2.513823781013489, "grad_norm": 1.0870771408081055, "lr": 6.064123335460796e-05, "s_per_step": 0.5868215465545654, "peak_gb": 45.225538048, "samples_seen": 192000, "time": 1791134346.748537}
242
+ {"step": 6025, "loss": 2.589872555732727, "grad_norm": 1.2455843687057495, "lr": 5.9848773973691574e-05, "s_per_step": 0.5989699459075928, "peak_gb": 45.225538048, "samples_seen": 192800, "time": 1791134361.7232323}
243
+ {"step": 6050, "loss": 2.5560085201263427, "grad_norm": 1.2540245056152344, "lr": 5.905930938001937e-05, "s_per_step": 0.5814158630371093, "peak_gb": 45.225538048, "samples_seen": 193600, "time": 1791134376.2590868}
244
+ {"step": 6075, "loss": 2.5238016700744628, "grad_norm": 1.5046987533569336, "lr": 5.827289845793256e-05, "s_per_step": 0.5851035404205323, "peak_gb": 45.225538048, "samples_seen": 194400, "time": 1791134390.8871312}
245
+ {"step": 6100, "loss": 2.528533492088318, "grad_norm": 1.174469232559204, "lr": 5.748959986400611e-05, "s_per_step": 0.5847522258758545, "peak_gb": 45.225538048, "samples_seen": 195200, "time": 1791134405.5064287}
246
+ {"step": 6125, "loss": 2.5303329944610597, "grad_norm": 1.1631888151168823, "lr": 5.670947202267356e-05, "s_per_step": 0.5879807376861572, "peak_gb": 45.225538048, "samples_seen": 196000, "time": 1791134420.2064354}
247
+ {"step": 6150, "loss": 2.5317347621917725, "grad_norm": 1.316127896308899, "lr": 5.593257312186937e-05, "s_per_step": 0.5837984085083008, "peak_gb": 45.225538048, "samples_seen": 196800, "time": 1791134434.8018806}
248
+ {"step": 6175, "loss": 2.434650685787201, "grad_norm": 1.1741480827331543, "lr": 5.5158961108688766e-05, "s_per_step": 0.5790462970733643, "peak_gb": 45.225538048, "samples_seen": 197600, "time": 1791134449.2784874}
249
+ {"step": 6200, "loss": 2.531181104183197, "grad_norm": 1.1735306978225708, "lr": 5.438869368506563e-05, "s_per_step": 0.584972219467163, "peak_gb": 45.225538048, "samples_seen": 198400, "time": 1791134463.9032767}
250
+ {"step": 6225, "loss": 2.527851085662842, "grad_norm": 1.0472126007080078, "lr": 5.36218283034686e-05, "s_per_step": 0.5841711044311524, "peak_gb": 45.225538048, "samples_seen": 199200, "time": 1791134478.5080554}
251
+ {"step": 6250, "loss": 2.4174211168289186, "grad_norm": 1.081107258796692, "lr": 5.285842216261594e-05, "s_per_step": 0.5814971351623535, "peak_gb": 45.225538048, "samples_seen": 200000, "time": 1791134493.045984}
252
+ {"step": 6275, "loss": 2.438464779853821, "grad_norm": 1.2638477087020874, "lr": 5.209853220320898e-05, "s_per_step": 0.5986836051940918, "peak_gb": 45.225538048, "samples_seen": 200800, "time": 1791134508.0135267}
253
+ {"step": 6300, "loss": 2.541532530784607, "grad_norm": 1.1554198265075684, "lr": 5.134221510368531e-05, "s_per_step": 0.5808733463287353, "peak_gb": 45.225538048, "samples_seen": 201600, "time": 1791134522.535798}
254
+ {"step": 6325, "loss": 2.5302062940597536, "grad_norm": 1.0484192371368408, "lr": 5.058952727599113e-05, "s_per_step": 0.5848421478271484, "peak_gb": 45.225538048, "samples_seen": 202400, "time": 1791134537.1572866}
255
+ {"step": 6350, "loss": 2.5093615889549254, "grad_norm": 1.3501886129379272, "lr": 4.984052486137364e-05, "s_per_step": 0.5831260204315185, "peak_gb": 45.225538048, "samples_seen": 203200, "time": 1791134551.7359152}
256
+ {"step": 6375, "loss": 2.4862116837501524, "grad_norm": 1.3331636190414429, "lr": 4.909526372619356e-05, "s_per_step": 0.5831469821929932, "peak_gb": 45.225538048, "samples_seen": 204000, "time": 1791134566.3150558}
257
+ {"step": 6400, "loss": 2.553971781730652, "grad_norm": 1.2492870092391968, "lr": 4.835379945775823e-05, "s_per_step": 0.5849734497070312, "peak_gb": 45.225538048, "samples_seen": 204800, "time": 1791134580.9397955}
258
+ {"step": 6425, "loss": 2.5170671367645263, "grad_norm": 1.2050851583480835, "lr": 4.7616187360175476e-05, "s_per_step": 0.5841898059844971, "peak_gb": 45.225538048, "samples_seen": 205600, "time": 1791134595.5449216}
259
+ {"step": 6450, "loss": 2.5319773864746096, "grad_norm": 1.1314105987548828, "lr": 4.6882482450228574e-05, "s_per_step": 0.5848858451843262, "peak_gb": 45.225538048, "samples_seen": 206400, "time": 1791134610.1675096}
260
+ {"step": 6475, "loss": 2.4931352043151858, "grad_norm": 1.1107985973358154, "lr": 4.615273945327272e-05, "s_per_step": 0.5831267642974853, "peak_gb": 45.225538048, "samples_seen": 207200, "time": 1791134624.7460885}
261
+ {"step": 6500, "loss": 2.5409064221382143, "grad_norm": 1.199941873550415, "lr": 4.542701279915318e-05, "s_per_step": 0.5819946575164795, "peak_gb": 45.225538048, "samples_seen": 208000, "time": 1791134639.2963989}
262
+ {"step": 6525, "loss": 2.49187207698822, "grad_norm": 1.4973266124725342, "lr": 4.470535661814538e-05, "s_per_step": 0.6058605766296387, "peak_gb": 45.225538048, "samples_seen": 208800, "time": 1791134654.4432883}
263
+ {"step": 6550, "loss": 2.490929820537567, "grad_norm": 1.222715139389038, "lr": 4.398782473691767e-05, "s_per_step": 0.5849639797210693, "peak_gb": 45.225538048, "samples_seen": 209600, "time": 1791134669.0677845}
264
+ {"step": 6575, "loss": 2.478930926322937, "grad_norm": 1.1837711334228516, "lr": 4.327447067451636e-05, "s_per_step": 0.5865503215789795, "peak_gb": 45.225538048, "samples_seen": 210400, "time": 1791134683.731912}
265
+ {"step": 6600, "loss": 2.505696654319763, "grad_norm": 1.1841992139816284, "lr": 4.256534763837391e-05, "s_per_step": 0.5857251071929932, "peak_gb": 45.225538048, "samples_seen": 211200, "time": 1791134698.375444}
266
+ {"step": 6625, "loss": 2.509692039489746, "grad_norm": 1.251827597618103, "lr": 4.186050852034028e-05, "s_per_step": 0.5832119750976562, "peak_gb": 45.225538048, "samples_seen": 212000, "time": 1791134712.9561343}
267
+ {"step": 6650, "loss": 2.4981483721733095, "grad_norm": 1.3467367887496948, "lr": 4.1160005892737896e-05, "s_per_step": 0.5832837581634521, "peak_gb": 45.225538048, "samples_seen": 212800, "time": 1791134727.5386372}
268
+ {"step": 6675, "loss": 2.487072801589966, "grad_norm": 1.149620771408081, "lr": 4.046389200444035e-05, "s_per_step": 0.5866239070892334, "peak_gb": 45.225538048, "samples_seen": 213600, "time": 1791134742.2046556}
269
+ {"step": 6700, "loss": 2.464690637588501, "grad_norm": 1.290765643119812, "lr": 3.9772218776975314e-05, "s_per_step": 0.5862217617034912, "peak_gb": 45.225538048, "samples_seen": 214400, "time": 1791134756.8606098}
270
+ {"step": 6725, "loss": 2.4712146949768066, "grad_norm": 1.1348451375961304, "lr": 3.908503780065186e-05, "s_per_step": 0.5837146186828613, "peak_gb": 45.225538048, "samples_seen": 215200, "time": 1791134771.4539366}
271
+ {"step": 6750, "loss": 2.4894593477249147, "grad_norm": 1.0581586360931396, "lr": 3.840240033071231e-05, "s_per_step": 0.5876280879974365, "peak_gb": 45.225538048, "samples_seen": 216000, "time": 1791134786.1450136}
272
+ {"step": 6775, "loss": 2.516192195415497, "grad_norm": 1.3426604270935059, "lr": 3.772435728350944e-05, "s_per_step": 0.6016309452056885, "peak_gb": 45.225538048, "samples_seen": 216800, "time": 1791134801.1862383}
273
+ {"step": 6800, "loss": 2.5307986450195314, "grad_norm": 1.3854995965957642, "lr": 3.70509592327086e-05, "s_per_step": 0.5847000885009765, "peak_gb": 45.225538048, "samples_seen": 217600, "time": 1791134815.8041387}
274
+ {"step": 6825, "loss": 2.5088115167617797, "grad_norm": 1.2087575197219849, "lr": 3.638225640551563e-05, "s_per_step": 0.5863976192474365, "peak_gb": 45.225538048, "samples_seen": 218400, "time": 1791134830.4645412}
275
+ {"step": 6850, "loss": 2.4794558453559876, "grad_norm": 1.3083394765853882, "lr": 3.571829867893038e-05, "s_per_step": 0.5870515918731689, "peak_gb": 45.225538048, "samples_seen": 219200, "time": 1791134845.1412776}
276
+ {"step": 6875, "loss": 2.495806655883789, "grad_norm": 1.250901222229004, "lr": 3.50591355760267e-05, "s_per_step": 0.5858700180053711, "peak_gb": 45.225538048, "samples_seen": 220000, "time": 1791134859.7884114}
277
+ {"step": 6900, "loss": 2.472851881980896, "grad_norm": 1.4390413761138916, "lr": 3.440481626225849e-05, "s_per_step": 0.5833972072601319, "peak_gb": 45.225538048, "samples_seen": 220800, "time": 1791134874.3737664}
278
+ {"step": 6925, "loss": 2.4779190325737, "grad_norm": 1.14964759349823, "lr": 3.3755389541792606e-05, "s_per_step": 0.5836262893676758, "peak_gb": 45.225538048, "samples_seen": 221600, "time": 1791134888.9648018}
279
+ {"step": 6950, "loss": 2.4924482202529905, "grad_norm": 1.1170302629470825, "lr": 3.311090385386867e-05, "s_per_step": 0.5860085487365723, "peak_gb": 45.225538048, "samples_seen": 222400, "time": 1791134903.6154745}
280
+ {"step": 6975, "loss": 2.449902744293213, "grad_norm": 1.2713782787322998, "lr": 3.247140726918607e-05, "s_per_step": 0.5829716205596924, "peak_gb": 45.225538048, "samples_seen": 223200, "time": 1791134918.190196}
281
+ {"step": 7000, "loss": 2.5060114574432375, "grad_norm": 1.2075129747390747, "lr": 3.183694748631856e-05, "s_per_step": 0.5825292396545411, "peak_gb": 45.225538048, "samples_seen": 224000, "time": 1791134932.753835}
282
+ {"step": 7025, "loss": 2.5190657806396484, "grad_norm": 1.2806220054626465, "lr": 3.1207571828156415e-05, "s_per_step": 0.6018135643005371, "peak_gb": 45.225538048, "samples_seen": 224800, "time": 1791134947.799628}
283
+ {"step": 7050, "loss": 2.488545150756836, "grad_norm": 1.1726572513580322, "lr": 3.0583327238376826e-05, "s_per_step": 0.5843643283843994, "peak_gb": 45.225538048, "samples_seen": 225600, "time": 1791134962.4091752}
284
+ {"step": 7075, "loss": 2.45016818523407, "grad_norm": 1.3082247972488403, "lr": 2.9964260277942414e-05, "s_per_step": 0.5856728076934814, "peak_gb": 45.225538048, "samples_seen": 226400, "time": 1791134977.0514216}
285
+ {"step": 7100, "loss": 2.4964564561843874, "grad_norm": 1.2286127805709839, "lr": 2.9350417121628438e-05, "s_per_step": 0.5862992858886719, "peak_gb": 45.225538048, "samples_seen": 227200, "time": 1791134991.7094052}
286
+ {"step": 7125, "loss": 2.4161909675598143, "grad_norm": 1.1147135496139526, "lr": 2.8741843554578563e-05, "s_per_step": 0.5842507934570312, "peak_gb": 45.225538048, "samples_seen": 228000, "time": 1791135006.3160708}
287
+ {"step": 7150, "loss": 2.5295400643348693, "grad_norm": 1.2771178483963013, "lr": 2.8138584968890036e-05, "s_per_step": 0.5852422046661377, "peak_gb": 45.225538048, "samples_seen": 228800, "time": 1791135020.9474854}
288
+ {"step": 7175, "loss": 2.5087658619880675, "grad_norm": 1.2442432641983032, "lr": 2.7540686360227917e-05, "s_per_step": 0.5861521339416504, "peak_gb": 45.225538048, "samples_seen": 229600, "time": 1791135035.6017551}
289
+ {"step": 7200, "loss": 2.4834299731254577, "grad_norm": 1.3997399806976318, "lr": 2.6948192324468925e-05, "s_per_step": 0.5873078346252442, "peak_gb": 45.225538048, "samples_seen": 230400, "time": 1791135050.2848673}
290
+ {"step": 7225, "loss": 2.44143901348114, "grad_norm": 1.4422211647033691, "lr": 2.636114705437519e-05, "s_per_step": 0.5871801567077637, "peak_gb": 45.225538048, "samples_seen": 231200, "time": 1791135064.9647627}
291
+ {"step": 7250, "loss": 2.470247166156769, "grad_norm": 1.3490147590637207, "lr": 2.5779594336297975e-05, "s_per_step": 0.5863297176361084, "peak_gb": 45.225538048, "samples_seen": 232000, "time": 1791135079.6234357}
292
+ {"step": 7275, "loss": 2.471605281829834, "grad_norm": 1.129075288772583, "lr": 2.5203577546911773e-05, "s_per_step": 0.599454345703125, "peak_gb": 45.225538048, "samples_seen": 232800, "time": 1791135094.6101875}
293
+ {"step": 7300, "loss": 2.440209083557129, "grad_norm": 1.0634057521820068, "lr": 2.463313964997894e-05, "s_per_step": 0.5864029026031494, "peak_gb": 45.225538048, "samples_seen": 233600, "time": 1791135109.2706358}
294
+ {"step": 7325, "loss": 2.5060981726646423, "grad_norm": 1.265033483505249, "lr": 2.4068323193145126e-05, "s_per_step": 0.5880957889556885, "peak_gb": 45.225538048, "samples_seen": 234400, "time": 1791135123.973469}
295
+ {"step": 7350, "loss": 2.529538559913635, "grad_norm": 1.1425158977508545, "lr": 2.350917030476578e-05, "s_per_step": 0.5846072196960449, "peak_gb": 45.225538048, "samples_seen": 235200, "time": 1791135138.5890818}
296
+ {"step": 7375, "loss": 2.4967665529251097, "grad_norm": 1.387877345085144, "lr": 2.295572269076374e-05, "s_per_step": 0.5827033042907714, "peak_gb": 45.225538048, "samples_seen": 236000, "time": 1791135153.1570456}
297
+ {"step": 7400, "loss": 2.515732476711273, "grad_norm": 1.2702908515930176, "lr": 2.2408021631518726e-05, "s_per_step": 0.5825530815124512, "peak_gb": 45.225538048, "samples_seen": 236800, "time": 1791135167.7212825}
298
+ {"step": 7425, "loss": 2.454144196510315, "grad_norm": 1.0984094142913818, "lr": 2.1866107978788153e-05, "s_per_step": 0.5797280502319336, "peak_gb": 45.225538048, "samples_seen": 237600, "time": 1791135182.2148888}
299
+ {"step": 7450, "loss": 2.4906524801254273, "grad_norm": 1.2278966903686523, "lr": 2.1330022152660156e-05, "s_per_step": 0.5791896152496337, "peak_gb": 45.225538048, "samples_seen": 238400, "time": 1791135196.69504}
300
+ {"step": 7475, "loss": 2.5189996767044067, "grad_norm": 1.3802945613861084, "lr": 2.0799804138538738e-05, "s_per_step": 0.5806455993652344, "peak_gb": 45.225538048, "samples_seen": 239200, "time": 1791135211.211557}
301
+ {"step": 7500, "loss": 2.4968285632133482, "grad_norm": 1.1902079582214355, "lr": 2.027549348416138e-05, "s_per_step": 0.5817130088806153, "peak_gb": 45.225538048, "samples_seen": 240000, "time": 1791135225.754808}
302
+ {"step": 7525, "loss": 2.5208859515190123, "grad_norm": 1.1940761804580688, "lr": 1.9757129296649146e-05, "s_per_step": 0.5999154567718505, "peak_gb": 45.225538048, "samples_seen": 240800, "time": 1791135240.753182}
303
+ {"step": 7550, "loss": 2.5416905546188353, "grad_norm": 1.4812296628952026, "lr": 1.924475023958998e-05, "s_per_step": 0.5843287658691406, "peak_gb": 45.225538048, "samples_seen": 241600, "time": 1791135255.3618777}
304
+ {"step": 7575, "loss": 2.4698385620117187, "grad_norm": 1.176838994026184, "lr": 1.87383945301547e-05, "s_per_step": 0.5848553848266601, "peak_gb": 45.225538048, "samples_seen": 242400, "time": 1791135269.9837594}
305
+ {"step": 7600, "loss": 2.4738434529304505, "grad_norm": 1.2223095893859863, "lr": 1.8238099936246555e-05, "s_per_step": 0.5825335311889649, "peak_gb": 45.225538048, "samples_seen": 243200, "time": 1791135284.547529}
306
+ {"step": 7625, "loss": 2.4790508365631103, "grad_norm": 1.1244370937347412, "lr": 1.774390377368418e-05, "s_per_step": 0.5838665103912354, "peak_gb": 45.225538048, "samples_seen": 244000, "time": 1791135299.144589}
307
+ {"step": 7650, "loss": 2.5447232913970947, "grad_norm": 1.2436165809631348, "lr": 1.7255842903418274e-05, "s_per_step": 0.5841799449920654, "peak_gb": 45.225538048, "samples_seen": 244800, "time": 1791135313.7496023}
308
+ {"step": 7675, "loss": 2.4780035495758055, "grad_norm": 1.200567603111267, "lr": 1.677395372878231e-05, "s_per_step": 0.5839977931976318, "peak_gb": 45.225538048, "samples_seen": 245600, "time": 1791135328.3500307}
309
+ {"step": 7700, "loss": 2.4953444504737856, "grad_norm": 1.3506298065185547, "lr": 1.629827219277713e-05, "s_per_step": 0.5840514087677002, "peak_gb": 45.225538048, "samples_seen": 246400, "time": 1791135342.9517403}
310
+ {"step": 7725, "loss": 2.4795714807510376, "grad_norm": 1.1271368265151978, "lr": 1.5828833775390227e-05, "s_per_step": 0.5836558628082276, "peak_gb": 45.225538048, "samples_seen": 247200, "time": 1791135357.5435355}
311
+ {"step": 7750, "loss": 2.4116058373451232, "grad_norm": 1.1426550149917603, "lr": 1.5365673490949285e-05, "s_per_step": 0.5836513805389404, "peak_gb": 45.225538048, "samples_seen": 248000, "time": 1791135372.135241}
312
+ {"step": 7775, "loss": 2.4785056114196777, "grad_norm": 1.1472325325012207, "lr": 1.4908825885510514e-05, "s_per_step": 0.5973025321960449, "peak_gb": 45.225538048, "samples_seen": 248800, "time": 1791135387.0681865}
313
+ {"step": 7800, "loss": 2.4868078422546387, "grad_norm": 1.2887753248214722, "lr": 1.4458325034281993e-05, "s_per_step": 0.580522985458374, "peak_gb": 45.225538048, "samples_seen": 249600, "time": 1791135401.5817068}
314
+ {"step": 7825, "loss": 2.4639941549301145, "grad_norm": 1.1615179777145386, "lr": 1.4014204539082044e-05, "s_per_step": 0.5852510643005371, "peak_gb": 45.225538048, "samples_seen": 250400, "time": 1791135416.213432}
315
+ {"step": 7850, "loss": 2.480881841182709, "grad_norm": 1.1630935668945312, "lr": 1.3576497525832998e-05, "s_per_step": 0.5810826301574707, "peak_gb": 45.225538048, "samples_seen": 251200, "time": 1791135430.7409315}
316
+ {"step": 7875, "loss": 2.467122640609741, "grad_norm": 1.288510799407959, "lr": 1.314523664209033e-05, "s_per_step": 0.5820274829864502, "peak_gb": 45.225538048, "samples_seen": 252000, "time": 1791135445.292001}
317
+ {"step": 7900, "loss": 2.4884210443496704, "grad_norm": 1.2538615465164185, "lr": 1.2720454054607633e-05, "s_per_step": 0.5825617122650146, "peak_gb": 45.225538048, "samples_seen": 252800, "time": 1791135459.8564053}
318
+ {"step": 7925, "loss": 2.480919179916382, "grad_norm": 1.403016448020935, "lr": 1.2302181446937333e-05, "s_per_step": 0.5786206722259521, "peak_gb": 45.225538048, "samples_seen": 253600, "time": 1791135474.3223386}
319
+ {"step": 7950, "loss": 2.439384446144104, "grad_norm": 1.1538753509521484, "lr": 1.1890450017067467e-05, "s_per_step": 0.5839206409454346, "peak_gb": 45.225538048, "samples_seen": 254400, "time": 1791135488.920715}
320
+ {"step": 7975, "loss": 2.484659490585327, "grad_norm": 1.5197055339813232, "lr": 1.1485290475094745e-05, "s_per_step": 0.5797516345977783, "peak_gb": 45.225538048, "samples_seen": 255200, "time": 1791135503.414908}
321
+ {"step": 8000, "loss": 2.4812649726867675, "grad_norm": 1.2070398330688477, "lr": 1.1086733040933938e-05, "s_per_step": 0.5830123424530029, "peak_gb": 45.225538048, "samples_seen": 256000, "time": 1791135517.9906447}
322
+ {"step": 8025, "loss": 2.453145351409912, "grad_norm": 1.3291114568710327, "lr": 1.0694807442063836e-05, "s_per_step": 0.5979341888427734, "peak_gb": 45.225538048, "samples_seen": 256800, "time": 1791135532.9394002}
323
+ {"step": 8050, "loss": 2.4227629256248475, "grad_norm": 1.178061604499817, "lr": 1.0309542911309932e-05, "s_per_step": 0.583064832687378, "peak_gb": 45.225538048, "samples_seen": 257600, "time": 1791135547.5164373}
324
+ {"step": 8075, "loss": 2.4514348554611205, "grad_norm": 1.5888311862945557, "lr": 9.930968184664046e-06, "s_per_step": 0.5811198616027832, "peak_gb": 45.225538048, "samples_seen": 258400, "time": 1791135562.0448437}
325
+ {"step": 8100, "loss": 2.4358947181701662, "grad_norm": 1.1722838878631592, "lr": 9.559111499140939e-06, "s_per_step": 0.579391222000122, "peak_gb": 45.225538048, "samples_seen": 259200, "time": 1791135576.530029}
326
+ {"step": 8125, "loss": 2.470351276397705, "grad_norm": 1.2379904985427856, "lr": 9.194000590672202e-06, "s_per_step": 0.5833250331878662, "peak_gb": 45.225538048, "samples_seen": 260000, "time": 1791135591.1136327}
327
+ {"step": 8150, "loss": 2.5454572772979738, "grad_norm": 1.0861706733703613, "lr": 8.835662692037516e-06, "s_per_step": 0.5829557609558106, "peak_gb": 45.225538048, "samples_seen": 260800, "time": 1791135605.6879003}
328
+ {"step": 8175, "loss": 2.5356476593017576, "grad_norm": 1.193375825881958, "lr": 8.484124530833349e-06, "s_per_step": 0.5813982677459717, "peak_gb": 45.225538048, "samples_seen": 261600, "time": 1791135620.2232423}
329
+ {"step": 8200, "loss": 2.5145446348190306, "grad_norm": 1.31069815158844, "lr": 8.139412327479512e-06, "s_per_step": 0.5830599784851074, "peak_gb": 45.225538048, "samples_seen": 262400, "time": 1791135634.8001153}
330
+ {"step": 8225, "loss": 2.51434268951416, "grad_norm": 1.2565644979476929, "lr": 7.801551793263307e-06, "s_per_step": 0.5842070770263672, "peak_gb": 45.225538048, "samples_seen": 263200, "time": 1791135649.4057364}
331
+ {"step": 8250, "loss": 2.4891091346740724, "grad_norm": 1.3129713535308838, "lr": 7.470568128421918e-06, "s_per_step": 0.5833453273773194, "peak_gb": 45.225538048, "samples_seen": 264000, "time": 1791135663.9898171}
332
+ {"step": 8275, "loss": 2.4999825406074523, "grad_norm": 1.2650532722473145, "lr": 7.146486020262688e-06, "s_per_step": 0.5951438903808594, "peak_gb": 45.225538048, "samples_seen": 264800, "time": 1791135678.8688512}
333
+ {"step": 8300, "loss": 2.50071989774704, "grad_norm": 1.366416335105896, "lr": 6.82932964132178e-06, "s_per_step": 0.585099573135376, "peak_gb": 45.225538048, "samples_seen": 265600, "time": 1791135693.4967763}
334
+ {"step": 8325, "loss": 2.4906865310668946, "grad_norm": 1.2222192287445068, "lr": 6.519122647561249e-06, "s_per_step": 0.5849275970458985, "peak_gb": 45.225538048, "samples_seen": 266400, "time": 1791135708.120377}
335
+ {"step": 8350, "loss": 2.482330822944641, "grad_norm": 1.2209337949752808, "lr": 6.215888176604501e-06, "s_per_step": 0.5821888160705566, "peak_gb": 45.225538048, "samples_seen": 267200, "time": 1791135722.6755044}
336
+ {"step": 8375, "loss": 2.455729749202728, "grad_norm": 1.3347740173339844, "lr": 5.919648846010584e-06, "s_per_step": 0.5881471824645996, "peak_gb": 45.225538048, "samples_seen": 268000, "time": 1791135737.3796217}
337
+ {"step": 8400, "loss": 2.4725853204727173, "grad_norm": 1.408981442451477, "lr": 5.6304267515872035e-06, "s_per_step": 0.5856296157836914, "peak_gb": 45.225538048, "samples_seen": 268800, "time": 1791135752.0207672}
338
+ {"step": 8425, "loss": 2.4187474822998047, "grad_norm": 1.3314207792282104, "lr": 5.3482434657425755e-06, "s_per_step": 0.5818990707397461, "peak_gb": 45.225538048, "samples_seen": 269600, "time": 1791135766.5687203}
339
+ {"step": 8450, "loss": 2.4801942539215087, "grad_norm": 1.0783225297927856, "lr": 5.073120035876466e-06, "s_per_step": 0.5826784038543701, "peak_gb": 45.225538048, "samples_seen": 270400, "time": 1791135781.136184}
340
+ {"step": 8475, "loss": 2.4995450139045716, "grad_norm": 1.3838938474655151, "lr": 4.805076982810275e-06, "s_per_step": 0.5815659427642822, "peak_gb": 45.225538048, "samples_seen": 271200, "time": 1791135795.6757793}
341
+ {"step": 8500, "loss": 2.4431419372558594, "grad_norm": 1.1133414506912231, "lr": 4.544134299256453e-06, "s_per_step": 0.5855115985870362, "peak_gb": 45.225538048, "samples_seen": 272000, "time": 1791135810.3140042}
342
+ {"step": 8525, "loss": 2.4860039710998536, "grad_norm": 1.1502230167388916, "lr": 4.290311448327267e-06, "s_per_step": 0.6029162120819092, "peak_gb": 45.225538048, "samples_seen": 272800, "time": 1791135825.3872917}
343
+ {"step": 8550, "loss": 2.502565712928772, "grad_norm": 1.4622975587844849, "lr": 4.043627362083124e-06, "s_per_step": 0.5819817352294921, "peak_gb": 45.225538048, "samples_seen": 273600, "time": 1791135839.9373028}
344
+ {"step": 8575, "loss": 2.4554481315612793, "grad_norm": 1.2948318719863892, "lr": 3.8041004401204392e-06, "s_per_step": 0.5855245685577393, "peak_gb": 45.225538048, "samples_seen": 274400, "time": 1791135854.575897}
345
+ {"step": 8600, "loss": 2.473366265296936, "grad_norm": 1.3055177927017212, "lr": 3.571748548199283e-06, "s_per_step": 0.5834248447418213, "peak_gb": 45.225538048, "samples_seen": 275200, "time": 1791135869.1619527}
346
+ {"step": 8625, "loss": 2.4664463996887207, "grad_norm": 1.3097848892211914, "lr": 3.346589016910795e-06, "s_per_step": 0.5833245754241944, "peak_gb": 45.225538048, "samples_seen": 276000, "time": 1791135883.7455816}
347
+ {"step": 8650, "loss": 2.5279615044593813, "grad_norm": 1.281739354133606, "lr": 3.1286386403845402e-06, "s_per_step": 0.5869896125793457, "peak_gb": 45.225538048, "samples_seen": 276800, "time": 1791135898.4208338}
348
+ {"step": 8675, "loss": 2.4256000685691834, "grad_norm": 1.1321176290512085, "lr": 2.917913675035888e-06, "s_per_step": 0.5836698818206787, "peak_gb": 45.225538048, "samples_seen": 277600, "time": 1791135913.0130074}
349
+ {"step": 8700, "loss": 2.520475058555603, "grad_norm": 1.046059250831604, "lr": 2.7144298383534607e-06, "s_per_step": 0.5819922924041748, "peak_gb": 45.225538048, "samples_seen": 278400, "time": 1791135927.5632734}
350
+ {"step": 8725, "loss": 2.555812520980835, "grad_norm": 1.1831002235412598, "lr": 2.5182023077268024e-06, "s_per_step": 0.5796582889556885, "peak_gb": 45.225538048, "samples_seen": 279200, "time": 1791135942.0551524}
351
+ {"step": 8750, "loss": 2.4607612800598146, "grad_norm": 1.2153393030166626, "lr": 2.3292457193143657e-06, "s_per_step": 0.5847400951385499, "peak_gb": 45.225538048, "samples_seen": 280000, "time": 1791135956.6740828}
352
+ {"step": 8775, "loss": 2.4432231545448304, "grad_norm": 1.3403053283691406, "lr": 2.1475741669517936e-06, "s_per_step": 0.5985502052307129, "peak_gb": 45.225538048, "samples_seen": 280800, "time": 1791135971.638234}
353
+ {"step": 8800, "loss": 2.45842068195343, "grad_norm": 1.3938454389572144, "lr": 1.973201201100705e-06, "s_per_step": 0.5829646301269531, "peak_gb": 45.225538048, "samples_seen": 281600, "time": 1791135986.2127783}
354
+ {"step": 8825, "loss": 2.4700878167152407, "grad_norm": 1.2799893617630005, "lr": 1.806139827838027e-06, "s_per_step": 0.5821155929565429, "peak_gb": 45.225538048, "samples_seen": 282400, "time": 1791136000.766119}
355
+ {"step": 8850, "loss": 2.505449166297913, "grad_norm": 1.1758151054382324, "lr": 1.646402507885858e-06, "s_per_step": 0.5805652141571045, "peak_gb": 45.225538048, "samples_seen": 283200, "time": 1791136015.2806501}
356
+ {"step": 8875, "loss": 2.4752099657058717, "grad_norm": 1.1823129653930664, "lr": 1.4940011556820676e-06, "s_per_step": 0.5807376194000244, "peak_gb": 45.225538048, "samples_seen": 284000, "time": 1791136029.7995193}
357
+ {"step": 8900, "loss": 2.477540431022644, "grad_norm": 1.2883774042129517, "lr": 1.3489471384916518e-06, "s_per_step": 0.5786253547668457, "peak_gb": 45.225538048, "samples_seen": 284800, "time": 1791136044.2656045}
358
+ {"step": 8925, "loss": 2.472158603668213, "grad_norm": 1.2201075553894043, "lr": 1.2112512755588335e-06, "s_per_step": 0.5808994483947754, "peak_gb": 45.225538048, "samples_seen": 285600, "time": 1791136058.7884433}
359
+ {"step": 8950, "loss": 2.4544294118881225, "grad_norm": 1.2970439195632935, "lr": 1.0809238373000962e-06, "s_per_step": 0.581709451675415, "peak_gb": 45.225538048, "samples_seen": 286400, "time": 1791136073.3315735}
360
+ {"step": 8975, "loss": 2.400626153945923, "grad_norm": 1.2315267324447632, "lr": 9.579745445381539e-07, "s_per_step": 0.579955415725708, "peak_gb": 45.225538048, "samples_seen": 287200, "time": 1791136087.8308325}
361
+ {"step": 9000, "loss": 2.4409583902359007, "grad_norm": 1.1493794918060303, "lr": 8.424125677768846e-07, "s_per_step": 0.5824070739746093, "peak_gb": 45.225538048, "samples_seen": 288000, "time": 1791136102.3914094}
362
+ {"step": 9025, "loss": 2.4901751470565796, "grad_norm": 1.405742883682251, "lr": 7.342465265172904e-07, "s_per_step": 0.5998865604400635, "peak_gb": 45.225538048, "samples_seen": 288800, "time": 1791136117.3889866}
363
+ {"step": 9050, "loss": 2.45655734539032, "grad_norm": 1.2151944637298584, "lr": 6.334844886146552e-07, "s_per_step": 0.5862552833557129, "peak_gb": 45.225538048, "samples_seen": 289600, "time": 1791136132.045783}
364
+ {"step": 9075, "loss": 2.440645921230316, "grad_norm": 1.2591733932495117, "lr": 5.401339696767371e-07, "s_per_step": 0.5819112586975098, "peak_gb": 45.225538048, "samples_seen": 290400, "time": 1791136146.593972}
365
+ {"step": 9100, "loss": 2.4351733326911926, "grad_norm": 1.4817209243774414, "lr": 4.542019325032176e-07, "s_per_step": 0.5811293697357178, "peak_gb": 45.225538048, "samples_seen": 291200, "time": 1791136161.1227205}
366
+ {"step": 9125, "loss": 2.499671845436096, "grad_norm": 1.075323224067688, "lr": 3.756947865663274e-07, "s_per_step": 0.580502290725708, "peak_gb": 45.225538048, "samples_seen": 292000, "time": 1791136175.635759}
367
+ {"step": 9150, "loss": 2.4365021347999574, "grad_norm": 1.199285626411438, "lr": 3.0461838753281793e-07, "s_per_step": 0.5829291343688965, "peak_gb": 45.225538048, "samples_seen": 292800, "time": 1791136190.209385}
368
+ {"step": 9175, "loss": 2.5040885043144225, "grad_norm": 1.390062689781189, "lr": 2.4097803682719967e-07, "s_per_step": 0.5805082321166992, "peak_gb": 45.225538048, "samples_seen": 293600, "time": 1791136204.7224927}
369
+ {"step": 9200, "loss": 2.4622223567962647, "grad_norm": 1.1260004043579102, "lr": 1.847784812362696e-07, "s_per_step": 0.5807758331298828, "peak_gb": 45.225538048, "samples_seen": 294400, "time": 1791136219.2422497}
370
+ {"step": 9225, "loss": 2.5410248279571532, "grad_norm": 1.2212775945663452, "lr": 1.360239125551388e-07, "s_per_step": 0.5833808612823487, "peak_gb": 45.225538048, "samples_seen": 295200, "time": 1791136233.8271465}
371
+ {"step": 9250, "loss": 2.4439662075042725, "grad_norm": 1.3457050323486328, "lr": 9.471796727449356e-08, "s_per_step": 0.5828397941589355, "peak_gb": 45.225538048, "samples_seen": 296000, "time": 1791136248.3985698}
372
+ {"step": 9275, "loss": 2.402722330093384, "grad_norm": 1.2421631813049316, "lr": 6.086372630945692e-08, "s_per_step": 0.5969734477996826, "peak_gb": 45.225538048, "samples_seen": 296800, "time": 1791136263.3232737}
373
+ {"step": 9300, "loss": 2.4314382100105285, "grad_norm": 1.442682147026062, "lr": 3.446371476966137e-08, "s_per_step": 0.5828660678863525, "peak_gb": 45.225538048, "samples_seen": 297600, "time": 1791136277.895351}
374
+ {"step": 9325, "loss": 2.46381530046463, "grad_norm": 1.3629335165023804, "lr": 1.5519901771032798e-08, "s_per_step": 0.5814544200897217, "peak_gb": 45.225538048, "samples_seen": 298400, "time": 1791136292.4321828}
375
+ {"step": 9350, "loss": 2.448145937919617, "grad_norm": 1.2016350030899048, "lr": 4.033700288830211e-09, "s_per_step": 0.5831421184539795, "peak_gb": 45.225538048, "samples_seen": 299200, "time": 1791136307.0111272}
376
+ {"step": 9375, "loss": 2.453929500579834, "grad_norm": 1.3293850421905518, "lr": 5.967052318922584e-12, "s_per_step": 0.5822473049163819, "peak_gb": 45.225538048, "samples_seen": 300000, "time": 1791136321.5678217}
logs/stage2_train_log.jsonl ADDED
@@ -0,0 +1,197 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"step": 1, "loss": 2.1018713116645813, "grad_norm": 3.935774087905884, "lr": 1.3698630136986302e-06, "s_per_step": 5.9358556270599365, "peak_gb": 28.923105792, "samples_seen": 32, "time": 1791137293.9179838}
2
+ {"step": 25, "loss": 1.8595843004683654, "grad_norm": 0.8364684581756592, "lr": 3.424657534246575e-05, "s_per_step": 3.0314967930316925, "peak_gb": 32.089119744, "samples_seen": 800, "time": 1791137366.67442}
3
+ {"step": 50, "loss": 1.4337403464317322, "grad_norm": 0.641377329826355, "lr": 6.84931506849315e-05, "s_per_step": 2.7474777412414553, "peak_gb": 32.089119744, "samples_seen": 1600, "time": 1791137435.3618677}
4
+ {"step": 75, "loss": 1.3591693234443665, "grad_norm": 0.46830347180366516, "lr": 0.00010273972602739728, "s_per_step": 2.675526475906372, "peak_gb": 32.18382848, "samples_seen": 2400, "time": 1791137502.2504642}
5
+ {"step": 100, "loss": 1.3316362035274505, "grad_norm": 0.6638643145561218, "lr": 0.000136986301369863, "s_per_step": 2.702469034194946, "peak_gb": 34.275917312, "samples_seen": 3200, "time": 1791137569.8126426}
6
+ {"step": 125, "loss": 1.3198813837766648, "grad_norm": 0.5527892112731934, "lr": 0.00017123287671232877, "s_per_step": 2.699006910324097, "peak_gb": 34.275917312, "samples_seen": 4000, "time": 1791137637.2882571}
7
+ {"step": 150, "loss": 1.3081526547670363, "grad_norm": 0.5911878347396851, "lr": 0.00019999980323762506, "s_per_step": 2.6959174919128417, "peak_gb": 34.275917312, "samples_seen": 4800, "time": 1791137704.6866806}
8
+ {"step": 175, "loss": 1.33544615983963, "grad_norm": 0.5037556290626526, "lr": 0.000199982860294911, "s_per_step": 2.5416007328033445, "peak_gb": 34.275917312, "samples_seen": 5600, "time": 1791137768.2271826}
9
+ {"step": 200, "loss": 1.30276043176651, "grad_norm": 0.6148776412010193, "lr": 0.00019993859454180504, "s_per_step": 2.7514454460144044, "peak_gb": 34.275917312, "samples_seen": 6400, "time": 1791137837.0137653}
10
+ {"step": 225, "loss": 1.3027793610095977, "grad_norm": 0.5531913042068481, "lr": 0.00019986701807502827, "s_per_step": 2.6572889137268065, "peak_gb": 34.275917312, "samples_seen": 7200, "time": 1791137903.4464118}
11
+ {"step": 250, "loss": 1.3026298081874848, "grad_norm": 0.5650343298912048, "lr": 0.00019976815045463558, "s_per_step": 2.581279354095459, "peak_gb": 34.275917312, "samples_seen": 8000, "time": 1791137967.9797318}
12
+ {"step": 275, "loss": 1.2859191417694091, "grad_norm": 0.5168231725692749, "lr": 0.00019964201869867025, "s_per_step": 2.7289039039611818, "peak_gb": 34.275917312, "samples_seen": 8800, "time": 1791138036.2027671}
13
+ {"step": 300, "loss": 1.3106050491333008, "grad_norm": 0.5257666707038879, "lr": 0.0001994886572757806, "s_per_step": 2.456781349182129, "peak_gb": 34.275917312, "samples_seen": 9600, "time": 1791138097.6227736}
14
+ {"step": 325, "loss": 1.285406960248947, "grad_norm": 0.5512126088142395, "lr": 0.00019930810809580063, "s_per_step": 2.6005754470825195, "peak_gb": 34.275917312, "samples_seen": 10400, "time": 1791138162.6376011}
15
+ {"step": 350, "loss": 1.2840537762641906, "grad_norm": 0.592390239238739, "lr": 0.00019910042049829714, "s_per_step": 2.6476001262664797, "peak_gb": 34.275917312, "samples_seen": 11200, "time": 1791138228.8280518}
16
+ {"step": 375, "loss": 1.310118284225464, "grad_norm": 0.4393337070941925, "lr": 0.00019886565123908637, "s_per_step": 2.5076206874847413, "peak_gb": 34.275917312, "samples_seen": 12000, "time": 1791138291.5190346}
17
+ {"step": 400, "loss": 1.3007353556156158, "grad_norm": 0.572140097618103, "lr": 0.0001986038644747241, "s_per_step": 2.54843355178833, "peak_gb": 34.275917312, "samples_seen": 12800, "time": 1791138355.230326}
18
+ {"step": 425, "loss": 1.3010186761617661, "grad_norm": 0.5119914412498474, "lr": 0.00019831513174497332, "s_per_step": 2.463373165130615, "peak_gb": 34.454321664, "samples_seen": 13600, "time": 1791138416.8150551}
19
+ {"step": 450, "loss": 1.2758523070812224, "grad_norm": 0.46384191513061523, "lr": 0.00019799953195325412, "s_per_step": 2.629770526885986, "peak_gb": 34.454321664, "samples_seen": 14400, "time": 1791138482.5597115}
20
+ {"step": 475, "loss": 1.2857936775684358, "grad_norm": 0.5037622451782227, "lr": 0.0001976571513450814, "s_per_step": 2.5591455364227294, "peak_gb": 34.454321664, "samples_seen": 15200, "time": 1791138546.5387669}
21
+ {"step": 500, "loss": 1.296420032978058, "grad_norm": 0.5816029906272888, "lr": 0.00019728808348449617, "s_per_step": 2.425617971420288, "peak_gb": 34.454321664, "samples_seen": 16000, "time": 1791138607.1797073}
22
+ {"step": 525, "loss": 1.2788980293273926, "grad_norm": 0.5487325191497803, "lr": 0.0001968924292284968, "s_per_step": 2.719025077819824, "peak_gb": 34.454321664, "samples_seen": 16800, "time": 1791138675.1558762}
23
+ {"step": 550, "loss": 1.2836597490310668, "grad_norm": 0.4508974850177765, "lr": 0.0001964702966994773, "s_per_step": 2.5938339710235594, "peak_gb": 35.16160256, "samples_seen": 17600, "time": 1791138740.0021875}
24
+ {"step": 575, "loss": 1.272069708108902, "grad_norm": 0.4157426059246063, "lr": 0.00019602180125568024, "s_per_step": 2.557501964569092, "peak_gb": 35.16160256, "samples_seen": 18400, "time": 1791138803.9402196}
25
+ {"step": 600, "loss": 1.2735493952035903, "grad_norm": 0.4135555326938629, "lr": 0.00019554706545967223, "s_per_step": 2.681543912887573, "peak_gb": 35.16160256, "samples_seen": 19200, "time": 1791138871.1860206}
26
+ {"step": 625, "loss": 1.2743602776527405, "grad_norm": 0.527656078338623, "lr": 0.00019504621904485055, "s_per_step": 2.574559545516968, "peak_gb": 35.16160256, "samples_seen": 20000, "time": 1791138935.5505657}
27
+ {"step": 650, "loss": 1.2700351548194886, "grad_norm": 0.47019869089126587, "lr": 0.00019451939887999044, "s_per_step": 2.5688095760345457, "peak_gb": 35.16160256, "samples_seen": 20800, "time": 1791138999.7712135}
28
+ {"step": 675, "loss": 1.278431007862091, "grad_norm": 0.49949219822883606, "lr": 0.0001939667489318421, "s_per_step": 2.5981833267211916, "peak_gb": 35.16160256, "samples_seen": 21600, "time": 1791139064.7262437}
29
+ {"step": 700, "loss": 1.2902556908130647, "grad_norm": 0.4944266378879547, "lr": 0.00019338842022578826, "s_per_step": 2.3775853538513183, "peak_gb": 35.16160256, "samples_seen": 22400, "time": 1791139124.16638}
30
+ {"step": 725, "loss": 1.2667275273799896, "grad_norm": 0.3541528880596161, "lr": 0.00019278457080457287, "s_per_step": 2.660867967605591, "peak_gb": 35.16160256, "samples_seen": 23200, "time": 1791139190.6885371}
31
+ {"step": 750, "loss": 1.2670336496829986, "grad_norm": 0.4926910996437073, "lr": 0.00019215536568511165, "s_per_step": 2.5423004722595213, "peak_gb": 35.16160256, "samples_seen": 24000, "time": 1791139254.246462}
32
+ {"step": 775, "loss": 1.2607650971412658, "grad_norm": 0.5327316522598267, "lr": 0.0001915009768133975, "s_per_step": 2.632803087234497, "peak_gb": 35.16160256, "samples_seen": 24800, "time": 1791139320.0670488}
33
+ {"step": 800, "loss": 1.298940641283989, "grad_norm": 0.5895273685455322, "lr": 0.00019082158301751158, "s_per_step": 2.3035424613952635, "peak_gb": 35.16160256, "samples_seen": 25600, "time": 1791139377.6561027}
34
+ {"step": 825, "loss": 1.2994678497314454, "grad_norm": 0.5701431632041931, "lr": 0.00019011736995875442, "s_per_step": 2.4659962368011477, "peak_gb": 35.16160256, "samples_seen": 26400, "time": 1791139439.3064585}
35
+ {"step": 850, "loss": 1.2705088353157044, "grad_norm": 0.4701496958732605, "lr": 0.0001893885300809091, "s_per_step": 2.506269245147705, "peak_gb": 35.16160256, "samples_seen": 27200, "time": 1791139501.96361}
36
+ {"step": 875, "loss": 1.2747311300039292, "grad_norm": 0.5782403349876404, "lr": 0.00018863526255765125, "s_per_step": 2.4462522220611573, "peak_gb": 35.16160256, "samples_seen": 28000, "time": 1791139563.1204047}
37
+ {"step": 900, "loss": 1.2845921885967255, "grad_norm": 0.5439372062683105, "lr": 0.00018785777323811998, "s_per_step": 2.414919261932373, "peak_gb": 35.16160256, "samples_seen": 28800, "time": 1791139623.4938784}
38
+ {"step": 925, "loss": 1.2421265226602554, "grad_norm": 0.4805165231227875, "lr": 0.00018705627459066429, "s_per_step": 2.59889892578125, "peak_gb": 35.16160256, "samples_seen": 29600, "time": 1791139688.4667766}
39
+ {"step": 950, "loss": 1.256026518344879, "grad_norm": 0.4744698405265808, "lr": 0.00018623098564478095, "s_per_step": 2.4729893398284912, "peak_gb": 35.16160256, "samples_seen": 30400, "time": 1791139750.3143578}
40
+ {"step": 975, "loss": 1.2527462810277938, "grad_norm": 0.45646801590919495, "lr": 0.00018538213193125915, "s_per_step": 2.529736156463623, "peak_gb": 35.16160256, "samples_seen": 31200, "time": 1791139813.5582156}
41
+ {"step": 1000, "loss": 1.2478597414493562, "grad_norm": 0.6265877485275269, "lr": 0.00018450994542054856, "s_per_step": 2.5998341178894044, "peak_gb": 35.16160256, "samples_seen": 32000, "time": 1791139878.5545025}
42
+ {"step": 1025, "loss": 1.2808118069171905, "grad_norm": 0.4756571054458618, "lr": 0.0001836146644593677, "s_per_step": 2.5343000221252443, "peak_gb": 35.16160256, "samples_seen": 32800, "time": 1791139941.9124794}
43
+ {"step": 1050, "loss": 1.2551424139738083, "grad_norm": 0.6432344913482666, "lr": 0.00018269653370556972, "s_per_step": 2.568102979660034, "peak_gb": 35.16160256, "samples_seen": 33600, "time": 1791140006.1155303}
44
+ {"step": 1075, "loss": 1.2568443739414215, "grad_norm": 0.5057801008224487, "lr": 0.0001817558040612835, "s_per_step": 2.6142832756042482, "peak_gb": 35.16160256, "samples_seen": 34400, "time": 1791140071.4731367}
45
+ {"step": 1100, "loss": 1.2462696027755737, "grad_norm": 0.537843644618988, "lr": 0.00018079273260434846, "s_per_step": 2.5834513759613036, "peak_gb": 35.16160256, "samples_seen": 35200, "time": 1791140136.05996}
46
+ {"step": 1125, "loss": 1.2574056977033614, "grad_norm": 0.5526939630508423, "lr": 0.00017980758251806152, "s_per_step": 2.4060669994354247, "peak_gb": 35.16160256, "samples_seen": 36000, "time": 1791140196.2120588}
47
+ {"step": 1150, "loss": 1.2388004493713378, "grad_norm": 0.5050171613693237, "lr": 0.00017880062301925582, "s_per_step": 2.5375908184051514, "peak_gb": 35.16160256, "samples_seen": 36800, "time": 1791140259.6522965}
48
+ {"step": 1175, "loss": 1.263253116607666, "grad_norm": 0.4916113317012787, "lr": 0.00017777212928473044, "s_per_step": 2.5061941528320313, "peak_gb": 35.16160256, "samples_seen": 37600, "time": 1791140322.3075662}
49
+ {"step": 1200, "loss": 1.2475608468055726, "grad_norm": 0.5672100186347961, "lr": 0.00017672238237605146, "s_per_step": 2.3854447269439696, "peak_gb": 35.16160256, "samples_seen": 38400, "time": 1791140381.9441311}
50
+ {"step": 1225, "loss": 1.2634010112285614, "grad_norm": 0.45946162939071655, "lr": 0.0001756516691627449, "s_per_step": 2.559474468231201, "peak_gb": 35.16160256, "samples_seen": 39200, "time": 1791140445.9314318}
51
+ {"step": 1250, "loss": 1.2636224496364594, "grad_norm": 0.5660274028778076, "lr": 0.00017456028224390258, "s_per_step": 2.3614064121246336, "peak_gb": 35.16160256, "samples_seen": 40000, "time": 1791140504.967005}
52
+ {"step": 1275, "loss": 1.2519952070713043, "grad_norm": 0.630153477191925, "lr": 0.00017344851986822186, "s_per_step": 2.604766302108765, "peak_gb": 35.16160256, "samples_seen": 40800, "time": 1791140570.0865934}
53
+ {"step": 1300, "loss": 1.248630976676941, "grad_norm": 0.603481650352478, "lr": 0.000172316685852502, "s_per_step": 2.5879988288879394, "peak_gb": 35.16160256, "samples_seen": 41600, "time": 1791140634.7870479}
54
+ {"step": 1325, "loss": 1.2556267881393433, "grad_norm": 0.5644108057022095, "lr": 0.00017116508949861846, "s_per_step": 2.4682638454437256, "peak_gb": 35.16160256, "samples_seen": 42400, "time": 1791140696.4940681}
55
+ {"step": 1350, "loss": 1.2743791633844375, "grad_norm": 0.552125871181488, "lr": 0.00016999404550899856, "s_per_step": 2.4177276992797854, "peak_gb": 35.16160256, "samples_seen": 43200, "time": 1791140756.9376965}
56
+ {"step": 1375, "loss": 1.2419914174079896, "grad_norm": 0.5084569454193115, "lr": 0.00016880387390062117, "s_per_step": 2.5251627922058106, "peak_gb": 35.16160256, "samples_seen": 44000, "time": 1791140820.067213}
57
+ {"step": 1400, "loss": 1.2473990058898925, "grad_norm": 0.48340925574302673, "lr": 0.00016759489991756402, "s_per_step": 2.5814351272583007, "peak_gb": 35.16160256, "samples_seen": 44800, "time": 1791140884.6035194}
58
+ {"step": 1425, "loss": 1.2643068969249724, "grad_norm": 0.6564701795578003, "lr": 0.0001663674539421228, "s_per_step": 2.4565196990966798, "peak_gb": 35.16160256, "samples_seen": 45600, "time": 1791140946.0169346}
59
+ {"step": 1450, "loss": 1.2657874739170074, "grad_norm": 0.5559526681900024, "lr": 0.0001651218714045258, "s_per_step": 2.38125319480896, "peak_gb": 35.16160256, "samples_seen": 46400, "time": 1791141005.5487158}
60
+ {"step": 1475, "loss": 1.2623559749126434, "grad_norm": 0.633904218673706, "lr": 0.00016385849269126923, "s_per_step": 2.3729316806793213, "peak_gb": 35.16160256, "samples_seen": 47200, "time": 1791141064.8724935}
61
+ {"step": 1500, "loss": 1.2323721277713775, "grad_norm": 0.5630310773849487, "lr": 0.00016257766305209824, "s_per_step": 2.5675553035736085, "peak_gb": 35.16160256, "samples_seen": 48000, "time": 1791141129.06184}
62
+ {"step": 1525, "loss": 1.2649996304512023, "grad_norm": 0.5584118962287903, "lr": 0.0001612797325056588, "s_per_step": 2.5546645450592043, "peak_gb": 35.16160256, "samples_seen": 48800, "time": 1791141192.9288673}
63
+ {"step": 1550, "loss": 1.2378357803821565, "grad_norm": 0.5218235850334167, "lr": 0.00015996505574384623, "s_per_step": 2.511321334838867, "peak_gb": 35.16160256, "samples_seen": 49600, "time": 1791141255.712335}
64
+ {"step": 1575, "loss": 1.2511804521083831, "grad_norm": 0.6376656889915466, "lr": 0.00015863399203487695, "s_per_step": 2.461472635269165, "peak_gb": 35.16160256, "samples_seen": 50400, "time": 1791141317.2496617}
65
+ {"step": 1600, "loss": 1.2562532049417496, "grad_norm": 0.5086262822151184, "lr": 0.00015728690512510944, "s_per_step": 2.506492853164673, "peak_gb": 35.16160256, "samples_seen": 51200, "time": 1791141379.912437}
66
+ {"step": 1625, "loss": 1.2379037940502167, "grad_norm": 0.5947778224945068, "lr": 0.00015592416313964136, "s_per_step": 2.4982398319244385, "peak_gb": 35.16160256, "samples_seen": 52000, "time": 1791141442.3688512}
67
+ {"step": 1650, "loss": 1.237431339621544, "grad_norm": 0.5146723985671997, "lr": 0.0001545461384817104, "s_per_step": 2.4878829956054687, "peak_gb": 35.16160256, "samples_seen": 52800, "time": 1791141504.5663276}
68
+ {"step": 1675, "loss": 1.2310943824052811, "grad_norm": 0.5532039403915405, "lr": 0.00015315320773092562, "s_per_step": 2.601818952560425, "peak_gb": 35.16160256, "samples_seen": 53600, "time": 1791141569.612219}
69
+ {"step": 1700, "loss": 1.233572164773941, "grad_norm": 0.491245836019516, "lr": 0.00015174575154035776, "s_per_step": 2.525400400161743, "peak_gb": 35.16160256, "samples_seen": 54400, "time": 1791141632.7476742}
70
+ {"step": 1725, "loss": 1.2353800696134567, "grad_norm": 0.5083756446838379, "lr": 0.00015032415453251628, "s_per_step": 2.500525407791138, "peak_gb": 35.16160256, "samples_seen": 55200, "time": 1791141695.2612443}
71
+ {"step": 1750, "loss": 1.2415151298046112, "grad_norm": 0.4848071336746216, "lr": 0.00014888880519424167, "s_per_step": 2.497219123840332, "peak_gb": 35.16160256, "samples_seen": 56000, "time": 1791141757.6921887}
72
+ {"step": 1775, "loss": 1.25896546125412, "grad_norm": 0.5702372193336487, "lr": 0.00014744009577054174, "s_per_step": 2.511766996383667, "peak_gb": 35.16160256, "samples_seen": 56800, "time": 1791141820.4869192}
73
+ {"step": 1800, "loss": 1.2329586499929428, "grad_norm": 0.5330091714859009, "lr": 0.0001459784221574008, "s_per_step": 2.550053768157959, "peak_gb": 35.34578432, "samples_seen": 57600, "time": 1791141884.2394202}
74
+ {"step": 1825, "loss": 1.2399318504333496, "grad_norm": 0.5269576907157898, "lr": 0.00014450418379359145, "s_per_step": 2.3936056327819824, "peak_gb": 35.34578432, "samples_seen": 58400, "time": 1791141944.0801225}
75
+ {"step": 1850, "loss": 1.1970552122592926, "grad_norm": 0.5154424905776978, "lr": 0.00014301778355151762, "s_per_step": 2.6440818309783936, "peak_gb": 35.34578432, "samples_seen": 59200, "time": 1791142010.1826928}
76
+ {"step": 1875, "loss": 1.2424454951286317, "grad_norm": 0.5999359488487244, "lr": 0.00014151962762711998, "s_per_step": 2.516843738555908, "peak_gb": 35.709385216, "samples_seen": 60000, "time": 1791142073.1044393}
77
+ {"step": 1900, "loss": 1.2413146513700486, "grad_norm": 0.5478770732879639, "lr": 0.0001400101254288724, "s_per_step": 2.479934730529785, "peak_gb": 35.709385216, "samples_seen": 60800, "time": 1791142135.1033275}
78
+ {"step": 1925, "loss": 1.2434303241968154, "grad_norm": 0.6004431247711182, "lr": 0.00013848968946590133, "s_per_step": 2.4324647426605224, "peak_gb": 35.709385216, "samples_seen": 61600, "time": 1791142195.9162219}
79
+ {"step": 1950, "loss": 1.2015808641910553, "grad_norm": 0.5266046524047852, "lr": 0.000136958735235257, "s_per_step": 2.506597595214844, "peak_gb": 35.709385216, "samples_seen": 62400, "time": 1791142258.5817263}
80
+ {"step": 1975, "loss": 1.2200964736938475, "grad_norm": 0.5610483884811401, "lr": 0.00013541768110836863, "s_per_step": 2.5384847927093506, "peak_gb": 35.709385216, "samples_seen": 63200, "time": 1791142322.0443137}
81
+ {"step": 2000, "loss": 1.2384649211168288, "grad_norm": 0.6361256241798401, "lr": 0.00013386694821671414, "s_per_step": 2.4207293224334716, "peak_gb": 35.709385216, "samples_seen": 64000, "time": 1791142382.5630329}
82
+ {"step": 2025, "loss": 1.214077221751213, "grad_norm": 0.6302129030227661, "lr": 0.0001323069603367351, "s_per_step": 2.6198851585388185, "peak_gb": 35.709385216, "samples_seen": 64800, "time": 1791142448.0606163}
83
+ {"step": 2050, "loss": 1.2055979889631272, "grad_norm": 0.509576141834259, "lr": 0.00013073814377402974, "s_per_step": 2.627358932495117, "peak_gb": 35.709385216, "samples_seen": 65600, "time": 1791142513.7450488}
84
+ {"step": 2075, "loss": 1.2103803342580794, "grad_norm": 0.5196846723556519, "lr": 0.00012916092724685383, "s_per_step": 2.598076095581055, "peak_gb": 35.709385216, "samples_seen": 66400, "time": 1791142578.6973796}
85
+ {"step": 2100, "loss": 1.2209484434127809, "grad_norm": 0.5177134871482849, "lr": 0.00012757574176896312, "s_per_step": 2.5112220001220704, "peak_gb": 35.709385216, "samples_seen": 67200, "time": 1791142641.47846}
86
+ {"step": 2125, "loss": 1.2264458799362183, "grad_norm": 0.5397101640701294, "lr": 0.0001259830205318278, "s_per_step": 2.3952220344543456, "peak_gb": 35.709385216, "samples_seen": 68000, "time": 1791142701.3594959}
87
+ {"step": 2150, "loss": 1.209760498404503, "grad_norm": 0.5998483300209045, "lr": 0.00012438319878625226, "s_per_step": 2.5326212215423585, "peak_gb": 35.709385216, "samples_seen": 68800, "time": 1791142764.675523}
88
+ {"step": 2175, "loss": 1.2117489975690843, "grad_norm": 0.5707188844680786, "lr": 0.00012277671372343195, "s_per_step": 2.4900473594665526, "peak_gb": 35.709385216, "samples_seen": 69600, "time": 1791142826.9271343}
89
+ {"step": 2200, "loss": 1.2348184549808503, "grad_norm": 0.5396585464477539, "lr": 0.00012116400435547992, "s_per_step": 2.492477035522461, "peak_gb": 35.709385216, "samples_seen": 70400, "time": 1791142889.2394576}
90
+ {"step": 2225, "loss": 1.2173430824279785, "grad_norm": 0.39253613352775574, "lr": 0.00011954551139545588, "s_per_step": 2.49932147026062, "peak_gb": 35.709385216, "samples_seen": 71200, "time": 1791142951.7229047}
91
+ {"step": 2250, "loss": 1.2054871225357056, "grad_norm": 0.5584589838981628, "lr": 0.00011792167713693036, "s_per_step": 2.4481615734100344, "peak_gb": 35.709385216, "samples_seen": 72000, "time": 1791143012.9273942}
92
+ {"step": 2275, "loss": 1.2060042411088943, "grad_norm": 0.5810075402259827, "lr": 0.0001162929453331168, "s_per_step": 2.66895037651062, "peak_gb": 35.709385216, "samples_seen": 72800, "time": 1791143079.6516283}
93
+ {"step": 2300, "loss": 1.2204527378082275, "grad_norm": 0.5374050140380859, "lr": 0.00011465976107560521, "s_per_step": 2.4989648342132567, "peak_gb": 35.709385216, "samples_seen": 73600, "time": 1791143142.12621}
94
+ {"step": 2325, "loss": 1.2322663950920105, "grad_norm": 0.5682536959648132, "lr": 0.00011302257067272952, "s_per_step": 2.3952228927612307, "peak_gb": 35.709385216, "samples_seen": 74400, "time": 1791143202.0071743}
95
+ {"step": 2350, "loss": 1.1958238542079926, "grad_norm": 0.4648407995700836, "lr": 0.00011138182152760286, "s_per_step": 2.6400107955932617, "peak_gb": 35.709385216, "samples_seen": 75200, "time": 1791143268.0078988}
96
+ {"step": 2375, "loss": 1.1953755247592925, "grad_norm": 0.4725061058998108, "lr": 0.00010973796201585335, "s_per_step": 2.4706161880493163, "peak_gb": 35.709385216, "samples_seen": 76000, "time": 1791143329.774484}
97
+ {"step": 2400, "loss": 1.2082021981477737, "grad_norm": 0.5769761800765991, "lr": 0.00010809144136309454, "s_per_step": 2.4323978424072266, "peak_gb": 35.709385216, "samples_seen": 76800, "time": 1791143390.5849032}
98
+ {"step": 2425, "loss": 1.2111977207660676, "grad_norm": 0.5421669483184814, "lr": 0.000106442709522163, "s_per_step": 2.4998205661773683, "peak_gb": 35.709385216, "samples_seen": 77600, "time": 1791143453.0808678}
99
+ {"step": 2450, "loss": 1.2098423671722411, "grad_norm": 0.484115332365036, "lr": 0.00010479221705015759, "s_per_step": 2.5170133399963377, "peak_gb": 35.709385216, "samples_seen": 78400, "time": 1791143516.0066352}
100
+ {"step": 2475, "loss": 1.223331753015518, "grad_norm": 0.49125203490257263, "lr": 0.00010314041498531371, "s_per_step": 2.4067738342285154, "peak_gb": 35.709385216, "samples_seen": 79200, "time": 1791143576.176388}
101
+ {"step": 2500, "loss": 1.2110005378723145, "grad_norm": 0.5610998868942261, "lr": 0.00010148775472374543, "s_per_step": 2.5736419677734377, "peak_gb": 35.709385216, "samples_seen": 80000, "time": 1791143640.5179398}
102
+ {"step": 2525, "loss": 1.2331609898805618, "grad_norm": 0.5671659708023071, "lr": 9.983468789609067e-05, "s_per_step": 2.5041617870330812, "peak_gb": 35.709385216, "samples_seen": 80800, "time": 1791143703.1224139}
103
+ {"step": 2550, "loss": 1.214245287179947, "grad_norm": 0.4357548654079437, "lr": 9.818166624409161e-05, "s_per_step": 2.4907432842254638, "peak_gb": 35.709385216, "samples_seen": 81600, "time": 1791143765.391451}
104
+ {"step": 2575, "loss": 1.214305859208107, "grad_norm": 0.5925732254981995, "lr": 9.6529141497145e-05, "s_per_step": 2.4908134937286377, "peak_gb": 35.709385216, "samples_seen": 82400, "time": 1791143827.6623023}
105
+ {"step": 2600, "loss": 1.1959744757413864, "grad_norm": 0.5484082698822021, "lr": 9.487756524885599e-05, "s_per_step": 2.544646863937378, "peak_gb": 35.709385216, "samples_seen": 83200, "time": 1791143891.2789044}
106
+ {"step": 2625, "loss": 1.2088198286294938, "grad_norm": 0.5090175271034241, "lr": 9.322738883362871e-05, "s_per_step": 2.5154834175109864, "peak_gb": 35.709385216, "samples_seen": 84000, "time": 1791143954.166384}
107
+ {"step": 2650, "loss": 1.2010453754663468, "grad_norm": 0.5693520307540894, "lr": 9.157906320332812e-05, "s_per_step": 2.5574374485015867, "peak_gb": 35.709385216, "samples_seen": 84800, "time": 1791144018.102869}
108
+ {"step": 2675, "loss": 1.214320752620697, "grad_norm": 0.5067724585533142, "lr": 8.993303880404592e-05, "s_per_step": 2.516569652557373, "peak_gb": 35.709385216, "samples_seen": 85600, "time": 1791144081.0175457}
109
+ {"step": 2700, "loss": 1.1918057614564896, "grad_norm": 0.49570122361183167, "lr": 8.828976545300505e-05, "s_per_step": 2.6196873950958253, "peak_gb": 35.709385216, "samples_seen": 86400, "time": 1791144146.5105312}
110
+ {"step": 2725, "loss": 1.1909438091516495, "grad_norm": 0.49999526143074036, "lr": 8.664969221563594e-05, "s_per_step": 2.577509126663208, "peak_gb": 35.709385216, "samples_seen": 87200, "time": 1791144210.9487243}
111
+ {"step": 2750, "loss": 1.2046889424324037, "grad_norm": 0.4591648280620575, "lr": 8.501326728285814e-05, "s_per_step": 2.463228282928467, "peak_gb": 35.709385216, "samples_seen": 88000, "time": 1791144272.529925}
112
+ {"step": 2775, "loss": 1.187312572002411, "grad_norm": 0.4475753605365753, "lr": 8.338093784860098e-05, "s_per_step": 2.65875358581543, "peak_gb": 35.709385216, "samples_seen": 88800, "time": 1791144338.9991896}
113
+ {"step": 2800, "loss": 1.2058877384662627, "grad_norm": 0.43829211592674255, "lr": 8.175314998759658e-05, "s_per_step": 2.525215663909912, "peak_gb": 35.709385216, "samples_seen": 89600, "time": 1791144402.1299999}
114
+ {"step": 2825, "loss": 1.1879939883947372, "grad_norm": 0.4615590274333954, "lr": 8.013034853347902e-05, "s_per_step": 2.61760648727417, "peak_gb": 35.709385216, "samples_seen": 90400, "time": 1791144467.5705826}
115
+ {"step": 2850, "loss": 1.1962288463115691, "grad_norm": 0.5492228269577026, "lr": 7.851297695722228e-05, "s_per_step": 2.416270399093628, "peak_gb": 35.709385216, "samples_seen": 91200, "time": 1791144527.9778004}
116
+ {"step": 2875, "loss": 1.189599106311798, "grad_norm": 0.4402064085006714, "lr": 7.690147724595069e-05, "s_per_step": 2.5021799850463866, "peak_gb": 35.709385216, "samples_seen": 92000, "time": 1791144590.5327408}
117
+ {"step": 2900, "loss": 1.1737463009357452, "grad_norm": 0.41730257868766785, "lr": 7.529628978215513e-05, "s_per_step": 2.671381711959839, "peak_gb": 35.709385216, "samples_seen": 92800, "time": 1791144657.317731}
118
+ {"step": 2925, "loss": 1.1911302745342254, "grad_norm": 0.5390501022338867, "lr": 7.369785322334733e-05, "s_per_step": 2.4857681655883788, "peak_gb": 35.709385216, "samples_seen": 93600, "time": 1791144719.4624217}
119
+ {"step": 2950, "loss": 1.1944737762212754, "grad_norm": 0.5601727962493896, "lr": 7.210660438218596e-05, "s_per_step": 2.4794444179534914, "peak_gb": 35.709385216, "samples_seen": 94400, "time": 1791144781.4489458}
120
+ {"step": 2975, "loss": 1.1881670886278153, "grad_norm": 0.5373884439468384, "lr": 7.052297810710643e-05, "s_per_step": 2.483798608779907, "peak_gb": 35.709385216, "samples_seen": 95200, "time": 1791144843.5443158}
121
+ {"step": 3000, "loss": 1.2002181720733642, "grad_norm": 0.5904892683029175, "lr": 6.894740716348797e-05, "s_per_step": 2.475346622467041, "peak_gb": 35.709385216, "samples_seen": 96000, "time": 1791144905.4283826}
122
+ {"step": 3025, "loss": 1.1833700090646744, "grad_norm": 0.5156553387641907, "lr": 6.738032211538946e-05, "s_per_step": 2.6700209045410155, "peak_gb": 35.709385216, "samples_seen": 96800, "time": 1791144972.179436}
123
+ {"step": 3050, "loss": 1.1821686124801636, "grad_norm": 0.5594586730003357, "lr": 6.582215120788719e-05, "s_per_step": 2.4451071548461916, "peak_gb": 35.709385216, "samples_seen": 97600, "time": 1791145033.3075602}
124
+ {"step": 3075, "loss": 1.1715920561552047, "grad_norm": 0.44480738043785095, "lr": 6.427332025004628e-05, "s_per_step": 2.616603384017944, "peak_gb": 35.709385216, "samples_seen": 98400, "time": 1791145098.7230656}
125
+ {"step": 3100, "loss": 1.178504084944725, "grad_norm": 0.527042031288147, "lr": 6.273425249855752e-05, "s_per_step": 2.5396651363372804, "peak_gb": 35.709385216, "samples_seen": 99200, "time": 1791145162.2151375}
126
+ {"step": 3125, "loss": 1.1844954544305801, "grad_norm": 0.521228551864624, "lr": 6.120536854207215e-05, "s_per_step": 2.5937981700897215, "peak_gb": 35.709385216, "samples_seen": 100000, "time": 1791145227.0605233}
127
+ {"step": 3150, "loss": 1.197044386267662, "grad_norm": 0.5560638308525085, "lr": 5.968708618626539e-05, "s_per_step": 2.4499312114715575, "peak_gb": 35.709385216, "samples_seen": 100800, "time": 1791145288.309198}
128
+ {"step": 3175, "loss": 1.1869262939691543, "grad_norm": 0.5193493962287903, "lr": 5.817982033966059e-05, "s_per_step": 2.4811374282836915, "peak_gb": 35.709385216, "samples_seen": 101600, "time": 1791145350.3380346}
129
+ {"step": 3200, "loss": 1.17697389960289, "grad_norm": 0.5577865839004517, "lr": 5.668398290024517e-05, "s_per_step": 2.537874908447266, "peak_gb": 35.709385216, "samples_seen": 102400, "time": 1791145413.7853076}
130
+ {"step": 3225, "loss": 1.1980028331279755, "grad_norm": 0.41638582944869995, "lr": 5.5199982642909375e-05, "s_per_step": 2.443896884918213, "peak_gb": 35.709385216, "samples_seen": 103200, "time": 1791145474.8831582}
131
+ {"step": 3250, "loss": 1.2014775633811952, "grad_norm": 0.5623952150344849, "lr": 5.37282251077381e-05, "s_per_step": 2.4464002323150633, "peak_gb": 35.709385216, "samples_seen": 104000, "time": 1791145536.0436106}
132
+ {"step": 3275, "loss": 1.2081005316972733, "grad_norm": 0.5129700899124146, "lr": 5.2269112489187e-05, "s_per_step": 2.5506320571899415, "peak_gb": 35.709385216, "samples_seen": 104800, "time": 1791145599.8098936}
133
+ {"step": 3300, "loss": 1.1788170540332794, "grad_norm": 0.505996823310852, "lr": 5.082304352617291e-05, "s_per_step": 2.4294992446899415, "peak_gb": 35.709385216, "samples_seen": 105600, "time": 1791145660.547851}
134
+ {"step": 3325, "loss": 1.1952128916978837, "grad_norm": 0.48522287607192993, "lr": 4.939041339310858e-05, "s_per_step": 2.348098773956299, "peak_gb": 35.709385216, "samples_seen": 106400, "time": 1791145719.2507544}
135
+ {"step": 3350, "loss": 1.164835900068283, "grad_norm": 0.4022349715232849, "lr": 4.797161359191104e-05, "s_per_step": 2.542964315414429, "peak_gb": 35.709385216, "samples_seen": 107200, "time": 1791145782.8253438}
136
+ {"step": 3375, "loss": 1.1721830260753632, "grad_norm": 0.5090497136116028, "lr": 4.656703184501427e-05, "s_per_step": 2.52671706199646, "peak_gb": 35.709385216, "samples_seen": 108000, "time": 1791145845.9937546}
137
+ {"step": 3400, "loss": 1.1724821329116821, "grad_norm": 0.5728841423988342, "lr": 4.517705198941442e-05, "s_per_step": 2.4924769973754883, "peak_gb": 35.709385216, "samples_seen": 108800, "time": 1791145908.3061934}
138
+ {"step": 3425, "loss": 1.1960626763105393, "grad_norm": 0.562131941318512, "lr": 4.380205387177645e-05, "s_per_step": 2.408029260635376, "peak_gb": 35.709385216, "samples_seen": 109600, "time": 1791145968.5073981}
139
+ {"step": 3450, "loss": 1.1602804332971572, "grad_norm": 0.4601585268974304, "lr": 4.244241324463182e-05, "s_per_step": 2.5518965339660644, "peak_gb": 35.709385216, "samples_seen": 110400, "time": 1791146032.3052933}
140
+ {"step": 3475, "loss": 1.1917330080270767, "grad_norm": 0.5174102187156677, "lr": 4.109850166369465e-05, "s_per_step": 2.4686901664733885, "peak_gb": 35.709385216, "samples_seen": 111200, "time": 1791146094.0923746}
141
+ {"step": 3500, "loss": 1.176536915898323, "grad_norm": 0.46731457114219666, "lr": 3.977068638632486e-05, "s_per_step": 2.5151400184631347, "peak_gb": 35.709385216, "samples_seen": 112000, "time": 1791146156.9713702}
142
+ {"step": 3525, "loss": 1.1758270978927612, "grad_norm": 0.5187398195266724, "lr": 3.845933027116595e-05, "s_per_step": 2.5711546421051024, "peak_gb": 35.709385216, "samples_seen": 112800, "time": 1791146221.2507052}
143
+ {"step": 3550, "loss": 1.1746439105272293, "grad_norm": 0.46097496151924133, "lr": 3.716479167898476e-05, "s_per_step": 2.4746162128448486, "peak_gb": 35.709385216, "samples_seen": 113600, "time": 1791146283.1165752}
144
+ {"step": 3575, "loss": 1.1660949611663818, "grad_norm": 0.5945229530334473, "lr": 3.588742437474063e-05, "s_per_step": 2.5315122985839844, "peak_gb": 35.709385216, "samples_seen": 114400, "time": 1791146346.4049218}
145
+ {"step": 3600, "loss": 1.177633284330368, "grad_norm": 0.42318975925445557, "lr": 3.4627577430910066e-05, "s_per_step": 2.5272295093536377, "peak_gb": 35.709385216, "samples_seen": 115200, "time": 1791146409.5861077}
146
+ {"step": 3625, "loss": 1.1747868591547013, "grad_norm": 0.5484309792518616, "lr": 3.3385595132094104e-05, "s_per_step": 2.5058855533599855, "peak_gb": 35.709385216, "samples_seen": 116000, "time": 1791146472.2342918}
147
+ {"step": 3650, "loss": 1.1845440715551376, "grad_norm": 0.5164344906806946, "lr": 3.2161816880933996e-05, "s_per_step": 2.420202522277832, "peak_gb": 35.709385216, "samples_seen": 116800, "time": 1791146532.7397869}
148
+ {"step": 3675, "loss": 1.1623845690488814, "grad_norm": 0.44582146406173706, "lr": 3.095657710536086e-05, "s_per_step": 2.5247189235687255, "peak_gb": 35.709385216, "samples_seen": 117600, "time": 1791146595.8581657}
149
+ {"step": 3700, "loss": 1.177596697807312, "grad_norm": 0.6056506633758545, "lr": 2.977020516720501e-05, "s_per_step": 2.4129674339294436, "peak_gb": 35.709385216, "samples_seen": 118400, "time": 1791146656.1832361}
150
+ {"step": 3725, "loss": 1.1791141772270202, "grad_norm": 0.5385143160820007, "lr": 2.860302527218951e-05, "s_per_step": 2.466406307220459, "peak_gb": 35.709385216, "samples_seen": 119200, "time": 1791146717.8438177}
151
+ {"step": 3750, "loss": 1.1454178375005721, "grad_norm": 0.5121068954467773, "lr": 2.745535638133304e-05, "s_per_step": 2.5826826667785645, "peak_gb": 35.709385216, "samples_seen": 120000, "time": 1791146782.4113297}
152
+ {"step": 3775, "loss": 1.1689221292734147, "grad_norm": 0.5559164881706238, "lr": 2.632751212378568e-05, "s_per_step": 2.653624095916748, "peak_gb": 35.709385216, "samples_seen": 120800, "time": 1791146848.7524045}
153
+ {"step": 3800, "loss": 1.182214805483818, "grad_norm": 0.4873104691505432, "lr": 2.5219800711121976e-05, "s_per_step": 2.498158531188965, "peak_gb": 35.709385216, "samples_seen": 121600, "time": 1791146911.2068315}
154
+ {"step": 3825, "loss": 1.1555895429849625, "grad_norm": 0.576235294342041, "lr": 2.4132524853114456e-05, "s_per_step": 2.463836088180542, "peak_gb": 35.709385216, "samples_seen": 122400, "time": 1791146972.803223}
155
+ {"step": 3850, "loss": 1.163134110569954, "grad_norm": 0.49555864930152893, "lr": 2.306598167501064e-05, "s_per_step": 2.5268060398101806, "peak_gb": 35.709385216, "samples_seen": 123200, "time": 1791147035.973819}
156
+ {"step": 3875, "loss": 1.1696758967638017, "grad_norm": 0.519751250743866, "lr": 2.2020462636336137e-05, "s_per_step": 2.439229755401611, "peak_gb": 35.709385216, "samples_seen": 124000, "time": 1791147096.9549644}
157
+ {"step": 3900, "loss": 1.1512171399593354, "grad_norm": 0.42172372341156006, "lr": 2.0996253451246017e-05, "s_per_step": 2.601565704345703, "peak_gb": 35.709385216, "samples_seen": 124800, "time": 1791147161.9945152}
158
+ {"step": 3925, "loss": 1.1549856066703796, "grad_norm": 0.5235016942024231, "lr": 1.9993634010446448e-05, "s_per_step": 2.5241199016571043, "peak_gb": 35.709385216, "samples_seen": 125600, "time": 1791147225.0979981}
159
+ {"step": 3950, "loss": 1.1684985667467118, "grad_norm": 0.5432916283607483, "lr": 1.9012878304707403e-05, "s_per_step": 2.5513547325134276, "peak_gb": 35.709385216, "samples_seen": 126400, "time": 1791147288.8822782}
160
+ {"step": 3975, "loss": 1.1723533356189728, "grad_norm": 0.5285301804542542, "lr": 1.805425434998782e-05, "s_per_step": 2.436748571395874, "peak_gb": 35.709385216, "samples_seen": 127200, "time": 1791147349.801405}
161
+ {"step": 4000, "loss": 1.1497523707151414, "grad_norm": 0.44535908102989197, "lr": 1.711802411419384e-05, "s_per_step": 2.6349299812316893, "peak_gb": 35.709385216, "samples_seen": 128000, "time": 1791147415.675062}
162
+ {"step": 4025, "loss": 1.1554417186975479, "grad_norm": 0.4447822570800781, "lr": 1.6204443445589225e-05, "s_per_step": 2.693064069747925, "peak_gb": 35.709385216, "samples_seen": 128800, "time": 1791147483.002096}
163
+ {"step": 4050, "loss": 1.1771225923299788, "grad_norm": 0.5996832251548767, "lr": 1.5313762002878586e-05, "s_per_step": 2.506432514190674, "peak_gb": 35.709385216, "samples_seen": 129600, "time": 1791147545.6633337}
164
+ {"step": 4075, "loss": 1.156247045993805, "grad_norm": 0.5294703841209412, "lr": 1.444622318698191e-05, "s_per_step": 2.4499019813537597, "peak_gb": 35.709385216, "samples_seen": 130400, "time": 1791147606.9112895}
165
+ {"step": 4100, "loss": 1.1654343736171722, "grad_norm": 0.472944051027298, "lr": 1.3602064074519216e-05, "s_per_step": 2.5358735370635985, "peak_gb": 35.709385216, "samples_seen": 131200, "time": 1791147670.3085465}
166
+ {"step": 4125, "loss": 1.164423559308052, "grad_norm": 0.41716107726097107, "lr": 1.2781515353023365e-05, "s_per_step": 2.4599155235290526, "peak_gb": 35.709385216, "samples_seen": 132000, "time": 1791147731.8068397}
167
+ {"step": 4150, "loss": 1.1569818568229675, "grad_norm": 0.6027914881706238, "lr": 1.1984801257898925e-05, "s_per_step": 2.5184785079956056, "peak_gb": 35.709385216, "samples_seen": 132800, "time": 1791147794.76921}
168
+ {"step": 4175, "loss": 1.1682433980703353, "grad_norm": 0.5346322059631348, "lr": 1.1212139511144448e-05, "s_per_step": 2.3972896575927733, "peak_gb": 35.709385216, "samples_seen": 133600, "time": 1791147854.701882}
169
+ {"step": 4200, "loss": 1.1649791526794433, "grad_norm": 0.5470781326293945, "lr": 1.0463741261854288e-05, "s_per_step": 2.492423486709595, "peak_gb": 35.709385216, "samples_seen": 134400, "time": 1791147917.01288}
170
+ {"step": 4225, "loss": 1.1546978533267975, "grad_norm": 0.4576592743396759, "lr": 9.739811028516932e-06, "s_per_step": 2.527438831329346, "peak_gb": 35.709385216, "samples_seen": 135200, "time": 1791147980.1992562}
171
+ {"step": 4250, "loss": 1.1561020785570144, "grad_norm": 0.5178248286247253, "lr": 9.040546643125203e-06, "s_per_step": 2.517265033721924, "peak_gb": 35.709385216, "samples_seen": 136000, "time": 1791148043.1313138}
172
+ {"step": 4275, "loss": 1.1623137599229814, "grad_norm": 0.47821733355522156, "lr": 8.366139197113831e-06, "s_per_step": 2.5566812801361083, "peak_gb": 35.709385216, "samples_seen": 136800, "time": 1791148107.0487514}
173
+ {"step": 4300, "loss": 1.1758109939098358, "grad_norm": 0.5945451855659485, "lr": 7.716772989138744e-06, "s_per_step": 2.4145172882080077, "peak_gb": 35.709385216, "samples_seen": 137600, "time": 1791148167.4120862}
174
+ {"step": 4325, "loss": 1.1620000940561295, "grad_norm": 0.4968617260456085, "lr": 7.092625474713055e-06, "s_per_step": 2.42730583190918, "peak_gb": 35.709385216, "samples_seen": 138400, "time": 1791148228.0951617}
175
+ {"step": 4350, "loss": 1.156725025177002, "grad_norm": 0.45200902223587036, "lr": 6.493867217712868e-06, "s_per_step": 2.5541227149963377, "peak_gb": 35.709385216, "samples_seen": 139200, "time": 1791148291.9486508}
176
+ {"step": 4375, "loss": 1.1519554841518402, "grad_norm": 0.5528603792190552, "lr": 5.920661843766395e-06, "s_per_step": 2.648443021774292, "peak_gb": 35.889089024, "samples_seen": 140000, "time": 1791148358.1601393}
177
+ {"step": 4400, "loss": 1.1472780203819275, "grad_norm": 0.5185584425926208, "lr": 5.3731659955391865e-06, "s_per_step": 2.5665468978881836, "peak_gb": 35.889089024, "samples_seen": 140800, "time": 1791148422.3242114}
178
+ {"step": 4425, "loss": 1.16106074988842, "grad_norm": 0.47097915410995483, "lr": 4.8515292899276695e-06, "s_per_step": 2.5675136184692384, "peak_gb": 35.889089024, "samples_seen": 141600, "time": 1791148486.5124738}
179
+ {"step": 4450, "loss": 1.1584270763397218, "grad_norm": 0.5423069596290588, "lr": 4.355894277172545e-06, "s_per_step": 2.5209413051605223, "peak_gb": 35.889089024, "samples_seen": 142400, "time": 1791148549.5364525}
180
+ {"step": 4475, "loss": 1.167509251832962, "grad_norm": 0.4573025703430176, "lr": 3.886396401903381e-06, "s_per_step": 2.452349863052368, "peak_gb": 35.889089024, "samples_seen": 143200, "time": 1791148610.8456094}
181
+ {"step": 4500, "loss": 1.15239905834198, "grad_norm": 0.5477348566055298, "lr": 3.443163966125007e-06, "s_per_step": 2.4918809604644774, "peak_gb": 35.889089024, "samples_seen": 144000, "time": 1791148673.1431098}
182
+ {"step": 4525, "loss": 1.161824835538864, "grad_norm": 0.6060011386871338, "lr": 3.0263180941558332e-06, "s_per_step": 2.6219866466522217, "peak_gb": 35.889089024, "samples_seen": 144800, "time": 1791148738.6932275}
183
+ {"step": 4550, "loss": 1.1565506535768508, "grad_norm": 0.5005341172218323, "lr": 2.6359726995275004e-06, "s_per_step": 2.566774663925171, "peak_gb": 35.889089024, "samples_seen": 145600, "time": 1791148802.8630507}
184
+ {"step": 4575, "loss": 1.162327172756195, "grad_norm": 0.5122349858283997, "lr": 2.2722344538552597e-06, "s_per_step": 2.571699342727661, "peak_gb": 35.889089024, "samples_seen": 146400, "time": 1791148867.1559367}
185
+ {"step": 4600, "loss": 1.1596791630983352, "grad_norm": 0.5953731536865234, "lr": 1.9352027576872824e-06, "s_per_step": 2.364423170089722, "peak_gb": 35.889089024, "samples_seen": 147200, "time": 1791148926.266977}
186
+ {"step": 4625, "loss": 1.1516347283124924, "grad_norm": 0.49013832211494446, "lr": 1.6249697133409069e-06, "s_per_step": 2.5397282600402833, "peak_gb": 35.889089024, "samples_seen": 148000, "time": 1791148989.760581}
187
+ {"step": 4650, "loss": 1.1604260104894637, "grad_norm": 0.44168010354042053, "lr": 1.341620099733487e-06, "s_per_step": 2.495543918609619, "peak_gb": 35.889089024, "samples_seen": 148800, "time": 1791149052.149569}
188
+ {"step": 4675, "loss": 1.144905739426613, "grad_norm": 0.47312289476394653, "lr": 1.0852313492143551e-06, "s_per_step": 2.527244834899902, "peak_gb": 35.889089024, "samples_seen": 149600, "time": 1791149115.3311636}
189
+ {"step": 4700, "loss": 1.164122424721718, "grad_norm": 0.5258249044418335, "lr": 8.558735264045714e-07, "s_per_step": 2.392429494857788, "peak_gb": 35.889089024, "samples_seen": 150400, "time": 1791149175.1423135}
190
+ {"step": 4725, "loss": 1.1654922747612, "grad_norm": 0.5145270228385925, "lr": 6.536093090499517e-07, "s_per_step": 2.493706407546997, "peak_gb": 35.889089024, "samples_seen": 151200, "time": 1791149237.4853363}
191
+ {"step": 4750, "loss": 1.170558487176895, "grad_norm": 0.6493576169013977, "lr": 4.78493970892846e-07, "s_per_step": 2.3715320777893067, "peak_gb": 35.889089024, "samples_seen": 152000, "time": 1791149296.775865}
192
+ {"step": 4775, "loss": 1.166077768802643, "grad_norm": 0.5020484328269958, "lr": 3.3057536656722065e-07, "s_per_step": 2.589771041870117, "peak_gb": 35.889089024, "samples_seen": 152800, "time": 1791149361.5205963}
193
+ {"step": 4800, "loss": 1.1575500231981277, "grad_norm": 0.4154891073703766, "lr": 2.0989391852116458e-07, "s_per_step": 2.4963897895812988, "peak_gb": 35.889089024, "samples_seen": 153600, "time": 1791149423.930882}
194
+ {"step": 4825, "loss": 1.1519961988925933, "grad_norm": 0.5121576189994812, "lr": 1.1648260597041383e-07, "s_per_step": 2.559989347457886, "peak_gb": 35.889089024, "samples_seen": 154400, "time": 1791149487.9312465}
195
+ {"step": 4850, "loss": 1.1676404130458833, "grad_norm": 0.550052285194397, "lr": 5.036695588606088e-08, "s_per_step": 2.4459603118896482, "peak_gb": 35.889089024, "samples_seen": 155200, "time": 1791149549.080826}
196
+ {"step": 4875, "loss": 1.169189488887787, "grad_norm": 0.5803373456001282, "lr": 1.1565036018568176e-08, "s_per_step": 2.4609833908081056, "peak_gb": 35.889089024, "samples_seen": 156000, "time": 1791149610.6059558}
197
+ {"step": 4897, "loss": 1.1497614966197447, "grad_norm": 0.4358130097389221, "lr": 2.1862492483037954e-11, "s_per_step": 2.5117971788753164, "peak_gb": 35.889089024, "samples_seen": 156704, "time": 1791149665.8659236}
stage1/projector.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:43bfac0c0f58fb43a395d253206896f0dfb1c30ecd1cc90750bb13bfc91cbced
3
+ size 79724872
stage2/lora_adapter/adapter_config.json ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "NCAIR1/N-ATLaS",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "kasa_config": null,
16
+ "layer_replication": null,
17
+ "layers_pattern": null,
18
+ "layers_to_transform": null,
19
+ "loftq_config": {},
20
+ "lora_alpha": 128,
21
+ "lora_bias": false,
22
+ "lora_dropout": 0.05,
23
+ "lora_ga_config": null,
24
+ "megatron_config": null,
25
+ "megatron_core": "megatron.core",
26
+ "modules_to_save": null,
27
+ "monteclora_config": null,
28
+ "peft_type": "LORA",
29
+ "peft_version": "0.21.2",
30
+ "qalora_group_size": 16,
31
+ "r": 64,
32
+ "rank_pattern": {},
33
+ "revision": null,
34
+ "target_modules": [
35
+ "k_proj",
36
+ "o_proj",
37
+ "q_proj",
38
+ "up_proj",
39
+ "down_proj",
40
+ "gate_proj",
41
+ "v_proj"
42
+ ],
43
+ "target_parameters": null,
44
+ "task_type": "CAUSAL_LM",
45
+ "trainable_token_indices": null,
46
+ "use_bdlora": null,
47
+ "use_dora": false,
48
+ "use_qalora": false,
49
+ "use_rslora": false,
50
+ "velora_config": null
51
+ }
stage2/lora_adapter/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:06df60018c6f604a693251a4a95e96d00ab90a5431f7421f2b6a57bc0020cba4
3
+ size 671149168
stage2/projector.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e6b3dd6e3137450b00c0a11467c6476dc32300d41a68ea63bdb7aa562af51b44
3
+ size 79724872