Image-Text-to-Text
PEFT
Safetensors
vision-language
multimodal
llava
lora
siglip2
n-atlas
nigerian-languages
UncleanCode commited on
Commit
c8317c0
·
verified ·
1 Parent(s): 2a2540a

Upload via givemeanode export_data

Browse files
README.md CHANGED
@@ -36,6 +36,8 @@ following the two-stage LLaVA recipe:
36
  1. **Stage 1 — alignment:** only the projector is trained, on 300k image–caption pairs, so the LLM can "read" image features.
37
  2. **Stage 2 — visual instruction tuning:** the projector keeps training and LoRA adapters are added to N-ATLaS, on
38
  LLaVA-Instruct-150K conversations, so the model answers questions and holds conversations about images.
 
 
39
 
40
  This repository contains **only the trained parts** (projectors and LoRA adapter, ~0.9 GB).
41
  The base models are downloaded from their own repositories at load time, so their licenses and access conditions
@@ -59,6 +61,8 @@ The base models are downloaded from their own repositories at load time, so thei
59
  stage1/projector.safetensors # stage-1 projector (image captioning / alignment)
60
  stage2/projector.safetensors # stage-2 projector (use together with the LoRA adapter)
61
  stage2/lora_adapter/ # PEFT LoRA adapter for NCAIR1/N-ATLaS
 
 
62
  chat.py # standalone inference script (CLI + Python API)
63
  code/ # exact training, evaluation and launch scripts used
64
  eval/ # raw evaluation results (JSON) and report
@@ -74,7 +78,7 @@ pip install -U torch transformers peft safetensors pillow accelerate huggingface
74
  export HF_TOKEN=hf_... # a token from the account that was granted N-ATLaS access
75
  wget https://huggingface.co/FUTO-NIGERIA/AtlasVision/resolve/main/chat.py
76
 
77
- python chat.py --image photo.jpg --question "What is happening in this picture?"
78
  python chat.py --image photo.jpg --question "Kedu ihe dị na foto a?" # Igbo
79
  python chat.py --stage 1 --image photo.jpg # stage-1 captioner
80
  python chat.py --image photo.jpg --interactive # follow-up questions
@@ -88,7 +92,7 @@ From Python:
88
  ```python
89
  from chat import AtlasVision, load_image
90
 
91
- model = AtlasVision(stage=2) # downloads N-ATLaS, SigLIP2 and this repo's weights
92
  image = load_image("https://example.com/street.jpg")
93
  print(model.ask(image, "Describe this image in detail."))
94
  print(model.ask(image, "How many people are there?")) # follow-ups keep the conversation
@@ -109,17 +113,17 @@ Follow-up turns append `{answer}<|eot_id|><|start_header_id|>user<|end_header_id
109
 
110
  Both stages ran on a single NVIDIA H100 80GB.
111
 
112
- | | Stage 1 — alignment | Stage 2 — instruction tuning |
113
- |---|---|---|
114
- | Data | LLaVA-Pretrain (BLIP captions of LAION/CC/SBU), random 300k of 558k | LLaVA-Instruct-150K on COCO train2017 (156,712 train / 1,000 held out) |
115
- | Trained | projector | projector + LoRA |
116
- | Effective batch | 32 (16 × 2 accumulation) | 32 (8 × 4 accumulation) |
117
- | Learning rate | 2e-4, 3% warmup, cosine | LoRA 2e-4, projector 2e-5, 3% warmup, cosine |
118
- | Max text length | 128 tokens | 1,024 tokens |
119
- | Optimizer steps | 9375 | 4897 |
120
- | Loss (first log → last log) | 8.551 → 2.454 | 2.102 → 1.15 |
121
- | Wall-clock time | ~91 min | ~206 min |
122
- | Peak GPU memory | 45.2 GB | 35.9 GB (gradient checkpointing) |
123
 
124
  Other details: AdamW (no weight decay), gradient clipping at 1.0, bf16 autocast, loss only on assistant tokens,
125
  length-bucketed batches in stage 2. Per-step logs are in `logs/`.
@@ -139,28 +143,40 @@ Measured on 2,000 LLaVA-Pretrain captions that were **not** in the training subs
139
  Held-out loss matches the final training loss (no over-fitting), and pairing captions with the wrong images
140
  roughly doubles the loss — the language model is relying on the visual content, not guessing generic captions.
141
 
142
- ### Stage 1 vs stage 2
143
 
144
- | Metric | Stage 1 | Stage 2 |
145
- |---|---|---|
146
- | Held-out LLaVA-Instruct loss (500 unseen conversations, lower is better) | 2.2521 | **1.1271** |
147
 
148
- **POPE** (object hallucination; yes/no questions about COCO val2014 images, scored from the Yes/No token probabilities;
149
- balanced 50% yes / 50% no, so a yes-ratio near 0.5 is ideal):
150
 
151
- | POPE split | Stage 1: accuracy / F1 / yes-ratio | Stage 2: accuracy / F1 / yes-ratio | Stage 2: precision / recall |
152
- |---|---|---|---|
153
- | random | 0.500 / 0.200 / 0.13 | **0.543** / **0.686** / 0.95 | 0.522 / 0.997 |
154
- | popular | 0.503 / 0.201 / 0.12 | **0.522** / **0.676** / 0.98 | 0.511 / 0.997 |
155
- | adversarial | 0.502 / 0.201 / 0.12 | **0.514** / **0.672** / 0.98 | 0.507 / 0.997 |
156
-
157
- **How to read this: stage 2 has a strong "yes" bias.** It answers "yes" to 95%–98% of POPE
158
- questions, so it almost never misses an object that is present (recall ≈ 1.0) but also confirms most objects that are
159
- *absent*, leaving accuracy close to chance. Stage 1 shows the opposite bias (it was never trained to answer questions).
160
- This is a known effect of training on LLaVA-Instruct-150K alone: its conversations almost never contain questions whose
161
- answer is "no". The descriptive examples below are far stronger than these numbers suggest, but **do not use AtlasVision
162
- for yes/no verification of image content**. The usual fix is to add short-answer VQA data (VQAv2, GQA) and negative
163
- examples (e.g. LRV-Instruction) to stage 2, as LLaVA-1.5 does — it reaches POPE F1 ≈ 0.86 with 665k mixed samples at 336 px.
 
 
 
 
 
 
 
 
 
 
 
 
164
 
165
  ### Examples — stage 1 captions on unseen images
166
 
@@ -173,47 +189,38 @@ examples (e.g. LRV-Instruction) to stage 2, as LLaVA-1.5 does — it reaches POP
173
  | the beach in el nido national park, puerto puerto | the limestone cliffs and limestone islands in the background |
174
  | three pieces of paper with the words, democratic decentified dp controlled centralized ccp | a diagram showing the different types of democracy |
175
 
176
- ### Examples — "Describe this image in detail." (COCO val2014, unseen)
177
 
178
  **COCO_val2014_000000310196**
179
 
180
  - *Stage 1:* a skier in the snow on a mountain
181
- - *Stage 2:* The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.
182
 
183
- There are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together.
184
 
185
  **COCO_val2014_000000210789**
186
 
187
  - *Stage 1:* a woman and her child in the rain
188
- - *Stage 2:* The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.
189
 
190
- In the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene.
191
 
192
  **COCO_val2014_000000429109**
193
 
194
  - *Stage 1:* a bus and several other vehicles parked in front of a building
195
- - *Stage 2:* The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.
196
-
197
- Numerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.
198
-
199
- In addition to the vehicles, there are traffic lights at various points along the street
200
-
201
- **COCO_val2014_000000211674**
202
-
203
- - *Stage 1:* a bus with a red and white logo on it, carrying passengers
204
- - *Stage 2:* The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.
205
 
206
- In addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by.
207
 
208
- ### Questions in Nigerian languages (stage 2)
209
 
210
- The stage-2 instruction data is English only. The model understands the questions below (its answers match the image) but replies in English; adding translated instruction data would be needed for answers in these languages.
211
 
212
  | Language | Question | Answer |
213
  |---|---|---|
214
- | Igbo | Kedu ihe dị na foto a? | The image features a person skiing down a snow-covered slope, with the skier wearing red pants. |
215
- | Yoruba | Kí ni ó wà nínú àwòrán yìí? | The image features a person skiing down a snow-covered slope, with the skier wearing red pants. |
216
- | Hausa | Me ke cikin wannan hoton? | In the image, a person is skiing down a snow-covered slope. |
217
 
218
  ### Text-only check (no image)
219
 
@@ -228,9 +235,11 @@ Both answer fluently in Igbo, so the adapter keeps N-ATLaS's text abilities inta
228
 
229
  ## Limitations
230
 
231
- - **Hallucination and yes-bias.** Long descriptions are fluent but often add plausible details that are not in the
232
- image (exact counts, extra cars, a handbag), and stage 2 answers "yes" to most yes/no questions (see POPE above).
233
- Treat counts, small objects and yes/no answers as unreliable.
 
 
234
  - **English-only visual training.** All image–text training data is English. Answers to Hausa, Igbo and Yoruba questions
235
  come from N-ATLaS's own multilingual ability and are noticeably less reliable; they may switch to English.
236
  - **Low resolution.** Images are resized to 224×224, so small text, fine details and dense documents are hard.
 
36
  1. **Stage 1 — alignment:** only the projector is trained, on 300k image–caption pairs, so the LLM can "read" image features.
37
  2. **Stage 2 — visual instruction tuning:** the projector keeps training and LoRA adapters are added to N-ATLaS, on
38
  LLaVA-Instruct-150K conversations, so the model answers questions and holds conversations about images.
39
+ 3. **Stage 2b — de-biasing (recommended):** stage 2 is continued on a 150k mix of short-answer VQA data, which fixes a
40
+ severe "yes" bias and raises POPE accuracy from ~52% to ~78%. **Use `stage2b/` unless you are reproducing earlier results.**
41
 
42
  This repository contains **only the trained parts** (projectors and LoRA adapter, ~0.9 GB).
43
  The base models are downloaded from their own repositories at load time, so their licenses and access conditions
 
61
  stage1/projector.safetensors # stage-1 projector (image captioning / alignment)
62
  stage2/projector.safetensors # stage-2 projector (use together with the LoRA adapter)
63
  stage2/lora_adapter/ # PEFT LoRA adapter for NCAIR1/N-ATLaS
64
+ stage2b/projector.safetensors # RECOMMENDED: de-biased projector
65
+ stage2b/lora_adapter/ # RECOMMENDED: de-biased LoRA adapter
66
  chat.py # standalone inference script (CLI + Python API)
67
  code/ # exact training, evaluation and launch scripts used
68
  eval/ # raw evaluation results (JSON) and report
 
78
  export HF_TOKEN=hf_... # a token from the account that was granted N-ATLaS access
79
  wget https://huggingface.co/FUTO-NIGERIA/AtlasVision/resolve/main/chat.py
80
 
81
+ python chat.py --image photo.jpg --question "What is happening in this picture?" # add --stage 2b
82
  python chat.py --image photo.jpg --question "Kedu ihe dị na foto a?" # Igbo
83
  python chat.py --stage 1 --image photo.jpg # stage-1 captioner
84
  python chat.py --image photo.jpg --interactive # follow-up questions
 
92
  ```python
93
  from chat import AtlasVision, load_image
94
 
95
+ model = AtlasVision(stage="2b") # downloads N-ATLaS, SigLIP2 and this repo's weights
96
  image = load_image("https://example.com/street.jpg")
97
  print(model.ask(image, "Describe this image in detail."))
98
  print(model.ask(image, "How many people are there?")) # follow-ups keep the conversation
 
113
 
114
  Both stages ran on a single NVIDIA H100 80GB.
115
 
116
+ | | Stage 1 — alignment | Stage 2 — instruction tuning | Stage 2b — de-biasing |
117
+ |---|---|---|---|
118
+ | Data | LLaVA-Pretrain (BLIP captions of LAION/CC/SBU), random 300k of 558k | LLaVA-Instruct-150K on COCO train2017 (156,712 train / 1,000 held out) | LLaVA-1.5 mix: 110k short-answer VQA (VQAv2 / OK-VQA / A-OKVQA) + 40k LLaVA-Instruct replay, COCO images |
119
+ | Trained | projector | projector + LoRA | projector + LoRA (continued from stage 2) |
120
+ | Effective batch | 32 (16 × 2 accumulation) | 32 (8 × 4 accumulation) | 32 (8 × 4 accumulation) |
121
+ | Learning rate | 2e-4, 3% warmup, cosine | LoRA 2e-4, projector 2e-5, 3% warmup, cosine | LoRA 1e-4, projector 1e-5, 3% warmup, cosine |
122
+ | Max text length | 128 tokens | 1,024 tokens | 1,024 tokens |
123
+ | Optimizer steps | 9375 | 4897 | 4,681 |
124
+ | Loss (first log → last log) | 8.551 → 2.454 | 2.102 → 1.15 | 3.14 → 0.55 |
125
+ | Wall-clock time | ~91 min | ~206 min | ~172 min |
126
+ | Peak GPU memory | 45.2 GB | 35.9 GB (gradient checkpointing) | 37.3 GB |
127
 
128
  Other details: AdamW (no weight decay), gradient clipping at 1.0, bf16 autocast, loss only on assistant tokens,
129
  length-bucketed batches in stage 2. Per-step logs are in `logs/`.
 
143
  Held-out loss matches the final training loss (no over-fitting), and pairing captions with the wrong images
144
  roughly doubles the loss — the language model is relying on the visual content, not guessing generic captions.
145
 
146
+ ### Stage 1 vs stage 2 vs stage 2b
147
 
148
+ | Metric | Stage 1 | Stage 2 | Stage 2b (recommended) |
149
+ |---|---|---|---|
150
+ | Held-out LLaVA-Instruct loss (500 unseen conversations, lower is better) | 2.2521 | 1.1271 | **1.1602** |
151
 
152
+ **POPE** (object hallucination: yes/no questions about COCO val2014 images, scored from the Yes/No token
153
+ probabilities; the benchmark is balanced 50% yes / 50% no, so a yes-ratio near 0.5 is ideal):
154
 
155
+ | POPE split | Stage 2: accuracy / F1 / yes-ratio | Stage 2b: accuracy / F1 / yes-ratio |
156
+ |---|---|---|
157
+ | random | 0.543 / 0.686 / 0.95 | **0.822** / **0.832** / **0.56** |
158
+ | popular | 0.522 / 0.676 / 0.98 | **0.776** / **0.797** / **0.60** |
159
+ | adversarial | 0.514 / 0.672 / 0.98 | **0.752** / **0.780** / **0.63** |
160
+
161
+ #### What stage 2b fixed
162
+
163
+ Stage 2 answered **"yes" to 97–98%** of POPE questions. It confirmed almost every object it was asked about,
164
+ present or not, so accuracy sat at chance (51–54%) even though its descriptions were good. The cause is the
165
+ training data: LLaVA-Instruct-150K is made of long GPT-4-written conversations and contains almost no question
166
+ whose answer is "no".
167
+
168
+ Stage 2b continues stage 2 for one epoch on a 150k mix drawn from the public LLaVA-1.5 training set: 110k
169
+ short-answer VQA conversations (**57,804 "yes" vs 59,443 "no"** turns) plus 40k LLaVA-Instruct conversations
170
+ replayed so detailed description is not lost. Every POPE test image and every held-out conversation was removed
171
+ from this mix before training, so the numbers above are not contaminated; region/bounding-box tasks were dropped
172
+ as unsupported. Result: the yes-ratio fell to 0.56–0.63, accuracy rose by **24–28 points**, and the held-out
173
+ conversation loss improved as well (1.1271 → 1.1602),
174
+ so description quality did not regress. For reference, LLaVA-1.5-7B reports POPE F1 ≈ 0.86 from 665k samples at
175
+ 336 px; AtlasVision reaches 0.83 from 150k samples at 224 px.
176
+
177
+ Recall (0.88) still exceeds precision
178
+ (0.70) on the adversarial split, so a mild "yes" lean remains on the
179
+ hardest confusable pairs.
180
 
181
  ### Examples — stage 1 captions on unseen images
182
 
 
189
  | the beach in el nido national park, puerto puerto | the limestone cliffs and limestone islands in the background |
190
  | three pieces of paper with the words, democratic decentified dp controlled centralized ccp | a diagram showing the different types of democracy |
191
 
192
+ ### Examples — "Describe this image in detail." (stage 2b, unseen COCO val2014)
193
 
194
  **COCO_val2014_000000310196**
195
 
196
  - *Stage 1:* a skier in the snow on a mountain
197
+ - *Stage 2b:* The image features a snow-covered slope with a skier in the center, wearing red and black clothing. The skier is skillfully navigating down the hill, making their way through the snow.
198
 
199
+ In addition to the main skier, there are several other people scattered across the scene, some of them closer to the top of the slope while others are further away or near the bottom. These individuals may also be skiing or enjoying the snowy landscape.
200
 
201
  **COCO_val2014_000000210789**
202
 
203
  - *Stage 1:* a woman and her child in the rain
204
+ - *Stage 2b:* The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The little girl is wearing a pink dress, while the woman appears to be her mother. They are positioned close together, with the child on the left side of the frame and the woman on the right.
205
 
206
+ In addition to the two main subjects, there are several other people visible in the background, some of whom may also be holding umbrellas or trying to stay dry during the rainy weather. A car can be seen parked further back in the scene, likely waiting for its occupants
207
 
208
  **COCO_val2014_000000429109**
209
 
210
  - *Stage 1:* a bus and several other vehicles parked in front of a building
211
+ - *Stage 2b:* The image features a busy street with several buses and cars parked or driving along the road. There are three buses in total, one on the left side of the scene, another near the center, and the third bus further to the right. A car is also visible on the left side of the image.
 
 
 
 
 
 
 
 
 
212
 
213
+ In addition to the vehicles, there are multiple people walking around the area. Some individuals can be seen closer to the buses, while others are scattered throughout the scene. The presence of both vehicles and pedestrians suggests that this could be a popular transportation hub or a bustling city street.
214
 
215
+ ### Questions in Nigerian languages (stage 2b)
216
 
217
+ Stage 2b answers Igbo and Yoruba questions **in those languages** without any translation step (stage 2 always replied in English). Hausa still falls back to English, so the `lang=` cascade below remains the reliable route for all three.
218
 
219
  | Language | Question | Answer |
220
  |---|---|---|
221
+ | Igbo | Kedu ihe dị na foto a? | Na foto a, e nwere skier na snow. |
222
+ | Yoruba | Kí ni ó wà nínú àwòrán yìí? | Nínú àwòrán yìí, ó wà skier tí ń ski lórí òkè. |
223
+ | Hausa | Me ke cikin wannan hoton? | In the image, there is a skier wearing red and black clothing who is skiing down a snow-covered slope. |
224
 
225
  ### Text-only check (no image)
226
 
 
235
 
236
  ## Limitations
237
 
238
+ - **Hallucination.** Long descriptions are fluent but often add plausible details that are not in the image
239
+ (exact counts, extra people, a bicycle at the edge of frame). Treat counts and small objects as unreliable.
240
+ - **Residual yes-bias.** Stage 2b largely fixed stage 2's yes-bias, but recall still exceeds precision on POPE's
241
+ adversarial split, so yes/no answers about easily-confused objects lean positive. Stage 2 weights are kept for
242
+ reproducibility only — do not use them for yes/no verification.
243
  - **English-only visual training.** All image–text training data is English. Answers to Hausa, Igbo and Yoruba questions
244
  come from N-ATLaS's own multilingual ability and are noticeably less reliable; they may switch to English.
245
  - **Low resolution.** Images are resized to 224×224, so small text, fine details and dense documents are hard.
code/prepare_mix.py ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Build the de-biasing mix for stage 2b from LLaVA-1.5's public training mix (COCO-image parts only).
3
+
4
+ Keeps VQAv2 / OK-VQA / A-OKVQA style short-answer conversations (many "no" answers) plus a replay of
5
+ LLaVA-Instruct conversations, drops region/bounding-box tasks, and removes every POPE test image and the
6
+ 1,000 held-out stage-2 conversations so evaluation stays clean."""
7
+ import collections
8
+ import glob
9
+ import json
10
+ import os
11
+ import random
12
+ import re
13
+
14
+ import pyarrow.parquet as pq
15
+ import torch
16
+ from huggingface_hub import hf_hub_download
17
+
18
+ H = os.path.expanduser("~")
19
+ D = f"{H}/data/llava_instruct"
20
+ N_VQA = int(os.environ.get("N_VQA", 110_000))
21
+ N_LLAVA = int(os.environ.get("N_LLAVA", 40_000))
22
+ OUT = os.environ.get("MIX_OUT", f"{D}/mix_debias.json")
23
+
24
+ mix_path = os.environ.get("MIX_PATH") or hf_hub_download("liuhaotian/LLaVA-Instruct-150K", "llava_v1_5_mix665k.json",
25
+ repo_type="dataset", local_dir=D)
26
+ mix = json.load(open(mix_path))
27
+ print(f"mix665k: {len(mix):,} entries")
28
+
29
+ inst = json.load(open(f"{D}/llava_instruct_150k.json")) # same split as train_stage2.py (seed 42, last 1000)
30
+ perm = torch.randperm(len(inst), generator=torch.Generator().manual_seed(42)).tolist()
31
+ held_imgs = {inst[i]["image"] for i in perm[-int(os.environ.get("HELDOUT_N", 1000)):]}
32
+
33
+ pope_ids = set()
34
+ for f in glob.glob(f"{H}/data/pope/**/*.parquet", recursive=True):
35
+ for src in pq.read_table(f, columns=["image_source"]).column("image_source").to_pylist():
36
+ m = re.search(r"(\d+)$", str(src))
37
+ if m:
38
+ pope_ids.add(int(m.group(1)))
39
+ print(f"POPE images to exclude: {len(pope_ids):,} | held-out stage-2 images to exclude: {len(held_imgs):,}")
40
+
41
+ BBOX = re.compile(r"\[\s*\d?\.\d+\s*,\s*\d?\.\d+")
42
+ SHORT = ("single word or phrase", "option's letter", "Answer the question using a single word")
43
+ stats = collections.Counter()
44
+ vqa, llava = [], []
45
+ for e in mix:
46
+ img = e.get("image") or ""
47
+ if not img.startswith("coco/train2017/"):
48
+ stats["skip_not_coco"] += 1
49
+ continue
50
+ base = img.rsplit("/", 1)[-1]
51
+ if int(base.split(".")[0]) in pope_ids:
52
+ stats["skip_pope_image"] += 1
53
+ continue
54
+ if base in held_imgs:
55
+ stats["skip_heldout_image"] += 1
56
+ continue
57
+ text = " ".join(t["value"] for t in e["conversations"])
58
+ if BBOX.search(text):
59
+ stats["skip_region_task"] += 1
60
+ continue
61
+ item = {"id": str(e.get("id", base)), "image": base, "conversations": e["conversations"]}
62
+ (vqa if any(s in text for s in SHORT) else llava).append(item)
63
+
64
+ rng = random.Random(42)
65
+ rng.shuffle(vqa)
66
+ rng.shuffle(llava)
67
+ out = vqa[:N_VQA] + llava[:N_LLAVA]
68
+ rng.shuffle(out)
69
+ json.dump(out, open(OUT, "w"))
70
+
71
+ answers = collections.Counter()
72
+ for e in vqa[:N_VQA]:
73
+ for t in e["conversations"]:
74
+ if t["from"] == "gpt":
75
+ a = t["value"].strip().lower().rstrip(".")
76
+ answers["yes" if a == "yes" else "no" if a == "no" else "other"] += 1
77
+ print("filter stats:", dict(stats))
78
+ print(f"available: {len(vqa):,} short-answer + {len(llava):,} instruct | using {min(N_VQA, len(vqa)):,} + {min(N_LLAVA, len(llava)):,} = {len(out):,}")
79
+ print(f"short-answer turns: yes {answers['yes']:,} | no {answers['no']:,} | other {answers['other']:,}")
80
+ if len(out) < 1000 and not os.environ.get("ALLOW_SMALL"):
81
+ raise SystemExit(f"only {len(out)} examples - something is wrong with the filters")
82
+ print(f"wrote {OUT}")
code/run_stage2b.sh ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+ # Stage 2b: de-bias stage 2 with short-answer VQA data (continues from stage-2 LoRA + projector), then evaluate.
3
+ set -uo pipefail
4
+ cd "$HOME/atlas"
5
+ export PYTHONUNBUFFERED=1 HF_HOME="$HOME/.cache/huggingface" PATH="$HOME/.local/bin:$PATH"
6
+ echo "=== [1/4] build de-bias mix ==="
7
+ python3 prepare_mix.py || exit 11
8
+ export HF_HUB_OFFLINE=1
9
+ TRAIN_ENV="INSTRUCT_JSON=$HOME/data/llava_instruct/mix_debias.json INIT_FROM=$HOME/checkpoints/stage2/latest.pt LORA_LR=${LORA_LR:-1e-4} PROJ_LR=${PROJ_LR:-1e-5} HELDOUT=200"
10
+ echo "=== [2/4] smoke test (10 steps) ==="
11
+ rm -rf "$HOME/checkpoints/stage2b_smoke"
12
+ env $TRAIN_ENV MAX_STEPS=10 LOG_EVERY=2 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/stage2b_smoke" python3 train_stage2.py || { echo "SMOKE TEST FAILED"; exit 12; }
13
+ echo "=== [3/4] full stage-2b run ==="
14
+ env $TRAIN_ENV CKPT_DIR="$HOME/checkpoints/stage2b" python3 train_stage2.py || { echo "TRAINING FAILED"; exit 13; }
15
+ echo "=== [4/4] evaluation (same held-out set and POPE as stage 2) ==="
16
+ env -u INSTRUCT_JSON -u INIT_FROM -u HELDOUT CKPT_DIR="$HOME/checkpoints/stage2b" EVAL_DIR="$HOME/eval/stage2b" python3 eval_stage2.py || { echo "EVAL FAILED"; exit 14; }
17
+ cat "$HOME/eval/stage2b/report.md"
18
+ echo "=== run_stage2b.sh finished ==="
code/train_stage2.py CHANGED
@@ -172,9 +172,15 @@ def main():
172
  f"max_text_len={MAX_TEXT_LEN} workers={NUM_WORKERS} max_steps={MAX_STEPS or 'full'}")
173
 
174
  ck = torch.load(ckpt_path, map_location="cpu") if os.path.exists(ckpt_path) else None
 
 
 
 
 
175
  model, tok, pad_id, processor = build_stage2_model(
176
- lora_state=ck["lora_state_dict"] if ck else None,
177
- projector_state=ck["projector_state_dict"] if ck else None)
 
178
  lora_params = [p for n, p in model.llm.named_parameters() if p.requires_grad]
179
  proj_params = list(model.projector.parameters())
180
  log(f"trainable: LoRA {sum(p.numel() for p in lora_params):,} + projector {sum(p.numel() for p in proj_params):,}")
 
172
  f"max_text_len={MAX_TEXT_LEN} workers={NUM_WORKERS} max_steps={MAX_STEPS or 'full'}")
173
 
174
  ck = torch.load(ckpt_path, map_location="cpu") if os.path.exists(ckpt_path) else None
175
+ init = None
176
+ if ck is None and os.environ.get("INIT_FROM"): # continue from an earlier stage-2 run, fresh optimizer
177
+ init = torch.load(os.environ["INIT_FROM"], map_location="cpu")
178
+ log(f"initialising LoRA + projector from {os.environ['INIT_FROM']}")
179
+ src = ck or init
180
  model, tok, pad_id, processor = build_stage2_model(
181
+ lora_state=src["lora_state_dict"] if src else None,
182
+ projector_state=src["projector_state_dict"] if src else None)
183
+ del init
184
  lora_params = [p for n, p in model.llm.named_parameters() if p.requires_grad]
185
  proj_params = list(model.projector.parameters())
186
  log(f"trainable: LoRA {sum(p.numel() for p in lora_params):,} + projector {sum(p.numel() for p in proj_params):,}")
eval/stage2b/report.md ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Atlas-Vision evaluation
2
+
3
+ ## Scores (stage 1 → stage 2)
4
+
5
+ | Metric | Stage 1 | Stage 2 |
6
+ |---|---|---|
7
+ | Held-out instruct loss (lower is better) | 2.2521 | 1.1602 |
8
+ | POPE random: accuracy / F1 / yes-ratio | 0.4997 / 0.2003 / 0.1257 | 0.822 / 0.8318 / 0.558 |
9
+ | POPE popular: accuracy / F1 / yes-ratio | 0.5027 / 0.2013 / 0.1227 | 0.7757 / 0.7969 / 0.6043 |
10
+ | POPE adversarial: accuracy / F1 / yes-ratio | 0.5023 / 0.2012 / 0.123 | 0.752 / 0.78 / 0.6273 |
11
+
12
+ ## Detailed descriptions
13
+
14
+ **COCO_val2014_000000310196**
15
+
16
+ - Stage 1: a skier in the snow on a mountain
17
+ - Stage 2: The image features a snow-covered slope with a skier in the center, wearing red and black clothing. The skier is skillfully navigating down the hill, making their way through the snow.
18
+
19
+ In addition to the main skier, there are several other people scattered across the scene, some of them closer to the top of the slope while others are further away or near the bottom. These individuals may also be skiing or enjoying the snowy landscape.
20
+
21
+ **COCO_val2014_000000210789**
22
+
23
+ - Stage 1: a woman and her child in the rain
24
+ - Stage 2: The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The little girl is wearing a pink dress, while the woman appears to be her mother. They are positioned close together, with the child on the left side of the frame and the woman on the right.
25
+
26
+ In addition to the two main subjects, there are several other people visible in the background, some of whom may also be holding umbrellas or trying to stay dry during the rainy weather. A car can be seen parked further back in the scene, likely waiting for its occupants
27
+
28
+ **COCO_val2014_000000429109**
29
+
30
+ - Stage 1: a bus and several other vehicles parked in front of a building
31
+ - Stage 2: The image features a busy street with several buses and cars parked or driving along the road. There are three buses in total, one on the left side of the scene, another near the center, and the third bus further to the right. A car is also visible on the left side of the image.
32
+
33
+ In addition to the vehicles, there are multiple people walking around the area. Some individuals can be seen closer to the buses, while others are scattered throughout the scene. The presence of both vehicles and pedestrians suggests that this could be a popular transportation hub or a bustling city street.
34
+
35
+ **COCO_val2014_000000211674**
36
+
37
+ - Stage 1: a bus with a red and white logo on it, carrying passengers
38
+ - Stage 2: The image features a red double-decker bus with an advertisement for "Ruta Arzola" on its side. The bus is driving down the street, and there are several people visible in various positions around it. Some of these individuals can be seen standing near the front of the bus, while others are scattered along the sides or further back.
39
+
40
+ In addition to the bus, there are two cars present in the scene – one located behind the bus and another positioned slightly ahead of it. A bicycle can also be spotted close to the right edge of the image.
41
+
42
+ ## Questions in Nigerian languages (stage 2)
43
+
44
+ - **igbo** — Kedu ihe dị na foto a?
45
+ → Na foto a, e nwere skier na snow.
46
+ - **yoruba** — Kí ni ó wà nínú àwòrán yìí?
47
+ → Nínú àwòrán yìí, ó wà skier tí ń ski lórí òkè.
48
+ - **hausa** — Me ke cikin wannan hoton?
49
+ → In the image, there is a skier wearing red and black clothing who is skiing down a snow-covered slope.
50
+
51
+ ## Text-only check (no image)
52
+
53
+ Question: Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.
54
+
55
+ - Base N-ATLaS: Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
56
+ - With stage-2 LoRA: Positron bụ ụdị subatomic dị ka electron mana ọ nwere nnukwu ọdịiche - ọ nwere ụgwọ dị mma, ebe electron nwere ụgwọ na-adịghị mma. A na-akpọkwa positron antiparticle nke electron, n'ihi na ọ na-emegide njirimara niile nke electron ma nwee ibu yiri ya. Nchọpụta nke positrons mere site n'aka physicist Paul Dirac
eval/stage2b/stage2b_eval.json ADDED
@@ -0,0 +1,104 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "stage1": {
3
+ "heldout_instruct_loss": 2.2521,
4
+ "pope_random": {
5
+ "n": 3000,
6
+ "accuracy": 0.4997,
7
+ "precision": 0.4987,
8
+ "recall": 0.1253,
9
+ "f1": 0.2003,
10
+ "yes_ratio": 0.1257
11
+ },
12
+ "pope_popular": {
13
+ "n": 3000,
14
+ "accuracy": 0.5027,
15
+ "precision": 0.5109,
16
+ "recall": 0.1253,
17
+ "f1": 0.2013,
18
+ "yes_ratio": 0.1227
19
+ },
20
+ "pope_adversarial": {
21
+ "n": 3000,
22
+ "accuracy": 0.5023,
23
+ "precision": 0.5095,
24
+ "recall": 0.1253,
25
+ "f1": 0.2012,
26
+ "yes_ratio": 0.123
27
+ }
28
+ },
29
+ "stage2": {
30
+ "heldout_instruct_loss": 1.1602,
31
+ "pope_random": {
32
+ "n": 3000,
33
+ "accuracy": 0.822,
34
+ "precision": 0.7885,
35
+ "recall": 0.88,
36
+ "f1": 0.8318,
37
+ "yes_ratio": 0.558
38
+ },
39
+ "pope_popular": {
40
+ "n": 3000,
41
+ "accuracy": 0.7757,
42
+ "precision": 0.7281,
43
+ "recall": 0.88,
44
+ "f1": 0.7969,
45
+ "yes_ratio": 0.6043
46
+ },
47
+ "pope_adversarial": {
48
+ "n": 3000,
49
+ "accuracy": 0.752,
50
+ "precision": 0.7009,
51
+ "recall": 0.8793,
52
+ "f1": 0.78,
53
+ "yes_ratio": 0.6273
54
+ }
55
+ },
56
+ "descriptions": [
57
+ {
58
+ "image": "COCO_val2014_000000310196",
59
+ "stage1": "a skier in the snow on a mountain",
60
+ "stage2": "The image features a snow-covered slope with a skier in the center, wearing red and black clothing. The skier is skillfully navigating down the hill, making their way through the snow. \n\nIn addition to the main skier, there are several other people scattered across the scene, some of them closer to the top of the slope while others are further away or near the bottom. These individuals may also be skiing or enjoying the snowy landscape."
61
+ },
62
+ {
63
+ "image": "COCO_val2014_000000210789",
64
+ "stage1": "a woman and her child in the rain",
65
+ "stage2": "The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The little girl is wearing a pink dress, while the woman appears to be her mother. They are positioned close together, with the child on the left side of the frame and the woman on the right.\n\nIn addition to the two main subjects, there are several other people visible in the background, some of whom may also be holding umbrellas or trying to stay dry during the rainy weather. A car can be seen parked further back in the scene, likely waiting for its occupants"
66
+ },
67
+ {
68
+ "image": "COCO_val2014_000000429109",
69
+ "stage1": "a bus and several other vehicles parked in front of a building",
70
+ "stage2": "The image features a busy street with several buses and cars parked or driving along the road. There are three buses in total, one on the left side of the scene, another near the center, and the third bus further to the right. A car is also visible on the left side of the image.\n\nIn addition to the vehicles, there are multiple people walking around the area. Some individuals can be seen closer to the buses, while others are scattered throughout the scene. The presence of both vehicles and pedestrians suggests that this could be a popular transportation hub or a bustling city street."
71
+ },
72
+ {
73
+ "image": "COCO_val2014_000000211674",
74
+ "stage1": "a bus with a red and white logo on it, carrying passengers",
75
+ "stage2": "The image features a red double-decker bus with an advertisement for \"Ruta Arzola\" on its side. The bus is driving down the street, and there are several people visible in various positions around it. Some of these individuals can be seen standing near the front of the bus, while others are scattered along the sides or further back.\n\nIn addition to the bus, there are two cars present in the scene – one located behind the bus and another positioned slightly ahead of it. A bicycle can also be spotted close to the right edge of the image."
76
+ }
77
+ ],
78
+ "multilingual": [
79
+ {
80
+ "image": "COCO_val2014_000000310196",
81
+ "language": "igbo",
82
+ "question": "Kedu ihe dị na foto a?",
83
+ "stage2": "Na foto a, e nwere skier na snow."
84
+ },
85
+ {
86
+ "image": "COCO_val2014_000000310196",
87
+ "language": "yoruba",
88
+ "question": "Kí ni ó wà nínú àwòrán yìí?",
89
+ "stage2": "Nínú àwòrán yìí, ó wà skier tí ń ski lórí òkè."
90
+ },
91
+ {
92
+ "image": "COCO_val2014_000000310196",
93
+ "language": "hausa",
94
+ "question": "Me ke cikin wannan hoton?",
95
+ "stage2": "In the image, there is a skier wearing red and black clothing who is skiing down a snow-covered slope."
96
+ }
97
+ ],
98
+ "text_only": {
99
+ "question": "Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.",
100
+ "base_n_atlas": "Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us",
101
+ "with_stage2_lora": "Positron bụ ụdị subatomic dị ka electron mana ọ nwere nnukwu ọdịiche - ọ nwere ụgwọ dị mma, ebe electron nwere ụgwọ na-adịghị mma. A na-akpọkwa positron antiparticle nke electron, n'ihi na ọ na-emegide njirimara niile nke electron ma nwee ibu yiri ya. Nchọpụta nke positrons mere site n'aka physicist Paul Dirac"
102
+ },
103
+ "eval_minutes": 5.4
104
+ }
logs/stage2b_train_log.jsonl ADDED
@@ -0,0 +1,189 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"step": 1, "loss": 3.14205402135849, "grad_norm": 22.113000869750977, "lr": 7.142857142857143e-07, "s_per_step": 6.455075025558472, "peak_gb": 26.618435072, "samples_seen": 32, "time": 1791165515.6419272}
2
+ {"step": 25, "loss": 1.588310784039398, "grad_norm": 3.3672797679901123, "lr": 1.785714285714286e-05, "s_per_step": 2.4963032404581704, "peak_gb": 37.31657216, "samples_seen": 800, "time": 1791165575.553675}
3
+ {"step": 50, "loss": 0.7974538557231426, "grad_norm": 2.1865363121032715, "lr": 3.571428571428572e-05, "s_per_step": 2.3631107234954833, "peak_gb": 37.31657216, "samples_seen": 1600, "time": 1791165634.6318743}
4
+ {"step": 75, "loss": 0.8159684363007546, "grad_norm": 1.9037853479385376, "lr": 5.3571428571428575e-05, "s_per_step": 2.239762897491455, "peak_gb": 37.31657216, "samples_seen": 2400, "time": 1791165690.6263816}
5
+ {"step": 100, "loss": 0.7896525931358337, "grad_norm": 2.5986344814300537, "lr": 7.142857142857143e-05, "s_per_step": 2.3967845153808596, "peak_gb": 37.31657216, "samples_seen": 3200, "time": 1791165750.546467}
6
+ {"step": 125, "loss": 0.8200986498594284, "grad_norm": 1.9665552377700806, "lr": 8.92857142857143e-05, "s_per_step": 2.33407000541687, "peak_gb": 37.31657216, "samples_seen": 4000, "time": 1791165808.8986936}
7
+ {"step": 150, "loss": 0.8343269512057304, "grad_norm": 3.3930838108062744, "lr": 9.999903078446618e-05, "s_per_step": 2.1298329830169678, "peak_gb": 37.31657216, "samples_seen": 4800, "time": 1791165862.14493}
8
+ {"step": 175, "loss": 0.8642021375894546, "grad_norm": 2.676706552505493, "lr": 9.99861683318768e-05, "s_per_step": 2.233229761123657, "peak_gb": 37.31657216, "samples_seen": 5600, "time": 1791165917.9761162}
9
+ {"step": 200, "loss": 0.8286110037565231, "grad_norm": 2.648852586746216, "lr": 9.995835331149929e-05, "s_per_step": 2.396289863586426, "peak_gb": 37.31657216, "samples_seen": 6400, "time": 1791165977.8838441}
10
+ {"step": 225, "loss": 0.9085636857151985, "grad_norm": 2.160125970840454, "lr": 9.99155940437549e-05, "s_per_step": 2.1998683071136473, "peak_gb": 37.31657216, "samples_seen": 7200, "time": 1791166032.881024}
11
+ {"step": 250, "loss": 0.8117574036121369, "grad_norm": 2.7077982425689697, "lr": 9.9857903319399e-05, "s_per_step": 2.2049441051483156, "peak_gb": 37.31657216, "samples_seen": 8000, "time": 1791166088.0050786}
12
+ {"step": 275, "loss": 0.7880503736436367, "grad_norm": 1.907461166381836, "lr": 9.978529839569481e-05, "s_per_step": 2.2465640830993654, "peak_gb": 37.31657216, "samples_seen": 8800, "time": 1791166144.169653}
13
+ {"step": 300, "loss": 0.8249761319160461, "grad_norm": 2.7506635189056396, "lr": 9.969780099125133e-05, "s_per_step": 2.1263711261749267, "peak_gb": 37.31657216, "samples_seen": 9600, "time": 1791166197.3293953}
14
+ {"step": 325, "loss": 0.8113455653190613, "grad_norm": 1.875877022743225, "lr": 9.959543727952643e-05, "s_per_step": 2.335303964614868, "peak_gb": 37.31657216, "samples_seen": 10400, "time": 1791166255.7124696}
15
+ {"step": 350, "loss": 0.82724973320961, "grad_norm": 1.6174412965774536, "lr": 9.947823788099753e-05, "s_per_step": 2.136453084945679, "peak_gb": 37.31657216, "samples_seen": 11200, "time": 1791166309.1243026}
16
+ {"step": 375, "loss": 0.7504141560196876, "grad_norm": 2.193037271499634, "lr": 9.934623785400195e-05, "s_per_step": 2.084479150772095, "peak_gb": 37.31657216, "samples_seen": 12000, "time": 1791166361.2366986}
17
+ {"step": 400, "loss": 0.824563305824995, "grad_norm": 1.1731573343276978, "lr": 9.919947668424977e-05, "s_per_step": 2.275921335220337, "peak_gb": 37.31657216, "samples_seen": 12800, "time": 1791166418.1351547}
18
+ {"step": 425, "loss": 0.8523908746242523, "grad_norm": 2.243957996368408, "lr": 9.903799827301237e-05, "s_per_step": 2.191669044494629, "peak_gb": 37.31657216, "samples_seen": 13600, "time": 1791166472.9273403}
19
+ {"step": 450, "loss": 0.7732072618603706, "grad_norm": 2.2687370777130127, "lr": 9.886185092398996e-05, "s_per_step": 2.1319501781463623, "peak_gb": 37.31657216, "samples_seen": 14400, "time": 1791166526.2265835}
20
+ {"step": 475, "loss": 0.8263177201151848, "grad_norm": 2.527397632598877, "lr": 9.867108732886235e-05, "s_per_step": 2.2787140560150148, "peak_gb": 37.321700352, "samples_seen": 15200, "time": 1791166583.1949017}
21
+ {"step": 500, "loss": 0.756193850338459, "grad_norm": 2.1782448291778564, "lr": 9.846576455152708e-05, "s_per_step": 2.24827579498291, "peak_gb": 37.321700352, "samples_seen": 16000, "time": 1791166639.4022188}
22
+ {"step": 525, "loss": 0.8762797820568085, "grad_norm": 1.765979528427124, "lr": 9.824594401102962e-05, "s_per_step": 2.4363709831237794, "peak_gb": 37.321700352, "samples_seen": 16800, "time": 1791166700.3121493}
23
+ {"step": 550, "loss": 0.8164837975800038, "grad_norm": 2.792057752609253, "lr": 9.801169146319091e-05, "s_per_step": 2.295490322113037, "peak_gb": 37.325714432, "samples_seen": 17600, "time": 1791166757.6998796}
24
+ {"step": 575, "loss": 0.7905827209353447, "grad_norm": 2.4493868350982666, "lr": 9.776307698093747e-05, "s_per_step": 2.152802686691284, "peak_gb": 37.325714432, "samples_seen": 18400, "time": 1791166811.5203996}
25
+ {"step": 600, "loss": 0.8037899886071682, "grad_norm": 1.6931008100509644, "lr": 9.750017493334023e-05, "s_per_step": 2.2320249462127686, "peak_gb": 37.325714432, "samples_seen": 19200, "time": 1791166867.3215806}
26
+ {"step": 625, "loss": 0.7798363075405359, "grad_norm": 1.994132399559021, "lr": 9.722306396336825e-05, "s_per_step": 2.337219753265381, "peak_gb": 37.325714432, "samples_seen": 20000, "time": 1791166925.7525764}
27
+ {"step": 650, "loss": 0.7738844112306833, "grad_norm": 2.684321880340576, "lr": 9.693182696436385e-05, "s_per_step": 2.181212863922119, "peak_gb": 37.325714432, "samples_seen": 20800, "time": 1791166980.2833657}
28
+ {"step": 675, "loss": 0.8178048168122768, "grad_norm": 2.06518816947937, "lr": 9.662655105524643e-05, "s_per_step": 2.1410923194885254, "peak_gb": 37.325714432, "samples_seen": 21600, "time": 1791167033.8111975}
29
+ {"step": 700, "loss": 0.8055640901625156, "grad_norm": 1.5254127979278564, "lr": 9.630732755445221e-05, "s_per_step": 2.269847593307495, "peak_gb": 37.325714432, "samples_seen": 22400, "time": 1791167090.5579088}
30
+ {"step": 725, "loss": 0.8045347545295953, "grad_norm": 3.4738898277282715, "lr": 9.597425195261783e-05, "s_per_step": 2.2504861736297608, "peak_gb": 37.325714432, "samples_seen": 23200, "time": 1791167146.8205574}
31
+ {"step": 750, "loss": 0.817788716852665, "grad_norm": 1.5313327312469482, "lr": 9.562742388401568e-05, "s_per_step": 2.396076259613037, "peak_gb": 37.325714432, "samples_seen": 24000, "time": 1791167206.72293}
32
+ {"step": 775, "loss": 0.7898177513480187, "grad_norm": 2.4144198894500732, "lr": 9.526694709675015e-05, "s_per_step": 2.249171714782715, "peak_gb": 37.325714432, "samples_seen": 24800, "time": 1791167262.952708}
33
+ {"step": 800, "loss": 0.7381506878137588, "grad_norm": 1.8181664943695068, "lr": 9.489292942172278e-05, "s_per_step": 2.0362033462524414, "peak_gb": 37.325714432, "samples_seen": 25600, "time": 1791167313.858285}
34
+ {"step": 825, "loss": 0.7578190127015114, "grad_norm": 2.2866311073303223, "lr": 9.450548274037653e-05, "s_per_step": 1.91242262840271, "peak_gb": 37.325714432, "samples_seen": 26400, "time": 1791167361.6693754}
35
+ {"step": 850, "loss": 0.817485048621893, "grad_norm": 1.920059323310852, "lr": 9.41047229512281e-05, "s_per_step": 2.3170962715148926, "peak_gb": 37.325714432, "samples_seen": 27200, "time": 1791167419.59729}
36
+ {"step": 875, "loss": 0.7546365970373153, "grad_norm": 1.885245442390442, "lr": 9.369076993519888e-05, "s_per_step": 2.1289653301239015, "peak_gb": 37.325714432, "samples_seen": 28000, "time": 1791167472.821967}
37
+ {"step": 900, "loss": 0.7460251280665398, "grad_norm": 2.3024418354034424, "lr": 9.326374751975431e-05, "s_per_step": 2.1195631408691407, "peak_gb": 37.325714432, "samples_seen": 28800, "time": 1791167525.8115594}
38
+ {"step": 925, "loss": 0.8006729355454445, "grad_norm": 3.084848642349243, "lr": 9.2823783441863e-05, "s_per_step": 2.1862466812133787, "peak_gb": 37.325714432, "samples_seen": 29600, "time": 1791167580.468191}
39
+ {"step": 950, "loss": 0.7907543142139911, "grad_norm": 1.976046085357666, "lr": 9.237100930978614e-05, "s_per_step": 2.252327461242676, "peak_gb": 37.325714432, "samples_seen": 30400, "time": 1791167636.7768707}
40
+ {"step": 975, "loss": 0.7727618554607034, "grad_norm": 1.6060981750488281, "lr": 9.190556056370907e-05, "s_per_step": 2.1356379699707033, "peak_gb": 37.325714432, "samples_seen": 31200, "time": 1791167690.168348}
41
+ {"step": 1000, "loss": 0.7722805575281382, "grad_norm": 2.346635580062866, "lr": 9.142757643522645e-05, "s_per_step": 2.184465847015381, "peak_gb": 37.325714432, "samples_seen": 32000, "time": 1791167744.7805607}
42
+ {"step": 1025, "loss": 0.790591963082552, "grad_norm": 2.7288553714752197, "lr": 9.093719990569333e-05, "s_per_step": 2.320113706588745, "peak_gb": 37.325714432, "samples_seen": 32800, "time": 1791167802.7838862}
43
+ {"step": 1050, "loss": 0.8003257100284099, "grad_norm": 1.9871152639389038, "lr": 9.043457766345464e-05, "s_per_step": 2.3094374752044677, "peak_gb": 37.325714432, "samples_seen": 33600, "time": 1791167860.5202947}
44
+ {"step": 1075, "loss": 0.7964102986454964, "grad_norm": 1.7926714420318604, "lr": 8.991986005996554e-05, "s_per_step": 2.170173463821411, "peak_gb": 37.325714432, "samples_seen": 34400, "time": 1791167914.7752576}
45
+ {"step": 1100, "loss": 0.769957457780838, "grad_norm": 1.2739399671554565, "lr": 8.93932010648163e-05, "s_per_step": 2.0905527210235597, "peak_gb": 37.325714432, "samples_seen": 35200, "time": 1791167967.0396023}
46
+ {"step": 1125, "loss": 0.7743674084544182, "grad_norm": 1.7435359954833984, "lr": 8.885475821967478e-05, "s_per_step": 2.130764627456665, "peak_gb": 37.325714432, "samples_seen": 36000, "time": 1791168020.3092685}
47
+ {"step": 1150, "loss": 0.7376279538124799, "grad_norm": 2.089308261871338, "lr": 8.830469259116016e-05, "s_per_step": 2.06257776260376, "peak_gb": 37.325714432, "samples_seen": 36800, "time": 1791168071.8742328}
48
+ {"step": 1175, "loss": 0.7563618395477534, "grad_norm": 2.581925630569458, "lr": 8.774316872266266e-05, "s_per_step": 2.104277467727661, "peak_gb": 37.325714432, "samples_seen": 37600, "time": 1791168124.4817083}
49
+ {"step": 1200, "loss": 0.7164853295683861, "grad_norm": 1.9859364032745361, "lr": 8.717035458512277e-05, "s_per_step": 2.175741729736328, "peak_gb": 37.325714432, "samples_seen": 38400, "time": 1791168178.8757303}
50
+ {"step": 1225, "loss": 0.7512343572080136, "grad_norm": 1.4672244787216187, "lr": 8.658642152678558e-05, "s_per_step": 2.196865644454956, "peak_gb": 37.325714432, "samples_seen": 39200, "time": 1791168233.7978122}
51
+ {"step": 1250, "loss": 0.7883425006270408, "grad_norm": 0.9867988228797913, "lr": 8.599154422194459e-05, "s_per_step": 2.1413458442687987, "peak_gb": 37.325714432, "samples_seen": 40000, "time": 1791168287.3319008}
52
+ {"step": 1275, "loss": 0.7481986878812313, "grad_norm": 1.657426118850708, "lr": 8.53859006186907e-05, "s_per_step": 2.2622928619384766, "peak_gb": 37.325714432, "samples_seen": 40800, "time": 1791168343.8896875}
53
+ {"step": 1300, "loss": 0.7483823847025632, "grad_norm": 2.765939235687256, "lr": 8.476967188568188e-05, "s_per_step": 2.1760437870025635, "peak_gb": 37.325714432, "samples_seen": 41600, "time": 1791168398.291302}
54
+ {"step": 1325, "loss": 0.7789613246917725, "grad_norm": 1.4214611053466797, "lr": 8.414304235794941e-05, "s_per_step": 2.260581455230713, "peak_gb": 37.325714432, "samples_seen": 42400, "time": 1791168454.8063686}
55
+ {"step": 1350, "loss": 0.6945078286528588, "grad_norm": 0.8304738402366638, "lr": 8.3506199481757e-05, "s_per_step": 2.104763975143433, "peak_gb": 37.325714432, "samples_seen": 43200, "time": 1791168507.4259417}
56
+ {"step": 1375, "loss": 0.7696672566235065, "grad_norm": 1.0556106567382812, "lr": 8.285933375852925e-05, "s_per_step": 2.2171495151519776, "peak_gb": 37.325714432, "samples_seen": 44000, "time": 1791168562.8552532}
57
+ {"step": 1400, "loss": 0.7444813340902329, "grad_norm": 1.4957417249679565, "lr": 8.220263868786613e-05, "s_per_step": 2.1510196685791017, "peak_gb": 37.325714432, "samples_seen": 44800, "time": 1791168616.631196}
58
+ {"step": 1425, "loss": 0.7628245377540588, "grad_norm": 1.9097741842269897, "lr": 8.153631070966065e-05, "s_per_step": 2.2037828826904295, "peak_gb": 37.325714432, "samples_seen": 45600, "time": 1791168671.7262833}
59
+ {"step": 1450, "loss": 0.7726829506456852, "grad_norm": 4.09751558303833, "lr": 8.086054914533705e-05, "s_per_step": 2.278927917480469, "peak_gb": 37.325714432, "samples_seen": 46400, "time": 1791168728.7000833}
60
+ {"step": 1475, "loss": 0.7285840643197298, "grad_norm": 2.5535264015197754, "lr": 8.017555613822691e-05, "s_per_step": 2.2388983535766602, "peak_gb": 37.325714432, "samples_seen": 47200, "time": 1791168784.6731944}
61
+ {"step": 1500, "loss": 0.7646115978434682, "grad_norm": 1.643479347229004, "lr": 7.948153659310116e-05, "s_per_step": 2.2301154708862305, "peak_gb": 37.325714432, "samples_seen": 48000, "time": 1791168840.426597}
62
+ {"step": 1525, "loss": 0.8089976158738136, "grad_norm": 1.9452074766159058, "lr": 7.87786981148762e-05, "s_per_step": 2.326566572189331, "peak_gb": 37.325714432, "samples_seen": 48800, "time": 1791168898.5912657}
63
+ {"step": 1550, "loss": 0.7462999747693538, "grad_norm": 1.458425760269165, "lr": 7.806725094651197e-05, "s_per_step": 2.1147663688659666, "peak_gb": 37.325714432, "samples_seen": 49600, "time": 1791168951.4609642}
64
+ {"step": 1575, "loss": 0.7764386893808841, "grad_norm": 1.451231837272644, "lr": 7.734740790612136e-05, "s_per_step": 2.115509796142578, "peak_gb": 37.325714432, "samples_seen": 50400, "time": 1791169004.349144}
65
+ {"step": 1600, "loss": 0.7370309364795685, "grad_norm": 1.2360936403274536, "lr": 7.661938432330886e-05, "s_per_step": 2.2563931560516357, "peak_gb": 37.325714432, "samples_seen": 51200, "time": 1791169060.7593937}
66
+ {"step": 1625, "loss": 0.7613919779658318, "grad_norm": 1.3745328187942505, "lr": 7.588339797475826e-05, "s_per_step": 2.2605470848083495, "peak_gb": 37.325714432, "samples_seen": 52000, "time": 1791169117.27354}
67
+ {"step": 1650, "loss": 0.7241708636283875, "grad_norm": 1.1827912330627441, "lr": 7.513966901908811e-05, "s_per_step": 2.072138557434082, "peak_gb": 37.325714432, "samples_seen": 52800, "time": 1791169169.0774484}
68
+ {"step": 1675, "loss": 0.8018289273977279, "grad_norm": 1.4603354930877686, "lr": 7.438841993099486e-05, "s_per_step": 2.274630489349365, "peak_gb": 37.325714432, "samples_seen": 53600, "time": 1791169225.9437313}
69
+ {"step": 1700, "loss": 0.7257322013378144, "grad_norm": 1.382242202758789, "lr": 7.362987543470305e-05, "s_per_step": 2.2222759437561037, "peak_gb": 37.325714432, "samples_seen": 54400, "time": 1791169281.5011303}
70
+ {"step": 1725, "loss": 0.7535443095862866, "grad_norm": 1.816344141960144, "lr": 7.286426243674262e-05, "s_per_step": 2.216677312850952, "peak_gb": 37.325714432, "samples_seen": 55200, "time": 1791169336.9185019}
71
+ {"step": 1750, "loss": 0.6980357927083969, "grad_norm": 1.857561469078064, "lr": 7.209180995807341e-05, "s_per_step": 2.2018299293518067, "peak_gb": 37.325714432, "samples_seen": 56000, "time": 1791169391.9646761}
72
+ {"step": 1775, "loss": 0.7253469225764274, "grad_norm": 2.774467706680298, "lr": 7.131274906557725e-05, "s_per_step": 2.397851753234863, "peak_gb": 37.325714432, "samples_seen": 56800, "time": 1791169451.9115198}
73
+ {"step": 1800, "loss": 0.743451382741332, "grad_norm": 2.1614222526550293, "lr": 7.052731280293791e-05, "s_per_step": 2.1326423931121825, "peak_gb": 37.325714432, "samples_seen": 57600, "time": 1791169505.2281344}
74
+ {"step": 1825, "loss": 0.7323271352052688, "grad_norm": 1.8209247589111328, "lr": 6.973573612092985e-05, "s_per_step": 2.3023265171051026, "peak_gb": 37.328663552, "samples_seen": 58400, "time": 1791169562.786706}
75
+ {"step": 1850, "loss": 0.7592687951028347, "grad_norm": 1.0971518754959106, "lr": 6.893825580713633e-05, "s_per_step": 2.175807409286499, "peak_gb": 37.328663552, "samples_seen": 59200, "time": 1791169617.1823072}
76
+ {"step": 1875, "loss": 0.7327933664619922, "grad_norm": 1.8120750188827515, "lr": 6.813511041511828e-05, "s_per_step": 2.188905382156372, "peak_gb": 37.328663552, "samples_seen": 60000, "time": 1791169671.9053829}
77
+ {"step": 1900, "loss": 0.7617114392668008, "grad_norm": 1.5995821952819824, "lr": 6.732654019305469e-05, "s_per_step": 2.2518668460845945, "peak_gb": 37.328663552, "samples_seen": 60800, "time": 1791169728.2025328}
78
+ {"step": 1925, "loss": 0.7184512709081173, "grad_norm": 3.1466126441955566, "lr": 6.651278701187629e-05, "s_per_step": 2.1966624069213867, "peak_gb": 37.328663552, "samples_seen": 61600, "time": 1791169783.119614}
79
+ {"step": 1950, "loss": 0.7223995580524206, "grad_norm": 2.8622689247131348, "lr": 6.569409429291362e-05, "s_per_step": 2.0271696281433105, "peak_gb": 37.328663552, "samples_seen": 62400, "time": 1791169833.799278}
80
+ {"step": 1975, "loss": 0.7936395730078221, "grad_norm": 1.976616382598877, "lr": 6.487070693508147e-05, "s_per_step": 2.2634132957458495, "peak_gb": 37.328663552, "samples_seen": 63200, "time": 1791169890.385065}
81
+ {"step": 2000, "loss": 0.7677774439752102, "grad_norm": 1.5000005960464478, "lr": 6.404287124162117e-05, "s_per_step": 2.219316129684448, "peak_gb": 37.328663552, "samples_seen": 64000, "time": 1791169945.8683996}
82
+ {"step": 2025, "loss": 0.8423459057509899, "grad_norm": 2.1223344802856445, "lr": 6.321083484642303e-05, "s_per_step": 2.251362638473511, "peak_gb": 37.328663552, "samples_seen": 64800, "time": 1791170002.1529393}
83
+ {"step": 2050, "loss": 0.7075329820811749, "grad_norm": 2.1938531398773193, "lr": 6.237484663995047e-05, "s_per_step": 2.1452943515777587, "peak_gb": 37.328663552, "samples_seen": 65600, "time": 1791170055.7858298}
84
+ {"step": 2075, "loss": 0.8434611645340919, "grad_norm": 1.8508089780807495, "lr": 6.153515669478846e-05, "s_per_step": 2.182622308731079, "peak_gb": 37.328663552, "samples_seen": 66400, "time": 1791170110.3518124}
85
+ {"step": 2100, "loss": 0.7706048192456365, "grad_norm": 2.157609462738037, "lr": 6.069201619083826e-05, "s_per_step": 2.2786025047302245, "peak_gb": 37.328663552, "samples_seen": 67200, "time": 1791170167.3173573}
86
+ {"step": 2125, "loss": 0.7334290408343077, "grad_norm": 1.6534168720245361, "lr": 5.984567734018094e-05, "s_per_step": 2.185659351348877, "peak_gb": 37.328663552, "samples_seen": 68000, "time": 1791170221.9593232}
87
+ {"step": 2150, "loss": 0.763808610253036, "grad_norm": 2.686551809310913, "lr": 5.899639331163218e-05, "s_per_step": 2.268961925506592, "peak_gb": 37.328663552, "samples_seen": 68800, "time": 1791170278.6838627}
88
+ {"step": 2175, "loss": 0.74334836833179, "grad_norm": 1.4271022081375122, "lr": 5.814441815501079e-05, "s_per_step": 2.15900297164917, "peak_gb": 37.328663552, "samples_seen": 69600, "time": 1791170332.6593788}
89
+ {"step": 2200, "loss": 0.7046704458445311, "grad_norm": 1.7289416790008545, "lr": 5.729000672514379e-05, "s_per_step": 2.078843479156494, "peak_gb": 37.328663552, "samples_seen": 70400, "time": 1791170384.6309428}
90
+ {"step": 2225, "loss": 0.7402220770716668, "grad_norm": 1.6976007223129272, "lr": 5.643341460563063e-05, "s_per_step": 2.128084936141968, "peak_gb": 37.328663552, "samples_seen": 71200, "time": 1791170437.8334522}
91
+ {"step": 2250, "loss": 0.6715993209183216, "grad_norm": 2.1607041358947754, "lr": 5.557489803238933e-05, "s_per_step": 2.159433355331421, "peak_gb": 37.328663552, "samples_seen": 72000, "time": 1791170491.8197389}
92
+ {"step": 2275, "loss": 0.6767104347795247, "grad_norm": 1.7123639583587646, "lr": 5.471471381700778e-05, "s_per_step": 2.1726886463165282, "peak_gb": 37.328663552, "samples_seen": 72800, "time": 1791170546.1373775}
93
+ {"step": 2300, "loss": 0.747688923329115, "grad_norm": 2.3197970390319824, "lr": 5.385311926992242e-05, "s_per_step": 2.1769580173492433, "peak_gb": 37.328663552, "samples_seen": 73600, "time": 1791170600.5618634}
94
+ {"step": 2325, "loss": 0.7415474114567041, "grad_norm": 2.262953996658325, "lr": 5.2990372123447995e-05, "s_per_step": 2.1750038051605225, "peak_gb": 37.328663552, "samples_seen": 74400, "time": 1791170654.9374063}
95
+ {"step": 2350, "loss": 0.7086691360920667, "grad_norm": 2.3506317138671875, "lr": 5.212673045468114e-05, "s_per_step": 2.1435218715667723, "peak_gb": 37.328663552, "samples_seen": 75200, "time": 1791170708.5259638}
96
+ {"step": 2375, "loss": 0.6722931536287069, "grad_norm": 2.5476269721984863, "lr": 5.1262452608300514e-05, "s_per_step": 2.0527558517456055, "peak_gb": 37.328663552, "samples_seen": 76000, "time": 1791170759.8453696}
97
+ {"step": 2400, "loss": 0.6826995080709457, "grad_norm": 2.1311166286468506, "lr": 5.0397797119287236e-05, "s_per_step": 2.0880999755859375, "peak_gb": 37.328663552, "samples_seen": 76800, "time": 1791170812.048361}
98
+ {"step": 2425, "loss": 0.7050965564697981, "grad_norm": 3.221079111099243, "lr": 4.9533022635588216e-05, "s_per_step": 2.119127044677734, "peak_gb": 37.328663552, "samples_seen": 77600, "time": 1791170865.0271327}
99
+ {"step": 2450, "loss": 0.7206440759450197, "grad_norm": 2.5294783115386963, "lr": 4.8668387840745734e-05, "s_per_step": 2.2606944370269777, "peak_gb": 37.328663552, "samples_seen": 78400, "time": 1791170921.5450323}
100
+ {"step": 2475, "loss": 0.6436616685986519, "grad_norm": 2.3800606727600098, "lr": 4.7804151376516375e-05, "s_per_step": 2.2651154708862307, "peak_gb": 37.328663552, "samples_seen": 79200, "time": 1791170978.1734085}
101
+ {"step": 2500, "loss": 0.7039403146505356, "grad_norm": 1.7854039669036865, "lr": 4.694057176550241e-05, "s_per_step": 2.0506904697418213, "peak_gb": 37.328663552, "samples_seen": 80000, "time": 1791171029.4411094}
102
+ {"step": 2525, "loss": 0.7302389190718531, "grad_norm": 2.1441829204559326, "lr": 4.607790733381898e-05, "s_per_step": 2.298251438140869, "peak_gb": 37.328663552, "samples_seen": 80800, "time": 1791171086.8978634}
103
+ {"step": 2550, "loss": 0.655513388440013, "grad_norm": 2.3894569873809814, "lr": 4.521641613381981e-05, "s_per_step": 2.033922777175903, "peak_gb": 37.328663552, "samples_seen": 81600, "time": 1791171137.7463996}
104
+ {"step": 2575, "loss": 0.7391077170521021, "grad_norm": 2.9977188110351562, "lr": 4.435635586690507e-05, "s_per_step": 2.149828996658325, "peak_gb": 37.328663552, "samples_seen": 82400, "time": 1791171191.4925926}
105
+ {"step": 2600, "loss": 0.7200800219923258, "grad_norm": 2.7793188095092773, "lr": 4.349798380643398e-05, "s_per_step": 2.1879476928710937, "peak_gb": 37.328663552, "samples_seen": 83200, "time": 1791171246.1917589}
106
+ {"step": 2625, "loss": 0.7436163403838872, "grad_norm": 1.5204609632492065, "lr": 4.264155672076572e-05, "s_per_step": 2.4061502170562745, "peak_gb": 37.328663552, "samples_seen": 84000, "time": 1791171306.346018}
107
+ {"step": 2650, "loss": 0.660215564146638, "grad_norm": 1.1763237714767456, "lr": 4.178733079645106e-05, "s_per_step": 2.2049616050720213, "peak_gb": 37.328663552, "samples_seen": 84800, "time": 1791171361.4704888}
108
+ {"step": 2675, "loss": 0.6927475852519274, "grad_norm": 2.5386247634887695, "lr": 4.093556156159842e-05, "s_per_step": 2.145371465682983, "peak_gb": 37.328663552, "samples_seen": 85600, "time": 1791171415.1052473}
109
+ {"step": 2700, "loss": 0.7659448549896478, "grad_norm": 1.650169014930725, "lr": 4.008650380943658e-05, "s_per_step": 2.305666389465332, "peak_gb": 37.328663552, "samples_seen": 86400, "time": 1791171472.7473748}
110
+ {"step": 2725, "loss": 0.6683748767152429, "grad_norm": 2.2317709922790527, "lr": 3.924041152209739e-05, "s_per_step": 2.1367229557037355, "peak_gb": 37.328663552, "samples_seen": 87200, "time": 1791171526.165955}
111
+ {"step": 2750, "loss": 0.6803565273061395, "grad_norm": 2.578368902206421, "lr": 3.8397537794641006e-05, "s_per_step": 2.1723069763183593, "peak_gb": 37.328663552, "samples_seen": 88000, "time": 1791171580.4741504}
112
+ {"step": 2775, "loss": 0.6919519574567675, "grad_norm": 2.4252359867095947, "lr": 3.755813475934657e-05, "s_per_step": 2.2460842609405516, "peak_gb": 37.328663552, "samples_seen": 88800, "time": 1791171636.626775}
113
+ {"step": 2800, "loss": 0.7549696940928697, "grad_norm": 1.3772821426391602, "lr": 3.672245351029079e-05, "s_per_step": 2.140784168243408, "peak_gb": 37.328663552, "samples_seen": 89600, "time": 1791171690.1468506}
114
+ {"step": 2825, "loss": 0.6725568605586887, "grad_norm": 2.269684314727783, "lr": 3.5890744028237224e-05, "s_per_step": 2.1861842250823975, "peak_gb": 37.328663552, "samples_seen": 90400, "time": 1791171744.8019009}
115
+ {"step": 2850, "loss": 0.7570831072703004, "grad_norm": 1.3243885040283203, "lr": 3.506325510585843e-05, "s_per_step": 2.1999145126342774, "peak_gb": 37.328663552, "samples_seen": 91200, "time": 1791171799.8001847}
116
+ {"step": 2875, "loss": 0.7394591322168708, "grad_norm": 0.9553370475769043, "lr": 3.424023427331364e-05, "s_per_step": 2.3712642765045167, "peak_gb": 37.328663552, "samples_seen": 92000, "time": 1791171859.0828495}
117
+ {"step": 2900, "loss": 0.6674072152376175, "grad_norm": 1.9110374450683594, "lr": 3.342192772420396e-05, "s_per_step": 1.943366403579712, "peak_gb": 37.328663552, "samples_seen": 92800, "time": 1791171907.6675425}
118
+ {"step": 2925, "loss": 0.728705649226904, "grad_norm": 1.81485116481781, "lr": 3.260858024192766e-05, "s_per_step": 2.3112503147125243, "peak_gb": 37.328663552, "samples_seen": 93600, "time": 1791171965.4493177}
119
+ {"step": 2950, "loss": 0.6727638635039329, "grad_norm": 1.9227899312973022, "lr": 3.180043512645685e-05, "s_per_step": 2.1781687355041504, "peak_gb": 37.328663552, "samples_seen": 94400, "time": 1791172019.9040134}
120
+ {"step": 2975, "loss": 0.7250781350582838, "grad_norm": 1.9956176280975342, "lr": 3.099773412155837e-05, "s_per_step": 2.1951618003845215, "peak_gb": 37.328663552, "samples_seen": 95200, "time": 1791172074.7835524}
121
+ {"step": 3000, "loss": 0.650225435718894, "grad_norm": 1.4397475719451904, "lr": 3.0200717342479927e-05, "s_per_step": 2.1171874523162844, "peak_gb": 37.328663552, "samples_seen": 96000, "time": 1791172127.713813}
122
+ {"step": 3025, "loss": 0.6574885688722134, "grad_norm": 2.6953608989715576, "lr": 2.9409623204123326e-05, "s_per_step": 2.1706494808197023, "peak_gb": 37.328663552, "samples_seen": 96800, "time": 1791172181.9805422}
123
+ {"step": 3050, "loss": 0.7310007732175291, "grad_norm": 2.1912403106689453, "lr": 2.8624688349726657e-05, "s_per_step": 2.2939053535461427, "peak_gb": 37.328663552, "samples_seen": 97600, "time": 1791172239.3286743}
124
+ {"step": 3075, "loss": 0.6456812293455004, "grad_norm": 1.7668237686157227, "lr": 2.7846147580076e-05, "s_per_step": 2.030617341995239, "peak_gb": 37.328663552, "samples_seen": 98400, "time": 1791172290.0945995}
125
+ {"step": 3100, "loss": 0.7004510406404734, "grad_norm": 0.9194343090057373, "lr": 2.7074233783268676e-05, "s_per_step": 2.0785311794281007, "peak_gb": 37.328663552, "samples_seen": 99200, "time": 1791172342.0582764}
126
+ {"step": 3125, "loss": 0.7067461168020963, "grad_norm": 0.9615018367767334, "lr": 2.630917786504836e-05, "s_per_step": 2.172322187423706, "peak_gb": 37.328663552, "samples_seen": 100000, "time": 1791172396.3667948}
127
+ {"step": 3150, "loss": 0.6149680116400122, "grad_norm": 1.5655421018600464, "lr": 2.5551208679733385e-05, "s_per_step": 2.068291358947754, "peak_gb": 37.328663552, "samples_seen": 100800, "time": 1791172448.0745811}
128
+ {"step": 3175, "loss": 0.7133472035825252, "grad_norm": 2.5095202922821045, "lr": 2.4800552961758572e-05, "s_per_step": 2.2229059314727784, "peak_gb": 37.328663552, "samples_seen": 101600, "time": 1791172503.6477566}
129
+ {"step": 3200, "loss": 0.7128471711371094, "grad_norm": 5.073495388031006, "lr": 2.4057435257851175e-05, "s_per_step": 2.1217748641967775, "peak_gb": 37.328663552, "samples_seen": 102400, "time": 1791172556.6926906}
130
+ {"step": 3225, "loss": 0.681353373080492, "grad_norm": 0.7074483633041382, "lr": 2.3322077859861412e-05, "s_per_step": 2.2286937618255616, "peak_gb": 37.328663552, "samples_seen": 103200, "time": 1791172612.4104753}
131
+ {"step": 3250, "loss": 0.6933992094546556, "grad_norm": 1.5983896255493164, "lr": 2.2594700738267265e-05, "s_per_step": 2.2323473739624022, "peak_gb": 37.328663552, "samples_seen": 104000, "time": 1791172668.2196336}
132
+ {"step": 3275, "loss": 0.6521770719438791, "grad_norm": 1.4988881349563599, "lr": 2.1875521476373913e-05, "s_per_step": 2.350031909942627, "peak_gb": 37.328663552, "samples_seen": 104800, "time": 1791172726.9709315}
133
+ {"step": 3300, "loss": 0.6772003842145204, "grad_norm": 1.3674873113632202, "lr": 2.11647552052271e-05, "s_per_step": 2.1167491817474366, "peak_gb": 37.328663552, "samples_seen": 105600, "time": 1791172779.890159}
134
+ {"step": 3325, "loss": 0.6974349667225033, "grad_norm": 1.8985122442245483, "lr": 2.046261453926006e-05, "s_per_step": 2.188663167953491, "peak_gb": 37.328663552, "samples_seen": 106400, "time": 1791172834.6072176}
135
+ {"step": 3350, "loss": 0.6722285900078714, "grad_norm": 3.2266275882720947, "lr": 1.9769309512693395e-05, "s_per_step": 2.0479675102233887, "peak_gb": 37.328663552, "samples_seen": 107200, "time": 1791172885.8069012}
136
+ {"step": 3375, "loss": 0.6281475232355297, "grad_norm": 3.8666343688964844, "lr": 1.9085047516706533e-05, "s_per_step": 2.2480383014678953, "peak_gb": 37.328663552, "samples_seen": 108000, "time": 1791172942.0083659}
137
+ {"step": 3400, "loss": 0.6186308555305005, "grad_norm": 1.4103232622146606, "lr": 1.8410033237400098e-05, "s_per_step": 2.2519297409057617, "peak_gb": 37.328663552, "samples_seen": 108800, "time": 1791172998.3070562}
138
+ {"step": 3425, "loss": 0.6535533627495169, "grad_norm": 1.9440644979476929, "lr": 1.7744468594567244e-05, "s_per_step": 2.151489324569702, "peak_gb": 37.328663552, "samples_seen": 109600, "time": 1791173052.0947604}
139
+ {"step": 3450, "loss": 0.6736921292543411, "grad_norm": 4.172308921813965, "lr": 1.7088552681292513e-05, "s_per_step": 2.1378032398223876, "peak_gb": 37.328663552, "samples_seen": 110400, "time": 1791173105.5402908}
140
+ {"step": 3475, "loss": 0.7076615856215358, "grad_norm": 2.3574821949005127, "lr": 1.6442481704396416e-05, "s_per_step": 2.167111291885376, "peak_gb": 37.328663552, "samples_seen": 111200, "time": 1791173159.7185235}
141
+ {"step": 3500, "loss": 0.754842204619199, "grad_norm": 1.3139305114746094, "lr": 1.5806448925743172e-05, "s_per_step": 2.2147708892822267, "peak_gb": 37.328663552, "samples_seen": 112000, "time": 1791173215.0882668}
142
+ {"step": 3525, "loss": 0.6143901840061881, "grad_norm": 1.189170002937317, "lr": 1.5180644604429567e-05, "s_per_step": 2.2504481792449953, "peak_gb": 37.328663552, "samples_seen": 112800, "time": 1791173271.3500268}
143
+ {"step": 3550, "loss": 0.7468885008245707, "grad_norm": 1.0117112398147583, "lr": 1.456525593987194e-05, "s_per_step": 2.2883919048309327, "peak_gb": 37.328663552, "samples_seen": 113600, "time": 1791173328.5602643}
144
+ {"step": 3575, "loss": 0.6123304791096598, "grad_norm": 1.6147323846817017, "lr": 1.3960467015808465e-05, "s_per_step": 2.1810163593292238, "peak_gb": 37.328663552, "samples_seen": 114400, "time": 1791173383.0861204}
145
+ {"step": 3600, "loss": 0.6615970143675804, "grad_norm": 2.9963948726654053, "lr": 1.3366458745233385e-05, "s_per_step": 2.190250720977783, "peak_gb": 37.328663552, "samples_seen": 115200, "time": 1791173437.8429022}
146
+ {"step": 3625, "loss": 0.6804951827973127, "grad_norm": 1.5510998964309692, "lr": 1.2783408816279802e-05, "s_per_step": 2.1283248138427733, "peak_gb": 37.328663552, "samples_seen": 116000, "time": 1791173491.051459}
147
+ {"step": 3650, "loss": 0.6892811707407236, "grad_norm": 1.4326013326644897, "lr": 1.2211491639067125e-05, "s_per_step": 2.369460802078247, "peak_gb": 37.328663552, "samples_seen": 116800, "time": 1791173550.289685}
148
+ {"step": 3675, "loss": 0.6716245946101844, "grad_norm": 2.649146795272827, "lr": 1.1650878293528994e-05, "s_per_step": 2.1898504161834715, "peak_gb": 37.328663552, "samples_seen": 117600, "time": 1791173605.0363798}
149
+ {"step": 3700, "loss": 0.6472007501125335, "grad_norm": 2.400968074798584, "lr": 1.110173647823749e-05, "s_per_step": 2.089416027069092, "peak_gb": 37.328663552, "samples_seen": 118400, "time": 1791173657.2722385}
150
+ {"step": 3725, "loss": 0.6779031552001834, "grad_norm": 1.842848300933838, "lr": 1.0564230460238694e-05, "s_per_step": 2.1233923053741455, "peak_gb": 37.328663552, "samples_seen": 119200, "time": 1791173710.357557}
151
+ {"step": 3750, "loss": 0.6172893346473575, "grad_norm": 3.009672164916992, "lr": 1.0038521025914932e-05, "s_per_step": 2.111167163848877, "peak_gb": 37.328663552, "samples_seen": 120000, "time": 1791173763.137248}
152
+ {"step": 3775, "loss": 0.6523078301548958, "grad_norm": 0.9374409317970276, "lr": 9.524765432887944e-06, "s_per_step": 2.3036874389648436, "peak_gb": 37.328663552, "samples_seen": 120800, "time": 1791173820.7299206}
153
+ {"step": 3800, "loss": 0.6383235438168049, "grad_norm": 1.7281076908111572, "lr": 9.02311736297789e-06, "s_per_step": 2.1478301906585693, "peak_gb": 37.328663552, "samples_seen": 121600, "time": 1791173874.4261167}
154
+ {"step": 3825, "loss": 0.6730420371307991, "grad_norm": 1.6016979217529297, "lr": 8.533726876231828e-06, "s_per_step": 2.2305167865753175, "peak_gb": 37.328663552, "samples_seen": 122400, "time": 1791173930.1895103}
155
+ {"step": 3850, "loss": 0.6855003063753248, "grad_norm": 1.8208370208740234, "lr": 8.056740366035576e-06, "s_per_step": 2.341342372894287, "peak_gb": 37.328663552, "samples_seen": 123200, "time": 1791173988.7235923}
156
+ {"step": 3875, "loss": 0.6778232954442501, "grad_norm": 1.5814423561096191, "lr": 7.592300515322576e-06, "s_per_step": 2.280441083908081, "peak_gb": 37.328663552, "samples_seen": 124000, "time": 1791174045.7350547}
157
+ {"step": 3900, "loss": 0.7005690622329712, "grad_norm": 1.7448139190673828, "lr": 7.140546253892439e-06, "s_per_step": 2.2497168445587157, "peak_gb": 37.328663552, "samples_seen": 124800, "time": 1791174101.9784636}
158
+ {"step": 3925, "loss": 0.680451717376709, "grad_norm": 3.5512075424194336, "lr": 6.70161271685244e-06, "s_per_step": 2.2352783203125, "peak_gb": 37.328663552, "samples_seen": 125600, "time": 1791174157.8608978}
159
+ {"step": 3950, "loss": 0.6795122207701206, "grad_norm": 3.0180165767669678, "lr": 6.27563120419386e-06, "s_per_step": 2.179275674819946, "peak_gb": 37.328663552, "samples_seen": 126400, "time": 1791174212.3432853}
160
+ {"step": 3975, "loss": 0.6303972934558988, "grad_norm": 1.2422635555267334, "lr": 5.8627291415157435e-06, "s_per_step": 2.078404426574707, "peak_gb": 37.328663552, "samples_seen": 127200, "time": 1791174264.3039021}
161
+ {"step": 4000, "loss": 0.6791560095734894, "grad_norm": 1.8804198503494263, "lr": 5.463030041907597e-06, "s_per_step": 2.176313486099243, "peak_gb": 37.328663552, "samples_seen": 128000, "time": 1791174318.7121499}
162
+ {"step": 4025, "loss": 0.6513661769777537, "grad_norm": 2.495976448059082, "lr": 5.076653469002313e-06, "s_per_step": 2.1938330554962158, "peak_gb": 37.328663552, "samples_seen": 128800, "time": 1791174373.5584402}
163
+ {"step": 4050, "loss": 0.6535539354383946, "grad_norm": 1.9437261819839478, "lr": 4.70371500121074e-06, "s_per_step": 2.1953781414031983, "peak_gb": 37.328663552, "samples_seen": 129600, "time": 1791174428.4433782}
164
+ {"step": 4075, "loss": 0.7009203843027353, "grad_norm": 1.0737240314483643, "lr": 4.344326197148107e-06, "s_per_step": 2.248444013595581, "peak_gb": 37.328663552, "samples_seen": 130400, "time": 1791174484.654927}
165
+ {"step": 4100, "loss": 0.64098393403925, "grad_norm": 1.6634223461151123, "lr": 3.9985945622630915e-06, "s_per_step": 2.09239275932312, "peak_gb": 37.328663552, "samples_seen": 131200, "time": 1791174536.9651742}
166
+ {"step": 4125, "loss": 0.7195338121801614, "grad_norm": 1.6115185022354126, "lr": 3.6666235166793127e-06, "s_per_step": 2.3179955768585203, "peak_gb": 37.328663552, "samples_seen": 132000, "time": 1791174594.9155335}
167
+ {"step": 4150, "loss": 0.6215781589318067, "grad_norm": 1.4749220609664917, "lr": 3.3485123642587658e-06, "s_per_step": 2.1519798469543456, "peak_gb": 37.328663552, "samples_seen": 132800, "time": 1791174648.7154896}
168
+ {"step": 4175, "loss": 0.6063127079233527, "grad_norm": 2.4063498973846436, "lr": 3.0443562628967415e-06, "s_per_step": 2.164182243347168, "peak_gb": 37.328663552, "samples_seen": 133600, "time": 1791174702.8204815}
169
+ {"step": 4200, "loss": 0.6975668650865555, "grad_norm": 1.5435162782669067, "lr": 2.754246196056759e-06, "s_per_step": 2.2063780212402344, "peak_gb": 37.328663552, "samples_seen": 134400, "time": 1791174757.9804275}
170
+ {"step": 4225, "loss": 0.6374323242623359, "grad_norm": 1.643761396408081, "lr": 2.4782689455543796e-06, "s_per_step": 2.173592166900635, "peak_gb": 37.328663552, "samples_seen": 135200, "time": 1791174812.3206897}
171
+ {"step": 4250, "loss": 0.6485517076496035, "grad_norm": 2.015836238861084, "lr": 2.2165070655977727e-06, "s_per_step": 2.0835715389251708, "peak_gb": 37.328663552, "samples_seen": 136000, "time": 1791174864.4104404}
172
+ {"step": 4275, "loss": 0.6164555564941838, "grad_norm": 1.3733229637145996, "lr": 1.9690388580929363e-06, "s_per_step": 2.3116278553009035, "peak_gb": 37.328663552, "samples_seen": 136800, "time": 1791174922.2016754}
173
+ {"step": 4300, "loss": 0.6570157048292458, "grad_norm": 1.1992275714874268, "lr": 1.7359383492209613e-06, "s_per_step": 2.2723003482818602, "peak_gb": 37.328663552, "samples_seen": 137600, "time": 1791174979.0097623}
174
+ {"step": 4325, "loss": 0.6460098330676556, "grad_norm": 1.471946358680725, "lr": 1.5172752672942159e-06, "s_per_step": 2.1469297885894774, "peak_gb": 37.328663552, "samples_seen": 138400, "time": 1791175032.6836147}
175
+ {"step": 4350, "loss": 0.6363432590290904, "grad_norm": 1.199631690979004, "lr": 1.3131150218982923e-06, "s_per_step": 2.2307802295684813, "peak_gb": 37.328663552, "samples_seen": 139200, "time": 1791175088.4536717}
176
+ {"step": 4375, "loss": 0.6324924950860441, "grad_norm": 1.9240083694458008, "lr": 1.123518684325714e-06, "s_per_step": 2.1913782024383544, "peak_gb": 37.328663552, "samples_seen": 140000, "time": 1791175143.2387235}
177
+ {"step": 4400, "loss": 0.6856773376883939, "grad_norm": 2.5178141593933105, "lr": 9.485429693074754e-07, "s_per_step": 2.324142723083496, "peak_gb": 37.328663552, "samples_seen": 140800, "time": 1791175201.3428428}
178
+ {"step": 4425, "loss": 0.6213866775110364, "grad_norm": 0.6766404509544373, "lr": 7.882402180477033e-07, "s_per_step": 2.1214879989624023, "peak_gb": 37.328663552, "samples_seen": 141600, "time": 1791175254.3807}
179
+ {"step": 4450, "loss": 0.7287836332805455, "grad_norm": 0.6748529672622681, "lr": 6.426583825666133e-07, "s_per_step": 2.485296297073364, "peak_gb": 37.328663552, "samples_seen": 142400, "time": 1791175316.5137088}
180
+ {"step": 4475, "loss": 0.6739577786624431, "grad_norm": 1.455942153930664, "lr": 5.118410113564565e-07, "s_per_step": 2.2427228450775147, "peak_gb": 37.328663552, "samples_seen": 143200, "time": 1791175372.5823774}
181
+ {"step": 4500, "loss": 0.5849272518604994, "grad_norm": 1.208377718925476, "lr": 3.9582723635464004e-07, "s_per_step": 2.0806651401519773, "peak_gb": 37.328663552, "samples_seen": 144000, "time": 1791175424.5995016}
182
+ {"step": 4525, "loss": 0.6931665486842394, "grad_norm": 2.3075685501098633, "lr": 2.9465176123806284e-07, "s_per_step": 2.244974718093872, "peak_gb": 37.328663552, "samples_seen": 144800, "time": 1791175480.7244606}
183
+ {"step": 4550, "loss": 0.675754471779801, "grad_norm": 1.8862049579620361, "lr": 2.083448510420527e-07, "s_per_step": 2.2479200553894043, "peak_gb": 37.328663552, "samples_seen": 145600, "time": 1791175536.9230142}
184
+ {"step": 4575, "loss": 0.6882729256153106, "grad_norm": 1.836661458015442, "lr": 1.3693232310705295e-07, "s_per_step": 2.2444938945770265, "peak_gb": 37.328663552, "samples_seen": 146400, "time": 1791175593.0358894}
185
+ {"step": 4600, "loss": 0.6727710048668086, "grad_norm": 2.4301044940948486, "lr": 8.043553935577208e-08, "s_per_step": 2.264568119049072, "peak_gb": 37.328663552, "samples_seen": 147200, "time": 1791175649.6505373}
186
+ {"step": 4625, "loss": 0.6940591262280941, "grad_norm": 1.7329171895980835, "lr": 3.8871399903134265e-08, "s_per_step": 2.141450147628784, "peak_gb": 37.328663552, "samples_seen": 148000, "time": 1791175703.1872365}
187
+ {"step": 4650, "loss": 0.6567035659402609, "grad_norm": 1.8683308362960815, "lr": 1.2252338000839914e-08, "s_per_step": 2.156028919219971, "peak_gb": 37.328663552, "samples_seen": 148800, "time": 1791175757.0884378}
188
+ {"step": 4675, "loss": 0.6862510319799184, "grad_norm": 2.6522982120513916, "lr": 5.863163181796249e-10, "s_per_step": 2.03354868888855, "peak_gb": 37.328663552, "samples_seen": 149600, "time": 1791175807.9276795}
189
+ {"step": 4681, "loss": 0.5528123727999628, "grad_norm": 2.920032024383545, "lr": 1.1965662055635207e-11, "s_per_step": 1.9081807533899944, "peak_gb": 37.328663552, "samples_seen": 149792, "time": 1791175819.3772266}
stage2b/lora_adapter/adapter_config.json ADDED
@@ -0,0 +1,51 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "alora_invocation_tokens": null,
3
+ "alpha_pattern": {},
4
+ "arrow_config": null,
5
+ "auto_mapping": null,
6
+ "base_model_name_or_path": "NCAIR1/N-ATLaS",
7
+ "bias": "none",
8
+ "corda_config": null,
9
+ "ensure_weight_tying": false,
10
+ "eva_config": null,
11
+ "exclude_modules": null,
12
+ "fan_in_fan_out": false,
13
+ "inference_mode": true,
14
+ "init_lora_weights": true,
15
+ "kasa_config": null,
16
+ "layer_replication": null,
17
+ "layers_pattern": null,
18
+ "layers_to_transform": null,
19
+ "loftq_config": {},
20
+ "lora_alpha": 128,
21
+ "lora_bias": false,
22
+ "lora_dropout": 0.05,
23
+ "lora_ga_config": null,
24
+ "megatron_config": null,
25
+ "megatron_core": "megatron.core",
26
+ "modules_to_save": null,
27
+ "monteclora_config": null,
28
+ "peft_type": "LORA",
29
+ "peft_version": "0.21.2",
30
+ "qalora_group_size": 16,
31
+ "r": 64,
32
+ "rank_pattern": {},
33
+ "revision": null,
34
+ "target_modules": [
35
+ "k_proj",
36
+ "gate_proj",
37
+ "v_proj",
38
+ "o_proj",
39
+ "q_proj",
40
+ "down_proj",
41
+ "up_proj"
42
+ ],
43
+ "target_parameters": null,
44
+ "task_type": "CAUSAL_LM",
45
+ "trainable_token_indices": null,
46
+ "use_bdlora": null,
47
+ "use_dora": false,
48
+ "use_qalora": false,
49
+ "use_rslora": false,
50
+ "velora_config": null
51
+ }
stage2b/lora_adapter/adapter_model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4454825a3b6bc31070b68c9be66650d486b8f333eb9acfedb50f65858b9d3902
3
+ size 671149168
stage2b/projector.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ac1556194fdedc845bfe93c00440d549c20760b730898e516f5b85a96efaea8f
3
+ size 79724872