Image-Text-to-Text
PEFT
Safetensors
vision-language
multimodal
llava
lora
siglip2
n-atlas
nigerian-languages
Instructions to use Modularcomputing/AtlasVision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Modularcomputing/AtlasVision with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Upload via givemeanode export_data
Browse files- README.md +275 -1
- chat.py +156 -0
- code/eval_stage1.py +85 -0
- code/eval_stage2.py +231 -0
- code/infer.py +45 -0
- code/prepare.py +44 -0
- code/run.sh +47 -0
- code/run_stage2.sh +38 -0
- code/train.py +384 -0
- code/train_stage2.py +294 -0
- eval/report.md +58 -0
- eval/stage1_eval.json +68 -0
- eval/stage2_eval.json +104 -0
- logs/stage1_train_log.jsonl +376 -0
- logs/stage2_train_log.jsonl +197 -0
- stage1/projector.safetensors +3 -0
- stage2/lora_adapter/adapter_config.json +51 -0
- stage2/lora_adapter/adapter_model.safetensors +3 -0
- stage2/projector.safetensors +3 -0
README.md
CHANGED
|
@@ -1,3 +1,277 @@
|
|
| 1 |
---
|
| 2 |
-
license:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: awarri-open-source-research-and-innovation-license
|
| 4 |
+
license_link: https://huggingface.co/NCAIR1/N-ATLaS
|
| 5 |
+
base_model:
|
| 6 |
+
- NCAIR1/N-ATLaS
|
| 7 |
+
- google/siglip2-base-patch16-224
|
| 8 |
+
library_name: peft
|
| 9 |
+
pipeline_tag: image-text-to-text
|
| 10 |
+
language:
|
| 11 |
+
- en
|
| 12 |
+
- ha
|
| 13 |
+
- ig
|
| 14 |
+
- yo
|
| 15 |
+
datasets:
|
| 16 |
+
- liuhaotian/LLaVA-Pretrain
|
| 17 |
+
- liuhaotian/LLaVA-Instruct-150K
|
| 18 |
+
- lmms-lab/POPE
|
| 19 |
+
tags:
|
| 20 |
+
- vision-language
|
| 21 |
+
- multimodal
|
| 22 |
+
- llava
|
| 23 |
+
- lora
|
| 24 |
+
- siglip2
|
| 25 |
+
- n-atlas
|
| 26 |
+
- nigerian-languages
|
| 27 |
---
|
| 28 |
+
|
| 29 |
+
# AtlasVision
|
| 30 |
+
|
| 31 |
+
**AtlasVision** gives [N-ATLaS](https://huggingface.co/NCAIR1/N-ATLaS) — Nigeria's multilingual Llama-3-8B model for
|
| 32 |
+
English, Hausa, Igbo and Yoruba — the ability to see. It connects a frozen
|
| 33 |
+
[SigLIP2](https://huggingface.co/google/siglip2-base-patch16-224) image encoder to N-ATLaS through a trained MLP projector,
|
| 34 |
+
following the two-stage LLaVA recipe:
|
| 35 |
+
|
| 36 |
+
1. **Stage 1 — alignment:** only the projector is trained, on 300k image–caption pairs, so the LLM can "read" image features.
|
| 37 |
+
2. **Stage 2 — visual instruction tuning:** the projector keeps training and LoRA adapters are added to N-ATLaS, on
|
| 38 |
+
LLaVA-Instruct-150K conversations, so the model answers questions and holds conversations about images.
|
| 39 |
+
|
| 40 |
+
This repository contains **only the trained parts** (projectors and LoRA adapter, ~0.9 GB).
|
| 41 |
+
The base models are downloaded from their own repositories at load time, so their licenses and access conditions
|
| 42 |
+
(N-ATLaS is gated) continue to apply.
|
| 43 |
+
|
| 44 |
+
## Model summary
|
| 45 |
+
|
| 46 |
+
| | |
|
| 47 |
+
|---|---|
|
| 48 |
+
| Language model | `NCAIR1/N-ATLaS` (Llama-3 8B, frozen; LoRA in stage 2) |
|
| 49 |
+
| Vision encoder | `google/siglip2-base-patch16-224` (ViT-B/16, 224 px, 196 patch tokens, frozen) |
|
| 50 |
+
| Projector | 2-layer MLP 768 → 4096 → 4096 with GELU (19.9M params) |
|
| 51 |
+
| LoRA (stage 2) | r = 64, alpha = 128, dropout 0.05 on q/k/v/o/gate/up/down projections (167.8M params) |
|
| 52 |
+
| Image tokens | 196 per image, inserted after `<|begin_of_text|>` and the user header |
|
| 53 |
+
| Precision | bf16 weights and compute; projector and LoRA trained in fp32 |
|
| 54 |
+
| Languages | Training data is English; N-ATLaS contributes Hausa, Igbo and Yoruba |
|
| 55 |
+
|
| 56 |
+
## Repository contents
|
| 57 |
+
|
| 58 |
+
```
|
| 59 |
+
stage1/projector.safetensors # stage-1 projector (image captioning / alignment)
|
| 60 |
+
stage2/projector.safetensors # stage-2 projector (use together with the LoRA adapter)
|
| 61 |
+
stage2/lora_adapter/ # PEFT LoRA adapter for NCAIR1/N-ATLaS
|
| 62 |
+
chat.py # standalone inference script (CLI + Python API)
|
| 63 |
+
code/ # exact training, evaluation and launch scripts used
|
| 64 |
+
eval/ # raw evaluation results (JSON) and report
|
| 65 |
+
logs/ # per-step training logs (loss, grad norm, LR, speed)
|
| 66 |
+
```
|
| 67 |
+
|
| 68 |
+
## Quick start
|
| 69 |
+
|
| 70 |
+
Accept the conditions on the [N-ATLaS model page](https://huggingface.co/NCAIR1/N-ATLaS) first, then:
|
| 71 |
+
|
| 72 |
+
```bash
|
| 73 |
+
pip install -U torch transformers peft safetensors pillow accelerate huggingface_hub
|
| 74 |
+
export HF_TOKEN=hf_... # a token from the account that was granted N-ATLaS access
|
| 75 |
+
wget https://huggingface.co/FUTO-NIGERIA/AtlasVision/resolve/main/chat.py
|
| 76 |
+
|
| 77 |
+
python chat.py --image photo.jpg --question "What is happening in this picture?"
|
| 78 |
+
python chat.py --image photo.jpg --question "Kedu ihe dị na foto a?" # Igbo
|
| 79 |
+
python chat.py --stage 1 --image photo.jpg # stage-1 captioner
|
| 80 |
+
python chat.py --image photo.jpg --interactive # follow-up questions
|
| 81 |
+
python chat.py --image photo.jpg --load-in-4bit # ~8 GB GPU (pip install bitsandbytes)
|
| 82 |
+
```
|
| 83 |
+
|
| 84 |
+
bf16 needs about 18 GB of GPU memory (A100, L4, RTX 4090…). `--load-in-4bit` runs on ~8 GB GPUs such as a free Colab T4.
|
| 85 |
+
|
| 86 |
+
From Python:
|
| 87 |
+
|
| 88 |
+
```python
|
| 89 |
+
from chat import AtlasVision, load_image
|
| 90 |
+
|
| 91 |
+
model = AtlasVision(stage=2) # downloads N-ATLaS, SigLIP2 and this repo's weights
|
| 92 |
+
image = load_image("https://example.com/street.jpg")
|
| 93 |
+
print(model.ask(image, "Describe this image in detail."))
|
| 94 |
+
print(model.ask(image, "How many people are there?")) # follow-ups keep the conversation
|
| 95 |
+
```
|
| 96 |
+
|
| 97 |
+
### Prompt format
|
| 98 |
+
|
| 99 |
+
The image embeddings are spliced into the Llama-3 chat format:
|
| 100 |
+
|
| 101 |
+
```
|
| 102 |
+
<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n [196 image embeddings] {question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n
|
| 103 |
+
```
|
| 104 |
+
|
| 105 |
+
Follow-up turns append `{answer}<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n{question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n`.
|
| 106 |
+
`chat.py` builds this for you.
|
| 107 |
+
|
| 108 |
+
## Training
|
| 109 |
+
|
| 110 |
+
Both stages ran on a single NVIDIA H100 80GB.
|
| 111 |
+
|
| 112 |
+
| | Stage 1 — alignment | Stage 2 — instruction tuning |
|
| 113 |
+
|---|---|---|
|
| 114 |
+
| Data | LLaVA-Pretrain (BLIP captions of LAION/CC/SBU), random 300k of 558k | LLaVA-Instruct-150K on COCO train2017 (156,712 train / 1,000 held out) |
|
| 115 |
+
| Trained | projector | projector + LoRA |
|
| 116 |
+
| Effective batch | 32 (16 × 2 accumulation) | 32 (8 × 4 accumulation) |
|
| 117 |
+
| Learning rate | 2e-4, 3% warmup, cosine | LoRA 2e-4, projector 2e-5, 3% warmup, cosine |
|
| 118 |
+
| Max text length | 128 tokens | 1,024 tokens |
|
| 119 |
+
| Optimizer steps | 9375 | 4897 |
|
| 120 |
+
| Loss (first log → last log) | 8.551 → 2.454 | 2.102 → 1.15 |
|
| 121 |
+
| Wall-clock time | ~91 min | ~206 min |
|
| 122 |
+
| Peak GPU memory | 45.2 GB | 35.9 GB (gradient checkpointing) |
|
| 123 |
+
|
| 124 |
+
Other details: AdamW (no weight decay), gradient clipping at 1.0, bf16 autocast, loss only on assistant tokens,
|
| 125 |
+
length-bucketed batches in stage 2. Per-step logs are in `logs/`.
|
| 126 |
+
|
| 127 |
+
## Evaluation
|
| 128 |
+
|
| 129 |
+
### Stage 1 — does the model actually use the image?
|
| 130 |
+
|
| 131 |
+
Measured on 2,000 LLaVA-Pretrain captions that were **not** in the training subset (caption loss, lower is better):
|
| 132 |
+
|
| 133 |
+
| Setting | Held-out loss |
|
| 134 |
+
|---|---|
|
| 135 |
+
| Untrained (random) projector | 8.7249 |
|
| 136 |
+
| **Trained stage-1 projector** | **2.4523** |
|
| 137 |
+
| Trained projector, captions paired with the wrong images | 5.4332 |
|
| 138 |
+
|
| 139 |
+
Held-out loss matches the final training loss (no over-fitting), and pairing captions with the wrong images
|
| 140 |
+
roughly doubles the loss — the language model is relying on the visual content, not guessing generic captions.
|
| 141 |
+
|
| 142 |
+
### Stage 1 vs stage 2
|
| 143 |
+
|
| 144 |
+
| Metric | Stage 1 | Stage 2 |
|
| 145 |
+
|---|---|---|
|
| 146 |
+
| Held-out LLaVA-Instruct loss (500 unseen conversations, lower is better) | 2.2521 | **1.1271** |
|
| 147 |
+
|
| 148 |
+
**POPE** (object hallucination; yes/no questions about COCO val2014 images, scored from the Yes/No token probabilities;
|
| 149 |
+
balanced 50% yes / 50% no, so a yes-ratio near 0.5 is ideal):
|
| 150 |
+
|
| 151 |
+
| POPE split | Stage 1: accuracy / F1 / yes-ratio | Stage 2: accuracy / F1 / yes-ratio | Stage 2: precision / recall |
|
| 152 |
+
|---|---|---|---|
|
| 153 |
+
| random | 0.500 / 0.200 / 0.13 | **0.543** / **0.686** / 0.95 | 0.522 / 0.997 |
|
| 154 |
+
| popular | 0.503 / 0.201 / 0.12 | **0.522** / **0.676** / 0.98 | 0.511 / 0.997 |
|
| 155 |
+
| adversarial | 0.502 / 0.201 / 0.12 | **0.514** / **0.672** / 0.98 | 0.507 / 0.997 |
|
| 156 |
+
|
| 157 |
+
**How to read this: stage 2 has a strong "yes" bias.** It answers "yes" to 95%–98% of POPE
|
| 158 |
+
questions, so it almost never misses an object that is present (recall ≈ 1.0) but also confirms most objects that are
|
| 159 |
+
*absent*, leaving accuracy close to chance. Stage 1 shows the opposite bias (it was never trained to answer questions).
|
| 160 |
+
This is a known effect of training on LLaVA-Instruct-150K alone: its conversations almost never contain questions whose
|
| 161 |
+
answer is "no". The descriptive examples below are far stronger than these numbers suggest, but **do not use AtlasVision
|
| 162 |
+
for yes/no verification of image content**. The usual fix is to add short-answer VQA data (VQAv2, GQA) and negative
|
| 163 |
+
examples (e.g. LRV-Instruction) to stage 2, as LLaVA-1.5 does — it reaches POPE F1 ≈ 0.86 with 665k mixed samples at 336 px.
|
| 164 |
+
|
| 165 |
+
### Examples — stage 1 captions on unseen images
|
| 166 |
+
|
| 167 |
+
| Reference caption | AtlasVision stage 1 |
|
| 168 |
+
|---|---|
|
| 169 |
+
| cute homecoming prom dress with lace top and satin skirt | a - line sweetheart sweetheart lace tulle prom dress |
|
| 170 |
+
| a pair of silicone bracelets with a four codes message | a pair of silicone bracelets with the word code on them |
|
| 171 |
+
| a set of fingerprint icons in white on a black background | the iphone 6s and iphone 6s plus are shown in a black and white image |
|
| 172 |
+
| a business woman drawing modern concept of a website creation | a man writing the word modern web design on a whiteboard |
|
| 173 |
+
| the beach in el nido national park, puerto puerto | the limestone cliffs and limestone islands in the background |
|
| 174 |
+
| three pieces of paper with the words, democratic decentified dp controlled centralized ccp | a diagram showing the different types of democracy |
|
| 175 |
+
|
| 176 |
+
### Examples — "Describe this image in detail." (COCO val2014, unseen)
|
| 177 |
+
|
| 178 |
+
**COCO_val2014_000000310196**
|
| 179 |
+
|
| 180 |
+
- *Stage 1:* a skier in the snow on a mountain
|
| 181 |
+
- *Stage 2:* The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.
|
| 182 |
+
|
| 183 |
+
There are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together.
|
| 184 |
+
|
| 185 |
+
**COCO_val2014_000000210789**
|
| 186 |
+
|
| 187 |
+
- *Stage 1:* a woman and her child in the rain
|
| 188 |
+
- *Stage 2:* The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.
|
| 189 |
+
|
| 190 |
+
In the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene.
|
| 191 |
+
|
| 192 |
+
**COCO_val2014_000000429109**
|
| 193 |
+
|
| 194 |
+
- *Stage 1:* a bus and several other vehicles parked in front of a building
|
| 195 |
+
- *Stage 2:* The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.
|
| 196 |
+
|
| 197 |
+
Numerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.
|
| 198 |
+
|
| 199 |
+
In addition to the vehicles, there are traffic lights at various points along the street
|
| 200 |
+
|
| 201 |
+
**COCO_val2014_000000211674**
|
| 202 |
+
|
| 203 |
+
- *Stage 1:* a bus with a red and white logo on it, carrying passengers
|
| 204 |
+
- *Stage 2:* The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.
|
| 205 |
+
|
| 206 |
+
In addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by.
|
| 207 |
+
|
| 208 |
+
### Questions in Nigerian languages (stage 2)
|
| 209 |
+
|
| 210 |
+
The stage-2 instruction data is English only. The model understands the questions below (its answers match the image) but replies in English; adding translated instruction data would be needed for answers in these languages.
|
| 211 |
+
|
| 212 |
+
| Language | Question | Answer |
|
| 213 |
+
|---|---|---|
|
| 214 |
+
| Igbo | Kedu ihe dị na foto a? | The image features a person skiing down a snow-covered slope, with the skier wearing red pants. |
|
| 215 |
+
| Yoruba | Kí ni ó wà nínú àwòrán yìí? | The image features a person skiing down a snow-covered slope, with the skier wearing red pants. |
|
| 216 |
+
| Hausa | Me ke cikin wannan hoton? | In the image, a person is skiing down a snow-covered slope. |
|
| 217 |
+
|
| 218 |
+
### Text-only check (no image)
|
| 219 |
+
|
| 220 |
+
Does the stage-2 LoRA damage N-ATLaS's original text abilities? Same Igbo question, greedy decoding:
|
| 221 |
+
|
| 222 |
+
**Q:** Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.
|
| 223 |
+
|
| 224 |
+
- *Base N-ATLaS:* Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
|
| 225 |
+
- *With stage-2 LoRA:* Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu "positron" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej
|
| 226 |
+
|
| 227 |
+
Both answer fluently in Igbo, so the adapter keeps N-ATLaS's text abilities intact (responses are cut at 120 tokens).
|
| 228 |
+
|
| 229 |
+
## Limitations
|
| 230 |
+
|
| 231 |
+
- **Hallucination and yes-bias.** Long descriptions are fluent but often add plausible details that are not in the
|
| 232 |
+
image (exact counts, extra cars, a handbag), and stage 2 answers "yes" to most yes/no questions (see POPE above).
|
| 233 |
+
Treat counts, small objects and yes/no answers as unreliable.
|
| 234 |
+
- **English-only visual training.** All image–text training data is English. Answers to Hausa, Igbo and Yoruba questions
|
| 235 |
+
come from N-ATLaS's own multilingual ability and are noticeably less reliable; they may switch to English.
|
| 236 |
+
- **Low resolution.** Images are resized to 224×224, so small text, fine details and dense documents are hard.
|
| 237 |
+
- **One image per conversation**, no video, no bounding boxes or grounding.
|
| 238 |
+
- **Data biases.** Web captions (LAION/CC/SBU) and COCO carry Western-centric content; performance on Nigerian scenes,
|
| 239 |
+
people and text has not been measured and is likely weaker.
|
| 240 |
+
- **Not for high-stakes use** (medical, legal, identity, surveillance or safety-critical decisions).
|
| 241 |
+
|
| 242 |
+
## Licenses and terms
|
| 243 |
+
|
| 244 |
+
This repository contains weights trained on top of other people's models and data. Using it means complying with all of:
|
| 245 |
+
|
| 246 |
+
- **N-ATLaS**: Awarri Open-Source Research and Innovation License (attribution required; caps deployments at 1,000
|
| 247 |
+
active end-users per 30 days) and the Llama 3 Community License it inherits. Access is gated — accept the
|
| 248 |
+
conditions on the [model page](https://huggingface.co/NCAIR1/N-ATLaS).
|
| 249 |
+
- **SigLIP2**: Apache-2.0.
|
| 250 |
+
- **LLaVA-Instruct-150K** (stage 2 data): CC BY-NC 4.0, generated with GPT-4 and subject to OpenAI's terms —
|
| 251 |
+
**stage-2 weights are for non-commercial research use.**
|
| 252 |
+
- **LLaVA-Pretrain** (stage 1 data): captions/images from LAION, Conceptual Captions and SBU under their respective terms.
|
| 253 |
+
- **COCO images**: Flickr images under their individual Creative Commons licenses.
|
| 254 |
+
|
| 255 |
+
## Acknowledgements
|
| 256 |
+
|
| 257 |
+
- N-ATLaS by Awarri Technologies and the Federal Ministry of Communications, Innovation and Digital Economy (Nigeria),
|
| 258 |
+
as part of the Nigerian Languages AI Initiative.
|
| 259 |
+
- SigLIP2 by Google; LLaVA data and recipe by Liu et al.; POPE by Li et al.
|
| 260 |
+
- Trained on an H100 from San Francisco Compute (Autoresearch).
|
| 261 |
+
|
| 262 |
+
## Citation
|
| 263 |
+
|
| 264 |
+
```bibtex
|
| 265 |
+
@misc{atlasvision2026,
|
| 266 |
+
title = {AtlasVision: a vision-language extension of N-ATLaS},
|
| 267 |
+
author = {FUTO-NIGERIA},
|
| 268 |
+
year = {2026},
|
| 269 |
+
url = {https://huggingface.co/FUTO-NIGERIA/AtlasVision}
|
| 270 |
+
}
|
| 271 |
+
@inproceedings{liu2023llava,
|
| 272 |
+
title = {Visual Instruction Tuning},
|
| 273 |
+
author = {Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae},
|
| 274 |
+
booktitle = {NeurIPS},
|
| 275 |
+
year = {2023}
|
| 276 |
+
}
|
| 277 |
+
```
|
chat.py
ADDED
|
@@ -0,0 +1,156 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""AtlasVision inference: SigLIP2 vision encoder + N-ATLaS (Llama-3 8B) via a trained MLP projector.
|
| 3 |
+
|
| 4 |
+
pip install -U torch transformers peft safetensors pillow accelerate huggingface_hub
|
| 5 |
+
# optional for 4-bit on small GPUs (e.g. Colab T4): pip install bitsandbytes
|
| 6 |
+
export HF_TOKEN=hf_... # needs accepted access to the gated NCAIR1/N-ATLaS
|
| 7 |
+
|
| 8 |
+
python chat.py --image photo.jpg --question "What is happening in this picture?" # stage 2
|
| 9 |
+
python chat.py --stage 1 --image photo.jpg # stage 1 captioner
|
| 10 |
+
python chat.py --image https://example.com/cat.jpg --question "Kedu ihe dị na foto a?"
|
| 11 |
+
python chat.py --image photo.jpg --load-in-4bit # ~8 GB GPU
|
| 12 |
+
python chat.py --image photo.jpg --interactive # several questions
|
| 13 |
+
|
| 14 |
+
Needs ~18 GB of GPU memory in bf16 (A100, L4, RTX 4090...); --load-in-4bit fits ~8 GB GPUs such as a Colab T4.
|
| 15 |
+
"""
|
| 16 |
+
import argparse
|
| 17 |
+
import io
|
| 18 |
+
import os
|
| 19 |
+
import sys
|
| 20 |
+
|
| 21 |
+
import torch
|
| 22 |
+
import torch.nn as nn
|
| 23 |
+
from PIL import Image
|
| 24 |
+
|
| 25 |
+
REPO = "FUTO-NIGERIA/AtlasVision"
|
| 26 |
+
LLM = "NCAIR1/N-ATLaS"
|
| 27 |
+
VISION = "google/siglip2-base-patch16-224"
|
| 28 |
+
USER_HEADER = "<|start_header_id|>user<|end_header_id|>\n\n"
|
| 29 |
+
ASSIST_HEADER = "<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"
|
| 30 |
+
EOT = "<|eot_id|>"
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
class ProjectionMLP(nn.Module):
|
| 34 |
+
def __init__(self, vision_dim, text_dim):
|
| 35 |
+
super().__init__()
|
| 36 |
+
self.net = nn.Sequential(nn.Linear(vision_dim, text_dim), nn.GELU(), nn.Linear(text_dim, text_dim))
|
| 37 |
+
|
| 38 |
+
def forward(self, x):
|
| 39 |
+
return self.net(x)
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
def fetch(repo, filename):
|
| 43 |
+
if os.path.isdir(repo):
|
| 44 |
+
return os.path.join(repo, filename)
|
| 45 |
+
from huggingface_hub import hf_hub_download
|
| 46 |
+
return hf_hub_download(repo, filename)
|
| 47 |
+
|
| 48 |
+
|
| 49 |
+
def load_image(src):
|
| 50 |
+
if src.startswith(("http://", "https://")):
|
| 51 |
+
import urllib.request
|
| 52 |
+
with urllib.request.urlopen(src) as r:
|
| 53 |
+
return Image.open(io.BytesIO(r.read())).convert("RGB")
|
| 54 |
+
return Image.open(src).convert("RGB")
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
class AtlasVision:
|
| 58 |
+
def __init__(self, stage=2, repo=REPO, llm=LLM, vision=VISION, load_in_4bit=False, device=None):
|
| 59 |
+
from safetensors.torch import load_file
|
| 60 |
+
from transformers import AutoImageProcessor, AutoModel, AutoModelForCausalLM, AutoTokenizer
|
| 61 |
+
|
| 62 |
+
self.device = torch.device(device or ("cuda" if torch.cuda.is_available() else "cpu"))
|
| 63 |
+
cuda = self.device.type == "cuda"
|
| 64 |
+
self.dtype = torch.bfloat16 if (not cuda or torch.cuda.is_bf16_supported()) else torch.float16
|
| 65 |
+
|
| 66 |
+
self.tok = AutoTokenizer.from_pretrained(llm)
|
| 67 |
+
self.pad_id = self.tok.pad_token_id if self.tok.pad_token_id is not None else self.tok.eos_token_id
|
| 68 |
+
self.eot_id = self.tok.convert_tokens_to_ids(EOT)
|
| 69 |
+
self.processor = AutoImageProcessor.from_pretrained(vision)
|
| 70 |
+
|
| 71 |
+
full_vision = AutoModel.from_pretrained(vision, dtype=self.dtype)
|
| 72 |
+
self.vision = full_vision.vision_model.to(self.device).eval()
|
| 73 |
+
vision_dim = full_vision.config.vision_config.hidden_size
|
| 74 |
+
del full_vision
|
| 75 |
+
|
| 76 |
+
kw = {"dtype": self.dtype}
|
| 77 |
+
if load_in_4bit:
|
| 78 |
+
from transformers import BitsAndBytesConfig
|
| 79 |
+
kw["quantization_config"] = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
|
| 80 |
+
bnb_4bit_compute_dtype=self.dtype)
|
| 81 |
+
kw["device_map"] = {"": self.device.index or 0}
|
| 82 |
+
self.llm = AutoModelForCausalLM.from_pretrained(llm, **kw)
|
| 83 |
+
if not load_in_4bit:
|
| 84 |
+
self.llm.to(self.device)
|
| 85 |
+
text_dim = self.llm.config.hidden_size
|
| 86 |
+
|
| 87 |
+
if stage == 2:
|
| 88 |
+
from peft import PeftModel
|
| 89 |
+
if os.path.isdir(repo):
|
| 90 |
+
self.llm = PeftModel.from_pretrained(self.llm, os.path.join(repo, "stage2/lora_adapter"))
|
| 91 |
+
else:
|
| 92 |
+
self.llm = PeftModel.from_pretrained(self.llm, repo, subfolder="stage2/lora_adapter")
|
| 93 |
+
self.llm.eval()
|
| 94 |
+
|
| 95 |
+
self.projector = ProjectionMLP(vision_dim, text_dim)
|
| 96 |
+
self.projector.load_state_dict(load_file(fetch(repo, f"stage{stage}/projector.safetensors")))
|
| 97 |
+
self.projector.to(self.device, dtype=torch.float32).eval()
|
| 98 |
+
|
| 99 |
+
self.prefix = torch.tensor([self.tok(USER_HEADER, add_special_tokens=True).input_ids], device=self.device)
|
| 100 |
+
self.history = [] # (question, answer) turns about the current image
|
| 101 |
+
|
| 102 |
+
@torch.no_grad()
|
| 103 |
+
def ask(self, image, question, max_new_tokens=256, temperature=0.0):
|
| 104 |
+
pv = self.processor(images=image, return_tensors="pt").pixel_values.to(self.device, self.dtype)
|
| 105 |
+
img = self.projector(self.vision(pixel_values=pv).last_hidden_state.float()).to(self.dtype)
|
| 106 |
+
|
| 107 |
+
text = ""
|
| 108 |
+
for i, (q, a) in enumerate(self.history):
|
| 109 |
+
text += (q if i == 0 else USER_HEADER + q) + ASSIST_HEADER + a + EOT
|
| 110 |
+
text += (question if not self.history else USER_HEADER + question) + ASSIST_HEADER
|
| 111 |
+
ids = torch.tensor([self.tok(text, add_special_tokens=False).input_ids], device=self.device)
|
| 112 |
+
|
| 113 |
+
emb = self.llm.get_input_embeddings()
|
| 114 |
+
embeds = torch.cat([emb(self.prefix).to(self.dtype), img, emb(ids).to(self.dtype)], dim=1)
|
| 115 |
+
mask = torch.ones(embeds.shape[:2], dtype=torch.long, device=self.device)
|
| 116 |
+
gen = dict(max_new_tokens=max_new_tokens, repetition_penalty=1.1, eos_token_id=self.eot_id, pad_token_id=self.pad_id)
|
| 117 |
+
if temperature > 0:
|
| 118 |
+
gen.update(do_sample=True, temperature=temperature, top_p=0.9)
|
| 119 |
+
else:
|
| 120 |
+
gen.update(do_sample=False)
|
| 121 |
+
out = self.llm.generate(inputs_embeds=embeds, attention_mask=mask, **gen)
|
| 122 |
+
answer = self.tok.decode(out[0], skip_special_tokens=True).strip()
|
| 123 |
+
self.history.append((question, answer))
|
| 124 |
+
return answer
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
def main():
|
| 128 |
+
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
|
| 129 |
+
ap.add_argument("--image", required=True, help="path or URL")
|
| 130 |
+
ap.add_argument("--question", default=None)
|
| 131 |
+
ap.add_argument("--stage", type=int, default=2, choices=[1, 2])
|
| 132 |
+
ap.add_argument("--repo", default=REPO, help="HF repo id, or a local folder containing stage1/ and stage2/")
|
| 133 |
+
ap.add_argument("--llm", default=LLM)
|
| 134 |
+
ap.add_argument("--vision", default=VISION)
|
| 135 |
+
ap.add_argument("--load-in-4bit", action="store_true")
|
| 136 |
+
ap.add_argument("--max-new-tokens", type=int, default=256)
|
| 137 |
+
ap.add_argument("--temperature", type=float, default=0.0)
|
| 138 |
+
ap.add_argument("--interactive", action="store_true")
|
| 139 |
+
a = ap.parse_args()
|
| 140 |
+
|
| 141 |
+
question = a.question or ("Describe this image briefly." if a.stage == 1 else "Describe this image in detail.")
|
| 142 |
+
model = AtlasVision(a.stage, a.repo, a.llm, a.vision, a.load_in_4bit)
|
| 143 |
+
image = load_image(a.image)
|
| 144 |
+
print(f"\nQ: {question}\nA: {model.ask(image, question, a.max_new_tokens, a.temperature)}", flush=True)
|
| 145 |
+
while a.interactive:
|
| 146 |
+
try:
|
| 147 |
+
q = input("\nQ (empty to quit): ").strip()
|
| 148 |
+
except EOFError:
|
| 149 |
+
break
|
| 150 |
+
if not q:
|
| 151 |
+
break
|
| 152 |
+
print(f"A: {model.ask(image, q, a.max_new_tokens, a.temperature)}", flush=True)
|
| 153 |
+
|
| 154 |
+
|
| 155 |
+
if __name__ == "__main__":
|
| 156 |
+
sys.exit(main())
|
code/eval_stage1.py
ADDED
|
@@ -0,0 +1,85 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Stage-1 verification on samples NOT seen in training.
|
| 3 |
+
1) checkpoint integrity 2) held-out loss: random projector vs trained vs trained-with-shuffled-images
|
| 4 |
+
3) captions on held-out images 4) COCO zip readiness for stage 2"""
|
| 5 |
+
import json, os, time, zipfile
|
| 6 |
+
os.environ.setdefault("HF_HUB_OFFLINE", "1")
|
| 7 |
+
import torch
|
| 8 |
+
from torch.utils.data import DataLoader
|
| 9 |
+
import train as T
|
| 10 |
+
|
| 11 |
+
OUT = os.path.expanduser("~/eval"); os.makedirs(OUT, exist_ok=True)
|
| 12 |
+
res = {}
|
| 13 |
+
t0 = time.time()
|
| 14 |
+
|
| 15 |
+
# 1. checkpoint integrity
|
| 16 |
+
sd = torch.load(os.path.join(T.CKPT_DIR, "projector_final.pt"), map_location="cpu")
|
| 17 |
+
res["checkpoint"] = {k: list(v.shape) for k, v in sd.items()}
|
| 18 |
+
res["checkpoint_all_finite"] = all(torch.isfinite(v).all().item() for v in sd.values())
|
| 19 |
+
print("checkpoint:", res["checkpoint"], "finite:", res["checkpoint_all_finite"], flush=True)
|
| 20 |
+
|
| 21 |
+
model, tok, pad_id, proc = T.build_model()
|
| 22 |
+
ann = json.load(open(T.JSON_PATH))
|
| 23 |
+
perm = torch.randperm(len(ann), generator=torch.Generator().manual_seed(T.SEED)).tolist()
|
| 24 |
+
held = [ann[i] for i in perm[T.NUM_SAMPLES:T.NUM_SAMPLES + 2000]] # never trained on
|
| 25 |
+
prefix = T.find_zip_prefix(held, T.ZIP_PATH)
|
| 26 |
+
ds = T.LLaVAPretrainDataset(held, T.ZIP_PATH, prefix, proc, tok, T.MAX_TEXT_LEN)
|
| 27 |
+
dl = DataLoader(ds, batch_size=32, num_workers=8, collate_fn=T.make_collate(pad_id))
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
@torch.no_grad()
|
| 31 |
+
def val_loss(shuffle_images=False):
|
| 32 |
+
model.eval(); tot, n = 0.0, 0
|
| 33 |
+
for b in dl:
|
| 34 |
+
pv = b["pixel_values"].cuda()
|
| 35 |
+
if shuffle_images:
|
| 36 |
+
pv = pv.roll(1, dims=0) # each caption paired with a different image
|
| 37 |
+
with torch.autocast("cuda", dtype=T.DTYPE):
|
| 38 |
+
l = model(pv, b["input_ids"].cuda(), b["attention_mask"].cuda(), b["labels"].cuda()).float()
|
| 39 |
+
k = int((b["labels"] != -100).sum()); tot += l.item() * k; n += k
|
| 40 |
+
return round(tot / n, 4)
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
# 2. held-out loss
|
| 44 |
+
torch.manual_seed(0)
|
| 45 |
+
model.projector = T.ProjectionMLP(model.projector.net[0].in_features, model.projector.net[2].out_features).cuda()
|
| 46 |
+
res["heldout_loss_random_projector"] = val_loss()
|
| 47 |
+
model.projector.load_state_dict(sd)
|
| 48 |
+
res["heldout_loss_trained"] = val_loss()
|
| 49 |
+
res["heldout_loss_trained_shuffled_images"] = val_loss(shuffle_images=True)
|
| 50 |
+
print("held-out loss:", {k: v for k, v in res.items() if k.startswith("heldout")}, flush=True)
|
| 51 |
+
|
| 52 |
+
# 3. captions on held-out images
|
| 53 |
+
eot = tok.convert_tokens_to_ids(T.EOT)
|
| 54 |
+
caps = []
|
| 55 |
+
for item in held[:8]:
|
| 56 |
+
img, ok = None, True
|
| 57 |
+
with zipfile.ZipFile(T.ZIP_PATH) as zf, zf.open(prefix + item["image"]) as f:
|
| 58 |
+
from PIL import Image
|
| 59 |
+
img = Image.open(f).convert("RGB")
|
| 60 |
+
pv = proc(images=img, return_tensors="pt").pixel_values.cuda()
|
| 61 |
+
ids = torch.tensor([tok("Describe this image briefly." + T.ASSIST_HEADER, add_special_tokens=False).input_ids], device="cuda")
|
| 62 |
+
with torch.no_grad(), torch.autocast("cuda", dtype=T.DTYPE):
|
| 63 |
+
e, m, _ = model.build_inputs(pv, ids, torch.ones_like(ids))
|
| 64 |
+
out = model.llm.generate(inputs_embeds=e, attention_mask=m, max_new_tokens=50, do_sample=False,
|
| 65 |
+
eos_token_id=eot, pad_token_id=pad_id)
|
| 66 |
+
caps.append({"image": item["image"], "reference": item["conversations"][1]["value"],
|
| 67 |
+
"model": tok.decode(out[0], skip_special_tokens=True).strip()})
|
| 68 |
+
res["captions"] = caps
|
| 69 |
+
for c in caps:
|
| 70 |
+
print(f"\n[{c['image']}]\n ref: {c['reference']}\n model: {c['model']}", flush=True)
|
| 71 |
+
|
| 72 |
+
# 4. stage-2 data readiness
|
| 73 |
+
with zipfile.ZipFile(os.path.expanduser("~/data/coco/train2017.zip")) as zf:
|
| 74 |
+
names = set(zf.namelist())
|
| 75 |
+
inst = json.load(open(os.path.expanduser("~/data/llava_instruct/llava_instruct_150k.json")))
|
| 76 |
+
hits = sum(("train2017/" + a["image"]) in names for a in inst[:2000])
|
| 77 |
+
res["coco_zip_files"] = len(names)
|
| 78 |
+
res["instruct_conversations"] = len(inst)
|
| 79 |
+
res["instruct_images_found_of_2000"] = hits
|
| 80 |
+
print("\nstage-2 data:", res["coco_zip_files"], "files in COCO zip;", len(inst), "conversations;",
|
| 81 |
+
hits, "/2000 images found", flush=True)
|
| 82 |
+
|
| 83 |
+
res["eval_minutes"] = round((time.time() - t0) / 60, 1)
|
| 84 |
+
json.dump(res, open(os.path.join(OUT, "stage1_eval.json"), "w"), indent=1)
|
| 85 |
+
print("EVAL_DONE", flush=True)
|
code/eval_stage2.py
ADDED
|
@@ -0,0 +1,231 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Evaluate stage 1 vs stage 2 and write ~/eval/stage2_eval.json + ~/eval/report.md
|
| 3 |
+
1) held-out LLaVA-Instruct loss 2) POPE (random / popular / adversarial): acc, precision, recall, F1, yes-ratio
|
| 4 |
+
3) detailed descriptions on POPE images 4) image questions in Igbo / Yoruba / Hausa 5) text-only check of N-ATLaS
|
| 5 |
+
"""
|
| 6 |
+
import glob
|
| 7 |
+
import io
|
| 8 |
+
import json
|
| 9 |
+
import os
|
| 10 |
+
import time
|
| 11 |
+
|
| 12 |
+
os.environ.setdefault("HF_HUB_OFFLINE", "1")
|
| 13 |
+
import torch
|
| 14 |
+
from PIL import Image
|
| 15 |
+
from torch.utils.data import DataLoader, Subset
|
| 16 |
+
|
| 17 |
+
import train as T
|
| 18 |
+
import train_stage2 as S
|
| 19 |
+
|
| 20 |
+
HOME = os.path.expanduser("~")
|
| 21 |
+
OUT = T.env("EVAL_DIR", f"{HOME}/eval")
|
| 22 |
+
POPE_DIR = T.env("POPE_DIR", f"{HOME}/data/pope")
|
| 23 |
+
POPE_LIMIT = T.env("POPE_LIMIT", 0, int) # 0 = all questions per split
|
| 24 |
+
HELD_N = T.env("HELD_N", 500, int)
|
| 25 |
+
os.makedirs(OUT, exist_ok=True)
|
| 26 |
+
DEV, DT = T.DEVICE, T.DTYPE
|
| 27 |
+
t0 = time.time()
|
| 28 |
+
|
| 29 |
+
proj1 = S._load_proj(S.STAGE1_PROJECTOR)
|
| 30 |
+
proj2 = torch.load(os.path.join(S.CKPT_DIR, "projector_stage2.pt"), map_location="cpu")
|
| 31 |
+
from peft import PeftModel
|
| 32 |
+
|
| 33 |
+
model, tok, pad_id, processor = T.build_model()
|
| 34 |
+
model.llm = PeftModel.from_pretrained(model.llm, os.path.join(S.CKPT_DIR, "lora_adapter")).to(DEV)
|
| 35 |
+
model.llm.eval()
|
| 36 |
+
model.eval()
|
| 37 |
+
EOT = tok.convert_tokens_to_ids(T.EOT)
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
class Stage:
|
| 41 |
+
"""Context manager: stage 1 = stage-1 projector, adapters off; stage 2 = stage-2 projector, adapters on."""
|
| 42 |
+
def __init__(self, n):
|
| 43 |
+
self.n = n
|
| 44 |
+
|
| 45 |
+
def __enter__(self):
|
| 46 |
+
model.projector.load_state_dict(proj1 if self.n == 1 else proj2)
|
| 47 |
+
model.projector.to(DEV, dtype=torch.float32)
|
| 48 |
+
self.ctx = model.llm.disable_adapter() if self.n == 1 else None
|
| 49 |
+
if self.ctx:
|
| 50 |
+
self.ctx.__enter__()
|
| 51 |
+
|
| 52 |
+
def __exit__(self, *a):
|
| 53 |
+
if self.ctx:
|
| 54 |
+
self.ctx.__exit__(*a)
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
res = {"stage1": {}, "stage2": {}}
|
| 58 |
+
|
| 59 |
+
# ---------------- 1. held-out instruct loss ----------------
|
| 60 |
+
ann = json.load(open(S.INSTRUCT_JSON))
|
| 61 |
+
_, held = S.split(ann)
|
| 62 |
+
prefix = T.find_zip_prefix([ann[i] for i in held[:200]], S.COCO_ZIP)
|
| 63 |
+
ds = S.InstructDataset(ann, S.COCO_ZIP, prefix, processor, tok, S.MAX_TEXT_LEN)
|
| 64 |
+
dl = DataLoader(Subset(ds, held[:HELD_N]), batch_size=8, num_workers=T.NUM_WORKERS,
|
| 65 |
+
collate_fn=T.make_collate(pad_id))
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
@torch.no_grad()
|
| 69 |
+
def heldout_loss():
|
| 70 |
+
tot, n = 0.0, 0
|
| 71 |
+
for b in dl:
|
| 72 |
+
with torch.autocast(device_type=DEV.type, dtype=DT):
|
| 73 |
+
l = model(b["pixel_values"].to(DEV), b["input_ids"].to(DEV), b["attention_mask"].to(DEV),
|
| 74 |
+
b["labels"].to(DEV)).float()
|
| 75 |
+
k = int((b["labels"] != -100).sum())
|
| 76 |
+
tot, n = tot + l.item() * k, n + k
|
| 77 |
+
return round(tot / n, 4)
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
for s in (1, 2):
|
| 81 |
+
with Stage(s):
|
| 82 |
+
res[f"stage{s}"]["heldout_instruct_loss"] = heldout_loss()
|
| 83 |
+
T.log(f"held-out instruct loss: stage1 {res['stage1']['heldout_instruct_loss']} | "
|
| 84 |
+
f"stage2 {res['stage2']['heldout_instruct_loss']}")
|
| 85 |
+
|
| 86 |
+
# ---------------- 2. POPE ----------------
|
| 87 |
+
import pyarrow.parquet as pq
|
| 88 |
+
|
| 89 |
+
SUFFIX = " Answer the question using a single word or phrase."
|
| 90 |
+
yes_ids = sorted({tok(w, add_special_tokens=False).input_ids[0] for w in ("Yes", "yes", " Yes", " yes")})
|
| 91 |
+
no_ids = sorted({tok(w, add_special_tokens=False).input_ids[0] for w in ("No", "no", " No", " no")})
|
| 92 |
+
|
| 93 |
+
|
| 94 |
+
def to_image(cell):
|
| 95 |
+
if isinstance(cell, dict):
|
| 96 |
+
cell = cell.get("bytes") or open(cell["path"], "rb").read()
|
| 97 |
+
return Image.open(io.BytesIO(cell)).convert("RGB")
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
def pope_rows(split):
|
| 101 |
+
files = sorted(glob.glob(f"{POPE_DIR}/**/{split}-*.parquet", recursive=True))
|
| 102 |
+
if files:
|
| 103 |
+
rows = pq.read_table(files[0]).to_pylist()
|
| 104 |
+
else: # some versions ship one 'test' table with a 'category' column
|
| 105 |
+
rows = [r for f in sorted(glob.glob(f"{POPE_DIR}/**/test-*.parquet", recursive=True))
|
| 106 |
+
for r in pq.read_table(f).to_pylist() if r.get("category") == split]
|
| 107 |
+
return rows[:POPE_LIMIT] if POPE_LIMIT else rows
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
@torch.no_grad()
|
| 111 |
+
def pope_eval(rows, bs=32):
|
| 112 |
+
tp = fp = tn = fn = 0
|
| 113 |
+
for s in range(0, len(rows), bs):
|
| 114 |
+
chunk = rows[s:s + bs]
|
| 115 |
+
pv = torch.stack([processor(images=to_image(r["image"]), return_tensors="pt").pixel_values[0] for r in chunk]).to(DEV)
|
| 116 |
+
seqs = [tok(r["question"].strip() + SUFFIX + T.ASSIST_HEADER, add_special_tokens=False).input_ids for r in chunk]
|
| 117 |
+
L = max(map(len, seqs))
|
| 118 |
+
ids = torch.full((len(seqs), L), pad_id, dtype=torch.long)
|
| 119 |
+
mask = torch.zeros_like(ids)
|
| 120 |
+
for i, q in enumerate(seqs):
|
| 121 |
+
ids[i, :len(q)] = torch.tensor(q)
|
| 122 |
+
mask[i, :len(q)] = 1
|
| 123 |
+
ids, mask = ids.to(DEV), mask.to(DEV)
|
| 124 |
+
with torch.autocast(device_type=DEV.type, dtype=DT):
|
| 125 |
+
e, m, _ = model.build_inputs(pv, ids, mask)
|
| 126 |
+
logits = model.llm(inputs_embeds=e, attention_mask=m).logits
|
| 127 |
+
n_fixed = e.shape[1] - L
|
| 128 |
+
for i, r in enumerate(chunk):
|
| 129 |
+
last = logits[i, n_fixed + len(seqs[i]) - 1].float()
|
| 130 |
+
pred_yes = last[yes_ids].max() > last[no_ids].max()
|
| 131 |
+
gold_yes = str(r["answer"]).strip().lower().startswith("yes")
|
| 132 |
+
tp += pred_yes and gold_yes
|
| 133 |
+
fp += pred_yes and not gold_yes
|
| 134 |
+
tn += (not pred_yes) and (not gold_yes)
|
| 135 |
+
fn += (not pred_yes) and gold_yes
|
| 136 |
+
tp, fp, tn, fn = map(int, (tp, fp, tn, fn))
|
| 137 |
+
n = tp + fp + tn + fn
|
| 138 |
+
prec = tp / max(1, tp + fp)
|
| 139 |
+
rec = tp / max(1, tp + fn)
|
| 140 |
+
return {"n": n, "accuracy": round((tp + tn) / max(1, n), 4), "precision": round(prec, 4),
|
| 141 |
+
"recall": round(rec, 4), "f1": round(2 * prec * rec / max(1e-9, prec + rec), 4),
|
| 142 |
+
"yes_ratio": round((tp + fp) / max(1, n), 4)}
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
pope_samples = []
|
| 146 |
+
for split in ("random", "popular", "adversarial"):
|
| 147 |
+
rows = pope_rows(split)
|
| 148 |
+
if not rows:
|
| 149 |
+
T.log(f"POPE {split}: no data found")
|
| 150 |
+
continue
|
| 151 |
+
pope_samples = pope_samples or rows
|
| 152 |
+
for s in (1, 2):
|
| 153 |
+
with Stage(s):
|
| 154 |
+
res[f"stage{s}"][f"pope_{split}"] = pope_eval(rows)
|
| 155 |
+
T.log(f"POPE {split}: stage1 {res['stage1'][f'pope_{split}']} | stage2 {res['stage2'][f'pope_{split}']}")
|
| 156 |
+
|
| 157 |
+
|
| 158 |
+
# ---------------- 3/4. generations ----------------
|
| 159 |
+
@torch.no_grad()
|
| 160 |
+
def answer(image, question, max_new_tokens=120):
|
| 161 |
+
pv = processor(images=image, return_tensors="pt").pixel_values.to(DEV)
|
| 162 |
+
ids = torch.tensor([tok(question + T.ASSIST_HEADER, add_special_tokens=False).input_ids], device=DEV)
|
| 163 |
+
with torch.autocast(device_type=DEV.type, dtype=DT):
|
| 164 |
+
e, m, _ = model.build_inputs(pv, ids, torch.ones_like(ids))
|
| 165 |
+
out = model.llm.generate(inputs_embeds=e, attention_mask=m, max_new_tokens=max_new_tokens,
|
| 166 |
+
do_sample=False, eos_token_id=EOT, pad_token_id=pad_id, repetition_penalty=1.1)
|
| 167 |
+
return tok.decode(out[0], skip_special_tokens=True).strip()
|
| 168 |
+
|
| 169 |
+
|
| 170 |
+
seen, gen_images = set(), []
|
| 171 |
+
for r in pope_samples:
|
| 172 |
+
key = r.get("image_source") or r.get("question_id")
|
| 173 |
+
if key not in seen:
|
| 174 |
+
seen.add(key)
|
| 175 |
+
gen_images.append((str(key), to_image(r["image"])))
|
| 176 |
+
if len(gen_images) == 4:
|
| 177 |
+
break
|
| 178 |
+
|
| 179 |
+
res["descriptions"] = []
|
| 180 |
+
for key, img in gen_images:
|
| 181 |
+
row = {"image": key}
|
| 182 |
+
for s in (1, 2):
|
| 183 |
+
with Stage(s):
|
| 184 |
+
row[f"stage{s}"] = answer(img, "Describe this image in detail.")
|
| 185 |
+
res["descriptions"].append(row)
|
| 186 |
+
|
| 187 |
+
MULTI = {"igbo": "Kedu ihe dị na foto a?", "yoruba": "Kí ni ó wà nínú àwòrán yìí?", "hausa": "Me ke cikin wannan hoton?"}
|
| 188 |
+
res["multilingual"] = []
|
| 189 |
+
if gen_images:
|
| 190 |
+
with Stage(2):
|
| 191 |
+
for lang, q in MULTI.items():
|
| 192 |
+
res["multilingual"].append({"image": gen_images[0][0], "language": lang, "question": q,
|
| 193 |
+
"stage2": answer(gen_images[0][1], q)})
|
| 194 |
+
|
| 195 |
+
# ---------------- 5. text-only check ----------------
|
| 196 |
+
@torch.no_grad()
|
| 197 |
+
def text_only(q, max_new_tokens=120):
|
| 198 |
+
ids = tok(T.USER_HEADER + q + T.ASSIST_HEADER, add_special_tokens=True, return_tensors="pt").input_ids.to(DEV)
|
| 199 |
+
with torch.autocast(device_type=DEV.type, dtype=DT):
|
| 200 |
+
out = model.llm.generate(input_ids=ids, attention_mask=torch.ones_like(ids), max_new_tokens=max_new_tokens,
|
| 201 |
+
do_sample=False, eos_token_id=EOT, pad_token_id=pad_id, repetition_penalty=1.1)
|
| 202 |
+
return tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True).strip()
|
| 203 |
+
|
| 204 |
+
|
| 205 |
+
q_text = "Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo."
|
| 206 |
+
with Stage(1):
|
| 207 |
+
base_answer = text_only(q_text)
|
| 208 |
+
with Stage(2):
|
| 209 |
+
lora_answer = text_only(q_text)
|
| 210 |
+
res["text_only"] = {"question": q_text, "base_n_atlas": base_answer, "with_stage2_lora": lora_answer}
|
| 211 |
+
res["eval_minutes"] = round((time.time() - t0) / 60, 1)
|
| 212 |
+
json.dump(res, open(os.path.join(OUT, "stage2_eval.json"), "w"), indent=1, ensure_ascii=False)
|
| 213 |
+
|
| 214 |
+
# ---------------- report ----------------
|
| 215 |
+
L = ["# Atlas-Vision evaluation\n", "## Scores (stage 1 → stage 2)\n", "| Metric | Stage 1 | Stage 2 |", "|---|---|---|",
|
| 216 |
+
f"| Held-out instruct loss (lower is better) | {res['stage1']['heldout_instruct_loss']} | {res['stage2']['heldout_instruct_loss']} |"]
|
| 217 |
+
for split in ("random", "popular", "adversarial"):
|
| 218 |
+
if f"pope_{split}" in res["stage1"]:
|
| 219 |
+
a, b = res["stage1"][f"pope_{split}"], res["stage2"][f"pope_{split}"]
|
| 220 |
+
L.append(f"| POPE {split}: accuracy / F1 / yes-ratio | {a['accuracy']} / {a['f1']} / {a['yes_ratio']} | "
|
| 221 |
+
f"{b['accuracy']} / {b['f1']} / {b['yes_ratio']} |")
|
| 222 |
+
L.append("\n## Detailed descriptions\n")
|
| 223 |
+
for d in res["descriptions"]:
|
| 224 |
+
L += [f"**{d['image']}**\n", f"- Stage 1: {d['stage1']}", f"- Stage 2: {d['stage2']}\n"]
|
| 225 |
+
L.append("## Questions in Nigerian languages (stage 2)\n")
|
| 226 |
+
for m in res["multilingual"]:
|
| 227 |
+
L.append(f"- **{m['language']}** — {m['question']}\n → {m['stage2']}")
|
| 228 |
+
L += ["\n## Text-only check (no image)\n", f"Question: {q_text}\n",
|
| 229 |
+
f"- Base N-ATLaS: {base_answer}", f"- With stage-2 LoRA: {lora_answer}"]
|
| 230 |
+
open(os.path.join(OUT, "report.md"), "w").write("\n".join(L) + "\n")
|
| 231 |
+
T.log(f"EVAL_DONE in {res['eval_minutes']} min -> {OUT}/stage2_eval.json, {OUT}/report.md")
|
code/infer.py
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Caption images with the trained projector.
|
| 3 |
+
Usage: python3 infer.py [projector.pt] [image paths...] (default: 8 held-out images from the zip)"""
|
| 4 |
+
import json
|
| 5 |
+
import os
|
| 6 |
+
import sys
|
| 7 |
+
import zipfile
|
| 8 |
+
|
| 9 |
+
import torch
|
| 10 |
+
from PIL import Image
|
| 11 |
+
|
| 12 |
+
os.environ.setdefault("HF_HUB_OFFLINE", "1")
|
| 13 |
+
import train as T # reuses the exact model layout and prompt format used in training
|
| 14 |
+
|
| 15 |
+
ckpt = sys.argv[1] if len(sys.argv) > 1 else os.path.join(T.CKPT_DIR, "projector_final.pt")
|
| 16 |
+
model, tok, pad_id, processor = T.build_model()
|
| 17 |
+
state = torch.load(ckpt, map_location="cpu")
|
| 18 |
+
state = state.get("projector_state_dict", state)
|
| 19 |
+
model.projector.load_state_dict(state)
|
| 20 |
+
model.eval()
|
| 21 |
+
|
| 22 |
+
if len(sys.argv) > 2:
|
| 23 |
+
images = [(p, Image.open(p).convert("RGB")) for p in sys.argv[2:]]
|
| 24 |
+
else: # images NOT in the training subset
|
| 25 |
+
ann = json.load(open(T.JSON_PATH))
|
| 26 |
+
g = torch.Generator().manual_seed(T.SEED)
|
| 27 |
+
held_out = torch.randperm(len(ann), generator=g)[T.NUM_SAMPLES:T.NUM_SAMPLES + 8].tolist()
|
| 28 |
+
zf = zipfile.ZipFile(T.ZIP_PATH)
|
| 29 |
+
prefix = T.find_zip_prefix([ann[i] for i in held_out], T.ZIP_PATH, n_check=8)
|
| 30 |
+
images = []
|
| 31 |
+
for i in held_out:
|
| 32 |
+
with zf.open(prefix + ann[i]["image"]) as f:
|
| 33 |
+
images.append((f"{ann[i]['image']} | reference: {ann[i]['conversations'][1]['value']}",
|
| 34 |
+
Image.open(f).convert("RGB")))
|
| 35 |
+
|
| 36 |
+
prompt = "Describe this image briefly."
|
| 37 |
+
for name, img in images:
|
| 38 |
+
pv = processor(images=img, return_tensors="pt").pixel_values.to(T.DEVICE)
|
| 39 |
+
ids = torch.tensor([tok(prompt + T.ASSIST_HEADER, add_special_tokens=False).input_ids], device=T.DEVICE)
|
| 40 |
+
mask = torch.ones_like(ids)
|
| 41 |
+
with torch.no_grad(), torch.autocast(device_type=T.DEVICE.type, dtype=T.DTYPE):
|
| 42 |
+
embeds, full_mask, _ = model.build_inputs(pv, ids, mask)
|
| 43 |
+
out = model.llm.generate(inputs_embeds=embeds, attention_mask=full_mask, max_new_tokens=60,
|
| 44 |
+
do_sample=False, eos_token_id=tok.convert_tokens_to_ids(T.EOT), pad_token_id=pad_id)
|
| 45 |
+
print(f"\n{name}\n -> {tok.decode(out[0], skip_special_tokens=True).strip()}", flush=True)
|
code/prepare.py
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Download everything onto the node's persistent disk. Safe to re-run (skips finished files).
|
| 3 |
+
Needs HF_TOKEN in the environment for the gated N-ATLaS repo."""
|
| 4 |
+
import os
|
| 5 |
+
import sys
|
| 6 |
+
import time
|
| 7 |
+
|
| 8 |
+
from huggingface_hub import hf_hub_download, snapshot_download
|
| 9 |
+
|
| 10 |
+
HOME = os.path.expanduser("~")
|
| 11 |
+
DATA_DIR = f"{HOME}/data/llava_pretrain"
|
| 12 |
+
LLM_NAME = os.environ.get("LLM_NAME", "NCAIR1/N-ATLaS")
|
| 13 |
+
VISION_NAME = os.environ.get("VISION_NAME", "google/siglip2-base-patch16-224")
|
| 14 |
+
|
| 15 |
+
|
| 16 |
+
def log(msg):
|
| 17 |
+
print(f"[{time.strftime('%H:%M:%S')}] {msg}", flush=True)
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
if not os.environ.get("HF_TOKEN"):
|
| 21 |
+
sys.exit("HF_TOKEN is not set - needed for the gated N-ATLaS repo")
|
| 22 |
+
|
| 23 |
+
# 1. Fail fast if the token can't see the gated model
|
| 24 |
+
try:
|
| 25 |
+
hf_hub_download(LLM_NAME, "config.json")
|
| 26 |
+
log(f"gated access OK for {LLM_NAME}")
|
| 27 |
+
except Exception as e:
|
| 28 |
+
sys.exit(f"Cannot access {LLM_NAME}: {type(e).__name__}: {e}\n"
|
| 29 |
+
"Check that access was granted on the model page and the token can read gated repos.")
|
| 30 |
+
|
| 31 |
+
t = time.time()
|
| 32 |
+
snapshot_download(LLM_NAME, allow_patterns=["*.json", "*.safetensors", "tokenizer*", "*.txt", "*.model", "*.jinja"])
|
| 33 |
+
log(f"LLM downloaded ({time.time() - t:.0f}s)")
|
| 34 |
+
|
| 35 |
+
t = time.time()
|
| 36 |
+
snapshot_download(VISION_NAME)
|
| 37 |
+
log(f"vision encoder downloaded ({time.time() - t:.0f}s)")
|
| 38 |
+
|
| 39 |
+
t = time.time()
|
| 40 |
+
os.makedirs(DATA_DIR, exist_ok=True)
|
| 41 |
+
hf_hub_download("liuhaotian/LLaVA-Pretrain", "blip_laion_cc_sbu_558k.json", repo_type="dataset", local_dir=DATA_DIR)
|
| 42 |
+
hf_hub_download("liuhaotian/LLaVA-Pretrain", "images.zip", repo_type="dataset", local_dir=DATA_DIR)
|
| 43 |
+
size_gb = os.path.getsize(f"{DATA_DIR}/images.zip") / 1e9
|
| 44 |
+
log(f"dataset downloaded: images.zip {size_gb:.1f} GB ({time.time() - t:.0f}s)")
|
code/run.sh
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/bin/bash
|
| 2 |
+
# Idempotent: safe to re-run after a stop/wake. Resumes training from ~/checkpoints/stage1/latest.pt.
|
| 3 |
+
set -uo pipefail
|
| 4 |
+
cd "$HOME/atlas"
|
| 5 |
+
export PYTHONUNBUFFERED=1
|
| 6 |
+
export HF_HOME="$HOME/.cache/huggingface"
|
| 7 |
+
export HF_HUB_ENABLE_HF_TRANSFER=1
|
| 8 |
+
|
| 9 |
+
echo "=== [1/4] python packages ==="
|
| 10 |
+
PIPFLAGS="--user --break-system-packages"
|
| 11 |
+
python3 -c "import sys; sys.exit(0 if sys.prefix != sys.base_prefix else 1)" && PIPFLAGS=""
|
| 12 |
+
python3 -m pip install -q $PIPFLAGS -U "transformers>=4.56" accelerate "huggingface_hub>=0.34" hf_transfer pillow sentencepiece protobuf || exit 10
|
| 13 |
+
export PATH="$HOME/.local/bin:$PATH"
|
| 14 |
+
python3 -c "import torch, transformers; print('torch', torch.__version__, '| transformers', transformers.__version__, '| cuda', torch.cuda.is_available())"
|
| 15 |
+
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv
|
| 16 |
+
|
| 17 |
+
echo "=== [2/4] downloads ==="
|
| 18 |
+
python3 prepare.py || exit 11
|
| 19 |
+
df -h "$HOME" | tail -1
|
| 20 |
+
|
| 21 |
+
export HF_HUB_OFFLINE=1 # everything is cached now; training never needs the token
|
| 22 |
+
GC="${GRAD_CKPT:-0}"
|
| 23 |
+
if [ "${SKIP_SMOKE:-0}" != "1" ]; then
|
| 24 |
+
echo "=== [3/4] smoke test (20 steps) ==="
|
| 25 |
+
rm -rf "$HOME/checkpoints/smoke"
|
| 26 |
+
MAX_STEPS=20 LOG_EVERY=5 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/smoke" GRAD_CKPT=$GC python3 train.py
|
| 27 |
+
rc=$?
|
| 28 |
+
if [ $rc -eq 3 ] && [ "$GC" = "0" ]; then
|
| 29 |
+
echo "OOM without gradient checkpointing -> retrying with GRAD_CKPT=1"
|
| 30 |
+
GC=1
|
| 31 |
+
rm -rf "$HOME/checkpoints/smoke"
|
| 32 |
+
MAX_STEPS=20 LOG_EVERY=5 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/smoke" GRAD_CKPT=1 python3 train.py
|
| 33 |
+
rc=$?
|
| 34 |
+
fi
|
| 35 |
+
if [ $rc -ne 0 ]; then echo "SMOKE TEST FAILED (exit $rc)"; exit 12; fi
|
| 36 |
+
fi
|
| 37 |
+
|
| 38 |
+
echo "=== [4/4] full stage-1 run (GRAD_CKPT=$GC) ==="
|
| 39 |
+
GRAD_CKPT=$GC python3 train.py
|
| 40 |
+
rc=$?
|
| 41 |
+
if [ $rc -eq 3 ] && [ "$GC" = "0" ]; then
|
| 42 |
+
echo "OOM in full run -> resuming with GRAD_CKPT=1"
|
| 43 |
+
GRAD_CKPT=1 python3 train.py
|
| 44 |
+
rc=$?
|
| 45 |
+
fi
|
| 46 |
+
echo "=== run.sh finished (exit $rc) ==="
|
| 47 |
+
exit $rc
|
code/run_stage2.sh
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/bin/bash
|
| 2 |
+
# Stage 2: smoke test -> full LoRA instruction tuning -> evaluation. Safe to re-run (resumes).
|
| 3 |
+
set -uo pipefail
|
| 4 |
+
cd "$HOME/atlas"
|
| 5 |
+
export PYTHONUNBUFFERED=1 HF_HOME="$HOME/.cache/huggingface" HF_HUB_OFFLINE=1 PATH="$HOME/.local/bin:$PATH"
|
| 6 |
+
python3 -m pip install -q --user --break-system-packages peft pyarrow || exit 10
|
| 7 |
+
python3 -c "import peft, pyarrow; print('peft', peft.__version__, '| pyarrow', pyarrow.__version__)"
|
| 8 |
+
BS="${BATCH_SIZE:-8}"; GA="${GRAD_ACCUM:-4}"
|
| 9 |
+
|
| 10 |
+
if [ "${SKIP_SMOKE:-0}" != "1" ]; then
|
| 11 |
+
echo "=== [1/3] smoke test (10 steps) ==="
|
| 12 |
+
rm -rf "$HOME/checkpoints/stage2_smoke"
|
| 13 |
+
MAX_STEPS=10 LOG_EVERY=2 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/stage2_smoke" BATCH_SIZE=$BS GRAD_ACCUM=$GA python3 train_stage2.py
|
| 14 |
+
rc=$?
|
| 15 |
+
if [ $rc -eq 3 ]; then
|
| 16 |
+
echo "OOM -> retrying with BATCH_SIZE=4 GRAD_ACCUM=8"; BS=4; GA=8
|
| 17 |
+
rm -rf "$HOME/checkpoints/stage2_smoke"
|
| 18 |
+
MAX_STEPS=10 LOG_EVERY=2 SAVE_EVERY=1000000 CKPT_DIR="$HOME/checkpoints/stage2_smoke" BATCH_SIZE=$BS GRAD_ACCUM=$GA python3 train_stage2.py
|
| 19 |
+
rc=$?
|
| 20 |
+
fi
|
| 21 |
+
if [ $rc -ne 0 ]; then echo "SMOKE TEST FAILED (exit $rc)"; exit 12; fi
|
| 22 |
+
fi
|
| 23 |
+
|
| 24 |
+
echo "=== [2/3] full stage-2 run (BATCH_SIZE=$BS GRAD_ACCUM=$GA) ==="
|
| 25 |
+
BATCH_SIZE=$BS GRAD_ACCUM=$GA python3 train_stage2.py
|
| 26 |
+
rc=$?
|
| 27 |
+
if [ $rc -eq 3 ] && [ "$BS" -gt 4 ]; then
|
| 28 |
+
echo "OOM in full run -> continuing with BATCH_SIZE=4 GRAD_ACCUM=8"
|
| 29 |
+
BATCH_SIZE=4 GRAD_ACCUM=8 python3 train_stage2.py
|
| 30 |
+
rc=$?
|
| 31 |
+
fi
|
| 32 |
+
if [ $rc -ne 0 ]; then echo "TRAINING FAILED (exit $rc)"; exit $rc; fi
|
| 33 |
+
|
| 34 |
+
echo "=== [3/3] evaluation ==="
|
| 35 |
+
python3 eval_stage2.py
|
| 36 |
+
rc=$?
|
| 37 |
+
echo "=== run_stage2.sh finished (exit $rc) ==="
|
| 38 |
+
exit $rc
|
code/train.py
ADDED
|
@@ -0,0 +1,384 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Stage 1 (projector alignment): SigLIP2 vision tower + N-ATLaS LLM, only the MLP projector trains.
|
| 3 |
+
|
| 4 |
+
Fixes vs. the Colab notebook:
|
| 5 |
+
* LLM loaded in bf16 (notebook loaded fp16), bf16 autocast, fp32 projector, no GradScaler
|
| 6 |
+
* dynamic padding + attention mask (notebook padded every sample to 512 tokens; captions are ~10-25)
|
| 7 |
+
* image tokens placed after <BOS><user header>, not before BOS
|
| 8 |
+
* fixed random subset, cosine LR schedule, resume by samples seen (instant, batch-size independent)
|
| 9 |
+
* zip path check so missing images fail loudly instead of silently training on white images
|
| 10 |
+
All settings come from environment variables (see CONFIG below).
|
| 11 |
+
"""
|
| 12 |
+
import json
|
| 13 |
+
import math
|
| 14 |
+
import os
|
| 15 |
+
import sys
|
| 16 |
+
import time
|
| 17 |
+
import zipfile
|
| 18 |
+
|
| 19 |
+
import torch
|
| 20 |
+
import torch.nn as nn
|
| 21 |
+
from PIL import Image
|
| 22 |
+
from torch.utils.data import DataLoader, Dataset, Subset
|
| 23 |
+
from transformers import AutoImageProcessor, AutoModel, AutoModelForCausalLM, AutoTokenizer
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
def env(name, default, cast=str):
|
| 27 |
+
v = os.environ.get(name)
|
| 28 |
+
return cast(v) if v not in (None, "") else default
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
# ----------------------------- CONFIG -----------------------------
|
| 32 |
+
HOME = os.path.expanduser("~")
|
| 33 |
+
LLM_NAME = env("LLM_NAME", "NCAIR1/N-ATLaS")
|
| 34 |
+
VISION_NAME = env("VISION_NAME", "google/siglip2-base-patch16-224")
|
| 35 |
+
DATA_DIR = env("DATA_DIR", f"{HOME}/data/llava_pretrain")
|
| 36 |
+
JSON_PATH = env("JSON_PATH", f"{DATA_DIR}/blip_laion_cc_sbu_558k.json")
|
| 37 |
+
ZIP_PATH = env("ZIP_PATH", f"{DATA_DIR}/images.zip")
|
| 38 |
+
CKPT_DIR = env("CKPT_DIR", f"{HOME}/checkpoints/stage1")
|
| 39 |
+
NUM_SAMPLES = env("NUM_SAMPLES", 150_000, int)
|
| 40 |
+
BATCH_SIZE = env("BATCH_SIZE", 16, int)
|
| 41 |
+
GRAD_ACCUM = env("GRAD_ACCUM", 1, int)
|
| 42 |
+
LR = env("LR", 2e-4, float)
|
| 43 |
+
WARMUP_RATIO = env("WARMUP_RATIO", 0.03, float)
|
| 44 |
+
MAX_TEXT_LEN = env("MAX_TEXT_LEN", 128, int)
|
| 45 |
+
MAX_STEPS = env("MAX_STEPS", 0, int) # >0 caps optimizer steps (smoke tests)
|
| 46 |
+
GRAD_CKPT = env("GRAD_CKPT", 0, int)
|
| 47 |
+
LOG_EVERY = env("LOG_EVERY", 25, int)
|
| 48 |
+
SAVE_EVERY = env("SAVE_EVERY", 500, int)
|
| 49 |
+
SEED = env("SEED", 42, int)
|
| 50 |
+
NUM_WORKERS = env("NUM_WORKERS", max(1, int(float(os.environ.get("GMN_CPU_LIMIT", "8"))) - 2), int)
|
| 51 |
+
DEVICE = torch.device(env("DEVICE", "cuda" if torch.cuda.is_available() else "cpu"))
|
| 52 |
+
DTYPE = torch.bfloat16
|
| 53 |
+
|
| 54 |
+
USER_HEADER = "<|start_header_id|>user<|end_header_id|>\n\n"
|
| 55 |
+
ASSIST_HEADER = "<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"
|
| 56 |
+
EOT = "<|eot_id|>"
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
def log(msg):
|
| 60 |
+
print(f"[{time.strftime('%H:%M:%S')}] {msg}", flush=True)
|
| 61 |
+
|
| 62 |
+
|
| 63 |
+
# ----------------------------- DATA -----------------------------
|
| 64 |
+
def find_zip_prefix(annotations, zip_path, n_check=500):
|
| 65 |
+
"""Return the prefix to put before item['image'] inside the zip, verifying images exist."""
|
| 66 |
+
with zipfile.ZipFile(zip_path) as zf:
|
| 67 |
+
names = set(zf.namelist())
|
| 68 |
+
sample = [a["image"] for a in annotations[:n_check]]
|
| 69 |
+
prefix = ""
|
| 70 |
+
if sum(s in names for s in sample) < 0.95 * len(sample):
|
| 71 |
+
first = sample[0]
|
| 72 |
+
hits = [n for n in names if n.endswith("/" + first)]
|
| 73 |
+
if hits:
|
| 74 |
+
prefix = hits[0][: -len(first)]
|
| 75 |
+
found = sum((prefix + s) in names for s in sample)
|
| 76 |
+
log(f"zip check: {found}/{len(sample)} images found (prefix={prefix!r}, {len(names):,} files in zip)")
|
| 77 |
+
if found < 0.95 * len(sample):
|
| 78 |
+
raise SystemExit(f"Too many images missing from {zip_path}; example: {sample[0]!r}")
|
| 79 |
+
return prefix
|
| 80 |
+
|
| 81 |
+
|
| 82 |
+
class LLaVAPretrainDataset(Dataset):
|
| 83 |
+
def __init__(self, annotations, zip_path, zip_prefix, processor, tokenizer, max_text_len):
|
| 84 |
+
self.ann = annotations
|
| 85 |
+
self.zip_path = zip_path
|
| 86 |
+
self.zip_prefix = zip_prefix
|
| 87 |
+
self.processor = processor
|
| 88 |
+
self.tok = tokenizer
|
| 89 |
+
self.max_len = max_text_len
|
| 90 |
+
self.zip = None # opened lazily, once per dataloader worker
|
| 91 |
+
|
| 92 |
+
def __len__(self):
|
| 93 |
+
return len(self.ann)
|
| 94 |
+
|
| 95 |
+
def __getitem__(self, idx):
|
| 96 |
+
if self.zip is None:
|
| 97 |
+
self.zip = zipfile.ZipFile(self.zip_path)
|
| 98 |
+
item = self.ann[idx]
|
| 99 |
+
user, answer = "", ""
|
| 100 |
+
for turn in item["conversations"]:
|
| 101 |
+
if turn["from"] == "human":
|
| 102 |
+
user = turn["value"].replace("<image>", "").strip()
|
| 103 |
+
elif turn["from"] == "gpt":
|
| 104 |
+
answer = turn["value"].strip()
|
| 105 |
+
ok = True
|
| 106 |
+
try:
|
| 107 |
+
with self.zip.open(self.zip_prefix + item["image"]) as f:
|
| 108 |
+
image = Image.open(f).convert("RGB")
|
| 109 |
+
except Exception:
|
| 110 |
+
image = Image.new("RGB", (224, 224), "white")
|
| 111 |
+
ok = False
|
| 112 |
+
pixel_values = self.processor(images=image, return_tensors="pt").pixel_values[0]
|
| 113 |
+
|
| 114 |
+
q = self.tok(user + ASSIST_HEADER, add_special_tokens=False).input_ids
|
| 115 |
+
a = self.tok(answer + EOT, add_special_tokens=False).input_ids
|
| 116 |
+
ids = (q + a)[: self.max_len]
|
| 117 |
+
labels = ([-100] * len(q) + a)[: self.max_len]
|
| 118 |
+
return {"pixel_values": pixel_values, "input_ids": ids, "labels": labels, "ok": ok}
|
| 119 |
+
|
| 120 |
+
|
| 121 |
+
def make_collate(pad_id):
|
| 122 |
+
def collate(batch):
|
| 123 |
+
L = max(len(b["input_ids"]) for b in batch)
|
| 124 |
+
B = len(batch)
|
| 125 |
+
ids = torch.full((B, L), pad_id, dtype=torch.long)
|
| 126 |
+
labels = torch.full((B, L), -100, dtype=torch.long)
|
| 127 |
+
mask = torch.zeros((B, L), dtype=torch.long)
|
| 128 |
+
for i, b in enumerate(batch):
|
| 129 |
+
n = len(b["input_ids"])
|
| 130 |
+
ids[i, :n] = torch.tensor(b["input_ids"])
|
| 131 |
+
labels[i, :n] = torch.tensor(b["labels"])
|
| 132 |
+
mask[i, :n] = 1
|
| 133 |
+
return {
|
| 134 |
+
"pixel_values": torch.stack([b["pixel_values"] for b in batch]),
|
| 135 |
+
"input_ids": ids,
|
| 136 |
+
"labels": labels,
|
| 137 |
+
"attention_mask": mask,
|
| 138 |
+
"n_bad": sum(not b["ok"] for b in batch),
|
| 139 |
+
}
|
| 140 |
+
return collate
|
| 141 |
+
|
| 142 |
+
|
| 143 |
+
# ----------------------------- MODEL -----------------------------
|
| 144 |
+
class ProjectionMLP(nn.Module):
|
| 145 |
+
"""Same layout as the notebook (net.0 / net.2), so Colab checkpoints load."""
|
| 146 |
+
def __init__(self, vision_dim, text_dim):
|
| 147 |
+
super().__init__()
|
| 148 |
+
self.net = nn.Sequential(nn.Linear(vision_dim, text_dim), nn.GELU(), nn.Linear(text_dim, text_dim))
|
| 149 |
+
|
| 150 |
+
def forward(self, x):
|
| 151 |
+
return self.net(x)
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
class AtlasVision(nn.Module):
|
| 155 |
+
"""[<BOS> user-header] [image tokens] [user text, assistant header, answer]"""
|
| 156 |
+
def __init__(self, vision, llm, prefix_ids, vision_dim, text_dim):
|
| 157 |
+
super().__init__()
|
| 158 |
+
self.vision = vision
|
| 159 |
+
self.llm = llm
|
| 160 |
+
self.projector = ProjectionMLP(vision_dim, text_dim)
|
| 161 |
+
self.register_buffer("prefix_ids", torch.tensor(prefix_ids, dtype=torch.long)[None], persistent=False)
|
| 162 |
+
|
| 163 |
+
def build_inputs(self, pixel_values, input_ids, attention_mask, labels=None):
|
| 164 |
+
B = input_ids.size(0)
|
| 165 |
+
with torch.no_grad():
|
| 166 |
+
feats = self.vision(pixel_values=pixel_values.to(DTYPE)).last_hidden_state
|
| 167 |
+
img = self.projector(feats.float()).to(DTYPE)
|
| 168 |
+
emb = self.llm.get_input_embeddings()
|
| 169 |
+
pre = emb(self.prefix_ids.expand(B, -1)).to(DTYPE)
|
| 170 |
+
txt = emb(input_ids).to(DTYPE)
|
| 171 |
+
n_fixed = pre.size(1) + img.size(1)
|
| 172 |
+
inputs_embeds = torch.cat([pre, img, txt], dim=1)
|
| 173 |
+
mask = torch.cat([attention_mask.new_ones(B, n_fixed), attention_mask], dim=1)
|
| 174 |
+
full_labels = None
|
| 175 |
+
if labels is not None:
|
| 176 |
+
full_labels = torch.cat([labels.new_full((B, n_fixed), -100), labels], dim=1)
|
| 177 |
+
return inputs_embeds, mask, full_labels
|
| 178 |
+
|
| 179 |
+
def forward(self, pixel_values, input_ids, attention_mask, labels):
|
| 180 |
+
inputs_embeds, mask, full_labels = self.build_inputs(pixel_values, input_ids, attention_mask, labels)
|
| 181 |
+
return self.llm(inputs_embeds=inputs_embeds, attention_mask=mask, labels=full_labels, use_cache=False).loss
|
| 182 |
+
|
| 183 |
+
|
| 184 |
+
def load_pretrained(cls, name, **kw):
|
| 185 |
+
try:
|
| 186 |
+
return cls.from_pretrained(name, dtype=DTYPE, **kw)
|
| 187 |
+
except TypeError:
|
| 188 |
+
return cls.from_pretrained(name, torch_dtype=DTYPE, **kw)
|
| 189 |
+
|
| 190 |
+
|
| 191 |
+
def build_model():
|
| 192 |
+
tok = AutoTokenizer.from_pretrained(LLM_NAME)
|
| 193 |
+
pad_id = tok.pad_token_id if tok.pad_token_id is not None else tok.eos_token_id
|
| 194 |
+
processor = AutoImageProcessor.from_pretrained(VISION_NAME)
|
| 195 |
+
|
| 196 |
+
full_vision = load_pretrained(AutoModel, VISION_NAME)
|
| 197 |
+
vision = full_vision.vision_model
|
| 198 |
+
vision_dim = full_vision.config.vision_config.hidden_size
|
| 199 |
+
del full_vision # drops the unused SigLIP text tower
|
| 200 |
+
|
| 201 |
+
llm = load_pretrained(AutoModelForCausalLM, LLM_NAME)
|
| 202 |
+
text_dim = llm.config.hidden_size
|
| 203 |
+
llm.config.use_cache = False
|
| 204 |
+
|
| 205 |
+
vision.to(DEVICE).eval()
|
| 206 |
+
llm.to(DEVICE)
|
| 207 |
+
for p in vision.parameters():
|
| 208 |
+
p.requires_grad_(False)
|
| 209 |
+
for p in llm.parameters():
|
| 210 |
+
p.requires_grad_(False)
|
| 211 |
+
if GRAD_CKPT:
|
| 212 |
+
llm.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
|
| 213 |
+
llm.train() # HF only checkpoints in train mode; Llama has no dropout
|
| 214 |
+
else:
|
| 215 |
+
llm.eval()
|
| 216 |
+
|
| 217 |
+
prefix_ids = tok(USER_HEADER, add_special_tokens=True).input_ids
|
| 218 |
+
model = AtlasVision(vision, llm, prefix_ids, vision_dim, text_dim)
|
| 219 |
+
model.projector.to(DEVICE, dtype=torch.float32)
|
| 220 |
+
model.prefix_ids = model.prefix_ids.to(DEVICE)
|
| 221 |
+
return model, tok, pad_id, processor
|
| 222 |
+
|
| 223 |
+
|
| 224 |
+
# ----------------------------- TRAIN -----------------------------
|
| 225 |
+
def lr_at(step, total, warmup):
|
| 226 |
+
if step < warmup:
|
| 227 |
+
return LR * (step + 1) / warmup
|
| 228 |
+
progress = (step - warmup) / max(1, total - warmup)
|
| 229 |
+
return LR * 0.5 * (1 + math.cos(math.pi * min(1.0, progress)))
|
| 230 |
+
|
| 231 |
+
|
| 232 |
+
def save_checkpoint(path, model, opt, samples_seen, step, extra=None):
|
| 233 |
+
tmp = path + ".tmp"
|
| 234 |
+
torch.save({
|
| 235 |
+
"projector_state_dict": model.projector.state_dict(),
|
| 236 |
+
"optimizer_state_dict": opt.state_dict(),
|
| 237 |
+
"samples_seen": samples_seen,
|
| 238 |
+
"opt_step": step,
|
| 239 |
+
"config": {"LR": LR, "BATCH_SIZE": BATCH_SIZE, "GRAD_ACCUM": GRAD_ACCUM, "NUM_SAMPLES": NUM_SAMPLES,
|
| 240 |
+
"SEED": SEED, "LLM_NAME": LLM_NAME, "VISION_NAME": VISION_NAME, **(extra or {})},
|
| 241 |
+
}, tmp)
|
| 242 |
+
os.replace(tmp, path) # atomic: a stop mid-save never corrupts the last good checkpoint
|
| 243 |
+
|
| 244 |
+
|
| 245 |
+
def write_result(d):
|
| 246 |
+
path = os.environ.get("GMN_RESULT_PATH")
|
| 247 |
+
if path:
|
| 248 |
+
with open(path, "w") as f:
|
| 249 |
+
json.dump(d, f)
|
| 250 |
+
|
| 251 |
+
|
| 252 |
+
def main():
|
| 253 |
+
torch.manual_seed(SEED)
|
| 254 |
+
os.makedirs(CKPT_DIR, exist_ok=True)
|
| 255 |
+
ckpt_path = os.path.join(CKPT_DIR, "latest.pt")
|
| 256 |
+
log(f"device={DEVICE} batch={BATCH_SIZE}x{GRAD_ACCUM} lr={LR} samples={NUM_SAMPLES} "
|
| 257 |
+
f"grad_ckpt={GRAD_CKPT} workers={NUM_WORKERS} max_steps={MAX_STEPS or 'full'}")
|
| 258 |
+
|
| 259 |
+
model, tok, pad_id, processor = build_model()
|
| 260 |
+
n_train = sum(p.numel() for p in model.parameters() if p.requires_grad)
|
| 261 |
+
log(f"trainable params: {n_train:,} (projector only)")
|
| 262 |
+
|
| 263 |
+
with open(JSON_PATH) as f:
|
| 264 |
+
annotations = json.load(f)
|
| 265 |
+
g = torch.Generator().manual_seed(SEED)
|
| 266 |
+
keep = torch.randperm(len(annotations), generator=g)[:NUM_SAMPLES].tolist()
|
| 267 |
+
annotations = [annotations[i] for i in keep]
|
| 268 |
+
zip_prefix = find_zip_prefix(annotations, ZIP_PATH)
|
| 269 |
+
dataset = LLaVAPretrainDataset(annotations, ZIP_PATH, zip_prefix, processor, tok, MAX_TEXT_LEN)
|
| 270 |
+
|
| 271 |
+
eff_bs = BATCH_SIZE * GRAD_ACCUM
|
| 272 |
+
total_steps = math.ceil(len(dataset) / eff_bs)
|
| 273 |
+
if MAX_STEPS:
|
| 274 |
+
total_steps = min(total_steps, MAX_STEPS)
|
| 275 |
+
warmup = max(1, int(total_steps * WARMUP_RATIO))
|
| 276 |
+
|
| 277 |
+
params = [p for p in model.projector.parameters()]
|
| 278 |
+
opt = torch.optim.AdamW(params, lr=LR, weight_decay=0.0)
|
| 279 |
+
|
| 280 |
+
samples_seen, step = 0, 0
|
| 281 |
+
if os.path.exists(ckpt_path):
|
| 282 |
+
ck = torch.load(ckpt_path, map_location="cpu")
|
| 283 |
+
model.projector.load_state_dict(ck["projector_state_dict"])
|
| 284 |
+
opt.load_state_dict(ck["optimizer_state_dict"])
|
| 285 |
+
samples_seen = ck["samples_seen"]
|
| 286 |
+
step = samples_seen // eff_bs # works even if batch size changed
|
| 287 |
+
log(f"resumed: {samples_seen:,} samples seen -> optimizer step {step}/{total_steps}")
|
| 288 |
+
del ck
|
| 289 |
+
if step >= total_steps:
|
| 290 |
+
log("already complete")
|
| 291 |
+
return
|
| 292 |
+
|
| 293 |
+
order = torch.randperm(len(dataset), generator=torch.Generator().manual_seed(SEED + 1))[samples_seen:].tolist()
|
| 294 |
+
loader = DataLoader(
|
| 295 |
+
Subset(dataset, order), batch_size=BATCH_SIZE, shuffle=False, num_workers=NUM_WORKERS,
|
| 296 |
+
pin_memory=DEVICE.type == "cuda", collate_fn=make_collate(pad_id), drop_last=True,
|
| 297 |
+
persistent_workers=NUM_WORKERS > 0, prefetch_factor=4 if NUM_WORKERS > 0 else None,
|
| 298 |
+
)
|
| 299 |
+
log(f"schedule: {total_steps} optimizer steps, {warmup} warmup, starting at {step}")
|
| 300 |
+
|
| 301 |
+
model.projector.train()
|
| 302 |
+
if DEVICE.type == "cuda":
|
| 303 |
+
torch.cuda.reset_peak_memory_stats()
|
| 304 |
+
opt.zero_grad(set_to_none=True)
|
| 305 |
+
micro, loss_sum, loss_n, bad_imgs, skipped = 0, 0.0, 0, 0, 0
|
| 306 |
+
last_loss, t_window, steps_window = None, time.time(), 0
|
| 307 |
+
t_start = time.time()
|
| 308 |
+
log_f = open(os.path.join(CKPT_DIR, "train_log.jsonl"), "a")
|
| 309 |
+
|
| 310 |
+
for batch in loader:
|
| 311 |
+
if step >= total_steps:
|
| 312 |
+
break
|
| 313 |
+
try:
|
| 314 |
+
pv = batch["pixel_values"].to(DEVICE, non_blocking=True)
|
| 315 |
+
ids = batch["input_ids"].to(DEVICE, non_blocking=True)
|
| 316 |
+
mask = batch["attention_mask"].to(DEVICE, non_blocking=True)
|
| 317 |
+
labels = batch["labels"].to(DEVICE, non_blocking=True)
|
| 318 |
+
bad_imgs += batch["n_bad"]
|
| 319 |
+
with torch.autocast(device_type=DEVICE.type, dtype=DTYPE):
|
| 320 |
+
loss = model(pv, ids, mask, labels).float()
|
| 321 |
+
if not torch.isfinite(loss):
|
| 322 |
+
skipped += 1
|
| 323 |
+
log(f"non-finite loss at step {step}; dropping this accumulation window (skipped={skipped})")
|
| 324 |
+
opt.zero_grad(set_to_none=True)
|
| 325 |
+
micro = 0
|
| 326 |
+
continue
|
| 327 |
+
(loss / GRAD_ACCUM).backward()
|
| 328 |
+
except torch.cuda.OutOfMemoryError:
|
| 329 |
+
log("CUDA OOM - rerun with GRAD_CKPT=1 or a smaller BATCH_SIZE")
|
| 330 |
+
sys.exit(3)
|
| 331 |
+
|
| 332 |
+
loss_sum += loss.item()
|
| 333 |
+
loss_n += 1
|
| 334 |
+
micro += 1
|
| 335 |
+
if micro < GRAD_ACCUM:
|
| 336 |
+
continue
|
| 337 |
+
micro = 0
|
| 338 |
+
|
| 339 |
+
lr = lr_at(step, total_steps, warmup)
|
| 340 |
+
for grp in opt.param_groups:
|
| 341 |
+
grp["lr"] = lr
|
| 342 |
+
grad_norm = torch.nn.utils.clip_grad_norm_(params, max_norm=1.0).item()
|
| 343 |
+
opt.step()
|
| 344 |
+
opt.zero_grad(set_to_none=True)
|
| 345 |
+
step += 1
|
| 346 |
+
steps_window += 1
|
| 347 |
+
samples_seen += eff_bs
|
| 348 |
+
|
| 349 |
+
if step % LOG_EVERY == 0 or step == 1 or step == total_steps:
|
| 350 |
+
dt = time.time() - t_window
|
| 351 |
+
sps = dt / max(1, steps_window)
|
| 352 |
+
avg = loss_sum / max(1, loss_n)
|
| 353 |
+
last_loss = avg
|
| 354 |
+
peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
|
| 355 |
+
eta_h = (total_steps - step) * sps / 3600
|
| 356 |
+
log(f"step {step}/{total_steps} | loss {avg:.4f} | grad {grad_norm:.2f} | lr {lr:.2e} | "
|
| 357 |
+
f"{sps:.3f} s/step | {eff_bs / sps:.1f} img/s | ETA {eta_h:.2f} h | peak {peak:.1f} GB | "
|
| 358 |
+
f"bad_imgs {bad_imgs}")
|
| 359 |
+
log_f.write(json.dumps({"step": step, "loss": avg, "grad_norm": grad_norm, "lr": lr,
|
| 360 |
+
"s_per_step": sps, "peak_gb": peak, "samples_seen": samples_seen,
|
| 361 |
+
"time": time.time()}) + "\n")
|
| 362 |
+
log_f.flush()
|
| 363 |
+
loss_sum, loss_n, t_window, steps_window = 0.0, 0, time.time(), 0
|
| 364 |
+
|
| 365 |
+
if step % SAVE_EVERY == 0:
|
| 366 |
+
save_checkpoint(ckpt_path, model, opt, samples_seen, step)
|
| 367 |
+
log(f"checkpoint saved at step {step} ({samples_seen:,} samples)")
|
| 368 |
+
|
| 369 |
+
save_checkpoint(ckpt_path, model, opt, samples_seen, step)
|
| 370 |
+
done = step >= total_steps
|
| 371 |
+
if done and not MAX_STEPS:
|
| 372 |
+
torch.save(model.projector.state_dict(), os.path.join(CKPT_DIR, "projector_final.pt"))
|
| 373 |
+
log("saved projector_final.pt")
|
| 374 |
+
peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
|
| 375 |
+
elapsed = time.time() - t_start
|
| 376 |
+
log(f"finished: step {step}/{total_steps}, last avg loss {last_loss}, {elapsed / 60:.1f} min, "
|
| 377 |
+
f"peak {peak:.1f} GB, skipped {skipped}, bad_imgs {bad_imgs}")
|
| 378 |
+
write_result({"complete": done, "smoke": bool(MAX_STEPS), "step": step, "total_steps": total_steps,
|
| 379 |
+
"last_loss": last_loss, "peak_gb": round(peak, 1), "skipped": skipped, "bad_imgs": bad_imgs,
|
| 380 |
+
"grad_ckpt": GRAD_CKPT, "batch_size": BATCH_SIZE})
|
| 381 |
+
|
| 382 |
+
|
| 383 |
+
if __name__ == "__main__":
|
| 384 |
+
main()
|
code/train_stage2.py
ADDED
|
@@ -0,0 +1,294 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Stage 2 (visual instruction tuning): LLaVA-Instruct-150K on COCO images.
|
| 3 |
+
|
| 4 |
+
Starts from the stage-1 projector. Vision tower frozen; the LLM gets LoRA adapters;
|
| 5 |
+
the projector keeps training at a lower LR. Same prompt layout as stage 1:
|
| 6 |
+
[<BOS> user-header] [image tokens] [q1 <eot> assistant-header] a1 <eot> [user-header q2 <eot> assistant-header] a2 <eot> ...
|
| 7 |
+
Loss only on assistant answers (+ their <eot>). Gradient checkpointing on, length-bucketed batches,
|
| 8 |
+
resume by samples seen. All settings via environment variables.
|
| 9 |
+
"""
|
| 10 |
+
import json
|
| 11 |
+
import math
|
| 12 |
+
import os
|
| 13 |
+
import sys
|
| 14 |
+
import time
|
| 15 |
+
import zipfile
|
| 16 |
+
|
| 17 |
+
import torch
|
| 18 |
+
from PIL import Image
|
| 19 |
+
from torch.utils.data import DataLoader, Dataset, Subset
|
| 20 |
+
|
| 21 |
+
import train as T # stage-1 module: model layout, prompt strings, collate, zip check
|
| 22 |
+
|
| 23 |
+
env = T.env
|
| 24 |
+
HOME = os.path.expanduser("~")
|
| 25 |
+
INSTRUCT_JSON = env("INSTRUCT_JSON", f"{HOME}/data/llava_instruct/llava_instruct_150k.json")
|
| 26 |
+
COCO_ZIP = env("COCO_ZIP", f"{HOME}/data/coco/train2017.zip")
|
| 27 |
+
STAGE1_PROJECTOR = env("STAGE1_PROJECTOR", f"{HOME}/checkpoints/stage1/projector_final.pt")
|
| 28 |
+
CKPT_DIR = env("CKPT_DIR", f"{HOME}/checkpoints/stage2")
|
| 29 |
+
NUM_SAMPLES = env("NUM_SAMPLES", 0, int) # 0 = all except the held-out set
|
| 30 |
+
HELDOUT = env("HELDOUT", 1000, int)
|
| 31 |
+
BATCH_SIZE = env("BATCH_SIZE", 8, int)
|
| 32 |
+
GRAD_ACCUM = env("GRAD_ACCUM", 4, int)
|
| 33 |
+
LORA_LR = env("LORA_LR", 2e-4, float)
|
| 34 |
+
PROJ_LR = env("PROJ_LR", 2e-5, float)
|
| 35 |
+
LORA_R = env("LORA_R", 64, int)
|
| 36 |
+
LORA_ALPHA = env("LORA_ALPHA", 128, int)
|
| 37 |
+
WARMUP_RATIO = env("WARMUP_RATIO", 0.03, float)
|
| 38 |
+
MAX_TEXT_LEN = env("MAX_TEXT_LEN", 1024, int)
|
| 39 |
+
MAX_STEPS = env("MAX_STEPS", 0, int)
|
| 40 |
+
LOG_EVERY = env("LOG_EVERY", 25, int)
|
| 41 |
+
SAVE_EVERY = env("SAVE_EVERY", 250, int)
|
| 42 |
+
SEED = env("SEED", 42, int)
|
| 43 |
+
NUM_WORKERS = T.NUM_WORKERS
|
| 44 |
+
DEVICE, DTYPE, log = T.DEVICE, T.DTYPE, T.log
|
| 45 |
+
LORA_TARGETS = ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
# ----------------------------- DATA -----------------------------
|
| 49 |
+
def split(ann):
|
| 50 |
+
perm = torch.randperm(len(ann), generator=torch.Generator().manual_seed(SEED)).tolist()
|
| 51 |
+
held, train = perm[-HELDOUT:], perm[:-HELDOUT]
|
| 52 |
+
if NUM_SAMPLES:
|
| 53 |
+
train = train[:NUM_SAMPLES]
|
| 54 |
+
return train, held
|
| 55 |
+
|
| 56 |
+
|
| 57 |
+
def build_ids(tok, conversations, max_len):
|
| 58 |
+
ids, labels = [], []
|
| 59 |
+
first = True
|
| 60 |
+
for turn in conversations:
|
| 61 |
+
v = turn["value"].replace("<image>", "").strip()
|
| 62 |
+
if turn["from"] == "human":
|
| 63 |
+
text = (v if first else T.USER_HEADER + v) + T.ASSIST_HEADER
|
| 64 |
+
t = tok(text, add_special_tokens=False).input_ids
|
| 65 |
+
ids += t
|
| 66 |
+
labels += [-100] * len(t)
|
| 67 |
+
first = False
|
| 68 |
+
else:
|
| 69 |
+
t = tok(v + T.EOT, add_special_tokens=False).input_ids
|
| 70 |
+
ids += t
|
| 71 |
+
labels += t
|
| 72 |
+
return ids[:max_len], labels[:max_len]
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
class InstructDataset(Dataset):
|
| 76 |
+
def __init__(self, ann, zip_path, zip_prefix, processor, tok, max_len):
|
| 77 |
+
self.ann, self.zip_path, self.prefix = ann, zip_path, zip_prefix
|
| 78 |
+
self.processor, self.tok, self.max_len = processor, tok, max_len
|
| 79 |
+
self.zip = None
|
| 80 |
+
|
| 81 |
+
def __len__(self):
|
| 82 |
+
return len(self.ann)
|
| 83 |
+
|
| 84 |
+
def load_image(self, item):
|
| 85 |
+
if self.zip is None:
|
| 86 |
+
self.zip = zipfile.ZipFile(self.zip_path)
|
| 87 |
+
try:
|
| 88 |
+
with self.zip.open(self.prefix + item["image"]) as f:
|
| 89 |
+
return Image.open(f).convert("RGB"), True
|
| 90 |
+
except Exception:
|
| 91 |
+
return Image.new("RGB", (224, 224), "white"), False
|
| 92 |
+
|
| 93 |
+
def __getitem__(self, idx):
|
| 94 |
+
item = self.ann[idx]
|
| 95 |
+
image, ok = self.load_image(item)
|
| 96 |
+
pv = self.processor(images=image, return_tensors="pt").pixel_values[0]
|
| 97 |
+
ids, labels = build_ids(self.tok, item["conversations"], self.max_len)
|
| 98 |
+
return {"pixel_values": pv, "input_ids": ids, "labels": labels, "ok": ok}
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
def bucketed_order(indices, lengths, mb, seed):
|
| 102 |
+
"""Shuffle, sort by length inside chunks of 64 batches, keep only full batches, shuffle batches."""
|
| 103 |
+
g = torch.Generator().manual_seed(seed)
|
| 104 |
+
perm = [indices[i] for i in torch.randperm(len(indices), generator=g).tolist()]
|
| 105 |
+
chunk = mb * 64
|
| 106 |
+
batches = []
|
| 107 |
+
for s in range(0, len(perm), chunk):
|
| 108 |
+
c = sorted(perm[s:s + chunk], key=lambda i: lengths[i])
|
| 109 |
+
batches += [c[j:j + mb] for j in range(0, len(c), mb) if len(c[j:j + mb]) == mb]
|
| 110 |
+
order = torch.randperm(len(batches), generator=g).tolist()
|
| 111 |
+
return [i for b in order for i in batches[b]]
|
| 112 |
+
|
| 113 |
+
|
| 114 |
+
# ----------------------------- MODEL -----------------------------
|
| 115 |
+
def build_stage2_model(lora_state=None, projector_state=None, train_mode=True):
|
| 116 |
+
from peft import LoraConfig, get_peft_model, set_peft_model_state_dict
|
| 117 |
+
|
| 118 |
+
model, tok, pad_id, processor = T.build_model() # frozen vision + frozen LLM (bf16), fresh projector
|
| 119 |
+
proj = projector_state if projector_state is not None else _load_proj(STAGE1_PROJECTOR)
|
| 120 |
+
model.projector.load_state_dict(proj)
|
| 121 |
+
model.projector.to(DEVICE, dtype=torch.float32)
|
| 122 |
+
|
| 123 |
+
cfg = LoraConfig(r=LORA_R, lora_alpha=LORA_ALPHA, lora_dropout=0.05, target_modules=LORA_TARGETS,
|
| 124 |
+
bias="none", task_type="CAUSAL_LM")
|
| 125 |
+
model.llm = get_peft_model(model.llm, cfg)
|
| 126 |
+
if lora_state is not None:
|
| 127 |
+
set_peft_model_state_dict(model.llm, lora_state)
|
| 128 |
+
for n, p in model.llm.named_parameters(): # LoRA weights in fp32 for stable AdamW updates
|
| 129 |
+
if p.requires_grad:
|
| 130 |
+
p.data = p.data.float()
|
| 131 |
+
if train_mode:
|
| 132 |
+
model.llm.base_model.model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
|
| 133 |
+
model.llm.train()
|
| 134 |
+
else:
|
| 135 |
+
model.llm.eval()
|
| 136 |
+
model.vision.eval()
|
| 137 |
+
return model, tok, pad_id, processor
|
| 138 |
+
|
| 139 |
+
|
| 140 |
+
def _load_proj(path):
|
| 141 |
+
sd = torch.load(path, map_location="cpu")
|
| 142 |
+
return sd.get("projector_state_dict", sd)
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
def lr_scale(step, total, warmup):
|
| 146 |
+
if step < warmup:
|
| 147 |
+
return (step + 1) / warmup
|
| 148 |
+
progress = (step - warmup) / max(1, total - warmup)
|
| 149 |
+
return 0.5 * (1 + math.cos(math.pi * min(1.0, progress)))
|
| 150 |
+
|
| 151 |
+
|
| 152 |
+
def save_checkpoint(path, model, opt, samples_seen, step):
|
| 153 |
+
from peft import get_peft_model_state_dict
|
| 154 |
+
tmp = path + ".tmp"
|
| 155 |
+
torch.save({
|
| 156 |
+
"lora_state_dict": get_peft_model_state_dict(model.llm),
|
| 157 |
+
"projector_state_dict": model.projector.state_dict(),
|
| 158 |
+
"optimizer_state_dict": opt.state_dict(),
|
| 159 |
+
"samples_seen": samples_seen, "opt_step": step,
|
| 160 |
+
"config": {"LORA_R": LORA_R, "LORA_ALPHA": LORA_ALPHA, "LORA_LR": LORA_LR, "PROJ_LR": PROJ_LR,
|
| 161 |
+
"BATCH_SIZE": BATCH_SIZE, "GRAD_ACCUM": GRAD_ACCUM, "SEED": SEED},
|
| 162 |
+
}, tmp)
|
| 163 |
+
os.replace(tmp, path)
|
| 164 |
+
|
| 165 |
+
|
| 166 |
+
# ----------------------------- TRAIN -----------------------------
|
| 167 |
+
def main():
|
| 168 |
+
torch.manual_seed(SEED)
|
| 169 |
+
os.makedirs(CKPT_DIR, exist_ok=True)
|
| 170 |
+
ckpt_path = os.path.join(CKPT_DIR, "latest.pt")
|
| 171 |
+
log(f"stage2: batch={BATCH_SIZE}x{GRAD_ACCUM} lora_r={LORA_R} lora_lr={LORA_LR} proj_lr={PROJ_LR} "
|
| 172 |
+
f"max_text_len={MAX_TEXT_LEN} workers={NUM_WORKERS} max_steps={MAX_STEPS or 'full'}")
|
| 173 |
+
|
| 174 |
+
ck = torch.load(ckpt_path, map_location="cpu") if os.path.exists(ckpt_path) else None
|
| 175 |
+
model, tok, pad_id, processor = build_stage2_model(
|
| 176 |
+
lora_state=ck["lora_state_dict"] if ck else None,
|
| 177 |
+
projector_state=ck["projector_state_dict"] if ck else None)
|
| 178 |
+
lora_params = [p for n, p in model.llm.named_parameters() if p.requires_grad]
|
| 179 |
+
proj_params = list(model.projector.parameters())
|
| 180 |
+
log(f"trainable: LoRA {sum(p.numel() for p in lora_params):,} + projector {sum(p.numel() for p in proj_params):,}")
|
| 181 |
+
|
| 182 |
+
ann = json.load(open(INSTRUCT_JSON))
|
| 183 |
+
train_idx, _ = split(ann)
|
| 184 |
+
prefix = T.find_zip_prefix([ann[i] for i in train_idx[:2000]], COCO_ZIP)
|
| 185 |
+
dataset = InstructDataset(ann, COCO_ZIP, prefix, processor, tok, MAX_TEXT_LEN)
|
| 186 |
+
lengths = [sum(len(t["value"]) for t in a["conversations"]) for a in ann]
|
| 187 |
+
|
| 188 |
+
eff_bs = BATCH_SIZE * GRAD_ACCUM
|
| 189 |
+
order = bucketed_order(train_idx, lengths, BATCH_SIZE, SEED + 1)
|
| 190 |
+
total_steps = len(order) // eff_bs
|
| 191 |
+
if MAX_STEPS:
|
| 192 |
+
total_steps = min(total_steps, MAX_STEPS)
|
| 193 |
+
warmup = max(1, int(total_steps * WARMUP_RATIO))
|
| 194 |
+
|
| 195 |
+
opt = torch.optim.AdamW([{"params": lora_params, "lr": LORA_LR, "base_lr": LORA_LR},
|
| 196 |
+
{"params": proj_params, "lr": PROJ_LR, "base_lr": PROJ_LR}], weight_decay=0.0)
|
| 197 |
+
samples_seen, step = 0, 0
|
| 198 |
+
if ck:
|
| 199 |
+
opt.load_state_dict(ck["optimizer_state_dict"])
|
| 200 |
+
samples_seen = ck["samples_seen"]
|
| 201 |
+
step = samples_seen // eff_bs
|
| 202 |
+
log(f"resumed: {samples_seen:,} samples -> step {step}/{total_steps}")
|
| 203 |
+
del ck
|
| 204 |
+
if step >= total_steps:
|
| 205 |
+
log("already complete")
|
| 206 |
+
return
|
| 207 |
+
|
| 208 |
+
loader = DataLoader(
|
| 209 |
+
Subset(dataset, order[samples_seen:]), batch_size=BATCH_SIZE, shuffle=False, num_workers=NUM_WORKERS,
|
| 210 |
+
pin_memory=True, collate_fn=T.make_collate(pad_id), drop_last=True,
|
| 211 |
+
persistent_workers=NUM_WORKERS > 0, prefetch_factor=4 if NUM_WORKERS > 0 else None)
|
| 212 |
+
log(f"schedule: {total_steps} optimizer steps ({len(order):,} samples), {warmup} warmup, starting at {step}")
|
| 213 |
+
|
| 214 |
+
params = lora_params + proj_params
|
| 215 |
+
if DEVICE.type == "cuda":
|
| 216 |
+
torch.cuda.reset_peak_memory_stats()
|
| 217 |
+
opt.zero_grad(set_to_none=True)
|
| 218 |
+
micro, loss_sum, loss_n, bad, skipped = 0, 0.0, 0, 0, 0
|
| 219 |
+
last_loss, t_window, steps_window, t_start = None, time.time(), 0, time.time()
|
| 220 |
+
log_f = open(os.path.join(CKPT_DIR, "train_log.jsonl"), "a")
|
| 221 |
+
|
| 222 |
+
for batch in loader:
|
| 223 |
+
if step >= total_steps:
|
| 224 |
+
break
|
| 225 |
+
try:
|
| 226 |
+
bad += batch["n_bad"]
|
| 227 |
+
with torch.autocast(device_type=DEVICE.type, dtype=DTYPE):
|
| 228 |
+
loss = model(batch["pixel_values"].to(DEVICE, non_blocking=True),
|
| 229 |
+
batch["input_ids"].to(DEVICE, non_blocking=True),
|
| 230 |
+
batch["attention_mask"].to(DEVICE, non_blocking=True),
|
| 231 |
+
batch["labels"].to(DEVICE, non_blocking=True)).float()
|
| 232 |
+
if not torch.isfinite(loss):
|
| 233 |
+
skipped += 1
|
| 234 |
+
log(f"non-finite loss at step {step}; dropping accumulation window (skipped={skipped})")
|
| 235 |
+
opt.zero_grad(set_to_none=True)
|
| 236 |
+
micro = 0
|
| 237 |
+
continue
|
| 238 |
+
(loss / GRAD_ACCUM).backward()
|
| 239 |
+
except torch.cuda.OutOfMemoryError:
|
| 240 |
+
log("CUDA OOM - rerun with a smaller BATCH_SIZE (and larger GRAD_ACCUM)")
|
| 241 |
+
sys.exit(3)
|
| 242 |
+
|
| 243 |
+
loss_sum += loss.item()
|
| 244 |
+
loss_n += 1
|
| 245 |
+
micro += 1
|
| 246 |
+
if micro < GRAD_ACCUM:
|
| 247 |
+
continue
|
| 248 |
+
micro = 0
|
| 249 |
+
|
| 250 |
+
s = lr_scale(step, total_steps, warmup)
|
| 251 |
+
for grp in opt.param_groups:
|
| 252 |
+
grp["lr"] = grp["base_lr"] * s
|
| 253 |
+
grad_norm = torch.nn.utils.clip_grad_norm_(params, max_norm=1.0).item()
|
| 254 |
+
opt.step()
|
| 255 |
+
opt.zero_grad(set_to_none=True)
|
| 256 |
+
step += 1
|
| 257 |
+
steps_window += 1
|
| 258 |
+
samples_seen += eff_bs
|
| 259 |
+
|
| 260 |
+
if step % LOG_EVERY == 0 or step == 1 or step == total_steps:
|
| 261 |
+
dt = time.time() - t_window
|
| 262 |
+
sps = dt / max(1, steps_window)
|
| 263 |
+
avg = loss_sum / max(1, loss_n)
|
| 264 |
+
last_loss = avg
|
| 265 |
+
peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
|
| 266 |
+
log(f"step {step}/{total_steps} | loss {avg:.4f} | grad {grad_norm:.2f} | lr {LORA_LR * s:.2e} | "
|
| 267 |
+
f"{sps:.3f} s/step | {eff_bs / sps:.1f} conv/s | ETA {(total_steps - step) * sps / 3600:.2f} h | "
|
| 268 |
+
f"peak {peak:.1f} GB | bad_imgs {bad}")
|
| 269 |
+
log_f.write(json.dumps({"step": step, "loss": avg, "grad_norm": grad_norm, "lr": LORA_LR * s,
|
| 270 |
+
"s_per_step": sps, "peak_gb": peak, "samples_seen": samples_seen,
|
| 271 |
+
"time": time.time()}) + "\n")
|
| 272 |
+
log_f.flush()
|
| 273 |
+
loss_sum, loss_n, t_window, steps_window = 0.0, 0, time.time(), 0
|
| 274 |
+
|
| 275 |
+
if step % SAVE_EVERY == 0:
|
| 276 |
+
save_checkpoint(ckpt_path, model, opt, samples_seen, step)
|
| 277 |
+
log(f"checkpoint saved at step {step} ({samples_seen:,} samples)")
|
| 278 |
+
|
| 279 |
+
save_checkpoint(ckpt_path, model, opt, samples_seen, step)
|
| 280 |
+
done = step >= total_steps
|
| 281 |
+
if done and not MAX_STEPS:
|
| 282 |
+
model.llm.save_pretrained(os.path.join(CKPT_DIR, "lora_adapter"))
|
| 283 |
+
torch.save(model.projector.state_dict(), os.path.join(CKPT_DIR, "projector_stage2.pt"))
|
| 284 |
+
log("saved lora_adapter/ and projector_stage2.pt")
|
| 285 |
+
peak = torch.cuda.max_memory_allocated() / 1e9 if DEVICE.type == "cuda" else 0.0
|
| 286 |
+
log(f"finished: step {step}/{total_steps}, last avg loss {last_loss}, {(time.time() - t_start) / 60:.1f} min, "
|
| 287 |
+
f"peak {peak:.1f} GB, skipped {skipped}, bad_imgs {bad}")
|
| 288 |
+
T.write_result({"complete": done, "smoke": bool(MAX_STEPS), "step": step, "total_steps": total_steps,
|
| 289 |
+
"last_loss": last_loss, "peak_gb": round(peak, 1), "skipped": skipped, "bad_imgs": bad,
|
| 290 |
+
"batch_size": BATCH_SIZE})
|
| 291 |
+
|
| 292 |
+
|
| 293 |
+
if __name__ == "__main__":
|
| 294 |
+
main()
|
eval/report.md
ADDED
|
@@ -0,0 +1,58 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Atlas-Vision evaluation
|
| 2 |
+
|
| 3 |
+
## Scores (stage 1 → stage 2)
|
| 4 |
+
|
| 5 |
+
| Metric | Stage 1 | Stage 2 |
|
| 6 |
+
|---|---|---|
|
| 7 |
+
| Held-out instruct loss (lower is better) | 2.2521 | 1.1271 |
|
| 8 |
+
| POPE random: accuracy / F1 / yes-ratio | 0.4997 / 0.2003 / 0.1257 | 0.5427 / 0.6856 / 0.9547 |
|
| 9 |
+
| POPE popular: accuracy / F1 / yes-ratio | 0.5027 / 0.2013 / 0.1227 | 0.522 / 0.676 / 0.9753 |
|
| 10 |
+
| POPE adversarial: accuracy / F1 / yes-ratio | 0.5023 / 0.2012 / 0.123 | 0.5137 / 0.6722 / 0.9837 |
|
| 11 |
+
|
| 12 |
+
## Detailed descriptions
|
| 13 |
+
|
| 14 |
+
**COCO_val2014_000000310196**
|
| 15 |
+
|
| 16 |
+
- Stage 1: a skier in the snow on a mountain
|
| 17 |
+
- Stage 2: The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.
|
| 18 |
+
|
| 19 |
+
There are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together.
|
| 20 |
+
|
| 21 |
+
**COCO_val2014_000000210789**
|
| 22 |
+
|
| 23 |
+
- Stage 1: a woman and her child in the rain
|
| 24 |
+
- Stage 2: The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.
|
| 25 |
+
|
| 26 |
+
In the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene.
|
| 27 |
+
|
| 28 |
+
**COCO_val2014_000000429109**
|
| 29 |
+
|
| 30 |
+
- Stage 1: a bus and several other vehicles parked in front of a building
|
| 31 |
+
- Stage 2: The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.
|
| 32 |
+
|
| 33 |
+
Numerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.
|
| 34 |
+
|
| 35 |
+
In addition to the vehicles, there are traffic lights at various points along the street
|
| 36 |
+
|
| 37 |
+
**COCO_val2014_000000211674**
|
| 38 |
+
|
| 39 |
+
- Stage 1: a bus with a red and white logo on it, carrying passengers
|
| 40 |
+
- Stage 2: The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.
|
| 41 |
+
|
| 42 |
+
In addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by.
|
| 43 |
+
|
| 44 |
+
## Questions in Nigerian languages (stage 2)
|
| 45 |
+
|
| 46 |
+
- **igbo** — Kedu ihe dị na foto a?
|
| 47 |
+
→ The image features a person skiing down a snow-covered slope, with the skier wearing red pants.
|
| 48 |
+
- **yoruba** — Kí ni ó wà nínú àwòrán yìí?
|
| 49 |
+
→ The image features a person skiing down a snow-covered slope, with the skier wearing red pants.
|
| 50 |
+
- **hausa** — Me ke cikin wannan hoton?
|
| 51 |
+
→ In the image, a person is skiing down a snow-covered slope.
|
| 52 |
+
|
| 53 |
+
## Text-only check (no image)
|
| 54 |
+
|
| 55 |
+
Question: Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.
|
| 56 |
+
|
| 57 |
+
- Base N-ATLaS: Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
|
| 58 |
+
- With stage-2 LoRA: Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu "positron" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej
|
eval/stage1_eval.json
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"checkpoint": {
|
| 3 |
+
"net.0.weight": [
|
| 4 |
+
4096,
|
| 5 |
+
768
|
| 6 |
+
],
|
| 7 |
+
"net.0.bias": [
|
| 8 |
+
4096
|
| 9 |
+
],
|
| 10 |
+
"net.2.weight": [
|
| 11 |
+
4096,
|
| 12 |
+
4096
|
| 13 |
+
],
|
| 14 |
+
"net.2.bias": [
|
| 15 |
+
4096
|
| 16 |
+
]
|
| 17 |
+
},
|
| 18 |
+
"checkpoint_all_finite": true,
|
| 19 |
+
"heldout_loss_random_projector": 8.7249,
|
| 20 |
+
"heldout_loss_trained": 2.4523,
|
| 21 |
+
"heldout_loss_trained_shuffled_images": 5.4332,
|
| 22 |
+
"captions": [
|
| 23 |
+
{
|
| 24 |
+
"image": "00390/003907752.jpg",
|
| 25 |
+
"reference": "cute homecoming prom dress with lace top and satin skirt",
|
| 26 |
+
"model": "a - line sweetheart sweetheart lace tulle prom dress"
|
| 27 |
+
},
|
| 28 |
+
{
|
| 29 |
+
"image": "00193/001937181.jpg",
|
| 30 |
+
"reference": "a pair of silicone bracelets with a four codes message",
|
| 31 |
+
"model": "a pair of silicone bracelets with the word code on them"
|
| 32 |
+
},
|
| 33 |
+
{
|
| 34 |
+
"image": "00530/005308076.jpg",
|
| 35 |
+
"reference": "a set of fingerprint icons in white on a black background",
|
| 36 |
+
"model": "the iphone 6s and iphone 6s plus are shown in a black and white image"
|
| 37 |
+
},
|
| 38 |
+
{
|
| 39 |
+
"image": "00399/003998740.jpg",
|
| 40 |
+
"reference": "a business woman drawing modern concept of a website creation",
|
| 41 |
+
"model": "a man writing the word modern web design on a whiteboard"
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"image": "00391/003917454.jpg",
|
| 45 |
+
"reference": "the beach in el nido national park, puerto puerto",
|
| 46 |
+
"model": "the limestone cliffs and limestone islands in the background"
|
| 47 |
+
},
|
| 48 |
+
{
|
| 49 |
+
"image": "00077/000776787.jpg",
|
| 50 |
+
"reference": "three pieces of paper with the words, democratic decentified dp controlled centralized ccp",
|
| 51 |
+
"model": "a diagram showing the different types of democracy"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"image": "00099/000999958.jpg",
|
| 55 |
+
"reference": "the beginer's guide to social security book",
|
| 56 |
+
"model": "a beginner's guide to social security"
|
| 57 |
+
},
|
| 58 |
+
{
|
| 59 |
+
"image": "00446/004469100.jpg",
|
| 60 |
+
"reference": "a flag of bolivia, with a smiley face water bottle",
|
| 61 |
+
"model": "a red and yellow flag with a red and yellow flag on it water bottle"
|
| 62 |
+
}
|
| 63 |
+
],
|
| 64 |
+
"coco_zip_files": 118288,
|
| 65 |
+
"instruct_conversations": 157712,
|
| 66 |
+
"instruct_images_found_of_2000": 2000,
|
| 67 |
+
"eval_minutes": 1.9
|
| 68 |
+
}
|
eval/stage2_eval.json
ADDED
|
@@ -0,0 +1,104 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"stage1": {
|
| 3 |
+
"heldout_instruct_loss": 2.2521,
|
| 4 |
+
"pope_random": {
|
| 5 |
+
"n": 3000,
|
| 6 |
+
"accuracy": 0.4997,
|
| 7 |
+
"precision": 0.4987,
|
| 8 |
+
"recall": 0.1253,
|
| 9 |
+
"f1": 0.2003,
|
| 10 |
+
"yes_ratio": 0.1257
|
| 11 |
+
},
|
| 12 |
+
"pope_popular": {
|
| 13 |
+
"n": 3000,
|
| 14 |
+
"accuracy": 0.5027,
|
| 15 |
+
"precision": 0.5109,
|
| 16 |
+
"recall": 0.1253,
|
| 17 |
+
"f1": 0.2013,
|
| 18 |
+
"yes_ratio": 0.1227
|
| 19 |
+
},
|
| 20 |
+
"pope_adversarial": {
|
| 21 |
+
"n": 3000,
|
| 22 |
+
"accuracy": 0.5023,
|
| 23 |
+
"precision": 0.5095,
|
| 24 |
+
"recall": 0.1253,
|
| 25 |
+
"f1": 0.2012,
|
| 26 |
+
"yes_ratio": 0.123
|
| 27 |
+
}
|
| 28 |
+
},
|
| 29 |
+
"stage2": {
|
| 30 |
+
"heldout_instruct_loss": 1.1271,
|
| 31 |
+
"pope_random": {
|
| 32 |
+
"n": 3000,
|
| 33 |
+
"accuracy": 0.5427,
|
| 34 |
+
"precision": 0.5223,
|
| 35 |
+
"recall": 0.9973,
|
| 36 |
+
"f1": 0.6856,
|
| 37 |
+
"yes_ratio": 0.9547
|
| 38 |
+
},
|
| 39 |
+
"pope_popular": {
|
| 40 |
+
"n": 3000,
|
| 41 |
+
"accuracy": 0.522,
|
| 42 |
+
"precision": 0.5113,
|
| 43 |
+
"recall": 0.9973,
|
| 44 |
+
"f1": 0.676,
|
| 45 |
+
"yes_ratio": 0.9753
|
| 46 |
+
},
|
| 47 |
+
"pope_adversarial": {
|
| 48 |
+
"n": 3000,
|
| 49 |
+
"accuracy": 0.5137,
|
| 50 |
+
"precision": 0.5069,
|
| 51 |
+
"recall": 0.9973,
|
| 52 |
+
"f1": 0.6722,
|
| 53 |
+
"yes_ratio": 0.9837
|
| 54 |
+
}
|
| 55 |
+
},
|
| 56 |
+
"descriptions": [
|
| 57 |
+
{
|
| 58 |
+
"image": "COCO_val2014_000000310196",
|
| 59 |
+
"stage1": "a skier in the snow on a mountain",
|
| 60 |
+
"stage2": "The image features a snowboarder wearing a red jacket, standing on top of a snowy hill. The person is in the process of skiing down the slope, with their skis visible beneath them. They are surrounded by a beautiful landscape that includes trees and mountains in the background.\n\nThere are several other people scattered throughout the scene, some closer to the foreground while others are further away. These individuals might be fellow snowboarders or skiers enjoying the winter sports activities together."
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"image": "COCO_val2014_000000210789",
|
| 64 |
+
"stage1": "a woman and her child in the rain",
|
| 65 |
+
"stage2": "The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The woman is on the left side of the scene, while the little girl stands next to her on the right. They are positioned close together, with the woman's umbrella covering them both.\n\nIn the background, there are two cars parked behind them, one closer to the left edge of the frame and another further back towards the center. A handbag can be seen placed near the woman, possibly belonging to her or someone else present in the scene."
|
| 66 |
+
},
|
| 67 |
+
{
|
| 68 |
+
"image": "COCO_val2014_000000429109",
|
| 69 |
+
"stage1": "a bus and several other vehicles parked in front of a building",
|
| 70 |
+
"stage2": "The image depicts a busy city street with several buses and cars parked or driving along the road. There are three buses in total, one of which is a large red bus occupying most of the scene, while the other two are smaller and positioned closer to the right side of the image.\n\nNumerous cars can be seen throughout the scene, some parked on the left side of the street and others driving down the road. A person is also visible near the center of the image, possibly waiting for public transportation or walking by.\n\nIn addition to the vehicles, there are traffic lights at various points along the street"
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"image": "COCO_val2014_000000211674",
|
| 74 |
+
"stage1": "a bus with a red and white logo on it, carrying passengers",
|
| 75 |
+
"stage2": "The image features a red double-decker bus driving down the street, with people on both levels of the bus. There are at least 12 passengers visible in the scene, some sitting and others standing, enjoying their ride. The bus is filled to capacity, indicating that it's a popular mode of transportation for these individuals.\n\nIn addition to the bus, there are two cars parked or moving along the street, one closer to the left side and another further back towards the right. A person can be seen walking near the center of the scene, possibly waiting to board the bus or simply passing by."
|
| 76 |
+
}
|
| 77 |
+
],
|
| 78 |
+
"multilingual": [
|
| 79 |
+
{
|
| 80 |
+
"image": "COCO_val2014_000000310196",
|
| 81 |
+
"language": "igbo",
|
| 82 |
+
"question": "Kedu ihe dị na foto a?",
|
| 83 |
+
"stage2": "The image features a person skiing down a snow-covered slope, with the skier wearing red pants."
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"image": "COCO_val2014_000000310196",
|
| 87 |
+
"language": "yoruba",
|
| 88 |
+
"question": "Kí ni ó wà nínú àwòrán yìí?",
|
| 89 |
+
"stage2": "The image features a person skiing down a snow-covered slope, with the skier wearing red pants."
|
| 90 |
+
},
|
| 91 |
+
{
|
| 92 |
+
"image": "COCO_val2014_000000310196",
|
| 93 |
+
"language": "hausa",
|
| 94 |
+
"question": "Me ke cikin wannan hoton?",
|
| 95 |
+
"stage2": "In the image, a person is skiing down a snow-covered slope."
|
| 96 |
+
}
|
| 97 |
+
],
|
| 98 |
+
"text_only": {
|
| 99 |
+
"question": "Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.",
|
| 100 |
+
"base_n_atlas": "Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us",
|
| 101 |
+
"with_stage2_lora": "Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu \"positron\" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej"
|
| 102 |
+
},
|
| 103 |
+
"eval_minutes": 5.2
|
| 104 |
+
}
|
logs/stage1_train_log.jsonl
ADDED
|
@@ -0,0 +1,376 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"step": 1, "loss": 8.55087947845459, "grad_norm": 87.72406768798828, "lr": 7.117437722419929e-07, "s_per_step": 8.396198749542236, "peak_gb": 41.247548928, "samples_seen": 32, "time": 1791130833.9630904}
|
| 2 |
+
{"step": 25, "loss": 6.397367646296819, "grad_norm": 15.911425590515137, "lr": 1.779359430604982e-05, "s_per_step": 0.6239869793256124, "peak_gb": 42.138478592, "samples_seen": 800, "time": 1791130848.9393055}
|
| 3 |
+
{"step": 50, "loss": 4.635109643936158, "grad_norm": 7.2413105964660645, "lr": 3.558718861209964e-05, "s_per_step": 0.5889950180053711, "peak_gb": 42.138478592, "samples_seen": 1600, "time": 1791130863.664765}
|
| 4 |
+
{"step": 75, "loss": 4.093637528419495, "grad_norm": 4.90704870223999, "lr": 5.338078291814947e-05, "s_per_step": 0.5850510597229004, "peak_gb": 42.138478592, "samples_seen": 2400, "time": 1791130878.2915204}
|
| 5 |
+
{"step": 100, "loss": 3.8673372650146485, "grad_norm": 4.5811944007873535, "lr": 7.117437722419929e-05, "s_per_step": 0.623746509552002, "peak_gb": 44.713082368, "samples_seen": 3200, "time": 1791130893.8860688}
|
| 6 |
+
{"step": 125, "loss": 3.740880389213562, "grad_norm": 3.184147834777832, "lr": 8.896797153024912e-05, "s_per_step": 0.5818255233764649, "peak_gb": 44.713082368, "samples_seen": 4000, "time": 1791130908.432502}
|
| 7 |
+
{"step": 150, "loss": 3.7184673833847044, "grad_norm": 2.9738147258758545, "lr": 0.00010676156583629893, "s_per_step": 0.576975965499878, "peak_gb": 44.713082368, "samples_seen": 4800, "time": 1791130922.857381}
|
| 8 |
+
{"step": 175, "loss": 3.5475265645980834, "grad_norm": 2.332338333129883, "lr": 0.00012455516014234878, "s_per_step": 0.5761029434204101, "peak_gb": 44.713082368, "samples_seen": 5600, "time": 1791130937.2603934}
|
| 9 |
+
{"step": 200, "loss": 3.5738428688049315, "grad_norm": 1.9974111318588257, "lr": 0.00014234875444839857, "s_per_step": 0.580996150970459, "peak_gb": 44.713082368, "samples_seen": 6400, "time": 1791130951.7857678}
|
| 10 |
+
{"step": 225, "loss": 3.5230879592895508, "grad_norm": 1.8189729452133179, "lr": 0.00016014234875444842, "s_per_step": 0.5787319278717041, "peak_gb": 44.713082368, "samples_seen": 7200, "time": 1791130966.2545114}
|
| 11 |
+
{"step": 250, "loss": 3.4794790887832643, "grad_norm": 2.1197030544281006, "lr": 0.00017793594306049823, "s_per_step": 0.5812134552001953, "peak_gb": 44.713082368, "samples_seen": 8000, "time": 1791130980.7852197}
|
| 12 |
+
{"step": 275, "loss": 3.5135861301422118, "grad_norm": 2.010641098022461, "lr": 0.00019572953736654805, "s_per_step": 0.5902993679046631, "peak_gb": 44.713082368, "samples_seen": 8800, "time": 1791130995.5431538}
|
| 13 |
+
{"step": 300, "loss": 3.332580523490906, "grad_norm": 1.7257816791534424, "lr": 0.00019999806668125934, "s_per_step": 0.5805612754821777, "peak_gb": 44.713082368, "samples_seen": 9600, "time": 1791131010.0576127}
|
| 14 |
+
{"step": 325, "loss": 3.331582636833191, "grad_norm": 1.817513346672058, "lr": 0.0001999889671230344, "s_per_step": 0.5767327976226807, "peak_gb": 44.713082368, "samples_seen": 10400, "time": 1791131024.4763641}
|
| 15 |
+
{"step": 350, "loss": 3.3034389925003054, "grad_norm": 1.4937862157821655, "lr": 0.00019997240961861618, "s_per_step": 0.5783243465423584, "peak_gb": 44.713082368, "samples_seen": 11200, "time": 1791131038.9348876}
|
| 16 |
+
{"step": 375, "loss": 3.310972599983215, "grad_norm": 1.4323818683624268, "lr": 0.00019994839540299076, "s_per_step": 0.5792804718017578, "peak_gb": 44.713082368, "samples_seen": 12000, "time": 1791131053.417312}
|
| 17 |
+
{"step": 400, "loss": 3.287726469039917, "grad_norm": 1.5991160869598389, "lr": 0.000199916926267323, "s_per_step": 0.578808069229126, "peak_gb": 44.713082368, "samples_seen": 12800, "time": 1791131067.8879645}
|
| 18 |
+
{"step": 425, "loss": 3.2084068727493285, "grad_norm": 1.271116018295288, "lr": 0.00019987800455882306, "s_per_step": 0.5803878498077393, "peak_gb": 44.713082368, "samples_seen": 13600, "time": 1791131082.3980415}
|
| 19 |
+
{"step": 450, "loss": 3.2198647975921633, "grad_norm": 1.286621332168579, "lr": 0.00019983163318057137, "s_per_step": 0.5836947727203369, "peak_gb": 44.713082368, "samples_seen": 14400, "time": 1791131096.991321}
|
| 20 |
+
{"step": 475, "loss": 3.1718136644363404, "grad_norm": 1.3426055908203125, "lr": 0.00019977781559130187, "s_per_step": 0.5814045429229736, "peak_gb": 44.713082368, "samples_seen": 15200, "time": 1791131111.52718}
|
| 21 |
+
{"step": 500, "loss": 3.176059513092041, "grad_norm": 1.3492554426193237, "lr": 0.00019971655580514438, "s_per_step": 0.5872831630706787, "peak_gb": 44.713082368, "samples_seen": 16000, "time": 1791131126.2099152}
|
| 22 |
+
{"step": 525, "loss": 3.1693886518478394, "grad_norm": 1.2558445930480957, "lr": 0.00019964785839132488, "s_per_step": 0.6033973217010498, "peak_gb": 44.713082368, "samples_seen": 16800, "time": 1791131141.295349}
|
| 23 |
+
{"step": 550, "loss": 3.14515998840332, "grad_norm": 1.2257124185562134, "lr": 0.0001995717284738248, "s_per_step": 0.5826546669006347, "peak_gb": 44.713082368, "samples_seen": 17600, "time": 1791131155.8622203}
|
| 24 |
+
{"step": 575, "loss": 3.12627272605896, "grad_norm": 1.2098281383514404, "lr": 0.000199488171730999, "s_per_step": 0.582458963394165, "peak_gb": 44.713082368, "samples_seen": 18400, "time": 1791131170.424219}
|
| 25 |
+
{"step": 600, "loss": 3.135997772216797, "grad_norm": 1.3937653303146362, "lr": 0.0001993971943951519, "s_per_step": 0.5858418941497803, "peak_gb": 44.713082368, "samples_seen": 19200, "time": 1791131185.0707946}
|
| 26 |
+
{"step": 625, "loss": 3.0999442672729494, "grad_norm": 1.0516718626022339, "lr": 0.000199298803252073, "s_per_step": 0.5833398532867432, "peak_gb": 44.713082368, "samples_seen": 20000, "time": 1791131199.654984}
|
| 27 |
+
{"step": 650, "loss": 3.080396270751953, "grad_norm": 1.1961040496826172, "lr": 0.00019919300564053046, "s_per_step": 0.6152436542510986, "peak_gb": 44.713082368, "samples_seen": 20800, "time": 1791131215.037142}
|
| 28 |
+
{"step": 675, "loss": 3.0702759647369384, "grad_norm": 1.0449596643447876, "lr": 0.00019907980945172384, "s_per_step": 0.5828781223297119, "peak_gb": 44.713082368, "samples_seen": 21600, "time": 1791131229.609633}
|
| 29 |
+
{"step": 700, "loss": 3.0974017333984376, "grad_norm": 1.0658111572265625, "lr": 0.00019895922312869555, "s_per_step": 0.5841869640350342, "peak_gb": 44.713082368, "samples_seen": 22400, "time": 1791131244.2148454}
|
| 30 |
+
{"step": 725, "loss": 3.074553599357605, "grad_norm": 0.9470886588096619, "lr": 0.00019883125566570094, "s_per_step": 0.5814198112487793, "peak_gb": 44.713082368, "samples_seen": 23200, "time": 1791131258.7509072}
|
| 31 |
+
{"step": 750, "loss": 3.1141643714904785, "grad_norm": 1.1800023317337036, "lr": 0.00019869591660753763, "s_per_step": 0.5784324645996094, "peak_gb": 44.713082368, "samples_seen": 24000, "time": 1791131273.212172}
|
| 32 |
+
{"step": 775, "loss": 3.1024985694885254, "grad_norm": 1.0973631143569946, "lr": 0.00019855321604883352, "s_per_step": 0.5959490013122558, "peak_gb": 44.713082368, "samples_seen": 24800, "time": 1791131288.1113417}
|
| 33 |
+
{"step": 800, "loss": 3.095311689376831, "grad_norm": 1.2262718677520752, "lr": 0.00019840316463329378, "s_per_step": 0.5826352882385254, "peak_gb": 44.713082368, "samples_seen": 25600, "time": 1791131302.6778867}
|
| 34 |
+
{"step": 825, "loss": 3.0253290700912476, "grad_norm": 1.2020397186279297, "lr": 0.000198245773552907, "s_per_step": 0.5797882652282715, "peak_gb": 44.713082368, "samples_seen": 26400, "time": 1791131317.173062}
|
| 35 |
+
{"step": 850, "loss": 3.0069773197174072, "grad_norm": 1.1738593578338623, "lr": 0.00019808105454711055, "s_per_step": 0.5793154048919678, "peak_gb": 44.713082368, "samples_seen": 27200, "time": 1791131331.6564252}
|
| 36 |
+
{"step": 875, "loss": 2.975392165184021, "grad_norm": 0.9737287163734436, "lr": 0.0001979090199019147, "s_per_step": 0.5790924072265625, "peak_gb": 44.713082368, "samples_seen": 28000, "time": 1791131346.1341693}
|
| 37 |
+
{"step": 900, "loss": 2.9498213052749636, "grad_norm": 1.2708204984664917, "lr": 0.0001977296824489864, "s_per_step": 0.5816666603088378, "peak_gb": 44.713082368, "samples_seen": 28800, "time": 1791131360.676415}
|
| 38 |
+
{"step": 925, "loss": 2.9142259979248046, "grad_norm": 1.089045763015747, "lr": 0.00019754305556469223, "s_per_step": 0.5844484996795655, "peak_gb": 44.713082368, "samples_seen": 29600, "time": 1791131375.2880468}
|
| 39 |
+
{"step": 950, "loss": 3.0302752017974854, "grad_norm": 1.205527663230896, "lr": 0.0001973491531691006, "s_per_step": 0.5822625255584717, "peak_gb": 44.713082368, "samples_seen": 30400, "time": 1791131389.8449948}
|
| 40 |
+
{"step": 975, "loss": 2.9161105060577395, "grad_norm": 1.1548945903778076, "lr": 0.00019714798972494347, "s_per_step": 0.5808561515808105, "peak_gb": 44.713082368, "samples_seen": 31200, "time": 1791131404.3668756}
|
| 41 |
+
{"step": 1000, "loss": 2.919966125488281, "grad_norm": 1.1152652502059937, "lr": 0.00019693958023653767, "s_per_step": 0.5795290184020996, "peak_gb": 44.713082368, "samples_seen": 32000, "time": 1791131418.8555546}
|
| 42 |
+
{"step": 1025, "loss": 2.952225284576416, "grad_norm": 1.1843825578689575, "lr": 0.00019672394024866576, "s_per_step": 0.5974857711791992, "peak_gb": 44.713082368, "samples_seen": 32800, "time": 1791131433.7931862}
|
| 43 |
+
{"step": 1050, "loss": 2.9690034532547, "grad_norm": 1.1854422092437744, "lr": 0.00019650108584541654, "s_per_step": 0.5811002635955811, "peak_gb": 44.713082368, "samples_seen": 33600, "time": 1791131448.3211079}
|
| 44 |
+
{"step": 1075, "loss": 2.9967500257492063, "grad_norm": 1.2190947532653809, "lr": 0.00019627103364898538, "s_per_step": 0.580099172592163, "peak_gb": 44.713082368, "samples_seen": 34400, "time": 1791131462.8241384}
|
| 45 |
+
{"step": 1100, "loss": 2.990977711677551, "grad_norm": 1.0495576858520508, "lr": 0.00019603380081843449, "s_per_step": 0.5808154773712159, "peak_gb": 44.713082368, "samples_seen": 35200, "time": 1791131477.3449311}
|
| 46 |
+
{"step": 1125, "loss": 2.8920970916748048, "grad_norm": 0.9397656321525574, "lr": 0.0001957894050484129, "s_per_step": 0.5838562393188477, "peak_gb": 44.713082368, "samples_seen": 36000, "time": 1791131491.941773}
|
| 47 |
+
{"step": 1150, "loss": 2.914958839416504, "grad_norm": 1.3053033351898193, "lr": 0.00019553786456783686, "s_per_step": 0.582167854309082, "peak_gb": 44.713082368, "samples_seen": 36800, "time": 1791131506.4963646}
|
| 48 |
+
{"step": 1175, "loss": 2.9034262990951536, "grad_norm": 1.0504006147384644, "lr": 0.00019527919813853, "s_per_step": 0.5793894004821777, "peak_gb": 44.713082368, "samples_seen": 37600, "time": 1791131520.981513}
|
| 49 |
+
{"step": 1200, "loss": 2.9084231233596802, "grad_norm": 1.1124707460403442, "lr": 0.00019501342505382405, "s_per_step": 0.5810192966461182, "peak_gb": 44.713082368, "samples_seen": 38400, "time": 1791131535.5074165}
|
| 50 |
+
{"step": 1225, "loss": 2.827880344390869, "grad_norm": 1.4236536026000977, "lr": 0.00019474056513711977, "s_per_step": 0.5862373161315918, "peak_gb": 44.713082368, "samples_seen": 39200, "time": 1791131550.1637754}
|
| 51 |
+
{"step": 1250, "loss": 2.8731605625152588, "grad_norm": 1.1672555208206177, "lr": 0.00019446063874040834, "s_per_step": 0.5826034450531006, "peak_gb": 44.713082368, "samples_seen": 40000, "time": 1791131564.7292533}
|
| 52 |
+
{"step": 1275, "loss": 2.9100488948822023, "grad_norm": 1.0153155326843262, "lr": 0.00019417366674275333, "s_per_step": 0.6047325611114502, "peak_gb": 44.713082368, "samples_seen": 40800, "time": 1791131579.8479767}
|
| 53 |
+
{"step": 1300, "loss": 2.8539438247680664, "grad_norm": 1.0282107591629028, "lr": 0.00019387967054873356, "s_per_step": 0.5824565982818604, "peak_gb": 44.713082368, "samples_seen": 41600, "time": 1791131594.4098606}
|
| 54 |
+
{"step": 1325, "loss": 2.8969113874435424, "grad_norm": 1.3478723764419556, "lr": 0.00019357867208684621, "s_per_step": 0.582642650604248, "peak_gb": 44.713082368, "samples_seen": 42400, "time": 1791131608.9763298}
|
| 55 |
+
{"step": 1350, "loss": 2.8323376035690306, "grad_norm": 1.1149699687957764, "lr": 0.00019327069380787164, "s_per_step": 0.5824921321868897, "peak_gb": 44.713082368, "samples_seen": 43200, "time": 1791131623.5390694}
|
| 56 |
+
{"step": 1375, "loss": 2.8894362831115723, "grad_norm": 1.1063823699951172, "lr": 0.00019295575868319857, "s_per_step": 0.5840862846374512, "peak_gb": 44.713082368, "samples_seen": 44000, "time": 1791131638.1416295}
|
| 57 |
+
{"step": 1400, "loss": 2.834498815536499, "grad_norm": 1.1182633638381958, "lr": 0.00019263389020311082, "s_per_step": 0.5812054634094238, "peak_gb": 44.713082368, "samples_seen": 44800, "time": 1791131652.6722355}
|
| 58 |
+
{"step": 1425, "loss": 2.8067142009735107, "grad_norm": 1.203196406364441, "lr": 0.00019230511237503515, "s_per_step": 0.5813498687744141, "peak_gb": 44.713082368, "samples_seen": 45600, "time": 1791131667.2064037}
|
| 59 |
+
{"step": 1450, "loss": 2.800572829246521, "grad_norm": 1.1607813835144043, "lr": 0.00019196944972175062, "s_per_step": 0.5796218490600586, "peak_gb": 44.713082368, "samples_seen": 46400, "time": 1791131681.697381}
|
| 60 |
+
{"step": 1475, "loss": 2.824986810684204, "grad_norm": 1.1562610864639282, "lr": 0.00019162692727955953, "s_per_step": 0.5842508506774903, "peak_gb": 44.713082368, "samples_seen": 47200, "time": 1791131696.3040268}
|
| 61 |
+
{"step": 1500, "loss": 2.8580026912689207, "grad_norm": 1.2550817728042603, "lr": 0.00019127757059642, "s_per_step": 0.5836690044403077, "peak_gb": 44.713082368, "samples_seen": 48000, "time": 1791131710.8980353}
|
| 62 |
+
{"step": 1525, "loss": 2.7956640100479127, "grad_norm": 1.0501290559768677, "lr": 0.00019092140573004042, "s_per_step": 0.5987179946899414, "peak_gb": 44.713082368, "samples_seen": 48800, "time": 1791131725.866389}
|
| 63 |
+
{"step": 1550, "loss": 2.904049711227417, "grad_norm": 1.2289342880249023, "lr": 0.0001905584592459358, "s_per_step": 0.579942455291748, "peak_gb": 44.713082368, "samples_seen": 49600, "time": 1791131740.3653781}
|
| 64 |
+
{"step": 1575, "loss": 2.801865735054016, "grad_norm": 1.4003002643585205, "lr": 0.0001901887582154464, "s_per_step": 0.5801137828826904, "peak_gb": 44.713082368, "samples_seen": 50400, "time": 1791131754.8686228}
|
| 65 |
+
{"step": 1600, "loss": 2.81393018245697, "grad_norm": 1.0660154819488525, "lr": 0.00018981233021371843, "s_per_step": 0.5823390769958496, "peak_gb": 44.713082368, "samples_seen": 51200, "time": 1791131769.4274774}
|
| 66 |
+
{"step": 1625, "loss": 2.8056074666976927, "grad_norm": 1.1312650442123413, "lr": 0.00018942920331764746, "s_per_step": 0.5844465160369873, "peak_gb": 44.713082368, "samples_seen": 52000, "time": 1791131784.0390692}
|
| 67 |
+
{"step": 1650, "loss": 2.77773108959198, "grad_norm": 1.0975704193115234, "lr": 0.00018903940610378407, "s_per_step": 0.582863941192627, "peak_gb": 44.713082368, "samples_seen": 52800, "time": 1791131798.6111288}
|
| 68 |
+
{"step": 1675, "loss": 2.7603001546859742, "grad_norm": 1.241600513458252, "lr": 0.00018864296764620238, "s_per_step": 0.5836851119995117, "peak_gb": 44.713082368, "samples_seen": 53600, "time": 1791131813.2036917}
|
| 69 |
+
{"step": 1700, "loss": 2.7555297422409057, "grad_norm": 1.1494439840316772, "lr": 0.00018823991751433167, "s_per_step": 0.5850475883483887, "peak_gb": 44.713082368, "samples_seen": 54400, "time": 1791131827.8303254}
|
| 70 |
+
{"step": 1725, "loss": 2.8411933469772337, "grad_norm": 0.991641104221344, "lr": 0.00018783028577075065, "s_per_step": 0.5855899620056152, "peak_gb": 44.713082368, "samples_seen": 55200, "time": 1791131842.4704814}
|
| 71 |
+
{"step": 1750, "loss": 2.8297329354286194, "grad_norm": 1.3848960399627686, "lr": 0.00018741410296894528, "s_per_step": 0.5831822967529297, "peak_gb": 44.713082368, "samples_seen": 56000, "time": 1791131857.0504653}
|
| 72 |
+
{"step": 1775, "loss": 2.7489702224731447, "grad_norm": 1.1250214576721191, "lr": 0.0001869914001510298, "s_per_step": 0.6029328918457031, "peak_gb": 44.713082368, "samples_seen": 56800, "time": 1791131872.1245182}
|
| 73 |
+
{"step": 1800, "loss": 2.840885500907898, "grad_norm": 1.192891240119934, "lr": 0.00018656220884543143, "s_per_step": 0.5853533840179443, "peak_gb": 44.713082368, "samples_seen": 57600, "time": 1791131886.7587817}
|
| 74 |
+
{"step": 1825, "loss": 2.7824712562561036, "grad_norm": 1.121734857559204, "lr": 0.00018612656106453871, "s_per_step": 0.5848813915252685, "peak_gb": 44.713082368, "samples_seen": 58400, "time": 1791131901.3812835}
|
| 75 |
+
{"step": 1850, "loss": 2.7347464036941527, "grad_norm": 1.163681149482727, "lr": 0.00018568448930231373, "s_per_step": 0.5818957328796387, "peak_gb": 44.713082368, "samples_seen": 59200, "time": 1791131915.9291348}
|
| 76 |
+
{"step": 1875, "loss": 2.7876648902893066, "grad_norm": 1.1300899982452393, "lr": 0.00018523602653186858, "s_per_step": 0.5813259029388428, "peak_gb": 44.713082368, "samples_seen": 60000, "time": 1791131930.462728}
|
| 77 |
+
{"step": 1900, "loss": 2.7792696571350097, "grad_norm": 1.1007814407348633, "lr": 0.00018478120620300578, "s_per_step": 0.5799147510528564, "peak_gb": 44.713082368, "samples_seen": 60800, "time": 1791131944.9613276}
|
| 78 |
+
{"step": 1925, "loss": 2.760815134048462, "grad_norm": 1.2840055227279663, "lr": 0.00018432006223972357, "s_per_step": 0.581923418045044, "peak_gb": 44.713082368, "samples_seen": 61600, "time": 1791131959.510147}
|
| 79 |
+
{"step": 1950, "loss": 2.72098069190979, "grad_norm": 1.1493003368377686, "lr": 0.00018385262903768547, "s_per_step": 0.5796582794189453, "peak_gb": 44.713082368, "samples_seen": 62400, "time": 1791131974.0020506}
|
| 80 |
+
{"step": 1975, "loss": 2.7809674453735354, "grad_norm": 1.0490471124649048, "lr": 0.00018337894146165467, "s_per_step": 0.5820777893066407, "peak_gb": 44.713082368, "samples_seen": 63200, "time": 1791131988.5544076}
|
| 81 |
+
{"step": 2000, "loss": 2.7531254482269287, "grad_norm": 1.0774011611938477, "lr": 0.00018289903484289383, "s_per_step": 0.5824406623840332, "peak_gb": 44.713082368, "samples_seen": 64000, "time": 1791132003.1158311}
|
| 82 |
+
{"step": 2025, "loss": 2.724808783531189, "grad_norm": 1.0720311403274536, "lr": 0.00018241294497652958, "s_per_step": 0.5992901039123535, "peak_gb": 44.713082368, "samples_seen": 64800, "time": 1791132018.0985038}
|
| 83 |
+
{"step": 2050, "loss": 2.790845193862915, "grad_norm": 1.2732620239257812, "lr": 0.00018192070811888272, "s_per_step": 0.5811270427703857, "peak_gb": 44.713082368, "samples_seen": 65600, "time": 1791132032.6270998}
|
| 84 |
+
{"step": 2075, "loss": 2.7737058067321776, "grad_norm": 1.0213253498077393, "lr": 0.00018142236098476398, "s_per_step": 0.5809852790832519, "peak_gb": 44.713082368, "samples_seen": 66400, "time": 1791132047.1521604}
|
| 85 |
+
{"step": 2100, "loss": 2.739048023223877, "grad_norm": 1.082443118095398, "lr": 0.00018091794074473537, "s_per_step": 0.5808903026580811, "peak_gb": 44.713082368, "samples_seen": 67200, "time": 1791132061.6748774}
|
| 86 |
+
{"step": 2125, "loss": 2.7185260105133056, "grad_norm": 1.1437773704528809, "lr": 0.00018040748502233802, "s_per_step": 0.5793789958953858, "peak_gb": 44.713082368, "samples_seen": 68000, "time": 1791132076.1597369}
|
| 87 |
+
{"step": 2150, "loss": 2.773498544692993, "grad_norm": 1.1243386268615723, "lr": 0.0001798910318912856, "s_per_step": 0.579986400604248, "peak_gb": 44.713082368, "samples_seen": 68800, "time": 1791132090.6597726}
|
| 88 |
+
{"step": 2175, "loss": 2.7247875928878784, "grad_norm": 1.1453449726104736, "lr": 0.00017936861987262475, "s_per_step": 0.6377030372619629, "peak_gb": 44.713082368, "samples_seen": 69600, "time": 1791132106.602996}
|
| 89 |
+
{"step": 2200, "loss": 2.758620800971985, "grad_norm": 1.0647848844528198, "lr": 0.0001788402879318618, "s_per_step": 0.583631067276001, "peak_gb": 44.713082368, "samples_seen": 70400, "time": 1791132121.194404}
|
| 90 |
+
{"step": 2225, "loss": 2.737944917678833, "grad_norm": 0.9973115921020508, "lr": 0.00017830607547605624, "s_per_step": 0.5805112552642823, "peak_gb": 44.713082368, "samples_seen": 71200, "time": 1791132135.7078857}
|
| 91 |
+
{"step": 2250, "loss": 2.7589354658126832, "grad_norm": 1.3257142305374146, "lr": 0.00017776602235088182, "s_per_step": 0.5818637657165527, "peak_gb": 44.713082368, "samples_seen": 72000, "time": 1791132150.255094}
|
| 92 |
+
{"step": 2275, "loss": 2.691265125274658, "grad_norm": 1.2526953220367432, "lr": 0.0001772201688376541, "s_per_step": 0.5970073795318603, "peak_gb": 44.713082368, "samples_seen": 72800, "time": 1791132165.180914}
|
| 93 |
+
{"step": 2300, "loss": 2.6809875345230103, "grad_norm": 1.1973172426223755, "lr": 0.00017666855565032642, "s_per_step": 0.581067304611206, "peak_gb": 44.713082368, "samples_seen": 73600, "time": 1791132179.7082095}
|
| 94 |
+
{"step": 2325, "loss": 2.7636598443984983, "grad_norm": 1.240551471710205, "lr": 0.00017611122393245275, "s_per_step": 0.5833901596069336, "peak_gb": 44.713082368, "samples_seen": 74400, "time": 1791132194.29364}
|
| 95 |
+
{"step": 2350, "loss": 2.6926284646987915, "grad_norm": 1.4566211700439453, "lr": 0.00017554821525411906, "s_per_step": 0.5827942657470703, "peak_gb": 44.713082368, "samples_seen": 75200, "time": 1791132208.8641186}
|
| 96 |
+
{"step": 2375, "loss": 2.7419660758972166, "grad_norm": 1.1245578527450562, "lr": 0.00017497957160884285, "s_per_step": 0.582745018005371, "peak_gb": 44.713082368, "samples_seen": 76000, "time": 1791132223.4334397}
|
| 97 |
+
{"step": 2400, "loss": 2.663253164291382, "grad_norm": 1.3107283115386963, "lr": 0.00017440533541044056, "s_per_step": 0.5869213008880615, "peak_gb": 44.713082368, "samples_seen": 76800, "time": 1791132238.1070802}
|
| 98 |
+
{"step": 2425, "loss": 2.719187321662903, "grad_norm": 1.1205676794052124, "lr": 0.00017382554948986445, "s_per_step": 0.5820108604431152, "peak_gb": 44.713082368, "samples_seen": 77600, "time": 1791132252.6579833}
|
| 99 |
+
{"step": 2450, "loss": 2.7780220317840576, "grad_norm": 1.3026337623596191, "lr": 0.00017324025709200766, "s_per_step": 0.5817010879516602, "peak_gb": 44.713082368, "samples_seen": 78400, "time": 1791132267.2012224}
|
| 100 |
+
{"step": 2475, "loss": 2.670451397895813, "grad_norm": 1.2014914751052856, "lr": 0.00017264950187247875, "s_per_step": 0.5823606967926025, "peak_gb": 44.713082368, "samples_seen": 79200, "time": 1791132281.7608585}
|
| 101 |
+
{"step": 2500, "loss": 2.716837215423584, "grad_norm": 1.10873281955719, "lr": 0.00017205332789434557, "s_per_step": 0.5853345394134521, "peak_gb": 44.713082368, "samples_seen": 80000, "time": 1791132296.3948615}
|
| 102 |
+
{"step": 2525, "loss": 2.7432607793807984, "grad_norm": 1.098906397819519, "lr": 0.00017145177962484857, "s_per_step": 0.5997240447998047, "peak_gb": 44.713082368, "samples_seen": 80800, "time": 1791132311.388538}
|
| 103 |
+
{"step": 2550, "loss": 2.72246383190155, "grad_norm": 1.1802773475646973, "lr": 0.00017084490193208435, "s_per_step": 0.5805338096618652, "peak_gb": 44.713082368, "samples_seen": 81600, "time": 1791132325.902455}
|
| 104 |
+
{"step": 2575, "loss": 2.7470813512802126, "grad_norm": 1.0650112628936768, "lr": 0.00017023274008165878, "s_per_step": 0.5804518604278565, "peak_gb": 44.713082368, "samples_seen": 82400, "time": 1791132340.4143882}
|
| 105 |
+
{"step": 2600, "loss": 2.733012390136719, "grad_norm": 1.1363471746444702, "lr": 0.00016961533973331081, "s_per_step": 0.5814594936370849, "peak_gb": 44.713082368, "samples_seen": 83200, "time": 1791132354.9514756}
|
| 106 |
+
{"step": 2625, "loss": 2.6893992137908938, "grad_norm": 1.1025614738464355, "lr": 0.00016899274693750696, "s_per_step": 0.5805016994476319, "peak_gb": 44.713082368, "samples_seen": 84000, "time": 1791132369.4646068}
|
| 107 |
+
{"step": 2650, "loss": 2.6821892070770263, "grad_norm": 1.081494688987732, "lr": 0.00016836500813200632, "s_per_step": 0.5812886905670166, "peak_gb": 44.713082368, "samples_seen": 84800, "time": 1791132383.9974716}
|
| 108 |
+
{"step": 2675, "loss": 2.6333597564697264, "grad_norm": 1.1299649477005005, "lr": 0.00016773217013839705, "s_per_step": 0.5820088291168213, "peak_gb": 44.713082368, "samples_seen": 85600, "time": 1791132398.5483217}
|
| 109 |
+
{"step": 2700, "loss": 2.6820573711395266, "grad_norm": 1.2370376586914062, "lr": 0.0001670942801586039, "s_per_step": 0.5822090339660645, "peak_gb": 44.713082368, "samples_seen": 86400, "time": 1791132413.1041813}
|
| 110 |
+
{"step": 2725, "loss": 2.695576796531677, "grad_norm": 1.1963305473327637, "lr": 0.0001664513857713677, "s_per_step": 0.5848300170898437, "peak_gb": 44.713082368, "samples_seen": 87200, "time": 1791132427.7255838}
|
| 111 |
+
{"step": 2750, "loss": 2.668781318664551, "grad_norm": 1.170694351196289, "lr": 0.00016580353492869635, "s_per_step": 0.5830868148803711, "peak_gb": 44.713082368, "samples_seen": 88000, "time": 1791132442.3036287}
|
| 112 |
+
{"step": 2775, "loss": 2.6674189043045042, "grad_norm": 1.4151184558868408, "lr": 0.00016515077595228837, "s_per_step": 0.6010197257995605, "peak_gb": 44.713082368, "samples_seen": 88800, "time": 1791132457.330008}
|
| 113 |
+
{"step": 2800, "loss": 2.716653389930725, "grad_norm": 1.140033483505249, "lr": 0.0001644931575299287, "s_per_step": 0.5838634395599365, "peak_gb": 44.713082368, "samples_seen": 89600, "time": 1791132471.9272482}
|
| 114 |
+
{"step": 2825, "loss": 2.7051201915740966, "grad_norm": 1.2430063486099243, "lr": 0.0001638307287118571, "s_per_step": 0.5820985507965087, "peak_gb": 44.713082368, "samples_seen": 90400, "time": 1791132486.4804387}
|
| 115 |
+
{"step": 2850, "loss": 2.6607730865478514, "grad_norm": 1.1695795059204102, "lr": 0.00016316353890710958, "s_per_step": 0.5837668800354003, "peak_gb": 44.713082368, "samples_seen": 91200, "time": 1791132501.0753098}
|
| 116 |
+
{"step": 2875, "loss": 2.6640982151031496, "grad_norm": 1.0638611316680908, "lr": 0.00016249163787983325, "s_per_step": 0.5824350929260254, "peak_gb": 44.713082368, "samples_seen": 92000, "time": 1791132515.6368423}
|
| 117 |
+
{"step": 2900, "loss": 2.6686540985107423, "grad_norm": 1.0542538166046143, "lr": 0.00016181507574557436, "s_per_step": 0.5836594104766846, "peak_gb": 44.713082368, "samples_seen": 92800, "time": 1791132530.229168}
|
| 118 |
+
{"step": 2925, "loss": 2.6495557260513305, "grad_norm": 1.2142537832260132, "lr": 0.00016113390296754036, "s_per_step": 0.5849769687652588, "peak_gb": 44.713082368, "samples_seen": 93600, "time": 1791132544.8542972}
|
| 119 |
+
{"step": 2950, "loss": 2.6701879692077637, "grad_norm": 1.2481223344802856, "lr": 0.00016044817035283605, "s_per_step": 0.5815197849273681, "peak_gb": 44.713082368, "samples_seen": 94400, "time": 1791132559.3929307}
|
| 120 |
+
{"step": 2975, "loss": 2.729030981063843, "grad_norm": 0.9954540729522705, "lr": 0.00015975792904867388, "s_per_step": 0.5838561344146729, "peak_gb": 44.713082368, "samples_seen": 95200, "time": 1791132573.989938}
|
| 121 |
+
{"step": 3000, "loss": 2.649655590057373, "grad_norm": 1.0716608762741089, "lr": 0.00015906323053855898, "s_per_step": 0.5867070293426514, "peak_gb": 44.713082368, "samples_seen": 96000, "time": 1791132588.6582751}
|
| 122 |
+
{"step": 3025, "loss": 2.6586125326156616, "grad_norm": 1.4092375040054321, "lr": 0.0001583641266384493, "s_per_step": 0.604687147140503, "peak_gb": 44.713082368, "samples_seen": 96800, "time": 1791132603.776329}
|
| 123 |
+
{"step": 3050, "loss": 2.6345360708236694, "grad_norm": 1.3088021278381348, "lr": 0.00015766066949289052, "s_per_step": 0.5844075584411621, "peak_gb": 44.713082368, "samples_seen": 97600, "time": 1791132618.3874304}
|
| 124 |
+
{"step": 3075, "loss": 2.697037835121155, "grad_norm": 1.2829176187515259, "lr": 0.00015695291157112693, "s_per_step": 0.5823062896728516, "peak_gb": 44.713082368, "samples_seen": 98400, "time": 1791132632.9458525}
|
| 125 |
+
{"step": 3100, "loss": 2.6461993885040282, "grad_norm": 1.1458219289779663, "lr": 0.0001562409056631878, "s_per_step": 0.5842306709289551, "peak_gb": 44.713082368, "samples_seen": 99200, "time": 1791132647.552255}
|
| 126 |
+
{"step": 3125, "loss": 2.6994865417480467, "grad_norm": 1.0527013540267944, "lr": 0.00015552470487594983, "s_per_step": 0.5832872867584229, "peak_gb": 44.713082368, "samples_seen": 100000, "time": 1791132662.1353507}
|
| 127 |
+
{"step": 3150, "loss": 2.6455965662002563, "grad_norm": 1.210248351097107, "lr": 0.00015480436262917615, "s_per_step": 0.5897402572631836, "peak_gb": 44.713082368, "samples_seen": 100800, "time": 1791132676.879246}
|
| 128 |
+
{"step": 3175, "loss": 2.7154378509521484, "grad_norm": 1.226425051689148, "lr": 0.0001540799326515317, "s_per_step": 0.5857205295562744, "peak_gb": 44.713082368, "samples_seen": 101600, "time": 1791132691.5226414}
|
| 129 |
+
{"step": 3200, "loss": 2.60671639919281, "grad_norm": 1.0465337038040161, "lr": 0.00015335146897657587, "s_per_step": 0.5847133922576905, "peak_gb": 44.713082368, "samples_seen": 102400, "time": 1791132706.1408818}
|
| 130 |
+
{"step": 3225, "loss": 2.5945321559906005, "grad_norm": 1.1408718824386597, "lr": 0.00015261902593873222, "s_per_step": 0.5843933010101319, "peak_gb": 44.713082368, "samples_seen": 103200, "time": 1791132720.7511504}
|
| 131 |
+
{"step": 3250, "loss": 2.627700114250183, "grad_norm": 1.0017207860946655, "lr": 0.00015188265816923585, "s_per_step": 0.5855321598052978, "peak_gb": 44.713082368, "samples_seen": 104000, "time": 1791132735.3898623}
|
| 132 |
+
{"step": 3275, "loss": 2.6447770547866822, "grad_norm": 1.066293478012085, "lr": 0.0001511424205920585, "s_per_step": 0.6015197467803955, "peak_gb": 44.713082368, "samples_seen": 104800, "time": 1791132750.4282053}
|
| 133 |
+
{"step": 3300, "loss": 2.6553041982650756, "grad_norm": 1.275646686553955, "lr": 0.00015039836841981185, "s_per_step": 0.5847999572753906, "peak_gb": 44.713082368, "samples_seen": 105600, "time": 1791132765.0488286}
|
| 134 |
+
{"step": 3325, "loss": 2.5941198587417604, "grad_norm": 1.0782554149627686, "lr": 0.00014965055714962952, "s_per_step": 0.5845184135437012, "peak_gb": 44.713082368, "samples_seen": 106400, "time": 1791132779.6623993}
|
| 135 |
+
{"step": 3350, "loss": 2.628866515159607, "grad_norm": 1.201255440711975, "lr": 0.0001488990425590276, "s_per_step": 0.5825023555755615, "peak_gb": 44.713082368, "samples_seen": 107200, "time": 1791132794.225617}
|
| 136 |
+
{"step": 3375, "loss": 2.6890606594085695, "grad_norm": 1.0845966339111328, "lr": 0.00014814388070174408, "s_per_step": 0.5846240043640136, "peak_gb": 44.713082368, "samples_seen": 108000, "time": 1791132808.841806}
|
| 137 |
+
{"step": 3400, "loss": 2.590783829689026, "grad_norm": 1.1155602931976318, "lr": 0.00014738512790355842, "s_per_step": 0.5839724159240722, "peak_gb": 44.713082368, "samples_seen": 108800, "time": 1791132823.4414928}
|
| 138 |
+
{"step": 3425, "loss": 2.6391691303253175, "grad_norm": 1.0816454887390137, "lr": 0.00014662284075808995, "s_per_step": 0.5850570201873779, "peak_gb": 44.713082368, "samples_seen": 109600, "time": 1791132838.0683677}
|
| 139 |
+
{"step": 3450, "loss": 2.6323579359054565, "grad_norm": 1.131683588027954, "lr": 0.00014585707612257672, "s_per_step": 0.5823802852630615, "peak_gb": 44.713082368, "samples_seen": 110400, "time": 1791132852.6284838}
|
| 140 |
+
{"step": 3475, "loss": 2.637928156852722, "grad_norm": 0.9959607720375061, "lr": 0.00014508789111363494, "s_per_step": 0.5853830051422119, "peak_gb": 44.713082368, "samples_seen": 111200, "time": 1791132867.2637053}
|
| 141 |
+
{"step": 3500, "loss": 2.65216420173645, "grad_norm": 1.3305087089538574, "lr": 0.00014431534310299838, "s_per_step": 0.5855358028411866, "peak_gb": 44.713082368, "samples_seen": 112000, "time": 1791132881.9027245}
|
| 142 |
+
{"step": 3525, "loss": 2.6313996505737305, "grad_norm": 1.2155920267105103, "lr": 0.00014353948971323942, "s_per_step": 0.6034837055206299, "peak_gb": 44.713082368, "samples_seen": 112800, "time": 1791132896.9904048}
|
| 143 |
+
{"step": 3550, "loss": 2.6454056882858277, "grad_norm": 1.2063369750976562, "lr": 0.00014276038881347105, "s_per_step": 0.5832828998565673, "peak_gb": 44.713082368, "samples_seen": 113600, "time": 1791132911.5728936}
|
| 144 |
+
{"step": 3575, "loss": 2.6533055901527405, "grad_norm": 1.2403618097305298, "lr": 0.00014197809851503049, "s_per_step": 0.5824703311920166, "peak_gb": 44.713082368, "samples_seen": 114400, "time": 1791132926.135257}
|
| 145 |
+
{"step": 3600, "loss": 2.645680537223816, "grad_norm": 1.168968915939331, "lr": 0.0001411926771671449, "s_per_step": 0.5845646953582764, "peak_gb": 44.713082368, "samples_seen": 115200, "time": 1791132940.7499871}
|
| 146 |
+
{"step": 3625, "loss": 2.5960699319839478, "grad_norm": 1.1931697130203247, "lr": 0.0001404041833525792, "s_per_step": 0.5853945255279541, "peak_gb": 44.713082368, "samples_seen": 116000, "time": 1791132955.3854115}
|
| 147 |
+
{"step": 3650, "loss": 2.6213397121429445, "grad_norm": 0.9856603741645813, "lr": 0.00013961267588326635, "s_per_step": 0.5852436065673828, "peak_gb": 44.713082368, "samples_seen": 116800, "time": 1791132970.017149}
|
| 148 |
+
{"step": 3675, "loss": 2.5998027467727662, "grad_norm": 1.1824589967727661, "lr": 0.0001388182137959211, "s_per_step": 0.5843393993377686, "peak_gb": 44.713082368, "samples_seen": 117600, "time": 1791132984.6262264}
|
| 149 |
+
{"step": 3700, "loss": 2.629755802154541, "grad_norm": 1.2840546369552612, "lr": 0.00013802085634763612, "s_per_step": 0.5829662704467773, "peak_gb": 44.713082368, "samples_seen": 118400, "time": 1791132999.2009673}
|
| 150 |
+
{"step": 3725, "loss": 2.6009298038482664, "grad_norm": 1.2750321626663208, "lr": 0.00013722066301146254, "s_per_step": 0.5825348472595215, "peak_gb": 44.713082368, "samples_seen": 119200, "time": 1791133013.764688}
|
| 151 |
+
{"step": 3750, "loss": 2.612462890148163, "grad_norm": 1.1066721677780151, "lr": 0.00013641769347197367, "s_per_step": 0.5850222778320312, "peak_gb": 44.713082368, "samples_seen": 120000, "time": 1791133028.3908353}
|
| 152 |
+
{"step": 3775, "loss": 2.591866898536682, "grad_norm": 1.233768343925476, "lr": 0.00013561200762081353, "s_per_step": 0.6027588081359864, "peak_gb": 44.713082368, "samples_seen": 120800, "time": 1791133043.4603734}
|
| 153 |
+
{"step": 3800, "loss": 2.578413248062134, "grad_norm": 1.108993649482727, "lr": 0.00013480366555222948, "s_per_step": 0.5819605350494385, "peak_gb": 44.713082368, "samples_seen": 121600, "time": 1791133058.0100117}
|
| 154 |
+
{"step": 3825, "loss": 2.547907509803772, "grad_norm": 0.9940484166145325, "lr": 0.00013399272755859002, "s_per_step": 0.5811294841766358, "peak_gb": 44.713082368, "samples_seen": 122400, "time": 1791133072.5388749}
|
| 155 |
+
{"step": 3850, "loss": 2.564617967605591, "grad_norm": 1.148879885673523, "lr": 0.00013317925412588777, "s_per_step": 0.5868311309814453, "peak_gb": 44.713082368, "samples_seen": 123200, "time": 1791133087.210223}
|
| 156 |
+
{"step": 3875, "loss": 2.587318305969238, "grad_norm": 1.0706918239593506, "lr": 0.00013236330592922784, "s_per_step": 0.5894640159606933, "peak_gb": 44.713082368, "samples_seen": 124000, "time": 1791133101.9471831}
|
| 157 |
+
{"step": 3900, "loss": 2.599490542411804, "grad_norm": 1.0891342163085938, "lr": 0.00013154494382830226, "s_per_step": 0.5840770053863525, "peak_gb": 44.713082368, "samples_seen": 124800, "time": 1791133116.5494475}
|
| 158 |
+
{"step": 3925, "loss": 2.628100709915161, "grad_norm": 1.1885673999786377, "lr": 0.00013072422886285063, "s_per_step": 0.5845713424682617, "peak_gb": 44.713082368, "samples_seen": 125600, "time": 1791133131.1641}
|
| 159 |
+
{"step": 3950, "loss": 2.6188903713226317, "grad_norm": 1.0750778913497925, "lr": 0.00012990122224810726, "s_per_step": 0.5842515659332276, "peak_gb": 44.713082368, "samples_seen": 126400, "time": 1791133145.7707798}
|
| 160 |
+
{"step": 3975, "loss": 2.577172713279724, "grad_norm": 1.1962721347808838, "lr": 0.00012907598537023528, "s_per_step": 0.5835914134979248, "peak_gb": 44.713082368, "samples_seen": 127200, "time": 1791133160.360931}
|
| 161 |
+
{"step": 4000, "loss": 2.5662565279006957, "grad_norm": 1.1467726230621338, "lr": 0.00012824857978174808, "s_per_step": 0.5806757926940918, "peak_gb": 44.713082368, "samples_seen": 128000, "time": 1791133174.8781514}
|
| 162 |
+
{"step": 4025, "loss": 2.6355027961730957, "grad_norm": 1.1590839624404907, "lr": 0.00012741906719691812, "s_per_step": 0.595490140914917, "peak_gb": 44.713082368, "samples_seen": 128800, "time": 1791133189.7657502}
|
| 163 |
+
{"step": 4050, "loss": 2.5592614030838012, "grad_norm": 0.9599524140357971, "lr": 0.0001265875094871738, "s_per_step": 0.5824095058441162, "peak_gb": 44.713082368, "samples_seen": 129600, "time": 1791133204.3263824}
|
| 164 |
+
{"step": 4075, "loss": 2.5762628889083863, "grad_norm": 0.9201562404632568, "lr": 0.0001257539686764847, "s_per_step": 0.5829432201385498, "peak_gb": 44.713082368, "samples_seen": 130400, "time": 1791133218.9003162}
|
| 165 |
+
{"step": 4100, "loss": 2.6393421840667726, "grad_norm": 1.352101445198059, "lr": 0.0001249185069367353, "s_per_step": 0.5822771739959717, "peak_gb": 44.713082368, "samples_seen": 131200, "time": 1791133233.4575992}
|
| 166 |
+
{"step": 4125, "loss": 2.545742371082306, "grad_norm": 1.2281466722488403, "lr": 0.00012408118658308783, "s_per_step": 0.5824446392059326, "peak_gb": 44.713082368, "samples_seen": 132000, "time": 1791133248.0190716}
|
| 167 |
+
{"step": 4150, "loss": 2.588506646156311, "grad_norm": 1.1411926746368408, "lr": 0.00012324207006933417, "s_per_step": 0.5804715347290039, "peak_gb": 44.713082368, "samples_seen": 132800, "time": 1791133262.5312164}
|
| 168 |
+
{"step": 4175, "loss": 2.5443468284606934, "grad_norm": 1.0073899030685425, "lr": 0.00012240121998323766, "s_per_step": 0.58156418800354, "peak_gb": 44.713082368, "samples_seen": 133600, "time": 1791133277.0706964}
|
| 169 |
+
{"step": 4200, "loss": 2.5501110649108885, "grad_norm": 1.205407977104187, "lr": 0.00012155869904186475, "s_per_step": 0.5811350154876709, "peak_gb": 44.713082368, "samples_seen": 134400, "time": 1791133291.5995052}
|
| 170 |
+
{"step": 4225, "loss": 2.602772789001465, "grad_norm": 1.2279781103134155, "lr": 0.00012071457008690715, "s_per_step": 0.5812279796600341, "peak_gb": 44.713082368, "samples_seen": 135200, "time": 1791133306.1305652}
|
| 171 |
+
{"step": 4250, "loss": 2.6113069915771483, "grad_norm": 1.1278338432312012, "lr": 0.00011986889607999463, "s_per_step": 0.580757417678833, "peak_gb": 44.713082368, "samples_seen": 136000, "time": 1791133320.6498582}
|
| 172 |
+
{"step": 4275, "loss": 2.566324009895325, "grad_norm": 1.2942112684249878, "lr": 0.00011902174009799881, "s_per_step": 0.5944104099273682, "peak_gb": 44.713082368, "samples_seen": 136800, "time": 1791133335.5104773}
|
| 173 |
+
{"step": 4300, "loss": 2.5078661108016966, "grad_norm": 1.2482675313949585, "lr": 0.00011817316532832833, "s_per_step": 0.5815795421600342, "peak_gb": 44.713082368, "samples_seen": 137600, "time": 1791133350.050364}
|
| 174 |
+
{"step": 4325, "loss": 2.6592298412322997, "grad_norm": 1.0717958211898804, "lr": 0.00011732323506421603, "s_per_step": 0.583767204284668, "peak_gb": 44.713082368, "samples_seen": 138400, "time": 1791133364.644909}
|
| 175 |
+
{"step": 4350, "loss": 2.6119984006881714, "grad_norm": 1.00667142868042, "lr": 0.00011647201269999792, "s_per_step": 0.5827497482299805, "peak_gb": 44.713082368, "samples_seen": 139200, "time": 1791133379.2140245}
|
| 176 |
+
{"step": 4375, "loss": 2.5818211317062376, "grad_norm": 1.1398861408233643, "lr": 0.0001156195617263847, "s_per_step": 0.6331087684631348, "peak_gb": 44.713082368, "samples_seen": 140000, "time": 1791133395.0420854}
|
| 177 |
+
{"step": 4400, "loss": 2.5258120822906496, "grad_norm": 1.2284903526306152, "lr": 0.00011476594572572632, "s_per_step": 0.5820397758483886, "peak_gb": 44.713082368, "samples_seen": 140800, "time": 1791133409.5934033}
|
| 178 |
+
{"step": 4425, "loss": 2.5755055570602416, "grad_norm": 1.141483187675476, "lr": 0.00011391122836726932, "s_per_step": 0.580828332901001, "peak_gb": 44.713082368, "samples_seen": 141600, "time": 1791133424.1144583}
|
| 179 |
+
{"step": 4450, "loss": 2.566441774368286, "grad_norm": 1.1366175413131714, "lr": 0.00011305547340240805, "s_per_step": 0.5789041805267334, "peak_gb": 44.713082368, "samples_seen": 142400, "time": 1791133438.5874844}
|
| 180 |
+
{"step": 4475, "loss": 2.561001148223877, "grad_norm": 1.0844303369522095, "lr": 0.00011219874465992948, "s_per_step": 0.5818092918395996, "peak_gb": 44.713082368, "samples_seen": 143200, "time": 1791133453.13306}
|
| 181 |
+
{"step": 4500, "loss": 2.5567925786972046, "grad_norm": 1.1574947834014893, "lr": 0.00011134110604125239, "s_per_step": 0.5787328243255615, "peak_gb": 44.713082368, "samples_seen": 144000, "time": 1791133467.6017637}
|
| 182 |
+
{"step": 4525, "loss": 2.590312509536743, "grad_norm": 1.257983684539795, "lr": 0.00011048262151566114, "s_per_step": 0.6001885604858398, "peak_gb": 44.713082368, "samples_seen": 144800, "time": 1791133482.606845}
|
| 183 |
+
{"step": 4550, "loss": 2.629988832473755, "grad_norm": 1.1604526042938232, "lr": 0.00010962335511553434, "s_per_step": 0.5829675579071045, "peak_gb": 44.713082368, "samples_seen": 145600, "time": 1791133497.1814847}
|
| 184 |
+
{"step": 4575, "loss": 2.5439812088012697, "grad_norm": 1.1712712049484253, "lr": 0.00010876337093156886, "s_per_step": 0.5811667346954346, "peak_gb": 44.713082368, "samples_seen": 146400, "time": 1791133511.711039}
|
| 185 |
+
{"step": 4600, "loss": 2.5177254247665406, "grad_norm": 1.1479181051254272, "lr": 0.00010790273310799933, "s_per_step": 0.5817756175994873, "peak_gb": 44.713082368, "samples_seen": 147200, "time": 1791133526.2557957}
|
| 186 |
+
{"step": 4625, "loss": 2.5491437196731566, "grad_norm": 1.2170847654342651, "lr": 0.00010704150583781388, "s_per_step": 0.5829072570800782, "peak_gb": 44.713082368, "samples_seen": 148000, "time": 1791133540.8288167}
|
| 187 |
+
{"step": 4650, "loss": 2.6045084619522094, "grad_norm": 1.189692497253418, "lr": 0.00010617975335796613, "s_per_step": 0.5828746032714843, "peak_gb": 44.713082368, "samples_seen": 148800, "time": 1791133555.4010704}
|
| 188 |
+
{"step": 4675, "loss": 2.576749534606934, "grad_norm": 1.1746689081192017, "lr": 0.00010531753994458382, "s_per_step": 0.5798631191253663, "peak_gb": 44.713082368, "samples_seen": 149600, "time": 1791133569.898002}
|
| 189 |
+
{"step": 4700, "loss": 2.5879501485824585, "grad_norm": 1.182418704032898, "lr": 0.00010445492990817471, "s_per_step": 0.58068115234375, "peak_gb": 44.713082368, "samples_seen": 150400, "time": 1791133584.415375}
|
| 190 |
+
{"step": 4725, "loss": 2.5357809329032897, "grad_norm": 1.4740160703659058, "lr": 0.00010359198758882973, "s_per_step": 0.5839321517944336, "peak_gb": 44.713082368, "samples_seen": 151200, "time": 1791133599.0140464}
|
| 191 |
+
{"step": 4750, "loss": 2.55805965423584, "grad_norm": 1.172712802886963, "lr": 0.00010272877735142407, "s_per_step": 0.5821543216705323, "peak_gb": 44.713082368, "samples_seen": 152000, "time": 1791133613.568319}
|
| 192 |
+
{"step": 4775, "loss": 2.5926496958732606, "grad_norm": 1.194968581199646, "lr": 0.00010186536358081623, "s_per_step": 0.5943687915802002, "peak_gb": 44.713082368, "samples_seen": 152800, "time": 1791133628.427965}
|
| 193 |
+
{"step": 4800, "loss": 2.5890002751350405, "grad_norm": 1.1471878290176392, "lr": 0.00010100181067704579, "s_per_step": 0.5851948833465577, "peak_gb": 44.713082368, "samples_seen": 153600, "time": 1791133643.0582457}
|
| 194 |
+
{"step": 4825, "loss": 2.5580673170089723, "grad_norm": 1.1443222761154175, "lr": 0.00010013818305053007, "s_per_step": 0.5806027603149414, "peak_gb": 44.713082368, "samples_seen": 154400, "time": 1791133657.573739}
|
| 195 |
+
{"step": 4850, "loss": 2.5510888147354125, "grad_norm": 1.1550148725509644, "lr": 9.927454511725963e-05, "s_per_step": 0.5874297904968262, "peak_gb": 44.947974656, "samples_seen": 155200, "time": 1791133672.2599056}
|
| 196 |
+
{"step": 4875, "loss": 2.5219507360458375, "grad_norm": 1.0392953157424927, "lr": 9.841096129399391e-05, "s_per_step": 0.5823034572601319, "peak_gb": 44.947974656, "samples_seen": 156000, "time": 1791133686.8179238}
|
| 197 |
+
{"step": 4900, "loss": 2.5420716381073, "grad_norm": 1.2059009075164795, "lr": 9.754749599345636e-05, "s_per_step": 0.585919017791748, "peak_gb": 44.947974656, "samples_seen": 156800, "time": 1791133701.4663377}
|
| 198 |
+
{"step": 4925, "loss": 2.5494388055801394, "grad_norm": 1.132745385169983, "lr": 9.668421361953004e-05, "s_per_step": 0.5843277454376221, "peak_gb": 44.947974656, "samples_seen": 157600, "time": 1791133716.0755186}
|
| 199 |
+
{"step": 4950, "loss": 2.534179072380066, "grad_norm": 1.1565256118774414, "lr": 9.582117856245404e-05, "s_per_step": 0.5862970352172852, "peak_gb": 44.947974656, "samples_seen": 158400, "time": 1791133730.7334623}
|
| 200 |
+
{"step": 4975, "loss": 2.575369110107422, "grad_norm": 1.273497223854065, "lr": 9.495845519402057e-05, "s_per_step": 0.582844295501709, "peak_gb": 44.947974656, "samples_seen": 159200, "time": 1791133745.3050113}
|
| 201 |
+
{"step": 5000, "loss": 2.5750187397003175, "grad_norm": 1.1561633348464966, "lr": 9.409610786277377e-05, "s_per_step": 0.5810141372680664, "peak_gb": 44.947974656, "samples_seen": 160000, "time": 1791133759.830802}
|
| 202 |
+
{"step": 5025, "loss": 2.560807399749756, "grad_norm": 1.3111895322799683, "lr": 9.323420088921002e-05, "s_per_step": 0.595534782409668, "peak_gb": 44.947974656, "samples_seen": 160800, "time": 1791133774.7197492}
|
| 203 |
+
{"step": 5050, "loss": 2.5090624332427978, "grad_norm": 1.128811240196228, "lr": 9.237279856098037e-05, "s_per_step": 0.5866300392150879, "peak_gb": 44.947974656, "samples_seen": 161600, "time": 1791133789.385984}
|
| 204 |
+
{"step": 5075, "loss": 2.5157794332504273, "grad_norm": 1.1408394575119019, "lr": 9.15119651280956e-05, "s_per_step": 0.5873813438415527, "peak_gb": 44.947974656, "samples_seen": 162400, "time": 1791133804.070975}
|
| 205 |
+
{"step": 5100, "loss": 2.4853475761413573, "grad_norm": 1.1142104864120483, "lr": 9.065176479813393e-05, "s_per_step": 0.5891704750061035, "peak_gb": 44.947974656, "samples_seen": 163200, "time": 1791133818.8006737}
|
| 206 |
+
{"step": 5125, "loss": 2.535648889541626, "grad_norm": 1.1073273420333862, "lr": 8.979226173145183e-05, "s_per_step": 0.5848058700561524, "peak_gb": 44.947974656, "samples_seen": 164000, "time": 1791133833.421258}
|
| 207 |
+
{"step": 5150, "loss": 2.582467722892761, "grad_norm": 1.2277833223342896, "lr": 8.893352003639856e-05, "s_per_step": 0.5909334659576416, "peak_gb": 44.947974656, "samples_seen": 164800, "time": 1791133848.1950076}
|
| 208 |
+
{"step": 5175, "loss": 2.624933614730835, "grad_norm": 1.0772781372070312, "lr": 8.807560376453437e-05, "s_per_step": 0.580198278427124, "peak_gb": 44.947974656, "samples_seen": 165600, "time": 1791133862.700406}
|
| 209 |
+
{"step": 5200, "loss": 2.5316876649856566, "grad_norm": 1.251028299331665, "lr": 8.721857690585314e-05, "s_per_step": 0.5918057537078858, "peak_gb": 44.947974656, "samples_seen": 166400, "time": 1791133877.495987}
|
| 210 |
+
{"step": 5225, "loss": 2.5712487649917604, "grad_norm": 1.1420977115631104, "lr": 8.636250338400954e-05, "s_per_step": 0.5819726657867431, "peak_gb": 44.947974656, "samples_seen": 167200, "time": 1791133892.0457838}
|
| 211 |
+
{"step": 5250, "loss": 2.5827745389938355, "grad_norm": 1.37429678440094, "lr": 8.550744705155087e-05, "s_per_step": 0.5890950107574463, "peak_gb": 44.947974656, "samples_seen": 168000, "time": 1791133906.7736533}
|
| 212 |
+
{"step": 5275, "loss": 2.493775963783264, "grad_norm": 1.1473252773284912, "lr": 8.465347168515478e-05, "s_per_step": 0.5988393306732178, "peak_gb": 44.947974656, "samples_seen": 168800, "time": 1791133921.745074}
|
| 213 |
+
{"step": 5300, "loss": 2.5703449940681455, "grad_norm": 1.315991997718811, "lr": 8.380064098087212e-05, "s_per_step": 0.5884280204772949, "peak_gb": 44.947974656, "samples_seen": 169600, "time": 1791133936.4563115}
|
| 214 |
+
{"step": 5325, "loss": 2.5584110116958616, "grad_norm": 1.3650532960891724, "lr": 8.294901854937599e-05, "s_per_step": 0.5825414657592773, "peak_gb": 44.947974656, "samples_seen": 170400, "time": 1791133951.0203662}
|
| 215 |
+
{"step": 5350, "loss": 2.5070259499549867, "grad_norm": 1.0850861072540283, "lr": 8.20986679112173e-05, "s_per_step": 0.5887926959991455, "peak_gb": 44.947974656, "samples_seen": 171200, "time": 1791133965.7406182}
|
| 216 |
+
{"step": 5375, "loss": 2.5562100505828855, "grad_norm": 1.1863794326782227, "lr": 8.124965249208671e-05, "s_per_step": 0.5837687969207763, "peak_gb": 44.947974656, "samples_seen": 172000, "time": 1791133980.3353226}
|
| 217 |
+
{"step": 5400, "loss": 2.5509732866287234, "grad_norm": 1.3163355588912964, "lr": 8.040203561808406e-05, "s_per_step": 0.5839274787902832, "peak_gb": 44.947974656, "samples_seen": 172800, "time": 1791133994.934007}
|
| 218 |
+
{"step": 5425, "loss": 2.518836889266968, "grad_norm": 1.0753092765808105, "lr": 7.955588051099495e-05, "s_per_step": 0.5834451007843018, "peak_gb": 44.947974656, "samples_seen": 173600, "time": 1791134009.5206103}
|
| 219 |
+
{"step": 5450, "loss": 2.486795573234558, "grad_norm": 1.2447965145111084, "lr": 7.87112502835751e-05, "s_per_step": 0.590032205581665, "peak_gb": 45.225538048, "samples_seen": 174400, "time": 1791134024.2720983}
|
| 220 |
+
{"step": 5475, "loss": 2.5397833967208863, "grad_norm": 1.2422091960906982, "lr": 7.786820793484306e-05, "s_per_step": 0.5841742324829101, "peak_gb": 45.225538048, "samples_seen": 175200, "time": 1791134038.8771572}
|
| 221 |
+
{"step": 5500, "loss": 2.582821946144104, "grad_norm": 1.1261637210845947, "lr": 7.702681634538104e-05, "s_per_step": 0.584408836364746, "peak_gb": 45.225538048, "samples_seen": 176000, "time": 1791134053.4881816}
|
| 222 |
+
{"step": 5525, "loss": 2.5332029485702514, "grad_norm": 1.3875887393951416, "lr": 7.618713827264505e-05, "s_per_step": 0.6076767444610596, "peak_gb": 45.225538048, "samples_seen": 176800, "time": 1791134068.680784}
|
| 223 |
+
{"step": 5550, "loss": 2.577248592376709, "grad_norm": 1.1470293998718262, "lr": 7.534923634628381e-05, "s_per_step": 0.582500295639038, "peak_gb": 45.225538048, "samples_seen": 177600, "time": 1791134083.2440093}
|
| 224 |
+
{"step": 5575, "loss": 2.521993110179901, "grad_norm": 1.2665389776229858, "lr": 7.451317306346738e-05, "s_per_step": 0.5894964313507081, "peak_gb": 45.225538048, "samples_seen": 178400, "time": 1791134097.9821842}
|
| 225 |
+
{"step": 5600, "loss": 2.538360843658447, "grad_norm": 1.314635992050171, "lr": 7.367901078422563e-05, "s_per_step": 0.5822025871276856, "peak_gb": 45.225538048, "samples_seen": 179200, "time": 1791134112.5379786}
|
| 226 |
+
{"step": 5625, "loss": 2.519887990951538, "grad_norm": 1.2479692697525024, "lr": 7.284681172679697e-05, "s_per_step": 0.5847748279571533, "peak_gb": 45.225538048, "samples_seen": 180000, "time": 1791134127.1580513}
|
| 227 |
+
{"step": 5650, "loss": 2.55150728225708, "grad_norm": 1.3254190683364868, "lr": 7.201663796298764e-05, "s_per_step": 0.5816287231445313, "peak_gb": 45.225538048, "samples_seen": 180800, "time": 1791134141.6998255}
|
| 228 |
+
{"step": 5675, "loss": 2.480019929409027, "grad_norm": 1.3282145261764526, "lr": 7.118855141354186e-05, "s_per_step": 0.5884755229949952, "peak_gb": 45.225538048, "samples_seen": 181600, "time": 1791134156.4120905}
|
| 229 |
+
{"step": 5700, "loss": 2.495427899360657, "grad_norm": 1.325265884399414, "lr": 7.036261384352343e-05, "s_per_step": 0.5823914813995361, "peak_gb": 45.225538048, "samples_seen": 182400, "time": 1791134170.972266}
|
| 230 |
+
{"step": 5725, "loss": 2.478832368850708, "grad_norm": 1.0613092184066772, "lr": 6.953888685770869e-05, "s_per_step": 0.5816488647460938, "peak_gb": 45.225538048, "samples_seen": 183200, "time": 1791134185.5139039}
|
| 231 |
+
{"step": 5750, "loss": 2.52920090675354, "grad_norm": 1.207517385482788, "lr": 6.871743189599156e-05, "s_per_step": 0.5833349800109864, "peak_gb": 45.225538048, "samples_seen": 184000, "time": 1791134200.0977366}
|
| 232 |
+
{"step": 5775, "loss": 2.5796267938613893, "grad_norm": 1.1419939994812012, "lr": 6.789831022880101e-05, "s_per_step": 0.5960377788543701, "peak_gb": 45.225538048, "samples_seen": 184800, "time": 1791134214.9990556}
|
| 233 |
+
{"step": 5800, "loss": 2.52916264295578, "grad_norm": 1.3384249210357666, "lr": 6.708158295253092e-05, "s_per_step": 0.5907851696014405, "peak_gb": 45.225538048, "samples_seen": 185600, "time": 1791134229.7691712}
|
| 234 |
+
{"step": 5825, "loss": 2.5286863899230956, "grad_norm": 1.351222276687622, "lr": 6.626731098498311e-05, "s_per_step": 0.5822148704528809, "peak_gb": 45.225538048, "samples_seen": 186400, "time": 1791134244.3249931}
|
| 235 |
+
{"step": 5850, "loss": 2.5227443981170654, "grad_norm": 1.1627928018569946, "lr": 6.545555506082355e-05, "s_per_step": 0.5858915615081787, "peak_gb": 45.225538048, "samples_seen": 187200, "time": 1791134258.9727561}
|
| 236 |
+
{"step": 5875, "loss": 2.534986820220947, "grad_norm": 1.2816951274871826, "lr": 6.464637572705237e-05, "s_per_step": 0.5801717948913574, "peak_gb": 45.225538048, "samples_seen": 188000, "time": 1791134273.4775524}
|
| 237 |
+
{"step": 5900, "loss": 2.489676251411438, "grad_norm": 1.4309545755386353, "lr": 6.383983333848773e-05, "s_per_step": 0.5923868274688721, "peak_gb": 45.225538048, "samples_seen": 188800, "time": 1791134288.2877378}
|
| 238 |
+
{"step": 5925, "loss": 2.4901873302459716, "grad_norm": 1.322837471961975, "lr": 6.303598805326419e-05, "s_per_step": 0.5796713924407959, "peak_gb": 45.225538048, "samples_seen": 189600, "time": 1791134302.7800236}
|
| 239 |
+
{"step": 5950, "loss": 2.53445716381073, "grad_norm": 1.3041024208068848, "lr": 6.223489982834558e-05, "s_per_step": 0.5908190441131592, "peak_gb": 45.225538048, "samples_seen": 190400, "time": 1791134317.550958}
|
| 240 |
+
{"step": 5975, "loss": 2.523839635848999, "grad_norm": 1.1598639488220215, "lr": 6.1436628415053e-05, "s_per_step": 0.5810464763641358, "peak_gb": 45.225538048, "samples_seen": 191200, "time": 1791134332.0775537}
|
| 241 |
+
{"step": 6000, "loss": 2.513823781013489, "grad_norm": 1.0870771408081055, "lr": 6.064123335460796e-05, "s_per_step": 0.5868215465545654, "peak_gb": 45.225538048, "samples_seen": 192000, "time": 1791134346.748537}
|
| 242 |
+
{"step": 6025, "loss": 2.589872555732727, "grad_norm": 1.2455843687057495, "lr": 5.9848773973691574e-05, "s_per_step": 0.5989699459075928, "peak_gb": 45.225538048, "samples_seen": 192800, "time": 1791134361.7232323}
|
| 243 |
+
{"step": 6050, "loss": 2.5560085201263427, "grad_norm": 1.2540245056152344, "lr": 5.905930938001937e-05, "s_per_step": 0.5814158630371093, "peak_gb": 45.225538048, "samples_seen": 193600, "time": 1791134376.2590868}
|
| 244 |
+
{"step": 6075, "loss": 2.5238016700744628, "grad_norm": 1.5046987533569336, "lr": 5.827289845793256e-05, "s_per_step": 0.5851035404205323, "peak_gb": 45.225538048, "samples_seen": 194400, "time": 1791134390.8871312}
|
| 245 |
+
{"step": 6100, "loss": 2.528533492088318, "grad_norm": 1.174469232559204, "lr": 5.748959986400611e-05, "s_per_step": 0.5847522258758545, "peak_gb": 45.225538048, "samples_seen": 195200, "time": 1791134405.5064287}
|
| 246 |
+
{"step": 6125, "loss": 2.5303329944610597, "grad_norm": 1.1631888151168823, "lr": 5.670947202267356e-05, "s_per_step": 0.5879807376861572, "peak_gb": 45.225538048, "samples_seen": 196000, "time": 1791134420.2064354}
|
| 247 |
+
{"step": 6150, "loss": 2.5317347621917725, "grad_norm": 1.316127896308899, "lr": 5.593257312186937e-05, "s_per_step": 0.5837984085083008, "peak_gb": 45.225538048, "samples_seen": 196800, "time": 1791134434.8018806}
|
| 248 |
+
{"step": 6175, "loss": 2.434650685787201, "grad_norm": 1.1741480827331543, "lr": 5.5158961108688766e-05, "s_per_step": 0.5790462970733643, "peak_gb": 45.225538048, "samples_seen": 197600, "time": 1791134449.2784874}
|
| 249 |
+
{"step": 6200, "loss": 2.531181104183197, "grad_norm": 1.1735306978225708, "lr": 5.438869368506563e-05, "s_per_step": 0.584972219467163, "peak_gb": 45.225538048, "samples_seen": 198400, "time": 1791134463.9032767}
|
| 250 |
+
{"step": 6225, "loss": 2.527851085662842, "grad_norm": 1.0472126007080078, "lr": 5.36218283034686e-05, "s_per_step": 0.5841711044311524, "peak_gb": 45.225538048, "samples_seen": 199200, "time": 1791134478.5080554}
|
| 251 |
+
{"step": 6250, "loss": 2.4174211168289186, "grad_norm": 1.081107258796692, "lr": 5.285842216261594e-05, "s_per_step": 0.5814971351623535, "peak_gb": 45.225538048, "samples_seen": 200000, "time": 1791134493.045984}
|
| 252 |
+
{"step": 6275, "loss": 2.438464779853821, "grad_norm": 1.2638477087020874, "lr": 5.209853220320898e-05, "s_per_step": 0.5986836051940918, "peak_gb": 45.225538048, "samples_seen": 200800, "time": 1791134508.0135267}
|
| 253 |
+
{"step": 6300, "loss": 2.541532530784607, "grad_norm": 1.1554198265075684, "lr": 5.134221510368531e-05, "s_per_step": 0.5808733463287353, "peak_gb": 45.225538048, "samples_seen": 201600, "time": 1791134522.535798}
|
| 254 |
+
{"step": 6325, "loss": 2.5302062940597536, "grad_norm": 1.0484192371368408, "lr": 5.058952727599113e-05, "s_per_step": 0.5848421478271484, "peak_gb": 45.225538048, "samples_seen": 202400, "time": 1791134537.1572866}
|
| 255 |
+
{"step": 6350, "loss": 2.5093615889549254, "grad_norm": 1.3501886129379272, "lr": 4.984052486137364e-05, "s_per_step": 0.5831260204315185, "peak_gb": 45.225538048, "samples_seen": 203200, "time": 1791134551.7359152}
|
| 256 |
+
{"step": 6375, "loss": 2.4862116837501524, "grad_norm": 1.3331636190414429, "lr": 4.909526372619356e-05, "s_per_step": 0.5831469821929932, "peak_gb": 45.225538048, "samples_seen": 204000, "time": 1791134566.3150558}
|
| 257 |
+
{"step": 6400, "loss": 2.553971781730652, "grad_norm": 1.2492870092391968, "lr": 4.835379945775823e-05, "s_per_step": 0.5849734497070312, "peak_gb": 45.225538048, "samples_seen": 204800, "time": 1791134580.9397955}
|
| 258 |
+
{"step": 6425, "loss": 2.5170671367645263, "grad_norm": 1.2050851583480835, "lr": 4.7616187360175476e-05, "s_per_step": 0.5841898059844971, "peak_gb": 45.225538048, "samples_seen": 205600, "time": 1791134595.5449216}
|
| 259 |
+
{"step": 6450, "loss": 2.5319773864746096, "grad_norm": 1.1314105987548828, "lr": 4.6882482450228574e-05, "s_per_step": 0.5848858451843262, "peak_gb": 45.225538048, "samples_seen": 206400, "time": 1791134610.1675096}
|
| 260 |
+
{"step": 6475, "loss": 2.4931352043151858, "grad_norm": 1.1107985973358154, "lr": 4.615273945327272e-05, "s_per_step": 0.5831267642974853, "peak_gb": 45.225538048, "samples_seen": 207200, "time": 1791134624.7460885}
|
| 261 |
+
{"step": 6500, "loss": 2.5409064221382143, "grad_norm": 1.199941873550415, "lr": 4.542701279915318e-05, "s_per_step": 0.5819946575164795, "peak_gb": 45.225538048, "samples_seen": 208000, "time": 1791134639.2963989}
|
| 262 |
+
{"step": 6525, "loss": 2.49187207698822, "grad_norm": 1.4973266124725342, "lr": 4.470535661814538e-05, "s_per_step": 0.6058605766296387, "peak_gb": 45.225538048, "samples_seen": 208800, "time": 1791134654.4432883}
|
| 263 |
+
{"step": 6550, "loss": 2.490929820537567, "grad_norm": 1.222715139389038, "lr": 4.398782473691767e-05, "s_per_step": 0.5849639797210693, "peak_gb": 45.225538048, "samples_seen": 209600, "time": 1791134669.0677845}
|
| 264 |
+
{"step": 6575, "loss": 2.478930926322937, "grad_norm": 1.1837711334228516, "lr": 4.327447067451636e-05, "s_per_step": 0.5865503215789795, "peak_gb": 45.225538048, "samples_seen": 210400, "time": 1791134683.731912}
|
| 265 |
+
{"step": 6600, "loss": 2.505696654319763, "grad_norm": 1.1841992139816284, "lr": 4.256534763837391e-05, "s_per_step": 0.5857251071929932, "peak_gb": 45.225538048, "samples_seen": 211200, "time": 1791134698.375444}
|
| 266 |
+
{"step": 6625, "loss": 2.509692039489746, "grad_norm": 1.251827597618103, "lr": 4.186050852034028e-05, "s_per_step": 0.5832119750976562, "peak_gb": 45.225538048, "samples_seen": 212000, "time": 1791134712.9561343}
|
| 267 |
+
{"step": 6650, "loss": 2.4981483721733095, "grad_norm": 1.3467367887496948, "lr": 4.1160005892737896e-05, "s_per_step": 0.5832837581634521, "peak_gb": 45.225538048, "samples_seen": 212800, "time": 1791134727.5386372}
|
| 268 |
+
{"step": 6675, "loss": 2.487072801589966, "grad_norm": 1.149620771408081, "lr": 4.046389200444035e-05, "s_per_step": 0.5866239070892334, "peak_gb": 45.225538048, "samples_seen": 213600, "time": 1791134742.2046556}
|
| 269 |
+
{"step": 6700, "loss": 2.464690637588501, "grad_norm": 1.290765643119812, "lr": 3.9772218776975314e-05, "s_per_step": 0.5862217617034912, "peak_gb": 45.225538048, "samples_seen": 214400, "time": 1791134756.8606098}
|
| 270 |
+
{"step": 6725, "loss": 2.4712146949768066, "grad_norm": 1.1348451375961304, "lr": 3.908503780065186e-05, "s_per_step": 0.5837146186828613, "peak_gb": 45.225538048, "samples_seen": 215200, "time": 1791134771.4539366}
|
| 271 |
+
{"step": 6750, "loss": 2.4894593477249147, "grad_norm": 1.0581586360931396, "lr": 3.840240033071231e-05, "s_per_step": 0.5876280879974365, "peak_gb": 45.225538048, "samples_seen": 216000, "time": 1791134786.1450136}
|
| 272 |
+
{"step": 6775, "loss": 2.516192195415497, "grad_norm": 1.3426604270935059, "lr": 3.772435728350944e-05, "s_per_step": 0.6016309452056885, "peak_gb": 45.225538048, "samples_seen": 216800, "time": 1791134801.1862383}
|
| 273 |
+
{"step": 6800, "loss": 2.5307986450195314, "grad_norm": 1.3854995965957642, "lr": 3.70509592327086e-05, "s_per_step": 0.5847000885009765, "peak_gb": 45.225538048, "samples_seen": 217600, "time": 1791134815.8041387}
|
| 274 |
+
{"step": 6825, "loss": 2.5088115167617797, "grad_norm": 1.2087575197219849, "lr": 3.638225640551563e-05, "s_per_step": 0.5863976192474365, "peak_gb": 45.225538048, "samples_seen": 218400, "time": 1791134830.4645412}
|
| 275 |
+
{"step": 6850, "loss": 2.4794558453559876, "grad_norm": 1.3083394765853882, "lr": 3.571829867893038e-05, "s_per_step": 0.5870515918731689, "peak_gb": 45.225538048, "samples_seen": 219200, "time": 1791134845.1412776}
|
| 276 |
+
{"step": 6875, "loss": 2.495806655883789, "grad_norm": 1.250901222229004, "lr": 3.50591355760267e-05, "s_per_step": 0.5858700180053711, "peak_gb": 45.225538048, "samples_seen": 220000, "time": 1791134859.7884114}
|
| 277 |
+
{"step": 6900, "loss": 2.472851881980896, "grad_norm": 1.4390413761138916, "lr": 3.440481626225849e-05, "s_per_step": 0.5833972072601319, "peak_gb": 45.225538048, "samples_seen": 220800, "time": 1791134874.3737664}
|
| 278 |
+
{"step": 6925, "loss": 2.4779190325737, "grad_norm": 1.14964759349823, "lr": 3.3755389541792606e-05, "s_per_step": 0.5836262893676758, "peak_gb": 45.225538048, "samples_seen": 221600, "time": 1791134888.9648018}
|
| 279 |
+
{"step": 6950, "loss": 2.4924482202529905, "grad_norm": 1.1170302629470825, "lr": 3.311090385386867e-05, "s_per_step": 0.5860085487365723, "peak_gb": 45.225538048, "samples_seen": 222400, "time": 1791134903.6154745}
|
| 280 |
+
{"step": 6975, "loss": 2.449902744293213, "grad_norm": 1.2713782787322998, "lr": 3.247140726918607e-05, "s_per_step": 0.5829716205596924, "peak_gb": 45.225538048, "samples_seen": 223200, "time": 1791134918.190196}
|
| 281 |
+
{"step": 7000, "loss": 2.5060114574432375, "grad_norm": 1.2075129747390747, "lr": 3.183694748631856e-05, "s_per_step": 0.5825292396545411, "peak_gb": 45.225538048, "samples_seen": 224000, "time": 1791134932.753835}
|
| 282 |
+
{"step": 7025, "loss": 2.5190657806396484, "grad_norm": 1.2806220054626465, "lr": 3.1207571828156415e-05, "s_per_step": 0.6018135643005371, "peak_gb": 45.225538048, "samples_seen": 224800, "time": 1791134947.799628}
|
| 283 |
+
{"step": 7050, "loss": 2.488545150756836, "grad_norm": 1.1726572513580322, "lr": 3.0583327238376826e-05, "s_per_step": 0.5843643283843994, "peak_gb": 45.225538048, "samples_seen": 225600, "time": 1791134962.4091752}
|
| 284 |
+
{"step": 7075, "loss": 2.45016818523407, "grad_norm": 1.3082247972488403, "lr": 2.9964260277942414e-05, "s_per_step": 0.5856728076934814, "peak_gb": 45.225538048, "samples_seen": 226400, "time": 1791134977.0514216}
|
| 285 |
+
{"step": 7100, "loss": 2.4964564561843874, "grad_norm": 1.2286127805709839, "lr": 2.9350417121628438e-05, "s_per_step": 0.5862992858886719, "peak_gb": 45.225538048, "samples_seen": 227200, "time": 1791134991.7094052}
|
| 286 |
+
{"step": 7125, "loss": 2.4161909675598143, "grad_norm": 1.1147135496139526, "lr": 2.8741843554578563e-05, "s_per_step": 0.5842507934570312, "peak_gb": 45.225538048, "samples_seen": 228000, "time": 1791135006.3160708}
|
| 287 |
+
{"step": 7150, "loss": 2.5295400643348693, "grad_norm": 1.2771178483963013, "lr": 2.8138584968890036e-05, "s_per_step": 0.5852422046661377, "peak_gb": 45.225538048, "samples_seen": 228800, "time": 1791135020.9474854}
|
| 288 |
+
{"step": 7175, "loss": 2.5087658619880675, "grad_norm": 1.2442432641983032, "lr": 2.7540686360227917e-05, "s_per_step": 0.5861521339416504, "peak_gb": 45.225538048, "samples_seen": 229600, "time": 1791135035.6017551}
|
| 289 |
+
{"step": 7200, "loss": 2.4834299731254577, "grad_norm": 1.3997399806976318, "lr": 2.6948192324468925e-05, "s_per_step": 0.5873078346252442, "peak_gb": 45.225538048, "samples_seen": 230400, "time": 1791135050.2848673}
|
| 290 |
+
{"step": 7225, "loss": 2.44143901348114, "grad_norm": 1.4422211647033691, "lr": 2.636114705437519e-05, "s_per_step": 0.5871801567077637, "peak_gb": 45.225538048, "samples_seen": 231200, "time": 1791135064.9647627}
|
| 291 |
+
{"step": 7250, "loss": 2.470247166156769, "grad_norm": 1.3490147590637207, "lr": 2.5779594336297975e-05, "s_per_step": 0.5863297176361084, "peak_gb": 45.225538048, "samples_seen": 232000, "time": 1791135079.6234357}
|
| 292 |
+
{"step": 7275, "loss": 2.471605281829834, "grad_norm": 1.129075288772583, "lr": 2.5203577546911773e-05, "s_per_step": 0.599454345703125, "peak_gb": 45.225538048, "samples_seen": 232800, "time": 1791135094.6101875}
|
| 293 |
+
{"step": 7300, "loss": 2.440209083557129, "grad_norm": 1.0634057521820068, "lr": 2.463313964997894e-05, "s_per_step": 0.5864029026031494, "peak_gb": 45.225538048, "samples_seen": 233600, "time": 1791135109.2706358}
|
| 294 |
+
{"step": 7325, "loss": 2.5060981726646423, "grad_norm": 1.265033483505249, "lr": 2.4068323193145126e-05, "s_per_step": 0.5880957889556885, "peak_gb": 45.225538048, "samples_seen": 234400, "time": 1791135123.973469}
|
| 295 |
+
{"step": 7350, "loss": 2.529538559913635, "grad_norm": 1.1425158977508545, "lr": 2.350917030476578e-05, "s_per_step": 0.5846072196960449, "peak_gb": 45.225538048, "samples_seen": 235200, "time": 1791135138.5890818}
|
| 296 |
+
{"step": 7375, "loss": 2.4967665529251097, "grad_norm": 1.387877345085144, "lr": 2.295572269076374e-05, "s_per_step": 0.5827033042907714, "peak_gb": 45.225538048, "samples_seen": 236000, "time": 1791135153.1570456}
|
| 297 |
+
{"step": 7400, "loss": 2.515732476711273, "grad_norm": 1.2702908515930176, "lr": 2.2408021631518726e-05, "s_per_step": 0.5825530815124512, "peak_gb": 45.225538048, "samples_seen": 236800, "time": 1791135167.7212825}
|
| 298 |
+
{"step": 7425, "loss": 2.454144196510315, "grad_norm": 1.0984094142913818, "lr": 2.1866107978788153e-05, "s_per_step": 0.5797280502319336, "peak_gb": 45.225538048, "samples_seen": 237600, "time": 1791135182.2148888}
|
| 299 |
+
{"step": 7450, "loss": 2.4906524801254273, "grad_norm": 1.2278966903686523, "lr": 2.1330022152660156e-05, "s_per_step": 0.5791896152496337, "peak_gb": 45.225538048, "samples_seen": 238400, "time": 1791135196.69504}
|
| 300 |
+
{"step": 7475, "loss": 2.5189996767044067, "grad_norm": 1.3802945613861084, "lr": 2.0799804138538738e-05, "s_per_step": 0.5806455993652344, "peak_gb": 45.225538048, "samples_seen": 239200, "time": 1791135211.211557}
|
| 301 |
+
{"step": 7500, "loss": 2.4968285632133482, "grad_norm": 1.1902079582214355, "lr": 2.027549348416138e-05, "s_per_step": 0.5817130088806153, "peak_gb": 45.225538048, "samples_seen": 240000, "time": 1791135225.754808}
|
| 302 |
+
{"step": 7525, "loss": 2.5208859515190123, "grad_norm": 1.1940761804580688, "lr": 1.9757129296649146e-05, "s_per_step": 0.5999154567718505, "peak_gb": 45.225538048, "samples_seen": 240800, "time": 1791135240.753182}
|
| 303 |
+
{"step": 7550, "loss": 2.5416905546188353, "grad_norm": 1.4812296628952026, "lr": 1.924475023958998e-05, "s_per_step": 0.5843287658691406, "peak_gb": 45.225538048, "samples_seen": 241600, "time": 1791135255.3618777}
|
| 304 |
+
{"step": 7575, "loss": 2.4698385620117187, "grad_norm": 1.176838994026184, "lr": 1.87383945301547e-05, "s_per_step": 0.5848553848266601, "peak_gb": 45.225538048, "samples_seen": 242400, "time": 1791135269.9837594}
|
| 305 |
+
{"step": 7600, "loss": 2.4738434529304505, "grad_norm": 1.2223095893859863, "lr": 1.8238099936246555e-05, "s_per_step": 0.5825335311889649, "peak_gb": 45.225538048, "samples_seen": 243200, "time": 1791135284.547529}
|
| 306 |
+
{"step": 7625, "loss": 2.4790508365631103, "grad_norm": 1.1244370937347412, "lr": 1.774390377368418e-05, "s_per_step": 0.5838665103912354, "peak_gb": 45.225538048, "samples_seen": 244000, "time": 1791135299.144589}
|
| 307 |
+
{"step": 7650, "loss": 2.5447232913970947, "grad_norm": 1.2436165809631348, "lr": 1.7255842903418274e-05, "s_per_step": 0.5841799449920654, "peak_gb": 45.225538048, "samples_seen": 244800, "time": 1791135313.7496023}
|
| 308 |
+
{"step": 7675, "loss": 2.4780035495758055, "grad_norm": 1.200567603111267, "lr": 1.677395372878231e-05, "s_per_step": 0.5839977931976318, "peak_gb": 45.225538048, "samples_seen": 245600, "time": 1791135328.3500307}
|
| 309 |
+
{"step": 7700, "loss": 2.4953444504737856, "grad_norm": 1.3506298065185547, "lr": 1.629827219277713e-05, "s_per_step": 0.5840514087677002, "peak_gb": 45.225538048, "samples_seen": 246400, "time": 1791135342.9517403}
|
| 310 |
+
{"step": 7725, "loss": 2.4795714807510376, "grad_norm": 1.1271368265151978, "lr": 1.5828833775390227e-05, "s_per_step": 0.5836558628082276, "peak_gb": 45.225538048, "samples_seen": 247200, "time": 1791135357.5435355}
|
| 311 |
+
{"step": 7750, "loss": 2.4116058373451232, "grad_norm": 1.1426550149917603, "lr": 1.5365673490949285e-05, "s_per_step": 0.5836513805389404, "peak_gb": 45.225538048, "samples_seen": 248000, "time": 1791135372.135241}
|
| 312 |
+
{"step": 7775, "loss": 2.4785056114196777, "grad_norm": 1.1472325325012207, "lr": 1.4908825885510514e-05, "s_per_step": 0.5973025321960449, "peak_gb": 45.225538048, "samples_seen": 248800, "time": 1791135387.0681865}
|
| 313 |
+
{"step": 7800, "loss": 2.4868078422546387, "grad_norm": 1.2887753248214722, "lr": 1.4458325034281993e-05, "s_per_step": 0.580522985458374, "peak_gb": 45.225538048, "samples_seen": 249600, "time": 1791135401.5817068}
|
| 314 |
+
{"step": 7825, "loss": 2.4639941549301145, "grad_norm": 1.1615179777145386, "lr": 1.4014204539082044e-05, "s_per_step": 0.5852510643005371, "peak_gb": 45.225538048, "samples_seen": 250400, "time": 1791135416.213432}
|
| 315 |
+
{"step": 7850, "loss": 2.480881841182709, "grad_norm": 1.1630935668945312, "lr": 1.3576497525832998e-05, "s_per_step": 0.5810826301574707, "peak_gb": 45.225538048, "samples_seen": 251200, "time": 1791135430.7409315}
|
| 316 |
+
{"step": 7875, "loss": 2.467122640609741, "grad_norm": 1.288510799407959, "lr": 1.314523664209033e-05, "s_per_step": 0.5820274829864502, "peak_gb": 45.225538048, "samples_seen": 252000, "time": 1791135445.292001}
|
| 317 |
+
{"step": 7900, "loss": 2.4884210443496704, "grad_norm": 1.2538615465164185, "lr": 1.2720454054607633e-05, "s_per_step": 0.5825617122650146, "peak_gb": 45.225538048, "samples_seen": 252800, "time": 1791135459.8564053}
|
| 318 |
+
{"step": 7925, "loss": 2.480919179916382, "grad_norm": 1.403016448020935, "lr": 1.2302181446937333e-05, "s_per_step": 0.5786206722259521, "peak_gb": 45.225538048, "samples_seen": 253600, "time": 1791135474.3223386}
|
| 319 |
+
{"step": 7950, "loss": 2.439384446144104, "grad_norm": 1.1538753509521484, "lr": 1.1890450017067467e-05, "s_per_step": 0.5839206409454346, "peak_gb": 45.225538048, "samples_seen": 254400, "time": 1791135488.920715}
|
| 320 |
+
{"step": 7975, "loss": 2.484659490585327, "grad_norm": 1.5197055339813232, "lr": 1.1485290475094745e-05, "s_per_step": 0.5797516345977783, "peak_gb": 45.225538048, "samples_seen": 255200, "time": 1791135503.414908}
|
| 321 |
+
{"step": 8000, "loss": 2.4812649726867675, "grad_norm": 1.2070398330688477, "lr": 1.1086733040933938e-05, "s_per_step": 0.5830123424530029, "peak_gb": 45.225538048, "samples_seen": 256000, "time": 1791135517.9906447}
|
| 322 |
+
{"step": 8025, "loss": 2.453145351409912, "grad_norm": 1.3291114568710327, "lr": 1.0694807442063836e-05, "s_per_step": 0.5979341888427734, "peak_gb": 45.225538048, "samples_seen": 256800, "time": 1791135532.9394002}
|
| 323 |
+
{"step": 8050, "loss": 2.4227629256248475, "grad_norm": 1.178061604499817, "lr": 1.0309542911309932e-05, "s_per_step": 0.583064832687378, "peak_gb": 45.225538048, "samples_seen": 257600, "time": 1791135547.5164373}
|
| 324 |
+
{"step": 8075, "loss": 2.4514348554611205, "grad_norm": 1.5888311862945557, "lr": 9.930968184664046e-06, "s_per_step": 0.5811198616027832, "peak_gb": 45.225538048, "samples_seen": 258400, "time": 1791135562.0448437}
|
| 325 |
+
{"step": 8100, "loss": 2.4358947181701662, "grad_norm": 1.1722838878631592, "lr": 9.559111499140939e-06, "s_per_step": 0.579391222000122, "peak_gb": 45.225538048, "samples_seen": 259200, "time": 1791135576.530029}
|
| 326 |
+
{"step": 8125, "loss": 2.470351276397705, "grad_norm": 1.2379904985427856, "lr": 9.194000590672202e-06, "s_per_step": 0.5833250331878662, "peak_gb": 45.225538048, "samples_seen": 260000, "time": 1791135591.1136327}
|
| 327 |
+
{"step": 8150, "loss": 2.5454572772979738, "grad_norm": 1.0861706733703613, "lr": 8.835662692037516e-06, "s_per_step": 0.5829557609558106, "peak_gb": 45.225538048, "samples_seen": 260800, "time": 1791135605.6879003}
|
| 328 |
+
{"step": 8175, "loss": 2.5356476593017576, "grad_norm": 1.193375825881958, "lr": 8.484124530833349e-06, "s_per_step": 0.5813982677459717, "peak_gb": 45.225538048, "samples_seen": 261600, "time": 1791135620.2232423}
|
| 329 |
+
{"step": 8200, "loss": 2.5145446348190306, "grad_norm": 1.31069815158844, "lr": 8.139412327479512e-06, "s_per_step": 0.5830599784851074, "peak_gb": 45.225538048, "samples_seen": 262400, "time": 1791135634.8001153}
|
| 330 |
+
{"step": 8225, "loss": 2.51434268951416, "grad_norm": 1.2565644979476929, "lr": 7.801551793263307e-06, "s_per_step": 0.5842070770263672, "peak_gb": 45.225538048, "samples_seen": 263200, "time": 1791135649.4057364}
|
| 331 |
+
{"step": 8250, "loss": 2.4891091346740724, "grad_norm": 1.3129713535308838, "lr": 7.470568128421918e-06, "s_per_step": 0.5833453273773194, "peak_gb": 45.225538048, "samples_seen": 264000, "time": 1791135663.9898171}
|
| 332 |
+
{"step": 8275, "loss": 2.4999825406074523, "grad_norm": 1.2650532722473145, "lr": 7.146486020262688e-06, "s_per_step": 0.5951438903808594, "peak_gb": 45.225538048, "samples_seen": 264800, "time": 1791135678.8688512}
|
| 333 |
+
{"step": 8300, "loss": 2.50071989774704, "grad_norm": 1.366416335105896, "lr": 6.82932964132178e-06, "s_per_step": 0.585099573135376, "peak_gb": 45.225538048, "samples_seen": 265600, "time": 1791135693.4967763}
|
| 334 |
+
{"step": 8325, "loss": 2.4906865310668946, "grad_norm": 1.2222192287445068, "lr": 6.519122647561249e-06, "s_per_step": 0.5849275970458985, "peak_gb": 45.225538048, "samples_seen": 266400, "time": 1791135708.120377}
|
| 335 |
+
{"step": 8350, "loss": 2.482330822944641, "grad_norm": 1.2209337949752808, "lr": 6.215888176604501e-06, "s_per_step": 0.5821888160705566, "peak_gb": 45.225538048, "samples_seen": 267200, "time": 1791135722.6755044}
|
| 336 |
+
{"step": 8375, "loss": 2.455729749202728, "grad_norm": 1.3347740173339844, "lr": 5.919648846010584e-06, "s_per_step": 0.5881471824645996, "peak_gb": 45.225538048, "samples_seen": 268000, "time": 1791135737.3796217}
|
| 337 |
+
{"step": 8400, "loss": 2.4725853204727173, "grad_norm": 1.408981442451477, "lr": 5.6304267515872035e-06, "s_per_step": 0.5856296157836914, "peak_gb": 45.225538048, "samples_seen": 268800, "time": 1791135752.0207672}
|
| 338 |
+
{"step": 8425, "loss": 2.4187474822998047, "grad_norm": 1.3314207792282104, "lr": 5.3482434657425755e-06, "s_per_step": 0.5818990707397461, "peak_gb": 45.225538048, "samples_seen": 269600, "time": 1791135766.5687203}
|
| 339 |
+
{"step": 8450, "loss": 2.4801942539215087, "grad_norm": 1.0783225297927856, "lr": 5.073120035876466e-06, "s_per_step": 0.5826784038543701, "peak_gb": 45.225538048, "samples_seen": 270400, "time": 1791135781.136184}
|
| 340 |
+
{"step": 8475, "loss": 2.4995450139045716, "grad_norm": 1.3838938474655151, "lr": 4.805076982810275e-06, "s_per_step": 0.5815659427642822, "peak_gb": 45.225538048, "samples_seen": 271200, "time": 1791135795.6757793}
|
| 341 |
+
{"step": 8500, "loss": 2.4431419372558594, "grad_norm": 1.1133414506912231, "lr": 4.544134299256453e-06, "s_per_step": 0.5855115985870362, "peak_gb": 45.225538048, "samples_seen": 272000, "time": 1791135810.3140042}
|
| 342 |
+
{"step": 8525, "loss": 2.4860039710998536, "grad_norm": 1.1502230167388916, "lr": 4.290311448327267e-06, "s_per_step": 0.6029162120819092, "peak_gb": 45.225538048, "samples_seen": 272800, "time": 1791135825.3872917}
|
| 343 |
+
{"step": 8550, "loss": 2.502565712928772, "grad_norm": 1.4622975587844849, "lr": 4.043627362083124e-06, "s_per_step": 0.5819817352294921, "peak_gb": 45.225538048, "samples_seen": 273600, "time": 1791135839.9373028}
|
| 344 |
+
{"step": 8575, "loss": 2.4554481315612793, "grad_norm": 1.2948318719863892, "lr": 3.8041004401204392e-06, "s_per_step": 0.5855245685577393, "peak_gb": 45.225538048, "samples_seen": 274400, "time": 1791135854.575897}
|
| 345 |
+
{"step": 8600, "loss": 2.473366265296936, "grad_norm": 1.3055177927017212, "lr": 3.571748548199283e-06, "s_per_step": 0.5834248447418213, "peak_gb": 45.225538048, "samples_seen": 275200, "time": 1791135869.1619527}
|
| 346 |
+
{"step": 8625, "loss": 2.4664463996887207, "grad_norm": 1.3097848892211914, "lr": 3.346589016910795e-06, "s_per_step": 0.5833245754241944, "peak_gb": 45.225538048, "samples_seen": 276000, "time": 1791135883.7455816}
|
| 347 |
+
{"step": 8650, "loss": 2.5279615044593813, "grad_norm": 1.281739354133606, "lr": 3.1286386403845402e-06, "s_per_step": 0.5869896125793457, "peak_gb": 45.225538048, "samples_seen": 276800, "time": 1791135898.4208338}
|
| 348 |
+
{"step": 8675, "loss": 2.4256000685691834, "grad_norm": 1.1321176290512085, "lr": 2.917913675035888e-06, "s_per_step": 0.5836698818206787, "peak_gb": 45.225538048, "samples_seen": 277600, "time": 1791135913.0130074}
|
| 349 |
+
{"step": 8700, "loss": 2.520475058555603, "grad_norm": 1.046059250831604, "lr": 2.7144298383534607e-06, "s_per_step": 0.5819922924041748, "peak_gb": 45.225538048, "samples_seen": 278400, "time": 1791135927.5632734}
|
| 350 |
+
{"step": 8725, "loss": 2.555812520980835, "grad_norm": 1.1831002235412598, "lr": 2.5182023077268024e-06, "s_per_step": 0.5796582889556885, "peak_gb": 45.225538048, "samples_seen": 279200, "time": 1791135942.0551524}
|
| 351 |
+
{"step": 8750, "loss": 2.4607612800598146, "grad_norm": 1.2153393030166626, "lr": 2.3292457193143657e-06, "s_per_step": 0.5847400951385499, "peak_gb": 45.225538048, "samples_seen": 280000, "time": 1791135956.6740828}
|
| 352 |
+
{"step": 8775, "loss": 2.4432231545448304, "grad_norm": 1.3403053283691406, "lr": 2.1475741669517936e-06, "s_per_step": 0.5985502052307129, "peak_gb": 45.225538048, "samples_seen": 280800, "time": 1791135971.638234}
|
| 353 |
+
{"step": 8800, "loss": 2.45842068195343, "grad_norm": 1.3938454389572144, "lr": 1.973201201100705e-06, "s_per_step": 0.5829646301269531, "peak_gb": 45.225538048, "samples_seen": 281600, "time": 1791135986.2127783}
|
| 354 |
+
{"step": 8825, "loss": 2.4700878167152407, "grad_norm": 1.2799893617630005, "lr": 1.806139827838027e-06, "s_per_step": 0.5821155929565429, "peak_gb": 45.225538048, "samples_seen": 282400, "time": 1791136000.766119}
|
| 355 |
+
{"step": 8850, "loss": 2.505449166297913, "grad_norm": 1.1758151054382324, "lr": 1.646402507885858e-06, "s_per_step": 0.5805652141571045, "peak_gb": 45.225538048, "samples_seen": 283200, "time": 1791136015.2806501}
|
| 356 |
+
{"step": 8875, "loss": 2.4752099657058717, "grad_norm": 1.1823129653930664, "lr": 1.4940011556820676e-06, "s_per_step": 0.5807376194000244, "peak_gb": 45.225538048, "samples_seen": 284000, "time": 1791136029.7995193}
|
| 357 |
+
{"step": 8900, "loss": 2.477540431022644, "grad_norm": 1.2883774042129517, "lr": 1.3489471384916518e-06, "s_per_step": 0.5786253547668457, "peak_gb": 45.225538048, "samples_seen": 284800, "time": 1791136044.2656045}
|
| 358 |
+
{"step": 8925, "loss": 2.472158603668213, "grad_norm": 1.2201075553894043, "lr": 1.2112512755588335e-06, "s_per_step": 0.5808994483947754, "peak_gb": 45.225538048, "samples_seen": 285600, "time": 1791136058.7884433}
|
| 359 |
+
{"step": 8950, "loss": 2.4544294118881225, "grad_norm": 1.2970439195632935, "lr": 1.0809238373000962e-06, "s_per_step": 0.581709451675415, "peak_gb": 45.225538048, "samples_seen": 286400, "time": 1791136073.3315735}
|
| 360 |
+
{"step": 8975, "loss": 2.400626153945923, "grad_norm": 1.2315267324447632, "lr": 9.579745445381539e-07, "s_per_step": 0.579955415725708, "peak_gb": 45.225538048, "samples_seen": 287200, "time": 1791136087.8308325}
|
| 361 |
+
{"step": 9000, "loss": 2.4409583902359007, "grad_norm": 1.1493794918060303, "lr": 8.424125677768846e-07, "s_per_step": 0.5824070739746093, "peak_gb": 45.225538048, "samples_seen": 288000, "time": 1791136102.3914094}
|
| 362 |
+
{"step": 9025, "loss": 2.4901751470565796, "grad_norm": 1.405742883682251, "lr": 7.342465265172904e-07, "s_per_step": 0.5998865604400635, "peak_gb": 45.225538048, "samples_seen": 288800, "time": 1791136117.3889866}
|
| 363 |
+
{"step": 9050, "loss": 2.45655734539032, "grad_norm": 1.2151944637298584, "lr": 6.334844886146552e-07, "s_per_step": 0.5862552833557129, "peak_gb": 45.225538048, "samples_seen": 289600, "time": 1791136132.045783}
|
| 364 |
+
{"step": 9075, "loss": 2.440645921230316, "grad_norm": 1.2591733932495117, "lr": 5.401339696767371e-07, "s_per_step": 0.5819112586975098, "peak_gb": 45.225538048, "samples_seen": 290400, "time": 1791136146.593972}
|
| 365 |
+
{"step": 9100, "loss": 2.4351733326911926, "grad_norm": 1.4817209243774414, "lr": 4.542019325032176e-07, "s_per_step": 0.5811293697357178, "peak_gb": 45.225538048, "samples_seen": 291200, "time": 1791136161.1227205}
|
| 366 |
+
{"step": 9125, "loss": 2.499671845436096, "grad_norm": 1.075323224067688, "lr": 3.756947865663274e-07, "s_per_step": 0.580502290725708, "peak_gb": 45.225538048, "samples_seen": 292000, "time": 1791136175.635759}
|
| 367 |
+
{"step": 9150, "loss": 2.4365021347999574, "grad_norm": 1.199285626411438, "lr": 3.0461838753281793e-07, "s_per_step": 0.5829291343688965, "peak_gb": 45.225538048, "samples_seen": 292800, "time": 1791136190.209385}
|
| 368 |
+
{"step": 9175, "loss": 2.5040885043144225, "grad_norm": 1.390062689781189, "lr": 2.4097803682719967e-07, "s_per_step": 0.5805082321166992, "peak_gb": 45.225538048, "samples_seen": 293600, "time": 1791136204.7224927}
|
| 369 |
+
{"step": 9200, "loss": 2.4622223567962647, "grad_norm": 1.1260004043579102, "lr": 1.847784812362696e-07, "s_per_step": 0.5807758331298828, "peak_gb": 45.225538048, "samples_seen": 294400, "time": 1791136219.2422497}
|
| 370 |
+
{"step": 9225, "loss": 2.5410248279571532, "grad_norm": 1.2212775945663452, "lr": 1.360239125551388e-07, "s_per_step": 0.5833808612823487, "peak_gb": 45.225538048, "samples_seen": 295200, "time": 1791136233.8271465}
|
| 371 |
+
{"step": 9250, "loss": 2.4439662075042725, "grad_norm": 1.3457050323486328, "lr": 9.471796727449356e-08, "s_per_step": 0.5828397941589355, "peak_gb": 45.225538048, "samples_seen": 296000, "time": 1791136248.3985698}
|
| 372 |
+
{"step": 9275, "loss": 2.402722330093384, "grad_norm": 1.2421631813049316, "lr": 6.086372630945692e-08, "s_per_step": 0.5969734477996826, "peak_gb": 45.225538048, "samples_seen": 296800, "time": 1791136263.3232737}
|
| 373 |
+
{"step": 9300, "loss": 2.4314382100105285, "grad_norm": 1.442682147026062, "lr": 3.446371476966137e-08, "s_per_step": 0.5828660678863525, "peak_gb": 45.225538048, "samples_seen": 297600, "time": 1791136277.895351}
|
| 374 |
+
{"step": 9325, "loss": 2.46381530046463, "grad_norm": 1.3629335165023804, "lr": 1.5519901771032798e-08, "s_per_step": 0.5814544200897217, "peak_gb": 45.225538048, "samples_seen": 298400, "time": 1791136292.4321828}
|
| 375 |
+
{"step": 9350, "loss": 2.448145937919617, "grad_norm": 1.2016350030899048, "lr": 4.033700288830211e-09, "s_per_step": 0.5831421184539795, "peak_gb": 45.225538048, "samples_seen": 299200, "time": 1791136307.0111272}
|
| 376 |
+
{"step": 9375, "loss": 2.453929500579834, "grad_norm": 1.3293850421905518, "lr": 5.967052318922584e-12, "s_per_step": 0.5822473049163819, "peak_gb": 45.225538048, "samples_seen": 300000, "time": 1791136321.5678217}
|
logs/stage2_train_log.jsonl
ADDED
|
@@ -0,0 +1,197 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"step": 1, "loss": 2.1018713116645813, "grad_norm": 3.935774087905884, "lr": 1.3698630136986302e-06, "s_per_step": 5.9358556270599365, "peak_gb": 28.923105792, "samples_seen": 32, "time": 1791137293.9179838}
|
| 2 |
+
{"step": 25, "loss": 1.8595843004683654, "grad_norm": 0.8364684581756592, "lr": 3.424657534246575e-05, "s_per_step": 3.0314967930316925, "peak_gb": 32.089119744, "samples_seen": 800, "time": 1791137366.67442}
|
| 3 |
+
{"step": 50, "loss": 1.4337403464317322, "grad_norm": 0.641377329826355, "lr": 6.84931506849315e-05, "s_per_step": 2.7474777412414553, "peak_gb": 32.089119744, "samples_seen": 1600, "time": 1791137435.3618677}
|
| 4 |
+
{"step": 75, "loss": 1.3591693234443665, "grad_norm": 0.46830347180366516, "lr": 0.00010273972602739728, "s_per_step": 2.675526475906372, "peak_gb": 32.18382848, "samples_seen": 2400, "time": 1791137502.2504642}
|
| 5 |
+
{"step": 100, "loss": 1.3316362035274505, "grad_norm": 0.6638643145561218, "lr": 0.000136986301369863, "s_per_step": 2.702469034194946, "peak_gb": 34.275917312, "samples_seen": 3200, "time": 1791137569.8126426}
|
| 6 |
+
{"step": 125, "loss": 1.3198813837766648, "grad_norm": 0.5527892112731934, "lr": 0.00017123287671232877, "s_per_step": 2.699006910324097, "peak_gb": 34.275917312, "samples_seen": 4000, "time": 1791137637.2882571}
|
| 7 |
+
{"step": 150, "loss": 1.3081526547670363, "grad_norm": 0.5911878347396851, "lr": 0.00019999980323762506, "s_per_step": 2.6959174919128417, "peak_gb": 34.275917312, "samples_seen": 4800, "time": 1791137704.6866806}
|
| 8 |
+
{"step": 175, "loss": 1.33544615983963, "grad_norm": 0.5037556290626526, "lr": 0.000199982860294911, "s_per_step": 2.5416007328033445, "peak_gb": 34.275917312, "samples_seen": 5600, "time": 1791137768.2271826}
|
| 9 |
+
{"step": 200, "loss": 1.30276043176651, "grad_norm": 0.6148776412010193, "lr": 0.00019993859454180504, "s_per_step": 2.7514454460144044, "peak_gb": 34.275917312, "samples_seen": 6400, "time": 1791137837.0137653}
|
| 10 |
+
{"step": 225, "loss": 1.3027793610095977, "grad_norm": 0.5531913042068481, "lr": 0.00019986701807502827, "s_per_step": 2.6572889137268065, "peak_gb": 34.275917312, "samples_seen": 7200, "time": 1791137903.4464118}
|
| 11 |
+
{"step": 250, "loss": 1.3026298081874848, "grad_norm": 0.5650343298912048, "lr": 0.00019976815045463558, "s_per_step": 2.581279354095459, "peak_gb": 34.275917312, "samples_seen": 8000, "time": 1791137967.9797318}
|
| 12 |
+
{"step": 275, "loss": 1.2859191417694091, "grad_norm": 0.5168231725692749, "lr": 0.00019964201869867025, "s_per_step": 2.7289039039611818, "peak_gb": 34.275917312, "samples_seen": 8800, "time": 1791138036.2027671}
|
| 13 |
+
{"step": 300, "loss": 1.3106050491333008, "grad_norm": 0.5257666707038879, "lr": 0.0001994886572757806, "s_per_step": 2.456781349182129, "peak_gb": 34.275917312, "samples_seen": 9600, "time": 1791138097.6227736}
|
| 14 |
+
{"step": 325, "loss": 1.285406960248947, "grad_norm": 0.5512126088142395, "lr": 0.00019930810809580063, "s_per_step": 2.6005754470825195, "peak_gb": 34.275917312, "samples_seen": 10400, "time": 1791138162.6376011}
|
| 15 |
+
{"step": 350, "loss": 1.2840537762641906, "grad_norm": 0.592390239238739, "lr": 0.00019910042049829714, "s_per_step": 2.6476001262664797, "peak_gb": 34.275917312, "samples_seen": 11200, "time": 1791138228.8280518}
|
| 16 |
+
{"step": 375, "loss": 1.310118284225464, "grad_norm": 0.4393337070941925, "lr": 0.00019886565123908637, "s_per_step": 2.5076206874847413, "peak_gb": 34.275917312, "samples_seen": 12000, "time": 1791138291.5190346}
|
| 17 |
+
{"step": 400, "loss": 1.3007353556156158, "grad_norm": 0.572140097618103, "lr": 0.0001986038644747241, "s_per_step": 2.54843355178833, "peak_gb": 34.275917312, "samples_seen": 12800, "time": 1791138355.230326}
|
| 18 |
+
{"step": 425, "loss": 1.3010186761617661, "grad_norm": 0.5119914412498474, "lr": 0.00019831513174497332, "s_per_step": 2.463373165130615, "peak_gb": 34.454321664, "samples_seen": 13600, "time": 1791138416.8150551}
|
| 19 |
+
{"step": 450, "loss": 1.2758523070812224, "grad_norm": 0.46384191513061523, "lr": 0.00019799953195325412, "s_per_step": 2.629770526885986, "peak_gb": 34.454321664, "samples_seen": 14400, "time": 1791138482.5597115}
|
| 20 |
+
{"step": 475, "loss": 1.2857936775684358, "grad_norm": 0.5037622451782227, "lr": 0.0001976571513450814, "s_per_step": 2.5591455364227294, "peak_gb": 34.454321664, "samples_seen": 15200, "time": 1791138546.5387669}
|
| 21 |
+
{"step": 500, "loss": 1.296420032978058, "grad_norm": 0.5816029906272888, "lr": 0.00019728808348449617, "s_per_step": 2.425617971420288, "peak_gb": 34.454321664, "samples_seen": 16000, "time": 1791138607.1797073}
|
| 22 |
+
{"step": 525, "loss": 1.2788980293273926, "grad_norm": 0.5487325191497803, "lr": 0.0001968924292284968, "s_per_step": 2.719025077819824, "peak_gb": 34.454321664, "samples_seen": 16800, "time": 1791138675.1558762}
|
| 23 |
+
{"step": 550, "loss": 1.2836597490310668, "grad_norm": 0.4508974850177765, "lr": 0.0001964702966994773, "s_per_step": 2.5938339710235594, "peak_gb": 35.16160256, "samples_seen": 17600, "time": 1791138740.0021875}
|
| 24 |
+
{"step": 575, "loss": 1.272069708108902, "grad_norm": 0.4157426059246063, "lr": 0.00019602180125568024, "s_per_step": 2.557501964569092, "peak_gb": 35.16160256, "samples_seen": 18400, "time": 1791138803.9402196}
|
| 25 |
+
{"step": 600, "loss": 1.2735493952035903, "grad_norm": 0.4135555326938629, "lr": 0.00019554706545967223, "s_per_step": 2.681543912887573, "peak_gb": 35.16160256, "samples_seen": 19200, "time": 1791138871.1860206}
|
| 26 |
+
{"step": 625, "loss": 1.2743602776527405, "grad_norm": 0.527656078338623, "lr": 0.00019504621904485055, "s_per_step": 2.574559545516968, "peak_gb": 35.16160256, "samples_seen": 20000, "time": 1791138935.5505657}
|
| 27 |
+
{"step": 650, "loss": 1.2700351548194886, "grad_norm": 0.47019869089126587, "lr": 0.00019451939887999044, "s_per_step": 2.5688095760345457, "peak_gb": 35.16160256, "samples_seen": 20800, "time": 1791138999.7712135}
|
| 28 |
+
{"step": 675, "loss": 1.278431007862091, "grad_norm": 0.49949219822883606, "lr": 0.0001939667489318421, "s_per_step": 2.5981833267211916, "peak_gb": 35.16160256, "samples_seen": 21600, "time": 1791139064.7262437}
|
| 29 |
+
{"step": 700, "loss": 1.2902556908130647, "grad_norm": 0.4944266378879547, "lr": 0.00019338842022578826, "s_per_step": 2.3775853538513183, "peak_gb": 35.16160256, "samples_seen": 22400, "time": 1791139124.16638}
|
| 30 |
+
{"step": 725, "loss": 1.2667275273799896, "grad_norm": 0.3541528880596161, "lr": 0.00019278457080457287, "s_per_step": 2.660867967605591, "peak_gb": 35.16160256, "samples_seen": 23200, "time": 1791139190.6885371}
|
| 31 |
+
{"step": 750, "loss": 1.2670336496829986, "grad_norm": 0.4926910996437073, "lr": 0.00019215536568511165, "s_per_step": 2.5423004722595213, "peak_gb": 35.16160256, "samples_seen": 24000, "time": 1791139254.246462}
|
| 32 |
+
{"step": 775, "loss": 1.2607650971412658, "grad_norm": 0.5327316522598267, "lr": 0.0001915009768133975, "s_per_step": 2.632803087234497, "peak_gb": 35.16160256, "samples_seen": 24800, "time": 1791139320.0670488}
|
| 33 |
+
{"step": 800, "loss": 1.298940641283989, "grad_norm": 0.5895273685455322, "lr": 0.00019082158301751158, "s_per_step": 2.3035424613952635, "peak_gb": 35.16160256, "samples_seen": 25600, "time": 1791139377.6561027}
|
| 34 |
+
{"step": 825, "loss": 1.2994678497314454, "grad_norm": 0.5701431632041931, "lr": 0.00019011736995875442, "s_per_step": 2.4659962368011477, "peak_gb": 35.16160256, "samples_seen": 26400, "time": 1791139439.3064585}
|
| 35 |
+
{"step": 850, "loss": 1.2705088353157044, "grad_norm": 0.4701496958732605, "lr": 0.0001893885300809091, "s_per_step": 2.506269245147705, "peak_gb": 35.16160256, "samples_seen": 27200, "time": 1791139501.96361}
|
| 36 |
+
{"step": 875, "loss": 1.2747311300039292, "grad_norm": 0.5782403349876404, "lr": 0.00018863526255765125, "s_per_step": 2.4462522220611573, "peak_gb": 35.16160256, "samples_seen": 28000, "time": 1791139563.1204047}
|
| 37 |
+
{"step": 900, "loss": 1.2845921885967255, "grad_norm": 0.5439372062683105, "lr": 0.00018785777323811998, "s_per_step": 2.414919261932373, "peak_gb": 35.16160256, "samples_seen": 28800, "time": 1791139623.4938784}
|
| 38 |
+
{"step": 925, "loss": 1.2421265226602554, "grad_norm": 0.4805165231227875, "lr": 0.00018705627459066429, "s_per_step": 2.59889892578125, "peak_gb": 35.16160256, "samples_seen": 29600, "time": 1791139688.4667766}
|
| 39 |
+
{"step": 950, "loss": 1.256026518344879, "grad_norm": 0.4744698405265808, "lr": 0.00018623098564478095, "s_per_step": 2.4729893398284912, "peak_gb": 35.16160256, "samples_seen": 30400, "time": 1791139750.3143578}
|
| 40 |
+
{"step": 975, "loss": 1.2527462810277938, "grad_norm": 0.45646801590919495, "lr": 0.00018538213193125915, "s_per_step": 2.529736156463623, "peak_gb": 35.16160256, "samples_seen": 31200, "time": 1791139813.5582156}
|
| 41 |
+
{"step": 1000, "loss": 1.2478597414493562, "grad_norm": 0.6265877485275269, "lr": 0.00018450994542054856, "s_per_step": 2.5998341178894044, "peak_gb": 35.16160256, "samples_seen": 32000, "time": 1791139878.5545025}
|
| 42 |
+
{"step": 1025, "loss": 1.2808118069171905, "grad_norm": 0.4756571054458618, "lr": 0.0001836146644593677, "s_per_step": 2.5343000221252443, "peak_gb": 35.16160256, "samples_seen": 32800, "time": 1791139941.9124794}
|
| 43 |
+
{"step": 1050, "loss": 1.2551424139738083, "grad_norm": 0.6432344913482666, "lr": 0.00018269653370556972, "s_per_step": 2.568102979660034, "peak_gb": 35.16160256, "samples_seen": 33600, "time": 1791140006.1155303}
|
| 44 |
+
{"step": 1075, "loss": 1.2568443739414215, "grad_norm": 0.5057801008224487, "lr": 0.0001817558040612835, "s_per_step": 2.6142832756042482, "peak_gb": 35.16160256, "samples_seen": 34400, "time": 1791140071.4731367}
|
| 45 |
+
{"step": 1100, "loss": 1.2462696027755737, "grad_norm": 0.537843644618988, "lr": 0.00018079273260434846, "s_per_step": 2.5834513759613036, "peak_gb": 35.16160256, "samples_seen": 35200, "time": 1791140136.05996}
|
| 46 |
+
{"step": 1125, "loss": 1.2574056977033614, "grad_norm": 0.5526939630508423, "lr": 0.00017980758251806152, "s_per_step": 2.4060669994354247, "peak_gb": 35.16160256, "samples_seen": 36000, "time": 1791140196.2120588}
|
| 47 |
+
{"step": 1150, "loss": 1.2388004493713378, "grad_norm": 0.5050171613693237, "lr": 0.00017880062301925582, "s_per_step": 2.5375908184051514, "peak_gb": 35.16160256, "samples_seen": 36800, "time": 1791140259.6522965}
|
| 48 |
+
{"step": 1175, "loss": 1.263253116607666, "grad_norm": 0.4916113317012787, "lr": 0.00017777212928473044, "s_per_step": 2.5061941528320313, "peak_gb": 35.16160256, "samples_seen": 37600, "time": 1791140322.3075662}
|
| 49 |
+
{"step": 1200, "loss": 1.2475608468055726, "grad_norm": 0.5672100186347961, "lr": 0.00017672238237605146, "s_per_step": 2.3854447269439696, "peak_gb": 35.16160256, "samples_seen": 38400, "time": 1791140381.9441311}
|
| 50 |
+
{"step": 1225, "loss": 1.2634010112285614, "grad_norm": 0.45946162939071655, "lr": 0.0001756516691627449, "s_per_step": 2.559474468231201, "peak_gb": 35.16160256, "samples_seen": 39200, "time": 1791140445.9314318}
|
| 51 |
+
{"step": 1250, "loss": 1.2636224496364594, "grad_norm": 0.5660274028778076, "lr": 0.00017456028224390258, "s_per_step": 2.3614064121246336, "peak_gb": 35.16160256, "samples_seen": 40000, "time": 1791140504.967005}
|
| 52 |
+
{"step": 1275, "loss": 1.2519952070713043, "grad_norm": 0.630153477191925, "lr": 0.00017344851986822186, "s_per_step": 2.604766302108765, "peak_gb": 35.16160256, "samples_seen": 40800, "time": 1791140570.0865934}
|
| 53 |
+
{"step": 1300, "loss": 1.248630976676941, "grad_norm": 0.603481650352478, "lr": 0.000172316685852502, "s_per_step": 2.5879988288879394, "peak_gb": 35.16160256, "samples_seen": 41600, "time": 1791140634.7870479}
|
| 54 |
+
{"step": 1325, "loss": 1.2556267881393433, "grad_norm": 0.5644108057022095, "lr": 0.00017116508949861846, "s_per_step": 2.4682638454437256, "peak_gb": 35.16160256, "samples_seen": 42400, "time": 1791140696.4940681}
|
| 55 |
+
{"step": 1350, "loss": 1.2743791633844375, "grad_norm": 0.552125871181488, "lr": 0.00016999404550899856, "s_per_step": 2.4177276992797854, "peak_gb": 35.16160256, "samples_seen": 43200, "time": 1791140756.9376965}
|
| 56 |
+
{"step": 1375, "loss": 1.2419914174079896, "grad_norm": 0.5084569454193115, "lr": 0.00016880387390062117, "s_per_step": 2.5251627922058106, "peak_gb": 35.16160256, "samples_seen": 44000, "time": 1791140820.067213}
|
| 57 |
+
{"step": 1400, "loss": 1.2473990058898925, "grad_norm": 0.48340925574302673, "lr": 0.00016759489991756402, "s_per_step": 2.5814351272583007, "peak_gb": 35.16160256, "samples_seen": 44800, "time": 1791140884.6035194}
|
| 58 |
+
{"step": 1425, "loss": 1.2643068969249724, "grad_norm": 0.6564701795578003, "lr": 0.0001663674539421228, "s_per_step": 2.4565196990966798, "peak_gb": 35.16160256, "samples_seen": 45600, "time": 1791140946.0169346}
|
| 59 |
+
{"step": 1450, "loss": 1.2657874739170074, "grad_norm": 0.5559526681900024, "lr": 0.0001651218714045258, "s_per_step": 2.38125319480896, "peak_gb": 35.16160256, "samples_seen": 46400, "time": 1791141005.5487158}
|
| 60 |
+
{"step": 1475, "loss": 1.2623559749126434, "grad_norm": 0.633904218673706, "lr": 0.00016385849269126923, "s_per_step": 2.3729316806793213, "peak_gb": 35.16160256, "samples_seen": 47200, "time": 1791141064.8724935}
|
| 61 |
+
{"step": 1500, "loss": 1.2323721277713775, "grad_norm": 0.5630310773849487, "lr": 0.00016257766305209824, "s_per_step": 2.5675553035736085, "peak_gb": 35.16160256, "samples_seen": 48000, "time": 1791141129.06184}
|
| 62 |
+
{"step": 1525, "loss": 1.2649996304512023, "grad_norm": 0.5584118962287903, "lr": 0.0001612797325056588, "s_per_step": 2.5546645450592043, "peak_gb": 35.16160256, "samples_seen": 48800, "time": 1791141192.9288673}
|
| 63 |
+
{"step": 1550, "loss": 1.2378357803821565, "grad_norm": 0.5218235850334167, "lr": 0.00015996505574384623, "s_per_step": 2.511321334838867, "peak_gb": 35.16160256, "samples_seen": 49600, "time": 1791141255.712335}
|
| 64 |
+
{"step": 1575, "loss": 1.2511804521083831, "grad_norm": 0.6376656889915466, "lr": 0.00015863399203487695, "s_per_step": 2.461472635269165, "peak_gb": 35.16160256, "samples_seen": 50400, "time": 1791141317.2496617}
|
| 65 |
+
{"step": 1600, "loss": 1.2562532049417496, "grad_norm": 0.5086262822151184, "lr": 0.00015728690512510944, "s_per_step": 2.506492853164673, "peak_gb": 35.16160256, "samples_seen": 51200, "time": 1791141379.912437}
|
| 66 |
+
{"step": 1625, "loss": 1.2379037940502167, "grad_norm": 0.5947778224945068, "lr": 0.00015592416313964136, "s_per_step": 2.4982398319244385, "peak_gb": 35.16160256, "samples_seen": 52000, "time": 1791141442.3688512}
|
| 67 |
+
{"step": 1650, "loss": 1.237431339621544, "grad_norm": 0.5146723985671997, "lr": 0.0001545461384817104, "s_per_step": 2.4878829956054687, "peak_gb": 35.16160256, "samples_seen": 52800, "time": 1791141504.5663276}
|
| 68 |
+
{"step": 1675, "loss": 1.2310943824052811, "grad_norm": 0.5532039403915405, "lr": 0.00015315320773092562, "s_per_step": 2.601818952560425, "peak_gb": 35.16160256, "samples_seen": 53600, "time": 1791141569.612219}
|
| 69 |
+
{"step": 1700, "loss": 1.233572164773941, "grad_norm": 0.491245836019516, "lr": 0.00015174575154035776, "s_per_step": 2.525400400161743, "peak_gb": 35.16160256, "samples_seen": 54400, "time": 1791141632.7476742}
|
| 70 |
+
{"step": 1725, "loss": 1.2353800696134567, "grad_norm": 0.5083756446838379, "lr": 0.00015032415453251628, "s_per_step": 2.500525407791138, "peak_gb": 35.16160256, "samples_seen": 55200, "time": 1791141695.2612443}
|
| 71 |
+
{"step": 1750, "loss": 1.2415151298046112, "grad_norm": 0.4848071336746216, "lr": 0.00014888880519424167, "s_per_step": 2.497219123840332, "peak_gb": 35.16160256, "samples_seen": 56000, "time": 1791141757.6921887}
|
| 72 |
+
{"step": 1775, "loss": 1.25896546125412, "grad_norm": 0.5702372193336487, "lr": 0.00014744009577054174, "s_per_step": 2.511766996383667, "peak_gb": 35.16160256, "samples_seen": 56800, "time": 1791141820.4869192}
|
| 73 |
+
{"step": 1800, "loss": 1.2329586499929428, "grad_norm": 0.5330091714859009, "lr": 0.0001459784221574008, "s_per_step": 2.550053768157959, "peak_gb": 35.34578432, "samples_seen": 57600, "time": 1791141884.2394202}
|
| 74 |
+
{"step": 1825, "loss": 1.2399318504333496, "grad_norm": 0.5269576907157898, "lr": 0.00014450418379359145, "s_per_step": 2.3936056327819824, "peak_gb": 35.34578432, "samples_seen": 58400, "time": 1791141944.0801225}
|
| 75 |
+
{"step": 1850, "loss": 1.1970552122592926, "grad_norm": 0.5154424905776978, "lr": 0.00014301778355151762, "s_per_step": 2.6440818309783936, "peak_gb": 35.34578432, "samples_seen": 59200, "time": 1791142010.1826928}
|
| 76 |
+
{"step": 1875, "loss": 1.2424454951286317, "grad_norm": 0.5999359488487244, "lr": 0.00014151962762711998, "s_per_step": 2.516843738555908, "peak_gb": 35.709385216, "samples_seen": 60000, "time": 1791142073.1044393}
|
| 77 |
+
{"step": 1900, "loss": 1.2413146513700486, "grad_norm": 0.5478770732879639, "lr": 0.0001400101254288724, "s_per_step": 2.479934730529785, "peak_gb": 35.709385216, "samples_seen": 60800, "time": 1791142135.1033275}
|
| 78 |
+
{"step": 1925, "loss": 1.2434303241968154, "grad_norm": 0.6004431247711182, "lr": 0.00013848968946590133, "s_per_step": 2.4324647426605224, "peak_gb": 35.709385216, "samples_seen": 61600, "time": 1791142195.9162219}
|
| 79 |
+
{"step": 1950, "loss": 1.2015808641910553, "grad_norm": 0.5266046524047852, "lr": 0.000136958735235257, "s_per_step": 2.506597595214844, "peak_gb": 35.709385216, "samples_seen": 62400, "time": 1791142258.5817263}
|
| 80 |
+
{"step": 1975, "loss": 1.2200964736938475, "grad_norm": 0.5610483884811401, "lr": 0.00013541768110836863, "s_per_step": 2.5384847927093506, "peak_gb": 35.709385216, "samples_seen": 63200, "time": 1791142322.0443137}
|
| 81 |
+
{"step": 2000, "loss": 1.2384649211168288, "grad_norm": 0.6361256241798401, "lr": 0.00013386694821671414, "s_per_step": 2.4207293224334716, "peak_gb": 35.709385216, "samples_seen": 64000, "time": 1791142382.5630329}
|
| 82 |
+
{"step": 2025, "loss": 1.214077221751213, "grad_norm": 0.6302129030227661, "lr": 0.0001323069603367351, "s_per_step": 2.6198851585388185, "peak_gb": 35.709385216, "samples_seen": 64800, "time": 1791142448.0606163}
|
| 83 |
+
{"step": 2050, "loss": 1.2055979889631272, "grad_norm": 0.509576141834259, "lr": 0.00013073814377402974, "s_per_step": 2.627358932495117, "peak_gb": 35.709385216, "samples_seen": 65600, "time": 1791142513.7450488}
|
| 84 |
+
{"step": 2075, "loss": 1.2103803342580794, "grad_norm": 0.5196846723556519, "lr": 0.00012916092724685383, "s_per_step": 2.598076095581055, "peak_gb": 35.709385216, "samples_seen": 66400, "time": 1791142578.6973796}
|
| 85 |
+
{"step": 2100, "loss": 1.2209484434127809, "grad_norm": 0.5177134871482849, "lr": 0.00012757574176896312, "s_per_step": 2.5112220001220704, "peak_gb": 35.709385216, "samples_seen": 67200, "time": 1791142641.47846}
|
| 86 |
+
{"step": 2125, "loss": 1.2264458799362183, "grad_norm": 0.5397101640701294, "lr": 0.0001259830205318278, "s_per_step": 2.3952220344543456, "peak_gb": 35.709385216, "samples_seen": 68000, "time": 1791142701.3594959}
|
| 87 |
+
{"step": 2150, "loss": 1.209760498404503, "grad_norm": 0.5998483300209045, "lr": 0.00012438319878625226, "s_per_step": 2.5326212215423585, "peak_gb": 35.709385216, "samples_seen": 68800, "time": 1791142764.675523}
|
| 88 |
+
{"step": 2175, "loss": 1.2117489975690843, "grad_norm": 0.5707188844680786, "lr": 0.00012277671372343195, "s_per_step": 2.4900473594665526, "peak_gb": 35.709385216, "samples_seen": 69600, "time": 1791142826.9271343}
|
| 89 |
+
{"step": 2200, "loss": 1.2348184549808503, "grad_norm": 0.5396585464477539, "lr": 0.00012116400435547992, "s_per_step": 2.492477035522461, "peak_gb": 35.709385216, "samples_seen": 70400, "time": 1791142889.2394576}
|
| 90 |
+
{"step": 2225, "loss": 1.2173430824279785, "grad_norm": 0.39253613352775574, "lr": 0.00011954551139545588, "s_per_step": 2.49932147026062, "peak_gb": 35.709385216, "samples_seen": 71200, "time": 1791142951.7229047}
|
| 91 |
+
{"step": 2250, "loss": 1.2054871225357056, "grad_norm": 0.5584589838981628, "lr": 0.00011792167713693036, "s_per_step": 2.4481615734100344, "peak_gb": 35.709385216, "samples_seen": 72000, "time": 1791143012.9273942}
|
| 92 |
+
{"step": 2275, "loss": 1.2060042411088943, "grad_norm": 0.5810075402259827, "lr": 0.0001162929453331168, "s_per_step": 2.66895037651062, "peak_gb": 35.709385216, "samples_seen": 72800, "time": 1791143079.6516283}
|
| 93 |
+
{"step": 2300, "loss": 1.2204527378082275, "grad_norm": 0.5374050140380859, "lr": 0.00011465976107560521, "s_per_step": 2.4989648342132567, "peak_gb": 35.709385216, "samples_seen": 73600, "time": 1791143142.12621}
|
| 94 |
+
{"step": 2325, "loss": 1.2322663950920105, "grad_norm": 0.5682536959648132, "lr": 0.00011302257067272952, "s_per_step": 2.3952228927612307, "peak_gb": 35.709385216, "samples_seen": 74400, "time": 1791143202.0071743}
|
| 95 |
+
{"step": 2350, "loss": 1.1958238542079926, "grad_norm": 0.4648407995700836, "lr": 0.00011138182152760286, "s_per_step": 2.6400107955932617, "peak_gb": 35.709385216, "samples_seen": 75200, "time": 1791143268.0078988}
|
| 96 |
+
{"step": 2375, "loss": 1.1953755247592925, "grad_norm": 0.4725061058998108, "lr": 0.00010973796201585335, "s_per_step": 2.4706161880493163, "peak_gb": 35.709385216, "samples_seen": 76000, "time": 1791143329.774484}
|
| 97 |
+
{"step": 2400, "loss": 1.2082021981477737, "grad_norm": 0.5769761800765991, "lr": 0.00010809144136309454, "s_per_step": 2.4323978424072266, "peak_gb": 35.709385216, "samples_seen": 76800, "time": 1791143390.5849032}
|
| 98 |
+
{"step": 2425, "loss": 1.2111977207660676, "grad_norm": 0.5421669483184814, "lr": 0.000106442709522163, "s_per_step": 2.4998205661773683, "peak_gb": 35.709385216, "samples_seen": 77600, "time": 1791143453.0808678}
|
| 99 |
+
{"step": 2450, "loss": 1.2098423671722411, "grad_norm": 0.484115332365036, "lr": 0.00010479221705015759, "s_per_step": 2.5170133399963377, "peak_gb": 35.709385216, "samples_seen": 78400, "time": 1791143516.0066352}
|
| 100 |
+
{"step": 2475, "loss": 1.223331753015518, "grad_norm": 0.49125203490257263, "lr": 0.00010314041498531371, "s_per_step": 2.4067738342285154, "peak_gb": 35.709385216, "samples_seen": 79200, "time": 1791143576.176388}
|
| 101 |
+
{"step": 2500, "loss": 1.2110005378723145, "grad_norm": 0.5610998868942261, "lr": 0.00010148775472374543, "s_per_step": 2.5736419677734377, "peak_gb": 35.709385216, "samples_seen": 80000, "time": 1791143640.5179398}
|
| 102 |
+
{"step": 2525, "loss": 1.2331609898805618, "grad_norm": 0.5671659708023071, "lr": 9.983468789609067e-05, "s_per_step": 2.5041617870330812, "peak_gb": 35.709385216, "samples_seen": 80800, "time": 1791143703.1224139}
|
| 103 |
+
{"step": 2550, "loss": 1.214245287179947, "grad_norm": 0.4357548654079437, "lr": 9.818166624409161e-05, "s_per_step": 2.4907432842254638, "peak_gb": 35.709385216, "samples_seen": 81600, "time": 1791143765.391451}
|
| 104 |
+
{"step": 2575, "loss": 1.214305859208107, "grad_norm": 0.5925732254981995, "lr": 9.6529141497145e-05, "s_per_step": 2.4908134937286377, "peak_gb": 35.709385216, "samples_seen": 82400, "time": 1791143827.6623023}
|
| 105 |
+
{"step": 2600, "loss": 1.1959744757413864, "grad_norm": 0.5484082698822021, "lr": 9.487756524885599e-05, "s_per_step": 2.544646863937378, "peak_gb": 35.709385216, "samples_seen": 83200, "time": 1791143891.2789044}
|
| 106 |
+
{"step": 2625, "loss": 1.2088198286294938, "grad_norm": 0.5090175271034241, "lr": 9.322738883362871e-05, "s_per_step": 2.5154834175109864, "peak_gb": 35.709385216, "samples_seen": 84000, "time": 1791143954.166384}
|
| 107 |
+
{"step": 2650, "loss": 1.2010453754663468, "grad_norm": 0.5693520307540894, "lr": 9.157906320332812e-05, "s_per_step": 2.5574374485015867, "peak_gb": 35.709385216, "samples_seen": 84800, "time": 1791144018.102869}
|
| 108 |
+
{"step": 2675, "loss": 1.214320752620697, "grad_norm": 0.5067724585533142, "lr": 8.993303880404592e-05, "s_per_step": 2.516569652557373, "peak_gb": 35.709385216, "samples_seen": 85600, "time": 1791144081.0175457}
|
| 109 |
+
{"step": 2700, "loss": 1.1918057614564896, "grad_norm": 0.49570122361183167, "lr": 8.828976545300505e-05, "s_per_step": 2.6196873950958253, "peak_gb": 35.709385216, "samples_seen": 86400, "time": 1791144146.5105312}
|
| 110 |
+
{"step": 2725, "loss": 1.1909438091516495, "grad_norm": 0.49999526143074036, "lr": 8.664969221563594e-05, "s_per_step": 2.577509126663208, "peak_gb": 35.709385216, "samples_seen": 87200, "time": 1791144210.9487243}
|
| 111 |
+
{"step": 2750, "loss": 1.2046889424324037, "grad_norm": 0.4591648280620575, "lr": 8.501326728285814e-05, "s_per_step": 2.463228282928467, "peak_gb": 35.709385216, "samples_seen": 88000, "time": 1791144272.529925}
|
| 112 |
+
{"step": 2775, "loss": 1.187312572002411, "grad_norm": 0.4475753605365753, "lr": 8.338093784860098e-05, "s_per_step": 2.65875358581543, "peak_gb": 35.709385216, "samples_seen": 88800, "time": 1791144338.9991896}
|
| 113 |
+
{"step": 2800, "loss": 1.2058877384662627, "grad_norm": 0.43829211592674255, "lr": 8.175314998759658e-05, "s_per_step": 2.525215663909912, "peak_gb": 35.709385216, "samples_seen": 89600, "time": 1791144402.1299999}
|
| 114 |
+
{"step": 2825, "loss": 1.1879939883947372, "grad_norm": 0.4615590274333954, "lr": 8.013034853347902e-05, "s_per_step": 2.61760648727417, "peak_gb": 35.709385216, "samples_seen": 90400, "time": 1791144467.5705826}
|
| 115 |
+
{"step": 2850, "loss": 1.1962288463115691, "grad_norm": 0.5492228269577026, "lr": 7.851297695722228e-05, "s_per_step": 2.416270399093628, "peak_gb": 35.709385216, "samples_seen": 91200, "time": 1791144527.9778004}
|
| 116 |
+
{"step": 2875, "loss": 1.189599106311798, "grad_norm": 0.4402064085006714, "lr": 7.690147724595069e-05, "s_per_step": 2.5021799850463866, "peak_gb": 35.709385216, "samples_seen": 92000, "time": 1791144590.5327408}
|
| 117 |
+
{"step": 2900, "loss": 1.1737463009357452, "grad_norm": 0.41730257868766785, "lr": 7.529628978215513e-05, "s_per_step": 2.671381711959839, "peak_gb": 35.709385216, "samples_seen": 92800, "time": 1791144657.317731}
|
| 118 |
+
{"step": 2925, "loss": 1.1911302745342254, "grad_norm": 0.5390501022338867, "lr": 7.369785322334733e-05, "s_per_step": 2.4857681655883788, "peak_gb": 35.709385216, "samples_seen": 93600, "time": 1791144719.4624217}
|
| 119 |
+
{"step": 2950, "loss": 1.1944737762212754, "grad_norm": 0.5601727962493896, "lr": 7.210660438218596e-05, "s_per_step": 2.4794444179534914, "peak_gb": 35.709385216, "samples_seen": 94400, "time": 1791144781.4489458}
|
| 120 |
+
{"step": 2975, "loss": 1.1881670886278153, "grad_norm": 0.5373884439468384, "lr": 7.052297810710643e-05, "s_per_step": 2.483798608779907, "peak_gb": 35.709385216, "samples_seen": 95200, "time": 1791144843.5443158}
|
| 121 |
+
{"step": 3000, "loss": 1.2002181720733642, "grad_norm": 0.5904892683029175, "lr": 6.894740716348797e-05, "s_per_step": 2.475346622467041, "peak_gb": 35.709385216, "samples_seen": 96000, "time": 1791144905.4283826}
|
| 122 |
+
{"step": 3025, "loss": 1.1833700090646744, "grad_norm": 0.5156553387641907, "lr": 6.738032211538946e-05, "s_per_step": 2.6700209045410155, "peak_gb": 35.709385216, "samples_seen": 96800, "time": 1791144972.179436}
|
| 123 |
+
{"step": 3050, "loss": 1.1821686124801636, "grad_norm": 0.5594586730003357, "lr": 6.582215120788719e-05, "s_per_step": 2.4451071548461916, "peak_gb": 35.709385216, "samples_seen": 97600, "time": 1791145033.3075602}
|
| 124 |
+
{"step": 3075, "loss": 1.1715920561552047, "grad_norm": 0.44480738043785095, "lr": 6.427332025004628e-05, "s_per_step": 2.616603384017944, "peak_gb": 35.709385216, "samples_seen": 98400, "time": 1791145098.7230656}
|
| 125 |
+
{"step": 3100, "loss": 1.178504084944725, "grad_norm": 0.527042031288147, "lr": 6.273425249855752e-05, "s_per_step": 2.5396651363372804, "peak_gb": 35.709385216, "samples_seen": 99200, "time": 1791145162.2151375}
|
| 126 |
+
{"step": 3125, "loss": 1.1844954544305801, "grad_norm": 0.521228551864624, "lr": 6.120536854207215e-05, "s_per_step": 2.5937981700897215, "peak_gb": 35.709385216, "samples_seen": 100000, "time": 1791145227.0605233}
|
| 127 |
+
{"step": 3150, "loss": 1.197044386267662, "grad_norm": 0.5560638308525085, "lr": 5.968708618626539e-05, "s_per_step": 2.4499312114715575, "peak_gb": 35.709385216, "samples_seen": 100800, "time": 1791145288.309198}
|
| 128 |
+
{"step": 3175, "loss": 1.1869262939691543, "grad_norm": 0.5193493962287903, "lr": 5.817982033966059e-05, "s_per_step": 2.4811374282836915, "peak_gb": 35.709385216, "samples_seen": 101600, "time": 1791145350.3380346}
|
| 129 |
+
{"step": 3200, "loss": 1.17697389960289, "grad_norm": 0.5577865839004517, "lr": 5.668398290024517e-05, "s_per_step": 2.537874908447266, "peak_gb": 35.709385216, "samples_seen": 102400, "time": 1791145413.7853076}
|
| 130 |
+
{"step": 3225, "loss": 1.1980028331279755, "grad_norm": 0.41638582944869995, "lr": 5.5199982642909375e-05, "s_per_step": 2.443896884918213, "peak_gb": 35.709385216, "samples_seen": 103200, "time": 1791145474.8831582}
|
| 131 |
+
{"step": 3250, "loss": 1.2014775633811952, "grad_norm": 0.5623952150344849, "lr": 5.37282251077381e-05, "s_per_step": 2.4464002323150633, "peak_gb": 35.709385216, "samples_seen": 104000, "time": 1791145536.0436106}
|
| 132 |
+
{"step": 3275, "loss": 1.2081005316972733, "grad_norm": 0.5129700899124146, "lr": 5.2269112489187e-05, "s_per_step": 2.5506320571899415, "peak_gb": 35.709385216, "samples_seen": 104800, "time": 1791145599.8098936}
|
| 133 |
+
{"step": 3300, "loss": 1.1788170540332794, "grad_norm": 0.505996823310852, "lr": 5.082304352617291e-05, "s_per_step": 2.4294992446899415, "peak_gb": 35.709385216, "samples_seen": 105600, "time": 1791145660.547851}
|
| 134 |
+
{"step": 3325, "loss": 1.1952128916978837, "grad_norm": 0.48522287607192993, "lr": 4.939041339310858e-05, "s_per_step": 2.348098773956299, "peak_gb": 35.709385216, "samples_seen": 106400, "time": 1791145719.2507544}
|
| 135 |
+
{"step": 3350, "loss": 1.164835900068283, "grad_norm": 0.4022349715232849, "lr": 4.797161359191104e-05, "s_per_step": 2.542964315414429, "peak_gb": 35.709385216, "samples_seen": 107200, "time": 1791145782.8253438}
|
| 136 |
+
{"step": 3375, "loss": 1.1721830260753632, "grad_norm": 0.5090497136116028, "lr": 4.656703184501427e-05, "s_per_step": 2.52671706199646, "peak_gb": 35.709385216, "samples_seen": 108000, "time": 1791145845.9937546}
|
| 137 |
+
{"step": 3400, "loss": 1.1724821329116821, "grad_norm": 0.5728841423988342, "lr": 4.517705198941442e-05, "s_per_step": 2.4924769973754883, "peak_gb": 35.709385216, "samples_seen": 108800, "time": 1791145908.3061934}
|
| 138 |
+
{"step": 3425, "loss": 1.1960626763105393, "grad_norm": 0.562131941318512, "lr": 4.380205387177645e-05, "s_per_step": 2.408029260635376, "peak_gb": 35.709385216, "samples_seen": 109600, "time": 1791145968.5073981}
|
| 139 |
+
{"step": 3450, "loss": 1.1602804332971572, "grad_norm": 0.4601585268974304, "lr": 4.244241324463182e-05, "s_per_step": 2.5518965339660644, "peak_gb": 35.709385216, "samples_seen": 110400, "time": 1791146032.3052933}
|
| 140 |
+
{"step": 3475, "loss": 1.1917330080270767, "grad_norm": 0.5174102187156677, "lr": 4.109850166369465e-05, "s_per_step": 2.4686901664733885, "peak_gb": 35.709385216, "samples_seen": 111200, "time": 1791146094.0923746}
|
| 141 |
+
{"step": 3500, "loss": 1.176536915898323, "grad_norm": 0.46731457114219666, "lr": 3.977068638632486e-05, "s_per_step": 2.5151400184631347, "peak_gb": 35.709385216, "samples_seen": 112000, "time": 1791146156.9713702}
|
| 142 |
+
{"step": 3525, "loss": 1.1758270978927612, "grad_norm": 0.5187398195266724, "lr": 3.845933027116595e-05, "s_per_step": 2.5711546421051024, "peak_gb": 35.709385216, "samples_seen": 112800, "time": 1791146221.2507052}
|
| 143 |
+
{"step": 3550, "loss": 1.1746439105272293, "grad_norm": 0.46097496151924133, "lr": 3.716479167898476e-05, "s_per_step": 2.4746162128448486, "peak_gb": 35.709385216, "samples_seen": 113600, "time": 1791146283.1165752}
|
| 144 |
+
{"step": 3575, "loss": 1.1660949611663818, "grad_norm": 0.5945229530334473, "lr": 3.588742437474063e-05, "s_per_step": 2.5315122985839844, "peak_gb": 35.709385216, "samples_seen": 114400, "time": 1791146346.4049218}
|
| 145 |
+
{"step": 3600, "loss": 1.177633284330368, "grad_norm": 0.42318975925445557, "lr": 3.4627577430910066e-05, "s_per_step": 2.5272295093536377, "peak_gb": 35.709385216, "samples_seen": 115200, "time": 1791146409.5861077}
|
| 146 |
+
{"step": 3625, "loss": 1.1747868591547013, "grad_norm": 0.5484309792518616, "lr": 3.3385595132094104e-05, "s_per_step": 2.5058855533599855, "peak_gb": 35.709385216, "samples_seen": 116000, "time": 1791146472.2342918}
|
| 147 |
+
{"step": 3650, "loss": 1.1845440715551376, "grad_norm": 0.5164344906806946, "lr": 3.2161816880933996e-05, "s_per_step": 2.420202522277832, "peak_gb": 35.709385216, "samples_seen": 116800, "time": 1791146532.7397869}
|
| 148 |
+
{"step": 3675, "loss": 1.1623845690488814, "grad_norm": 0.44582146406173706, "lr": 3.095657710536086e-05, "s_per_step": 2.5247189235687255, "peak_gb": 35.709385216, "samples_seen": 117600, "time": 1791146595.8581657}
|
| 149 |
+
{"step": 3700, "loss": 1.177596697807312, "grad_norm": 0.6056506633758545, "lr": 2.977020516720501e-05, "s_per_step": 2.4129674339294436, "peak_gb": 35.709385216, "samples_seen": 118400, "time": 1791146656.1832361}
|
| 150 |
+
{"step": 3725, "loss": 1.1791141772270202, "grad_norm": 0.5385143160820007, "lr": 2.860302527218951e-05, "s_per_step": 2.466406307220459, "peak_gb": 35.709385216, "samples_seen": 119200, "time": 1791146717.8438177}
|
| 151 |
+
{"step": 3750, "loss": 1.1454178375005721, "grad_norm": 0.5121068954467773, "lr": 2.745535638133304e-05, "s_per_step": 2.5826826667785645, "peak_gb": 35.709385216, "samples_seen": 120000, "time": 1791146782.4113297}
|
| 152 |
+
{"step": 3775, "loss": 1.1689221292734147, "grad_norm": 0.5559164881706238, "lr": 2.632751212378568e-05, "s_per_step": 2.653624095916748, "peak_gb": 35.709385216, "samples_seen": 120800, "time": 1791146848.7524045}
|
| 153 |
+
{"step": 3800, "loss": 1.182214805483818, "grad_norm": 0.4873104691505432, "lr": 2.5219800711121976e-05, "s_per_step": 2.498158531188965, "peak_gb": 35.709385216, "samples_seen": 121600, "time": 1791146911.2068315}
|
| 154 |
+
{"step": 3825, "loss": 1.1555895429849625, "grad_norm": 0.576235294342041, "lr": 2.4132524853114456e-05, "s_per_step": 2.463836088180542, "peak_gb": 35.709385216, "samples_seen": 122400, "time": 1791146972.803223}
|
| 155 |
+
{"step": 3850, "loss": 1.163134110569954, "grad_norm": 0.49555864930152893, "lr": 2.306598167501064e-05, "s_per_step": 2.5268060398101806, "peak_gb": 35.709385216, "samples_seen": 123200, "time": 1791147035.973819}
|
| 156 |
+
{"step": 3875, "loss": 1.1696758967638017, "grad_norm": 0.519751250743866, "lr": 2.2020462636336137e-05, "s_per_step": 2.439229755401611, "peak_gb": 35.709385216, "samples_seen": 124000, "time": 1791147096.9549644}
|
| 157 |
+
{"step": 3900, "loss": 1.1512171399593354, "grad_norm": 0.42172372341156006, "lr": 2.0996253451246017e-05, "s_per_step": 2.601565704345703, "peak_gb": 35.709385216, "samples_seen": 124800, "time": 1791147161.9945152}
|
| 158 |
+
{"step": 3925, "loss": 1.1549856066703796, "grad_norm": 0.5235016942024231, "lr": 1.9993634010446448e-05, "s_per_step": 2.5241199016571043, "peak_gb": 35.709385216, "samples_seen": 125600, "time": 1791147225.0979981}
|
| 159 |
+
{"step": 3950, "loss": 1.1684985667467118, "grad_norm": 0.5432916283607483, "lr": 1.9012878304707403e-05, "s_per_step": 2.5513547325134276, "peak_gb": 35.709385216, "samples_seen": 126400, "time": 1791147288.8822782}
|
| 160 |
+
{"step": 3975, "loss": 1.1723533356189728, "grad_norm": 0.5285301804542542, "lr": 1.805425434998782e-05, "s_per_step": 2.436748571395874, "peak_gb": 35.709385216, "samples_seen": 127200, "time": 1791147349.801405}
|
| 161 |
+
{"step": 4000, "loss": 1.1497523707151414, "grad_norm": 0.44535908102989197, "lr": 1.711802411419384e-05, "s_per_step": 2.6349299812316893, "peak_gb": 35.709385216, "samples_seen": 128000, "time": 1791147415.675062}
|
| 162 |
+
{"step": 4025, "loss": 1.1554417186975479, "grad_norm": 0.4447822570800781, "lr": 1.6204443445589225e-05, "s_per_step": 2.693064069747925, "peak_gb": 35.709385216, "samples_seen": 128800, "time": 1791147483.002096}
|
| 163 |
+
{"step": 4050, "loss": 1.1771225923299788, "grad_norm": 0.5996832251548767, "lr": 1.5313762002878586e-05, "s_per_step": 2.506432514190674, "peak_gb": 35.709385216, "samples_seen": 129600, "time": 1791147545.6633337}
|
| 164 |
+
{"step": 4075, "loss": 1.156247045993805, "grad_norm": 0.5294703841209412, "lr": 1.444622318698191e-05, "s_per_step": 2.4499019813537597, "peak_gb": 35.709385216, "samples_seen": 130400, "time": 1791147606.9112895}
|
| 165 |
+
{"step": 4100, "loss": 1.1654343736171722, "grad_norm": 0.472944051027298, "lr": 1.3602064074519216e-05, "s_per_step": 2.5358735370635985, "peak_gb": 35.709385216, "samples_seen": 131200, "time": 1791147670.3085465}
|
| 166 |
+
{"step": 4125, "loss": 1.164423559308052, "grad_norm": 0.41716107726097107, "lr": 1.2781515353023365e-05, "s_per_step": 2.4599155235290526, "peak_gb": 35.709385216, "samples_seen": 132000, "time": 1791147731.8068397}
|
| 167 |
+
{"step": 4150, "loss": 1.1569818568229675, "grad_norm": 0.6027914881706238, "lr": 1.1984801257898925e-05, "s_per_step": 2.5184785079956056, "peak_gb": 35.709385216, "samples_seen": 132800, "time": 1791147794.76921}
|
| 168 |
+
{"step": 4175, "loss": 1.1682433980703353, "grad_norm": 0.5346322059631348, "lr": 1.1212139511144448e-05, "s_per_step": 2.3972896575927733, "peak_gb": 35.709385216, "samples_seen": 133600, "time": 1791147854.701882}
|
| 169 |
+
{"step": 4200, "loss": 1.1649791526794433, "grad_norm": 0.5470781326293945, "lr": 1.0463741261854288e-05, "s_per_step": 2.492423486709595, "peak_gb": 35.709385216, "samples_seen": 134400, "time": 1791147917.01288}
|
| 170 |
+
{"step": 4225, "loss": 1.1546978533267975, "grad_norm": 0.4576592743396759, "lr": 9.739811028516932e-06, "s_per_step": 2.527438831329346, "peak_gb": 35.709385216, "samples_seen": 135200, "time": 1791147980.1992562}
|
| 171 |
+
{"step": 4250, "loss": 1.1561020785570144, "grad_norm": 0.5178248286247253, "lr": 9.040546643125203e-06, "s_per_step": 2.517265033721924, "peak_gb": 35.709385216, "samples_seen": 136000, "time": 1791148043.1313138}
|
| 172 |
+
{"step": 4275, "loss": 1.1623137599229814, "grad_norm": 0.47821733355522156, "lr": 8.366139197113831e-06, "s_per_step": 2.5566812801361083, "peak_gb": 35.709385216, "samples_seen": 136800, "time": 1791148107.0487514}
|
| 173 |
+
{"step": 4300, "loss": 1.1758109939098358, "grad_norm": 0.5945451855659485, "lr": 7.716772989138744e-06, "s_per_step": 2.4145172882080077, "peak_gb": 35.709385216, "samples_seen": 137600, "time": 1791148167.4120862}
|
| 174 |
+
{"step": 4325, "loss": 1.1620000940561295, "grad_norm": 0.4968617260456085, "lr": 7.092625474713055e-06, "s_per_step": 2.42730583190918, "peak_gb": 35.709385216, "samples_seen": 138400, "time": 1791148228.0951617}
|
| 175 |
+
{"step": 4350, "loss": 1.156725025177002, "grad_norm": 0.45200902223587036, "lr": 6.493867217712868e-06, "s_per_step": 2.5541227149963377, "peak_gb": 35.709385216, "samples_seen": 139200, "time": 1791148291.9486508}
|
| 176 |
+
{"step": 4375, "loss": 1.1519554841518402, "grad_norm": 0.5528603792190552, "lr": 5.920661843766395e-06, "s_per_step": 2.648443021774292, "peak_gb": 35.889089024, "samples_seen": 140000, "time": 1791148358.1601393}
|
| 177 |
+
{"step": 4400, "loss": 1.1472780203819275, "grad_norm": 0.5185584425926208, "lr": 5.3731659955391865e-06, "s_per_step": 2.5665468978881836, "peak_gb": 35.889089024, "samples_seen": 140800, "time": 1791148422.3242114}
|
| 178 |
+
{"step": 4425, "loss": 1.16106074988842, "grad_norm": 0.47097915410995483, "lr": 4.8515292899276695e-06, "s_per_step": 2.5675136184692384, "peak_gb": 35.889089024, "samples_seen": 141600, "time": 1791148486.5124738}
|
| 179 |
+
{"step": 4450, "loss": 1.1584270763397218, "grad_norm": 0.5423069596290588, "lr": 4.355894277172545e-06, "s_per_step": 2.5209413051605223, "peak_gb": 35.889089024, "samples_seen": 142400, "time": 1791148549.5364525}
|
| 180 |
+
{"step": 4475, "loss": 1.167509251832962, "grad_norm": 0.4573025703430176, "lr": 3.886396401903381e-06, "s_per_step": 2.452349863052368, "peak_gb": 35.889089024, "samples_seen": 143200, "time": 1791148610.8456094}
|
| 181 |
+
{"step": 4500, "loss": 1.15239905834198, "grad_norm": 0.5477348566055298, "lr": 3.443163966125007e-06, "s_per_step": 2.4918809604644774, "peak_gb": 35.889089024, "samples_seen": 144000, "time": 1791148673.1431098}
|
| 182 |
+
{"step": 4525, "loss": 1.161824835538864, "grad_norm": 0.6060011386871338, "lr": 3.0263180941558332e-06, "s_per_step": 2.6219866466522217, "peak_gb": 35.889089024, "samples_seen": 144800, "time": 1791148738.6932275}
|
| 183 |
+
{"step": 4550, "loss": 1.1565506535768508, "grad_norm": 0.5005341172218323, "lr": 2.6359726995275004e-06, "s_per_step": 2.566774663925171, "peak_gb": 35.889089024, "samples_seen": 145600, "time": 1791148802.8630507}
|
| 184 |
+
{"step": 4575, "loss": 1.162327172756195, "grad_norm": 0.5122349858283997, "lr": 2.2722344538552597e-06, "s_per_step": 2.571699342727661, "peak_gb": 35.889089024, "samples_seen": 146400, "time": 1791148867.1559367}
|
| 185 |
+
{"step": 4600, "loss": 1.1596791630983352, "grad_norm": 0.5953731536865234, "lr": 1.9352027576872824e-06, "s_per_step": 2.364423170089722, "peak_gb": 35.889089024, "samples_seen": 147200, "time": 1791148926.266977}
|
| 186 |
+
{"step": 4625, "loss": 1.1516347283124924, "grad_norm": 0.49013832211494446, "lr": 1.6249697133409069e-06, "s_per_step": 2.5397282600402833, "peak_gb": 35.889089024, "samples_seen": 148000, "time": 1791148989.760581}
|
| 187 |
+
{"step": 4650, "loss": 1.1604260104894637, "grad_norm": 0.44168010354042053, "lr": 1.341620099733487e-06, "s_per_step": 2.495543918609619, "peak_gb": 35.889089024, "samples_seen": 148800, "time": 1791149052.149569}
|
| 188 |
+
{"step": 4675, "loss": 1.144905739426613, "grad_norm": 0.47312289476394653, "lr": 1.0852313492143551e-06, "s_per_step": 2.527244834899902, "peak_gb": 35.889089024, "samples_seen": 149600, "time": 1791149115.3311636}
|
| 189 |
+
{"step": 4700, "loss": 1.164122424721718, "grad_norm": 0.5258249044418335, "lr": 8.558735264045714e-07, "s_per_step": 2.392429494857788, "peak_gb": 35.889089024, "samples_seen": 150400, "time": 1791149175.1423135}
|
| 190 |
+
{"step": 4725, "loss": 1.1654922747612, "grad_norm": 0.5145270228385925, "lr": 6.536093090499517e-07, "s_per_step": 2.493706407546997, "peak_gb": 35.889089024, "samples_seen": 151200, "time": 1791149237.4853363}
|
| 191 |
+
{"step": 4750, "loss": 1.170558487176895, "grad_norm": 0.6493576169013977, "lr": 4.78493970892846e-07, "s_per_step": 2.3715320777893067, "peak_gb": 35.889089024, "samples_seen": 152000, "time": 1791149296.775865}
|
| 192 |
+
{"step": 4775, "loss": 1.166077768802643, "grad_norm": 0.5020484328269958, "lr": 3.3057536656722065e-07, "s_per_step": 2.589771041870117, "peak_gb": 35.889089024, "samples_seen": 152800, "time": 1791149361.5205963}
|
| 193 |
+
{"step": 4800, "loss": 1.1575500231981277, "grad_norm": 0.4154891073703766, "lr": 2.0989391852116458e-07, "s_per_step": 2.4963897895812988, "peak_gb": 35.889089024, "samples_seen": 153600, "time": 1791149423.930882}
|
| 194 |
+
{"step": 4825, "loss": 1.1519961988925933, "grad_norm": 0.5121576189994812, "lr": 1.1648260597041383e-07, "s_per_step": 2.559989347457886, "peak_gb": 35.889089024, "samples_seen": 154400, "time": 1791149487.9312465}
|
| 195 |
+
{"step": 4850, "loss": 1.1676404130458833, "grad_norm": 0.550052285194397, "lr": 5.036695588606088e-08, "s_per_step": 2.4459603118896482, "peak_gb": 35.889089024, "samples_seen": 155200, "time": 1791149549.080826}
|
| 196 |
+
{"step": 4875, "loss": 1.169189488887787, "grad_norm": 0.5803373456001282, "lr": 1.1565036018568176e-08, "s_per_step": 2.4609833908081056, "peak_gb": 35.889089024, "samples_seen": 156000, "time": 1791149610.6059558}
|
| 197 |
+
{"step": 4897, "loss": 1.1497614966197447, "grad_norm": 0.4358130097389221, "lr": 2.1862492483037954e-11, "s_per_step": 2.5117971788753164, "peak_gb": 35.889089024, "samples_seen": 156704, "time": 1791149665.8659236}
|
stage1/projector.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:43bfac0c0f58fb43a395d253206896f0dfb1c30ecd1cc90750bb13bfc91cbced
|
| 3 |
+
size 79724872
|
stage2/lora_adapter/adapter_config.json
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "NCAIR1/N-ATLaS",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"kasa_config": null,
|
| 16 |
+
"layer_replication": null,
|
| 17 |
+
"layers_pattern": null,
|
| 18 |
+
"layers_to_transform": null,
|
| 19 |
+
"loftq_config": {},
|
| 20 |
+
"lora_alpha": 128,
|
| 21 |
+
"lora_bias": false,
|
| 22 |
+
"lora_dropout": 0.05,
|
| 23 |
+
"lora_ga_config": null,
|
| 24 |
+
"megatron_config": null,
|
| 25 |
+
"megatron_core": "megatron.core",
|
| 26 |
+
"modules_to_save": null,
|
| 27 |
+
"monteclora_config": null,
|
| 28 |
+
"peft_type": "LORA",
|
| 29 |
+
"peft_version": "0.21.2",
|
| 30 |
+
"qalora_group_size": 16,
|
| 31 |
+
"r": 64,
|
| 32 |
+
"rank_pattern": {},
|
| 33 |
+
"revision": null,
|
| 34 |
+
"target_modules": [
|
| 35 |
+
"k_proj",
|
| 36 |
+
"o_proj",
|
| 37 |
+
"q_proj",
|
| 38 |
+
"up_proj",
|
| 39 |
+
"down_proj",
|
| 40 |
+
"gate_proj",
|
| 41 |
+
"v_proj"
|
| 42 |
+
],
|
| 43 |
+
"target_parameters": null,
|
| 44 |
+
"task_type": "CAUSAL_LM",
|
| 45 |
+
"trainable_token_indices": null,
|
| 46 |
+
"use_bdlora": null,
|
| 47 |
+
"use_dora": false,
|
| 48 |
+
"use_qalora": false,
|
| 49 |
+
"use_rslora": false,
|
| 50 |
+
"velora_config": null
|
| 51 |
+
}
|
stage2/lora_adapter/adapter_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:06df60018c6f604a693251a4a95e96d00ab90a5431f7421f2b6a57bc0020cba4
|
| 3 |
+
size 671149168
|
stage2/projector.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e6b3dd6e3137450b00c0a11467c6476dc32300d41a68ea63bdb7aa562af51b44
|
| 3 |
+
size 79724872
|