Image-Text-to-Text
PEFT
Safetensors
vision-language
multimodal
llava
lora
siglip2
n-atlas
nigerian-languages
File size: 16,222 Bytes
afb7034
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
afb7034
2a2540a
 
 
 
 
 
 
 
 
 
 
c8317c0
 
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c8317c0
 
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c8317c0
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
c8317c0
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c8317c0
 
 
 
 
 
 
 
 
 
 
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c8317c0
2a2540a
c8317c0
 
 
2a2540a
c8317c0
 
2a2540a
c8317c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2a2540a
 
 
 
 
 
 
 
 
 
 
 
c8317c0
2a2540a
 
 
 
c8317c0
2a2540a
c8317c0
2a2540a
 
 
 
c8317c0
2a2540a
c8317c0
2a2540a
 
 
 
c8317c0
2a2540a
c8317c0
2a2540a
c8317c0
2a2540a
c8317c0
2a2540a
 
 
c8317c0
 
 
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
c8317c0
 
 
 
 
2a2540a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
---
license: other
license_name: awarri-open-source-research-and-innovation-license
license_link: https://huggingface.co/NCAIR1/N-ATLaS
base_model:
- NCAIR1/N-ATLaS
- google/siglip2-base-patch16-224
library_name: peft
pipeline_tag: image-text-to-text
language:
- en
- ha
- ig
- yo
datasets:
- liuhaotian/LLaVA-Pretrain
- liuhaotian/LLaVA-Instruct-150K
- lmms-lab/POPE
tags:
- vision-language
- multimodal
- llava
- lora
- siglip2
- n-atlas
- nigerian-languages
---

# AtlasVision

**AtlasVision** gives [N-ATLaS](https://huggingface.co/NCAIR1/N-ATLaS) — Nigeria's multilingual Llama-3-8B model for
English, Hausa, Igbo and Yoruba — the ability to see. It connects a frozen
[SigLIP2](https://huggingface.co/google/siglip2-base-patch16-224) image encoder to N-ATLaS through a trained MLP projector,
following the two-stage LLaVA recipe:

1. **Stage 1 — alignment:** only the projector is trained, on 300k image–caption pairs, so the LLM can "read" image features.
2. **Stage 2 — visual instruction tuning:** the projector keeps training and LoRA adapters are added to N-ATLaS, on
   LLaVA-Instruct-150K conversations, so the model answers questions and holds conversations about images.
3. **Stage 2b — de-biasing (recommended):** stage 2 is continued on a 150k mix of short-answer VQA data, which fixes a
   severe "yes" bias and raises POPE accuracy from ~52% to ~78%. **Use `stage2b/` unless you are reproducing earlier results.**

This repository contains **only the trained parts** (projectors and LoRA adapter, ~0.9 GB).
The base models are downloaded from their own repositories at load time, so their licenses and access conditions
(N-ATLaS is gated) continue to apply.

## Model summary

| | |
|---|---|
| Language model | `NCAIR1/N-ATLaS` (Llama-3 8B, frozen; LoRA in stage 2) |
| Vision encoder | `google/siglip2-base-patch16-224` (ViT-B/16, 224 px, 196 patch tokens, frozen) |
| Projector | 2-layer MLP 768 → 4096 → 4096 with GELU (19.9M params) |
| LoRA (stage 2) | r = 64, alpha = 128, dropout 0.05 on q/k/v/o/gate/up/down projections (167.8M params) |
| Image tokens | 196 per image, inserted after `<|begin_of_text|>` and the user header |
| Precision | bf16 weights and compute; projector and LoRA trained in fp32 |
| Languages | Training data is English; N-ATLaS contributes Hausa, Igbo and Yoruba |

## Repository contents

```
stage1/projector.safetensors          # stage-1 projector (image captioning / alignment)
stage2/projector.safetensors          # stage-2 projector (use together with the LoRA adapter)
stage2/lora_adapter/                  # PEFT LoRA adapter for NCAIR1/N-ATLaS
stage2b/projector.safetensors         # RECOMMENDED: de-biased projector
stage2b/lora_adapter/                 # RECOMMENDED: de-biased LoRA adapter
chat.py                               # standalone inference script (CLI + Python API)
code/                                 # exact training, evaluation and launch scripts used
eval/                                 # raw evaluation results (JSON) and report
logs/                                 # per-step training logs (loss, grad norm, LR, speed)
```

## Quick start

Accept the conditions on the [N-ATLaS model page](https://huggingface.co/NCAIR1/N-ATLaS) first, then:

```bash
pip install -U torch transformers peft safetensors pillow accelerate huggingface_hub
export HF_TOKEN=hf_...            # a token from the account that was granted N-ATLaS access
wget https://huggingface.co/FUTO-NIGERIA/AtlasVision/resolve/main/chat.py

python chat.py --image photo.jpg --question "What is happening in this picture?"   # add --stage 2b
python chat.py --image photo.jpg --question "Kedu ihe dị na foto a?"      # Igbo
python chat.py --stage 1 --image photo.jpg                                 # stage-1 captioner
python chat.py --image photo.jpg --interactive                             # follow-up questions
python chat.py --image photo.jpg --load-in-4bit                            # ~8 GB GPU (pip install bitsandbytes)
```

bf16 needs about 18 GB of GPU memory (A100, L4, RTX 4090…). `--load-in-4bit` runs on ~8 GB GPUs such as a free Colab T4.

From Python:

```python
from chat import AtlasVision, load_image

model = AtlasVision(stage="2b")                  # downloads N-ATLaS, SigLIP2 and this repo's weights
image = load_image("https://example.com/street.jpg")
print(model.ask(image, "Describe this image in detail."))
print(model.ask(image, "How many people are there?"))   # follow-ups keep the conversation
```

### Prompt format

The image embeddings are spliced into the Llama-3 chat format:

```
<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n [196 image embeddings] {question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n
```

Follow-up turns append `{answer}<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n{question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n`.
`chat.py` builds this for you.

## Training

Both stages ran on a single NVIDIA H100 80GB.

| | Stage 1 — alignment | Stage 2 — instruction tuning | Stage 2b — de-biasing |
|---|---|---|---|
| Data | LLaVA-Pretrain (BLIP captions of LAION/CC/SBU), random 300k of 558k | LLaVA-Instruct-150K on COCO train2017 (156,712 train / 1,000 held out) | LLaVA-1.5 mix: 110k short-answer VQA (VQAv2 / OK-VQA / A-OKVQA) + 40k LLaVA-Instruct replay, COCO images |
| Trained | projector | projector + LoRA | projector + LoRA (continued from stage 2) |
| Effective batch | 32 (16 × 2 accumulation) | 32 (8 × 4 accumulation) | 32 (8 × 4 accumulation) |
| Learning rate | 2e-4, 3% warmup, cosine | LoRA 2e-4, projector 2e-5, 3% warmup, cosine | LoRA 1e-4, projector 1e-5, 3% warmup, cosine |
| Max text length | 128 tokens | 1,024 tokens | 1,024 tokens |
| Optimizer steps | 9375 | 4897 | 4,681 |
| Loss (first log → last log) | 8.551 → 2.454 | 2.102 → 1.15 | 3.14 → 0.55 |
| Wall-clock time | ~91 min | ~206 min | ~172 min |
| Peak GPU memory | 45.2 GB | 35.9 GB (gradient checkpointing) | 37.3 GB |

Other details: AdamW (no weight decay), gradient clipping at 1.0, bf16 autocast, loss only on assistant tokens,
length-bucketed batches in stage 2. Per-step logs are in `logs/`.

## Evaluation

### Stage 1 — does the model actually use the image?

Measured on 2,000 LLaVA-Pretrain captions that were **not** in the training subset (caption loss, lower is better):

| Setting | Held-out loss |
|---|---|
| Untrained (random) projector | 8.7249 |
| **Trained stage-1 projector** | **2.4523** |
| Trained projector, captions paired with the wrong images | 5.4332 |

Held-out loss matches the final training loss (no over-fitting), and pairing captions with the wrong images
roughly doubles the loss — the language model is relying on the visual content, not guessing generic captions.

### Stage 1 vs stage 2 vs stage 2b

| Metric | Stage 1 | Stage 2 | Stage 2b (recommended) |
|---|---|---|---|
| Held-out LLaVA-Instruct loss (500 unseen conversations, lower is better) | 2.2521 | 1.1271 | **1.1602** |

**POPE** (object hallucination: yes/no questions about COCO val2014 images, scored from the Yes/No token
probabilities; the benchmark is balanced 50% yes / 50% no, so a yes-ratio near 0.5 is ideal):

| POPE split | Stage 2: accuracy / F1 / yes-ratio | Stage 2b: accuracy / F1 / yes-ratio |
|---|---|---|
| random | 0.543 / 0.686 / 0.95 | **0.822** / **0.832** / **0.56** |
| popular | 0.522 / 0.676 / 0.98 | **0.776** / **0.797** / **0.60** |
| adversarial | 0.514 / 0.672 / 0.98 | **0.752** / **0.780** / **0.63** |

#### What stage 2b fixed

Stage 2 answered **"yes" to 97–98%** of POPE questions. It confirmed almost every object it was asked about,
present or not, so accuracy sat at chance (51–54%) even though its descriptions were good. The cause is the
training data: LLaVA-Instruct-150K is made of long GPT-4-written conversations and contains almost no question
whose answer is "no".

Stage 2b continues stage 2 for one epoch on a 150k mix drawn from the public LLaVA-1.5 training set: 110k
short-answer VQA conversations (**57,804 "yes" vs 59,443 "no"** turns) plus 40k LLaVA-Instruct conversations
replayed so detailed description is not lost. Every POPE test image and every held-out conversation was removed
from this mix before training, so the numbers above are not contaminated; region/bounding-box tasks were dropped
as unsupported. Result: the yes-ratio fell to 0.56–0.63, accuracy rose by **24–28 points**, and the held-out
conversation loss improved as well (1.1271 → 1.1602),
so description quality did not regress. For reference, LLaVA-1.5-7B reports POPE F1 ≈ 0.86 from 665k samples at
336 px; AtlasVision reaches 0.83 from 150k samples at 224 px.

Recall (0.88) still exceeds precision
(0.70) on the adversarial split, so a mild "yes" lean remains on the
hardest confusable pairs.

### Examples — stage 1 captions on unseen images

| Reference caption | AtlasVision stage 1 |
|---|---|
| cute homecoming prom dress with lace top and satin skirt | a - line sweetheart sweetheart lace tulle prom dress |
| a pair of silicone bracelets with a four codes message | a pair of silicone bracelets with the word code on them |
| a set of fingerprint icons in white on a black background | the iphone 6s and iphone 6s plus are shown in a black and white image |
| a business woman drawing modern concept of a website creation | a man writing the word modern web design on a whiteboard |
| the beach in el nido national park, puerto puerto | the limestone cliffs and limestone islands in the background |
| three pieces of paper with the words, democratic decentified dp controlled centralized ccp | a diagram showing the different types of democracy |

### Examples — "Describe this image in detail." (stage 2b, unseen COCO val2014)

**COCO_val2014_000000310196**

- *Stage 1:* a skier in the snow on a mountain
- *Stage 2b:* The image features a snow-covered slope with a skier in the center, wearing red and black clothing. The skier is skillfully navigating down the hill, making their way through the snow. 

In addition to the main skier, there are several other people scattered across the scene, some of them closer to the top of the slope while others are further away or near the bottom. These individuals may also be skiing or enjoying the snowy landscape.

**COCO_val2014_000000210789**

- *Stage 1:* a woman and her child in the rain
- *Stage 2b:* The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The little girl is wearing a pink dress, while the woman appears to be her mother. They are positioned close together, with the child on the left side of the frame and the woman on the right.

In addition to the two main subjects, there are several other people visible in the background, some of whom may also be holding umbrellas or trying to stay dry during the rainy weather. A car can be seen parked further back in the scene, likely waiting for its occupants

**COCO_val2014_000000429109**

- *Stage 1:* a bus and several other vehicles parked in front of a building
- *Stage 2b:* The image features a busy street with several buses and cars parked or driving along the road. There are three buses in total, one on the left side of the scene, another near the center, and the third bus further to the right. A car is also visible on the left side of the image.

In addition to the vehicles, there are multiple people walking around the area. Some individuals can be seen closer to the buses, while others are scattered throughout the scene. The presence of both vehicles and pedestrians suggests that this could be a popular transportation hub or a bustling city street.

### Questions in Nigerian languages (stage 2b)

Stage 2b answers Igbo and Yoruba questions **in those languages** without any translation step (stage 2 always replied in English). Hausa still falls back to English, so the `lang=` cascade below remains the reliable route for all three.

| Language | Question | Answer |
|---|---|---|
| Igbo | Kedu ihe dị na foto a? | Na foto a, e nwere skier na snow. |
| Yoruba | Kí ni ó wà nínú àwòrán yìí? | Nínú àwòrán yìí, ó wà skier tí ń ski lórí òkè. |
| Hausa | Me ke cikin wannan hoton? | In the image, there is a skier wearing red and black clothing who is skiing down a snow-covered slope. |

### Text-only check (no image)

Does the stage-2 LoRA damage N-ATLaS's original text abilities? Same Igbo question, greedy decoding:

**Q:** Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.

- *Base N-ATLaS:* Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
- *With stage-2 LoRA:* Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu "positron" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej

Both answer fluently in Igbo, so the adapter keeps N-ATLaS's text abilities intact (responses are cut at 120 tokens).

## Limitations

- **Hallucination.** Long descriptions are fluent but often add plausible details that are not in the image
  (exact counts, extra people, a bicycle at the edge of frame). Treat counts and small objects as unreliable.
- **Residual yes-bias.** Stage 2b largely fixed stage 2's yes-bias, but recall still exceeds precision on POPE's
  adversarial split, so yes/no answers about easily-confused objects lean positive. Stage 2 weights are kept for
  reproducibility only — do not use them for yes/no verification.
- **English-only visual training.** All image–text training data is English. Answers to Hausa, Igbo and Yoruba questions
  come from N-ATLaS's own multilingual ability and are noticeably less reliable; they may switch to English.
- **Low resolution.** Images are resized to 224×224, so small text, fine details and dense documents are hard.
- **One image per conversation**, no video, no bounding boxes or grounding.
- **Data biases.** Web captions (LAION/CC/SBU) and COCO carry Western-centric content; performance on Nigerian scenes,
  people and text has not been measured and is likely weaker.
- **Not for high-stakes use** (medical, legal, identity, surveillance or safety-critical decisions).

## Licenses and terms

This repository contains weights trained on top of other people's models and data. Using it means complying with all of:

- **N-ATLaS**: Awarri Open-Source Research and Innovation License (attribution required; caps deployments at 1,000
  active end-users per 30 days) and the Llama 3 Community License it inherits. Access is gated — accept the
  conditions on the [model page](https://huggingface.co/NCAIR1/N-ATLaS).
- **SigLIP2**: Apache-2.0.
- **LLaVA-Instruct-150K** (stage 2 data): CC BY-NC 4.0, generated with GPT-4 and subject to OpenAI's terms —
  **stage-2 weights are for non-commercial research use.**
- **LLaVA-Pretrain** (stage 1 data): captions/images from LAION, Conceptual Captions and SBU under their respective terms.
- **COCO images**: Flickr images under their individual Creative Commons licenses.

## Acknowledgements

- N-ATLaS by Awarri Technologies and the Federal Ministry of Communications, Innovation and Digital Economy (Nigeria),
  as part of the Nigerian Languages AI Initiative.
- SigLIP2 by Google; LLaVA data and recipe by Liu et al.; POPE by Li et al.
- Trained on an H100 from San Francisco Compute (Autoresearch).

## Citation

```bibtex
@misc{atlasvision2026,
  title  = {AtlasVision: a vision-language extension of N-ATLaS},
  author = {FUTO-NIGERIA},
  year   = {2026},
  url    = {https://huggingface.co/FUTO-NIGERIA/AtlasVision}
}
@inproceedings{liu2023llava,
  title     = {Visual Instruction Tuning},
  author    = {Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae},
  booktitle = {NeurIPS},
  year      = {2023}
}
```