Image-Text-to-Text
PEFT
Safetensors
vision-language
multimodal
llava
lora
siglip2
n-atlas
nigerian-languages

AtlasVision

AtlasVision gives N-ATLaS — Nigeria's multilingual Llama-3-8B model for English, Hausa, Igbo and Yoruba — the ability to see. It connects a frozen SigLIP2 image encoder to N-ATLaS through a trained MLP projector, following the two-stage LLaVA recipe:

  1. Stage 1 — alignment: only the projector is trained, on 300k image–caption pairs, so the LLM can "read" image features.
  2. Stage 2 — visual instruction tuning: the projector keeps training and LoRA adapters are added to N-ATLaS, on LLaVA-Instruct-150K conversations, so the model answers questions and holds conversations about images.
  3. Stage 2b — de-biasing (recommended): stage 2 is continued on a 150k mix of short-answer VQA data, which fixes a severe "yes" bias and raises POPE accuracy from ~52% to ~78%. Use stage2b/ unless you are reproducing earlier results.

This repository contains only the trained parts (projectors and LoRA adapter, ~0.9 GB). The base models are downloaded from their own repositories at load time, so their licenses and access conditions (N-ATLaS is gated) continue to apply.

Model summary

Language model NCAIR1/N-ATLaS (Llama-3 8B, frozen; LoRA in stage 2)
Vision encoder google/siglip2-base-patch16-224 (ViT-B/16, 224 px, 196 patch tokens, frozen)
Projector 2-layer MLP 768 → 4096 → 4096 with GELU (19.9M params)
LoRA (stage 2) r = 64, alpha = 128, dropout 0.05 on q/k/v/o/gate/up/down projections (167.8M params)
Image tokens 196 per image, inserted after `<
Precision bf16 weights and compute; projector and LoRA trained in fp32
Languages Training data is English; N-ATLaS contributes Hausa, Igbo and Yoruba

Repository contents

stage1/projector.safetensors          # stage-1 projector (image captioning / alignment)
stage2/projector.safetensors          # stage-2 projector (use together with the LoRA adapter)
stage2/lora_adapter/                  # PEFT LoRA adapter for NCAIR1/N-ATLaS
stage2b/projector.safetensors         # RECOMMENDED: de-biased projector
stage2b/lora_adapter/                 # RECOMMENDED: de-biased LoRA adapter
chat.py                               # standalone inference script (CLI + Python API, incl. lang= cascade)
code/                                 # exact training, evaluation and launch scripts used
eval/                                 # raw evaluation results (JSON) and report
logs/                                 # per-step training logs (loss, grad norm, LR, speed)

Quick start

Accept the conditions on the N-ATLaS model page first, then:

pip install -U torch transformers peft safetensors pillow accelerate huggingface_hub
export HF_TOKEN=hf_...            # a token from the account that was granted N-ATLaS access
wget https://huggingface.co/Modularcomputing/AtlasVision/resolve/main/chat.py

python chat.py --image photo.jpg --question "What is happening in this picture?"   # add --stage 2b
python chat.py --image photo.jpg --question "Kedu ihe dị na foto a?"      # Igbo
python chat.py --stage 1 --image photo.jpg                                 # stage-1 captioner
python chat.py --image photo.jpg --interactive                             # follow-up questions
python chat.py --image photo.jpg --load-in-4bit                            # ~8 GB GPU (pip install bitsandbytes)

bf16 needs about 18 GB of GPU memory (A100, L4, RTX 4090…). --load-in-4bit runs on ~8 GB GPUs such as a free Colab T4.

From Python:

from chat import AtlasVision, load_image

model = AtlasVision(stage="2b")                  # downloads N-ATLaS, SigLIP2 and this repo's weights
image = load_image("https://example.com/street.jpg")
print(model.ask(image, "Describe this image in detail."))
print(model.ask(image, "How many people are there?"))   # follow-ups keep the conversation

Prompt format

The image embeddings are spliced into the Llama-3 chat format:

<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n [196 image embeddings] {question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n

Follow-up turns append {answer}<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n{question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n. chat.py builds this for you.

Training

Both stages ran on a single NVIDIA H100 80GB.

Stage 1 — alignment Stage 2 — instruction tuning Stage 2b — de-biasing
Data LLaVA-Pretrain (BLIP captions of LAION/CC/SBU), random 300k of 558k LLaVA-Instruct-150K on COCO train2017 (156,712 train / 1,000 held out) LLaVA-1.5 mix: 110k short-answer VQA (VQAv2 / OK-VQA / A-OKVQA) + 40k LLaVA-Instruct replay, COCO images
Trained projector projector + LoRA projector + LoRA (continued from stage 2)
Effective batch 32 (16 × 2 accumulation) 32 (8 × 4 accumulation) 32 (8 × 4 accumulation)
Learning rate 2e-4, 3% warmup, cosine LoRA 2e-4, projector 2e-5, 3% warmup, cosine LoRA 1e-4, projector 1e-5, 3% warmup, cosine
Max text length 128 tokens 1,024 tokens 1,024 tokens
Optimizer steps 9375 4897 4,681
Loss (first log → last log) 8.551 → 2.454 2.102 → 1.15 3.14 → 0.55
Wall-clock time ~91 min ~206 min ~172 min
Peak GPU memory 45.2 GB 35.9 GB (gradient checkpointing) 37.3 GB

Other details: AdamW (no weight decay), gradient clipping at 1.0, bf16 autocast, loss only on assistant tokens, length-bucketed batches in stage 2. Per-step logs are in logs/.

Evaluation

Stage 1 — does the model actually use the image?

Measured on 2,000 LLaVA-Pretrain captions that were not in the training subset (caption loss, lower is better):

Setting Held-out loss
Untrained (random) projector 8.7249
Trained stage-1 projector 2.4523
Trained projector, captions paired with the wrong images 5.4332

Held-out loss matches the final training loss (no over-fitting), and pairing captions with the wrong images roughly doubles the loss — the language model is relying on the visual content, not guessing generic captions.

Stage 1 vs stage 2 vs stage 2b

Metric Stage 1 Stage 2 Stage 2b (recommended)
Held-out LLaVA-Instruct loss (500 unseen conversations, lower is better) 2.2521 1.1271 1.1602

POPE (object hallucination: yes/no questions about COCO val2014 images, scored from the Yes/No token probabilities; the benchmark is balanced 50% yes / 50% no, so a yes-ratio near 0.5 is ideal):

POPE split Stage 2: accuracy / F1 / yes-ratio Stage 2b: accuracy / F1 / yes-ratio
random 0.543 / 0.686 / 0.95 0.822 / 0.832 / 0.56
popular 0.522 / 0.676 / 0.98 0.776 / 0.797 / 0.60
adversarial 0.514 / 0.672 / 0.98 0.752 / 0.780 / 0.63

What stage 2b fixed

Stage 2 answered "yes" to 97–98% of POPE questions. It confirmed almost every object it was asked about, present or not, so accuracy sat at chance (51–54%) even though its descriptions were good. The cause is the training data: LLaVA-Instruct-150K is made of long GPT-4-written conversations and contains almost no question whose answer is "no".

Stage 2b continues stage 2 for one epoch on a 150k mix drawn from the public LLaVA-1.5 training set: 110k short-answer VQA conversations (57,804 "yes" vs 59,443 "no" turns) plus 40k LLaVA-Instruct conversations replayed so detailed description is not lost. Every POPE test image and every held-out conversation was removed from this mix before training, so the numbers above are not contaminated; region/bounding-box tasks were dropped as unsupported. Result: the yes-ratio fell to 0.56–0.63 and accuracy rose by 24–28 points.

The held-out conversation loss rose slightly (1.1271 → 1.1602) — a small regression from the changed data mix. Description quality looks preserved in the examples below, but we have not measured that quantitatively.

Averaged over the three splits, stage 2b scores F1 0.80. We deliberately do not put this head-to-head with the number quoted for LLaVA-1.5-7B: that paper reports a single POPE figure without stating how it aggregates the splits, it covers the full POPE benchmark (COCO + A-OKVQA + GQA) while we evaluate the COCO split only, and independent reproductions of LLaVA-1.5 on COCO with greedy decoding land anywhere from F1 ≈ 0.78 to ≈ 0.87 depending on setup. Our setup also differs in scale (150k vs 665k samples), resolution (224 px vs 336 px) and base model, so treat these as our own baseline, not a ranking.

Recall (0.88) still exceeds precision (0.70) on the adversarial split, so a mild "yes" lean remains on the hardest confusable pairs.

Examples — stage 1 captions on unseen images

Reference caption AtlasVision stage 1
cute homecoming prom dress with lace top and satin skirt a - line sweetheart sweetheart lace tulle prom dress
a pair of silicone bracelets with a four codes message a pair of silicone bracelets with the word code on them
a set of fingerprint icons in white on a black background the iphone 6s and iphone 6s plus are shown in a black and white image
a business woman drawing modern concept of a website creation a man writing the word modern web design on a whiteboard
the beach in el nido national park, puerto puerto the limestone cliffs and limestone islands in the background
three pieces of paper with the words, democratic decentified dp controlled centralized ccp a diagram showing the different types of democracy

Examples — "Describe this image in detail." (stage 2b, unseen COCO val2014)

COCO_val2014_000000310196

  • Stage 1: a skier in the snow on a mountain
  • Stage 2b: The image features a snow-covered slope with a skier in the center, wearing red and black clothing. The skier is skillfully navigating down the hill, making their way through the snow.

In addition to the main skier, there are several other people scattered across the scene, some of them closer to the top of the slope while others are further away or near the bottom. These individuals may also be skiing or enjoying the snowy landscape.

COCO_val2014_000000210789

  • Stage 1: a woman and her child in the rain
  • Stage 2b: The image features a woman and a young girl standing in the rain, both holding umbrellas to protect themselves from getting wet. The little girl is wearing a pink dress, while the woman appears to be her mother. They are positioned close together, with the child on the left side of the frame and the woman on the right.

In addition to the two main subjects, there are several other people visible in the background, some of whom may also be holding umbrellas or trying to stay dry during the rainy weather. A car can be seen parked further back in the scene, likely waiting for its occupants

COCO_val2014_000000429109

  • Stage 1: a bus and several other vehicles parked in front of a building
  • Stage 2b: The image features a busy street with several buses and cars parked or driving along the road. There are three buses in total, one on the left side of the scene, another near the center, and the third bus further to the right. A car is also visible on the left side of the image.

In addition to the vehicles, there are multiple people walking around the area. Some individuals can be seen closer to the buses, while others are scattered throughout the scene. The presence of both vehicles and pedestrians suggests that this could be a popular transportation hub or a bustling city street.

Questions in Nigerian languages (stage 2b)

In our tests, stage 2b answered the Igbo and Yoruba questions in those languages without any translation step, where stage 2 always replied in English. This is one example per language, both short and containing English loanwords ("skier", "snow"), so treat it as an encouraging signal rather than a measured result — a proper multilingual evaluation is still to come. Hausa fell back to English. For reliable answers in all three languages, use the lang= cascade described below.

Language Question Answer
Igbo Kedu ihe dị na foto a? Na foto a, e nwere skier na snow.
Yoruba Kí ni ó wà nínú àwòrán yìí? Nínú àwòrán yìí, ó wà skier tí ń ski lórí òkè.
Hausa Me ke cikin wannan hoton? In the image, there is a skier wearing red and black clothing who is skiing down a snow-covered slope.

Text-only check (no image)

Does the stage-2 LoRA damage N-ATLaS's original text abilities? Same Igbo question, greedy decoding:

Q: Kedu ihe bụ positron? Kọwaa ya n'asụsụ Igbo.

  • Base N-ATLaS: Positron bụ eletrọn dị mma, ma ọ bụ antiparticle nke eletrọn. Ọ na-emegide eletrọn nkịtị site n'ịnwe mpempe ojii nwere otu ụgwọ ahụ mana ọdịiche dị na ntụziaka nke ụgwọ ha. Positrons na-ekere òkè dị mkpa na physics nuklia, ebe a na-emepụta ha mgbe ụfọdụ n'oge us
  • With stage-2 LoRA: Positron bụ akụkụ subatomic dị mma, nke a na-akpọkwa antiparticle nke electron. A na-eji okwu "positron" mee ka ọ pụta ìhè site n'aka physicist Paul Dirac na 1928. Positrons nwere njirimara yiri nke electrons, gụnyere ibu, ụgwọ, na spin, mana ha na-emegharịrị n'ihe gbasara mass na momentum. Mgbe positron na-ej

Both answer fluently in Igbo, so the adapter keeps N-ATLaS's text abilities intact (responses are cut at 120 tokens).

Answers in Igbo, Yoruba and Hausa (cascade)

The image training data is English, so AtlasVision reasons about images in English. chat.py can route any of the four languages through one loaded model, switching the LoRA adapter off to recover the original N-ATLaS for the translation steps:

  1. Question → English with plain N-ATLaS (adapter off).
  2. Look and answer in English with AtlasVision (adapter on, image tokens in).
  3. Answer → target language with plain N-ATLaS (adapter off).
from chat import AtlasVision, load_image

model = AtlasVision(stage="2b")
image = load_image("photo.jpg")

r = model.chat(image, "Mutane nawa ne a cikin hoton?", lang="ha")
print(r["answer"])            # Hausa
print(r["english_answer"])    # English draft, so you can check what was translated

print(model.describe(image, lang="ig"))        # Igbo description
print(model.translate("Good morning", "yo"))   # plain N-ATLaS translation, no image

From the command line: python chat.py --image photo.jpg --question "Kedu ihe di na foto a?" --lang ig.

Costs and caveats: each answer takes roughly 2–3× longer than English-only; translation quality is N-ATLaS's own; and any error in the English answer is carried into the translation, which is why english_answer is always returned. Pass translate_question=False to skip step 1 when the question is already short and simple.

Limitations

  • Hallucination. Long descriptions are fluent but often add plausible details that are not in the image (exact counts, extra people, a bicycle at the edge of frame). Treat counts and small objects as unreliable.
  • Residual yes-bias. Stage 2b largely fixed stage 2's yes-bias, but recall still exceeds precision on POPE's adversarial split, so yes/no answers about easily-confused objects lean positive. Stage 2 weights are kept for reproducibility only — do not use them for yes/no verification.
  • English-only visual training. All image–text training data is English. Answers to Hausa, Igbo and Yoruba questions come from N-ATLaS's own multilingual ability and are noticeably less reliable; they may switch to English.
  • Low resolution. Images are resized to 224×224, so small text, fine details and dense documents are hard.
  • One image per conversation, no video, no bounding boxes or grounding.
  • Data biases. Web captions (LAION/CC/SBU) and COCO carry Western-centric content; performance on Nigerian scenes, people and text has not been measured and is likely weaker.
  • Not for high-stakes use (medical, legal, identity, surveillance or safety-critical decisions).

Licenses and terms

This repository contains weights trained on top of other people's models and data. Using it means complying with all of:

  • N-ATLaS: Awarri Open-Source Research and Innovation License (attribution required; caps deployments at 1,000 active end-users per 30 days) and the Llama 3 Community License it inherits. Access is gated — accept the conditions on the model page.
  • SigLIP2: Apache-2.0.
  • LLaVA-Instruct-150K (stage 2 data): CC BY-NC 4.0, generated with GPT-4 and subject to OpenAI's terms — stage-2 and stage-2b weights are for non-commercial research use.
  • LLaVA-1.5 mix / VQAv2, OK-VQA, A-OKVQA (stage 2b data): drawn from llava_v1_5_mix665k, which inherits the CC BY-NC 4.0 terms above. The underlying short-answer sets carry their own licenses (VQAv2 annotations are CC BY 4.0 over COCO images; OK-VQA and A-OKVQA are research datasets released by their authors) — check each dataset's page before any redistribution or commercial use.
  • LLaVA-Pretrain (stage 1 data): captions/images from LAION, Conceptual Captions and SBU under their respective terms.
  • COCO images: Flickr images under their individual Creative Commons licenses.

Acknowledgements

  • N-ATLaS by Awarri Technologies and the Federal Ministry of Communications, Innovation and Digital Economy (Nigeria), as part of the Nigerian Languages AI Initiative.
  • SigLIP2 by Google; LLaVA data and recipe by Liu et al.; POPE by Li et al.
  • Trained on an H100 from San Francisco Compute (Autoresearch).

Citation

@misc{atlasvision2026,
  title  = {AtlasVision: a vision-language extension of N-ATLaS},
  author = {Modularcomputing},
  year   = {2026},
  url    = {https://huggingface.co/Modularcomputing/AtlasVision}
}
@inproceedings{liu2023llava,
  title     = {Visual Instruction Tuning},
  author    = {Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae},
  booktitle = {NeurIPS},
  year      = {2023}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for Modularcomputing/AtlasVision

Base model

NCAIR1/N-ATLaS
Adapter
(3)
this model

Datasets used to train Modularcomputing/AtlasVision