Vision Encoder MLLM Evaluation

Downstream evaluation checkpoints for comparing vision encoders in multimodal large language models (MLLMs).

The evaluation target is the vision encoder. Runs use different language backbones and are organized by backbone, vision encoder, and training stage. These are trained multimodal evaluation artifacts, rather than a standalone language-model release.

Contents

Language-backbone group Files Weight files Size (GiB)
qwen3base 35 8 10.46
smollm2 974 215 285.01
qwen3 978 214 305.64
qwen25 1173 268 326.19

Model Zoo

This Model Zoo covers the 70 visual tokenizers used in our paper: 43 language-supervised, 22 self-supervised, and 5 discrete tokenizers. The current release provides both pretraining and finetuning checkpoints for all 65 continuous tokenizers with each of the three main language backbones, plus 2 Qwen3-1.7B-Base runs. The 5 discrete tokenizers are listed separately with their checkpoint availability.

Training data and checkpoint types

  • Pretrain Data — LCS-558K: the image–text alignment dataset from LLaVA-Pretrain, using blip_laion_cc_sbu_558k.json. This dataset is used for MLLM projector training.
  • Finetuning Data — LLaVA-v1.5 mix665k (filtered): the LLaVA-v1.5 instruction mixture, using llava_v1_5_mix665k_drop_ge8kchars.json. The training configuration removes 395 examples with at least 8,000 characters of conversation text.
  • Download: projector downloads the pretraining mm_projector.bin; finetuned opens the finetuned checkpoint directory, including model weights, configuration, and language-tokenizer files. The frozen vision encoder must be supplied separately using the matching architecture and weights recorded in config.json. Pretraining projector weights alone are not instruction-tuned MLLMs.

Qwen2.5-1.5B-Instruct

Base LLM: Qwen/Qwen2.5-1.5B-Instruct. 65 tokenizers, each with pretraining and finetuning checkpoints.

Base LLM Vision Encoder / Tokenizer Stage 1 Pretrained weights Stage 2 Finetuned weights
Qwen2.5-1.5B-Instruct PE-Core-G/14 (448) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-G/14 (378) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-G/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-2.5B ViT-G/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-G/16 (384) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-G/16 (256) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-H/14 (378) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-2.5B ViT-H/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-v1.2 ViT-H/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct OpenAI CLIP ViT-L/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-L/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-2.5B ViT-L/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-400M ViT-L/14 (224) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 So400m/l14 (384) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 So400m/l14 (224) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 So400m/l16 (512) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 So400m/l16 (384) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 So400m/l16 (256) projector finetuned
Qwen2.5-1.5B-Instruct PE-Lang-L/14 (448) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-L/16 (512) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-L/16 (384) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-L/16 (256) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-M/16 (384) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-M/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2-mT5 ViT-M/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-B/16 (384) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-B/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-B/32 (384) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-B/32 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2-mT5 ViT-B/32 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-2.5B ViT-B/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-2.5B ViT-B/32 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-400M ViT-B/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP-400M ViT-B/32 (224) projector finetuned
Qwen2.5-1.5B-Instruct PE-Core-B/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-B/16 (512) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-B/16 (384) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-B/16 (256) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-B/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct SigLIP2 ViT-B/32 (256) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-S/16 (384) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2 ViT-S/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct MetaCLIP 2-mT5 ViT-S/16 (224) projector finetuned
Qwen2.5-1.5B-Instruct Web-SSL MAE 3B (224) projector finetuned
Qwen2.5-1.5B-Instruct DINOv2 ViT-G/14 projector finetuned
Qwen2.5-1.5B-Instruct Web-SSL DINO 1B (224) projector finetuned
Qwen2.5-1.5B-Instruct Web-SSL MAE 1B (224) projector finetuned
Qwen2.5-1.5B-Instruct Pixio ViT-H/16 projector finetuned
Qwen2.5-1.5B-Instruct I-JEPA ViT-H/14 projector finetuned
Qwen2.5-1.5B-Instruct DINOv2 ViT-L/14 projector finetuned
Qwen2.5-1.5B-Instruct DINOv3 ViT-L/16 projector finetuned
Qwen2.5-1.5B-Instruct Pixio ViT-L/16 projector finetuned
Qwen2.5-1.5B-Instruct RAEv2 DINOv3-L (k=7) projector finetuned
Qwen2.5-1.5B-Instruct Web-SSL MAE 300M (224) projector finetuned
Qwen2.5-1.5B-Instruct EUPE ConvNeXt-B projector finetuned
Qwen2.5-1.5B-Instruct DINOv2 ViT-B/14 projector finetuned
Qwen2.5-1.5B-Instruct DINO ViT-B/8 projector finetuned
Qwen2.5-1.5B-Instruct DINO ViT-B/16 projector finetuned
Qwen2.5-1.5B-Instruct EUPE ViT-B projector finetuned
Qwen2.5-1.5B-Instruct Pixio ViT-B/16 projector finetuned
Qwen2.5-1.5B-Instruct DINOv2 ViT-S/14 projector finetuned
Qwen2.5-1.5B-Instruct DINO ViT-S/16 projector finetuned
Qwen2.5-1.5B-Instruct EUPE ViT-S projector finetuned
Qwen2.5-1.5B-Instruct DINO ViT-S/8 projector finetuned
Qwen2.5-1.5B-Instruct EUPE ViT-T projector finetuned

Qwen3-1.7B

Base LLM: Qwen/Qwen3-1.7B. 65 tokenizers, each with pretraining and finetuning checkpoints.

Base LLM Vision Encoder / Tokenizer Stage 1 Pretrained weights Stage 2 Finetuned weights
Qwen3-1.7B PE-Core-G/14 (448) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-G/14 (378) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-G/14 (224) projector finetuned
Qwen3-1.7B MetaCLIP-2.5B ViT-G/14 (224) projector finetuned
Qwen3-1.7B SigLIP2 ViT-G/16 (384) projector finetuned
Qwen3-1.7B SigLIP2 ViT-G/16 (256) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-H/14 (378) projector finetuned
Qwen3-1.7B MetaCLIP-2.5B ViT-H/14 (224) projector finetuned
Qwen3-1.7B MetaCLIP-v1.2 ViT-H/14 (224) projector finetuned
Qwen3-1.7B OpenAI CLIP ViT-L/14 (224) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-L/14 (224) projector finetuned
Qwen3-1.7B MetaCLIP-2.5B ViT-L/14 (224) projector finetuned
Qwen3-1.7B MetaCLIP-400M ViT-L/14 (224) projector finetuned
Qwen3-1.7B SigLIP2 So400m/l14 (384) projector finetuned
Qwen3-1.7B SigLIP2 So400m/l14 (224) projector finetuned
Qwen3-1.7B SigLIP2 So400m/l16 (512) projector finetuned
Qwen3-1.7B SigLIP2 So400m/l16 (384) projector finetuned
Qwen3-1.7B SigLIP2 So400m/l16 (256) projector finetuned
Qwen3-1.7B PE-Lang-L/14 (448) projector finetuned
Qwen3-1.7B SigLIP2 ViT-L/16 (512) projector finetuned
Qwen3-1.7B SigLIP2 ViT-L/16 (384) projector finetuned
Qwen3-1.7B SigLIP2 ViT-L/16 (256) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-M/16 (384) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-M/16 (224) projector finetuned
Qwen3-1.7B MetaCLIP 2-mT5 ViT-M/16 (224) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-B/16 (384) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-B/16 (224) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-B/32 (384) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-B/32 (224) projector finetuned
Qwen3-1.7B MetaCLIP 2-mT5 ViT-B/32 (224) projector finetuned
Qwen3-1.7B MetaCLIP-2.5B ViT-B/16 (224) projector finetuned
Qwen3-1.7B MetaCLIP-2.5B ViT-B/32 (224) projector finetuned
Qwen3-1.7B MetaCLIP-400M ViT-B/16 (224) projector finetuned
Qwen3-1.7B MetaCLIP-400M ViT-B/32 (224) projector finetuned
Qwen3-1.7B PE-Core-B/16 (224) projector finetuned
Qwen3-1.7B SigLIP2 ViT-B/16 (512) projector finetuned
Qwen3-1.7B SigLIP2 ViT-B/16 (384) projector finetuned
Qwen3-1.7B SigLIP2 ViT-B/16 (256) projector finetuned
Qwen3-1.7B SigLIP2 ViT-B/16 (224) projector finetuned
Qwen3-1.7B SigLIP2 ViT-B/32 (256) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-S/16 (384) projector finetuned
Qwen3-1.7B MetaCLIP 2 ViT-S/16 (224) projector finetuned
Qwen3-1.7B MetaCLIP 2-mT5 ViT-S/16 (224) projector finetuned
Qwen3-1.7B Web-SSL MAE 3B (224) projector finetuned
Qwen3-1.7B DINOv2 ViT-G/14 projector finetuned
Qwen3-1.7B Web-SSL DINO 1B (224) projector finetuned
Qwen3-1.7B Web-SSL MAE 1B (224) projector finetuned
Qwen3-1.7B Pixio ViT-H/16 projector finetuned
Qwen3-1.7B I-JEPA ViT-H/14 projector finetuned
Qwen3-1.7B DINOv2 ViT-L/14 projector finetuned
Qwen3-1.7B DINOv3 ViT-L/16 projector finetuned
Qwen3-1.7B Pixio ViT-L/16 projector finetuned
Qwen3-1.7B RAEv2 DINOv3-L (k=7) projector finetuned
Qwen3-1.7B Web-SSL MAE 300M (224) projector finetuned
Qwen3-1.7B EUPE ConvNeXt-B projector finetuned
Qwen3-1.7B DINOv2 ViT-B/14 projector finetuned
Qwen3-1.7B DINO ViT-B/8 projector finetuned
Qwen3-1.7B DINO ViT-B/16 projector finetuned
Qwen3-1.7B EUPE ViT-B projector finetuned
Qwen3-1.7B Pixio ViT-B/16 projector finetuned
Qwen3-1.7B DINOv2 ViT-S/14 projector finetuned
Qwen3-1.7B DINO ViT-S/16 projector finetuned
Qwen3-1.7B EUPE ViT-S projector finetuned
Qwen3-1.7B DINO ViT-S/8 projector finetuned
Qwen3-1.7B EUPE ViT-T projector finetuned

SmolLM2-1.7B-Instruct

Base LLM: HuggingFaceTB/SmolLM2-1.7B-Instruct. 65 tokenizers, each with pretraining and finetuning checkpoints.

Base LLM Vision Encoder / Tokenizer Stage 1 Pretrained weights Stage 2 Finetuned weights
SmolLM2-1.7B-Instruct PE-Core-G/14 (448) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-G/14 (378) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-G/14 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-2.5B ViT-G/14 (224) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-G/16 (384) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-G/16 (256) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-H/14 (378) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-2.5B ViT-H/14 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-v1.2 ViT-H/14 (224) projector finetuned
SmolLM2-1.7B-Instruct OpenAI CLIP ViT-L/14 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-L/14 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-2.5B ViT-L/14 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-400M ViT-L/14 (224) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 So400m/l14 (384) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 So400m/l14 (224) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 So400m/l16 (512) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 So400m/l16 (384) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 So400m/l16 (256) projector finetuned
SmolLM2-1.7B-Instruct PE-Lang-L/14 (448) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-L/16 (512) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-L/16 (384) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-L/16 (256) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-M/16 (384) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-M/16 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2-mT5 ViT-M/16 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-B/16 (384) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-B/16 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-B/32 (384) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-B/32 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2-mT5 ViT-B/32 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-2.5B ViT-B/16 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-2.5B ViT-B/32 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-400M ViT-B/16 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP-400M ViT-B/32 (224) projector finetuned
SmolLM2-1.7B-Instruct PE-Core-B/16 (224) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-B/16 (512) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-B/16 (384) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-B/16 (256) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-B/16 (224) projector finetuned
SmolLM2-1.7B-Instruct SigLIP2 ViT-B/32 (256) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-S/16 (384) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2 ViT-S/16 (224) projector finetuned
SmolLM2-1.7B-Instruct MetaCLIP 2-mT5 ViT-S/16 (224) projector finetuned
SmolLM2-1.7B-Instruct Web-SSL MAE 3B (224) projector finetuned
SmolLM2-1.7B-Instruct DINOv2 ViT-G/14 projector finetuned
SmolLM2-1.7B-Instruct Web-SSL DINO 1B (224) projector finetuned
SmolLM2-1.7B-Instruct Web-SSL MAE 1B (224) projector finetuned
SmolLM2-1.7B-Instruct Pixio ViT-H/16 projector finetuned
SmolLM2-1.7B-Instruct I-JEPA ViT-H/14 projector finetuned
SmolLM2-1.7B-Instruct DINOv2 ViT-L/14 projector finetuned
SmolLM2-1.7B-Instruct DINOv3 ViT-L/16 projector finetuned
SmolLM2-1.7B-Instruct Pixio ViT-L/16 projector finetuned
SmolLM2-1.7B-Instruct RAEv2 DINOv3-L (k=7) projector finetuned
SmolLM2-1.7B-Instruct Web-SSL MAE 300M (224) projector finetuned
SmolLM2-1.7B-Instruct EUPE ConvNeXt-B projector finetuned
SmolLM2-1.7B-Instruct DINOv2 ViT-B/14 projector finetuned
SmolLM2-1.7B-Instruct DINO ViT-B/8 projector finetuned
SmolLM2-1.7B-Instruct DINO ViT-B/16 projector finetuned
SmolLM2-1.7B-Instruct EUPE ViT-B projector finetuned
SmolLM2-1.7B-Instruct Pixio ViT-B/16 projector finetuned
SmolLM2-1.7B-Instruct DINOv2 ViT-S/14 projector finetuned
SmolLM2-1.7B-Instruct DINO ViT-S/16 projector finetuned
SmolLM2-1.7B-Instruct EUPE ViT-S projector finetuned
SmolLM2-1.7B-Instruct DINO ViT-S/8 projector finetuned
SmolLM2-1.7B-Instruct EUPE ViT-T projector finetuned

Discrete tokenizers in the paper

Vision Tokenizer Checkpoint Availability
UniTok-Attn (256) Not released in this repository
UniAR-BSQ Not released in this repository
TokLIP-L (384) Not released in this repository
TokLIP-S (256) Not released in this repository
VILA-U (256) Not released in this repository

Download

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="336labs/VisionEncoder-MLLM-Eval",
    allow_patterns=["continuous/qwen3base/**", "FILES.json", "README.md"],
    local_dir="VisionEncoder-MLLM-Eval",
)

Choose the needed language-backbone group and vision-encoder run. Each run retains its original weights, configuration, tokenizer files, and available training metadata. Check the configuration in that subdirectory for the corresponding architecture.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support