A Strong Baseline for Evaluating Vision Encoders
in Multimodal Large Language Models
Yilin Yang1,* · Jun-Tao Tang2,* · Kengyi Wang3 · Siyuan Su3 · Gaoyong Luo4 · Mingda Chen1,†
1School of Artificial Intelligence, Shanghai Jiao Tong University
2Nanjing University ·
3Fudan University ·
4Independent Researcher
*Equal contribution. †Corresponding author.
Overview
This repository hosts the downstream MLLM evaluation checkpoints accompanying A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models.
The study evaluates 70 vision encoders with three main language backbones. The released checkpoints are organized by language backbone, vision encoder and training stage; links to the original frozen encoder weights appear alongside the corresponding MLLM weights below.
Model Zoo
This Model Zoo covers the 70 visual tokenizers used in our paper: 43 language-supervised, 22 self-supervised, and 5 discrete tokenizers. The current release provides both pretraining and finetuning checkpoints for all 65 continuous tokenizers with each of the three main language backbones, plus 2 Qwen3-1.7B-Base runs. The 5 discrete tokenizers are listed separately with their checkpoint availability.
Training data and checkpoint types
- Pretrain Data — LCS-558K: the image–text alignment dataset from LLaVA-Pretrain, using
blip_laion_cc_sbu_558k.json. This dataset is used for MLLM projector training. - Finetuning Data — LLaVA-v1.5 mix665k (filtered): the LLaVA-v1.5 instruction mixture, using
llava_v1_5_mix665k_drop_ge8kchars.json. The training configuration removes 395 examples with at least 8,000 characters of conversation text. - Download:
projectordownloads the pretrainingmm_projector.bin;finetunedopens the finetuned checkpoint directory, including model weights, configuration, and language-tokenizer files. The frozen vision encoder must be supplied separately using the matching architecture and weights recorded inconfig.json. Pretraining projector weights alone are not instruction-tuned MLLMs.
Vision encoder downloads
The Encoder weights column links to the original upstream vision encoder or visual tokenizer. Each encoder name links to its model or project page. Use these frozen encoder weights together with the corresponding Stage 1 projector or Stage 2 MLLM checkpoint, preserving the architecture, feature layer and preprocessing specified by that run's config.json.
- For Hugging Face models, retain the associated configuration and processor files. Web-SSL MAE 3B has sharded weights; its link opens the complete model repository.
- DINOv3 requires acceptance of the upstream model's terms and authentication. The RAEv2 DINOv3-L (k=7) entry uses the same DINOv3-L/16 backbone with the project's k=7 intermediate-layer readout; its download is the backbone used to construct that representation.
- TokLIP-S/L require both the encoder checkpoint and the linked VQ checkpoint. For UniAR and VILA-U, retain the config files alongside the weights in the respective
bsq_encoder/andvision_tower/directories.
Qwen2.5-1.5B-Instruct
Base LLM: Qwen/Qwen2.5-1.5B-Instruct. 65 tokenizers, each with pretraining and finetuning checkpoints.
Qwen3-1.7B
Base LLM: Qwen/Qwen3-1.7B. 65 tokenizers, each with pretraining and finetuning checkpoints.
SmolLM2-1.7B-Instruct
Base LLM: HuggingFaceTB/SmolLM2-1.7B-Instruct. 65 tokenizers, each with pretraining and finetuning checkpoints.
Discrete tokenizers in the paper
| Vision Tokenizer | Encoder weights | MLLM checkpoint availability |
|---|---|---|
| UniTok-Attn (256) | encoder | Not released in this repository |
| UniAR-BSQ | encoder | Not released in this repository |
| TokLIP-L (384) | encoder · VQ | Not released in this repository |
| TokLIP-S (256) | encoder · VQ | Not released in this repository |
| VILA-U (256) | encoder | Not released in this repository |
Download
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="336labs/VisionEncoder-to-MLLM-ModelZoo",
allow_patterns=["continuous/qwen3base/**", "FILES.json", "README.md"],
local_dir="VisionEncoder-to-MLLM-ModelZoo",
)
Choose the needed language-backbone group and vision-encoder run. Each run retains its original weights, configuration, tokenizer files, and available training metadata. Check the configuration in that subdirectory for the corresponding architecture.
Citation
If you use RAVEL or this model zoo, please cite:
@misc{yang2026strong,
title={A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models},
author={Yang, Yilin and Tang, Jun-Tao and Wang, Kengyi and Su, Siyuan and Luo, Gaoyong and Chen, Mingda},
year={2026},
eprint={2610.05413},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.05413}
}