δ-Vision Embedding Adapter
A layer-wise visual-memory adapter for Qwen3-VL-4B-Instruct. Low-rank residual MLPs predict each layer's visual memory from the initial visual embeddings. The frozen language model projects that memory into visual keys and values for text retrieval, preserving all visual tokens.
This repository contains adapter weights. The corresponding Qwen3-VL-4B-Instruct base model and the δ-Vision implementation are required to load them; this is not a standalone Transformers model or a PEFT adapter.
Checkpoint
- File:
qwen_embedding_adapter_step2000.pt - Mode:
embedding_adapter - Rank: 128 across all 36 layers
- Trainable parameters: 23,592,960
- Training: 2,000 steps on PixMo-Ask-Model-Anything, global batch size 32
- Objective: answer-token KL distillation under teacher forcing, teacher top-k 1,024, temperature 2
- Optimizer: AdamW, learning rate
5e-5 - Training data: huaXiaKyrie/pixmo-ama-train
Download and evaluate
From an installed δ-Vision checkout:
hf download huaXiaKyrie/delta-Vision-Embedding-Adapter \
qwen_embedding_adapter_step2000.pt \
--local-dir artifacts/experiments/delta_vision/checkpoints
bash scripts/eval_benchmark.sh \
--benchmark mmstar \
--model-kind qwen \
--model-path model/Qwen3-VL-4B-Instruct \
--run-dir artifacts/experiments/delta_vision \
--step 2000 --num-shards 1 --max-samples 1000 \
--out-dir artifacts/eval/delta_vision/mmstar
Prepare the base model and benchmark inputs first, following the project README. The checkpoint's saved training arguments contain historical paths; pass the local model and data paths explicitly when running evaluation.
Recorded evaluation
These are the preserved September 11–12, 2026 evaluation results for this checkpoint. Each benchmark uses 1,000 examples except RealWorldQA, which uses all 765. Scores are on a 0–100 scale; MME and POPE use question accuracy, and VQAv2 uses its soft score. The average is the unweighted mean of the nine scores.
| Benchmark | Score |
|---|---|
| mmstar | 54.40 |
| realworldqa | 63.40 |
| gqa | 55.40 |
| mmb | 84.10 |
| mmb-cn | 82.70 |
| mme | 80.20 |
| pope | 86.70 |
| sqa | 81.20 |
| vqav2 | 77.98 |
| Average | 74.01 |
Citation
If you find our work helpful, please kindly cite:
@misc{lei2026justmlpsefficientvisual,
title={Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models},
author={Jingdi Lei and Junxian Li and Di Zhang and Zhanqiu Zhang and Yiwen Guo and Soujanya Poria},
year={2026},
eprint={2609.34972},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.34972},
}
Model tree for huaXiaKyrie/delta-Vision-Embedding-Adapter
Base model
Qwen/Qwen3-VL-4B-Instruct