δ-Vision Embedding Adapter

A layer-wise visual-memory adapter for Qwen3-VL-4B-Instruct. Low-rank residual MLPs predict each layer's visual memory from the initial visual embeddings. The frozen language model projects that memory into visual keys and values for text retrieval, preserving all visual tokens.

This repository contains adapter weights. The corresponding Qwen3-VL-4B-Instruct base model and the δ-Vision implementation are required to load them; this is not a standalone Transformers model or a PEFT adapter.

Checkpoint

  • File: qwen_embedding_adapter_step2000.pt
  • Mode: embedding_adapter
  • Rank: 128 across all 36 layers
  • Trainable parameters: 23,592,960
  • Training: 2,000 steps on PixMo-Ask-Model-Anything, global batch size 32
  • Objective: answer-token KL distillation under teacher forcing, teacher top-k 1,024, temperature 2
  • Optimizer: AdamW, learning rate 5e-5
  • Training data: huaXiaKyrie/pixmo-ama-train

Download and evaluate

From an installed δ-Vision checkout:

hf download huaXiaKyrie/delta-Vision-Embedding-Adapter \
  qwen_embedding_adapter_step2000.pt \
  --local-dir artifacts/experiments/delta_vision/checkpoints

bash scripts/eval_benchmark.sh \
  --benchmark mmstar \
  --model-kind qwen \
  --model-path model/Qwen3-VL-4B-Instruct \
  --run-dir artifacts/experiments/delta_vision \
  --step 2000 --num-shards 1 --max-samples 1000 \
  --out-dir artifacts/eval/delta_vision/mmstar

Prepare the base model and benchmark inputs first, following the project README. The checkpoint's saved training arguments contain historical paths; pass the local model and data paths explicitly when running evaluation.

Recorded evaluation

These are the preserved September 11–12, 2026 evaluation results for this checkpoint. Each benchmark uses 1,000 examples except RealWorldQA, which uses all 765. Scores are on a 0–100 scale; MME and POPE use question accuracy, and VQAv2 uses its soft score. The average is the unweighted mean of the nine scores.

Benchmark Score
mmstar 54.40
realworldqa 63.40
gqa 55.40
mmb 84.10
mmb-cn 82.70
mme 80.20
pope 86.70
sqa 81.20
vqav2 77.98
Average 74.01

Citation

If you find our work helpful, please kindly cite:

@misc{lei2026justmlpsefficientvisual,
      title={Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models},
      author={Jingdi Lei and Junxian Li and Di Zhang and Zhanqiu Zhang and Yiwen Guo and Soujanya Poria},
      year={2026},
      eprint={2609.34972},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.34972},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for huaXiaKyrie/delta-Vision-Embedding-Adapter

Adapter
(210)
this model

Dataset used to train huaXiaKyrie/delta-Vision-Embedding-Adapter

Paper for huaXiaKyrie/delta-Vision-Embedding-Adapter