CA-VTR โ Official Checkpoints
Trained checkpoints for CA-VTR: Cross-Attention Vision-Token Refiller โ From Passive Refill to Active Retrieval in Decoder-Free MLLM Segmentation.
๐ Paper link coming soon ยท ๐ป Code: github.com/jackwang0108/CA-VTR
CA-VTR replaces the MLP-based refilling decoder of decoder-free MLLM segmenters with a single
cross-attention layer โ high-resolution tile features as queries, the MLLM's semantic
output tokens as keys/values โ and produces the mask by a plain dot product with the
[SEG] embedding. 33% fewer decoder parameters than SELF1E, better results on all
benchmarks.
Checkpoints
| Directory | Training config | Backbone | Note |
|---|---|---|---|
2b-vanilla/ |
4-class mixture, sample rates 1,1,1,1 (~562k samples, 1 epoch, lr 1e-4) | InternVL3-2B | main result |
2b-seg/ |
sample rates 6,20,6,1 (~2.05M samples, 1 epoch, lr 1e-4) | InternVL3-2B | SEG recipe |
8b-vanilla/ |
โ | InternVL3-8B | coming soon |
8b-seg/ |
โ | InternVL3-8B | coming soon |
All runs: seed 42, LoRA r=128 (vision + LLM), tile(896) + fusion(concat) + ฮฒ=0 + 2D-RoPE.
Key results (single seed 42)
| Model | RefCOCO val cIoU | RefCOCOโบ val cIoU | RefCOCOg val cIoU | ReasonSeg val cIoU |
|---|---|---|---|---|
CA-VTR-2B (2b-vanilla) |
80.9 | 75.2 | 78.0 | 74.5 |
CA-VTR-SEG-2B (2b-seg) |
84.2 | 79.4 | 81.2 | 67.7 |
Full tables (incl. gRefCOCO and open-vocabulary segmentation): docs/perf-baseline.md in the
code repository.
Usage
The checkpoints use CA-VTR's custom architecture (InternVL3SELF1E) and are loaded by the
CA-VTR codebase โ not by plain transformers:
git clone https://github.com/jackwang0108/CA-VTR && cd CA-VTR
conda env create -f environment.yml && conda activate cavtr
# download a checkpoint and wire it up as an experiment
hf download JackWang0107/CA-VTR --include "2b-vanilla/*" --local-dir ckpts
mkdir -p runs/main_2b_vanilla
ln -s "$(pwd)/ckpts/2b-vanilla" runs/main_2b_vanilla/ckpt_model
# evaluate on all Referring / GRES / Reasoning splits
EXP_NAME=main_2b_vanilla bash scripts/eval_all.sh
Alternatively, point --model_weights <checkpoint dir> at a downloaded directory. Dataset
preparation: docs/data_preparation.md in the code repository.
License
MIT โ see the code repository.