Masking-KD (Qwen3-VL-2B-Thinking)

Checkpoint for Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation (NeurIPS 2026).

Qwen3-VL-2B-Thinking distilled from Qwen3-VL-8B-Thinking with Masking-KD.

Paper: https://arxiv.org/abs/2605.11651 | Code: https://github.com/Seonghoon-Yu/Masking-KD

Training

Student / Teacher Qwen3-VL-2B-Thinking / Qwen3-VL-8B-Thinking
Data SeonghoonYu/Masking-KD-Rollouts: 19,387 correct greedy teacher rollouts on ViRL39K
Schedule 2 epochs (76 steps), global batch size 512, learning rate 1e-6

Usage

The model was trained with the reasoning instruction below appended to each question.

from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained("SeonghoonYu/Masking-KD", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("SeonghoonYu/Masking-KD")

question = "Find x."
instruction = (
    "You first think through the reasoning process as an internal monologue, enclosed within <think> </think> tags. "
    "Then, provide your final answer enclosed within \\boxed{}."
)
messages = [{
    "role": "user",
    "content": [
        {"type": "image", "image": "path/to/image.png"},
        {"type": "text", "text": f"{question}\n\n{instruction}"},
    ],
}]

inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt"
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
print(processor.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

To reproduce the benchmark evaluation (Geo3K, MathVista, We-Math, MMK12, MathVerse, LogicVista, MMMU-Pro):

git clone https://github.com/Seonghoon-Yu/Masking-KD && cd Masking-KD
python evaluation/prepare_data.py
bash scripts/eval.sh SeonghoonYu/Masking-KD

Citation

@inproceedings{yu2026hide,
  title     = {Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation},
  author    = {Yu, Seonghoon and Nam, Dongjun and Lee, Byung-Kwan and Son, Jeany},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}
Downloads last month
4
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SeonghoonYu/Masking-KD

Finetuned
(24)
this model

Dataset used to train SeonghoonYu/Masking-KD

Paper for SeonghoonYu/Masking-KD