--- license: apache-2.0 base_model: Qwen/Qwen3-VL-2B-Thinking datasets: - SeonghoonYu/Masking-KD-Rollouts pipeline_tag: image-text-to-text library_name: transformers tags: - knowledge-distillation - multimodal-reasoning - qwen3_vl --- # Masking-KD (Qwen3-VL-2B-Thinking) Checkpoint for **Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation** (NeurIPS 2026). [Qwen3-VL-2B-Thinking](https://huggingface.co/Qwen/Qwen3-VL-2B-Thinking) distilled from [Qwen3-VL-8B-Thinking](https://huggingface.co/Qwen/Qwen3-VL-8B-Thinking) with Masking-KD. Paper: https://arxiv.org/abs/2605.11651 | Code: https://github.com/Seonghoon-Yu/Masking-KD ## Training | | | |---|---| | Student / Teacher | Qwen3-VL-2B-Thinking / Qwen3-VL-8B-Thinking | | Data | [SeonghoonYu/Masking-KD-Rollouts](https://huggingface.co/datasets/SeonghoonYu/Masking-KD-Rollouts): 19,387 correct greedy teacher rollouts on ViRL39K | | Schedule | 2 epochs (76 steps), global batch size 512, learning rate 1e-6 | ## Usage The model was trained with the reasoning instruction below appended to each question. ```python from transformers import AutoModelForImageTextToText, AutoProcessor model = AutoModelForImageTextToText.from_pretrained("SeonghoonYu/Masking-KD", dtype="auto", device_map="auto") processor = AutoProcessor.from_pretrained("SeonghoonYu/Masking-KD") question = "Find x." instruction = ( "You first think through the reasoning process as an internal monologue, enclosed within tags. " "Then, provide your final answer enclosed within \\boxed{}." ) messages = [{ "role": "user", "content": [ {"type": "image", "image": "path/to/image.png"}, {"type": "text", "text": f"{question}\n\n{instruction}"}, ], }] inputs = processor.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt" ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=4096, do_sample=False) print(processor.decode(outputs[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)) ``` To reproduce the benchmark evaluation (Geo3K, MathVista, We-Math, MMK12, MathVerse, LogicVista, MMMU-Pro): ```bash git clone https://github.com/Seonghoon-Yu/Masking-KD && cd Masking-KD python evaluation/prepare_data.py bash scripts/eval.sh SeonghoonYu/Masking-KD ``` ## Citation ```bibtex @inproceedings{yu2026hide, title = {Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation}, author = {Yu, Seonghoon and Nam, Dongjun and Lee, Byung-Kwan and Son, Jeany}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026} } ```