--- title: RefCaptioner emoji: 🎬 colorFrom: indigo colorTo: green sdk: gradio sdk_version: 6.30.0 app_file: app.py short_description: Multi-reference image-grounded video captioning python_version: "3.10" startup_duration_timeout: 1h --- # RefCaptioner: Multi-Reference Image-Grounded Video Captioning Demo of [RefCaptioner](https://huggingface.co/TengfeiLiuCoder/RefCaptioner) (paper: [2607.28509](https://arxiv.org/abs/2607.28509)), an 8B vision-language model fine-tuned from Qwen3-VL-8B-Instruct for **multi-reference image-grounded video captioning**. Given a video and an ordered list of reference images, the model writes a fluent English video caption and places `` tags immediately after the visual phrases that each reference grounds. References that are distractors (not visible in the video) are omitted. **Inputs** - 🎥 one video - 🖼️ 1–6 reference images — upload order defines the tag mapping (``, ``, …) - optional: leave a reference slot empty to see the model reject a distractor **Output** - one grounded caption paragraph with phrase-level `` bindings (tags highlighted) The demo follows the authors' released inference recipe exactly (prompt `Prompt_1.0`, greedy decoding, 2 FPS sampling, up to 20 frames, 602112 max pixels per image/frame, thinking disabled). ## Links - 📄 Paper: [RefCaptioner: Multi-Reference Image-Grounded Video Captioning](https://arxiv.org/abs/2607.28509) - 🤗 Weights: [TengfeiLiuCoder/RefCaptioner](https://huggingface.co/TengfeiLiuCoder/RefCaptioner) - 💻 Code: [pkucs-Ltf/RefCaptioner](https://github.com/pkucs-Ltf/RefCaptioner) - 📊 Benchmark: [TengfeiLiuCoder/MRVBench](https://huggingface.co/datasets/TengfeiLiuCoder/MRVBench) Example media in this Space comes from the [`linoyts/repo-to-space-example-videos`](https://huggingface.co/datasets/linoyts/repo-to-space-example-videos) and [`linoyts/repo-to-space-example-inputs`](https://huggingface.co/datasets/linoyts/repo-to-space-example-inputs) datasets (free-to-use / CC0-style stock media).