refcaptioner / README.md
multimodalart's picture
multimodalart HF Staff
Upload folder using huggingface_hub
f3fc14a verified
|
Raw History Blame Contribute Delete
2.06 kB
---
title: RefCaptioner
emoji: 🎬
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.30.0
app_file: app.py
short_description: Multi-reference image-grounded video captioning
python_version: "3.10"
startup_duration_timeout: 1h
---
# RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Demo of [RefCaptioner](https://huggingface.co/TengfeiLiuCoder/RefCaptioner) (paper: [2607.28509](https://arxiv.org/abs/2607.28509)), an 8B vision-language model fine-tuned from Qwen3-VL-8B-Instruct for **multi-reference image-grounded video captioning**.
Given a video and an ordered list of reference images, the model writes a fluent English video caption and places `<Image_n>` tags immediately after the visual phrases that each reference grounds. References that are distractors (not visible in the video) are omitted.
**Inputs**
- πŸŽ₯ one video
- πŸ–ΌοΈ 1–6 reference images β€” upload order defines the tag mapping (`<Image_1>`, `<Image_2>`, …)
- optional: leave a reference slot empty to see the model reject a distractor
**Output**
- one grounded caption paragraph with phrase-level `<Image_n>` bindings (tags highlighted)
The demo follows the authors' released inference recipe exactly (prompt `Prompt_1.0`, greedy decoding, 2 FPS sampling, up to 20 frames, 602112 max pixels per image/frame, thinking disabled).
## Links
- πŸ“„ Paper: [RefCaptioner: Multi-Reference Image-Grounded Video Captioning](https://arxiv.org/abs/2607.28509)
- πŸ€— Weights: [TengfeiLiuCoder/RefCaptioner](https://huggingface.co/TengfeiLiuCoder/RefCaptioner)
- πŸ’» Code: [pkucs-Ltf/RefCaptioner](https://github.com/pkucs-Ltf/RefCaptioner)
- πŸ“Š Benchmark: [TengfeiLiuCoder/MRVBench](https://huggingface.co/datasets/TengfeiLiuCoder/MRVBench)
Example media in this Space comes from the [`linoyts/repo-to-space-example-videos`](https://huggingface.co/datasets/linoyts/repo-to-space-example-videos) and [`linoyts/repo-to-space-example-inputs`](https://huggingface.co/datasets/linoyts/repo-to-space-example-inputs) datasets (free-to-use / CC0-style stock media).