Spaces:
Running on Zero
Download README.md from hugging-apps/refcaptioner: direct link, hf CLI and curl.
- Browser
- Download file 2.06 kB
-
https://huggingface.co/spaces/hugging-apps/refcaptioner/resolve/main/README.md
- Command line
-
hf download hf://spaces/hugging-apps/refcaptioner/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/refcaptioner/resolve/main/README.md
title: RefCaptioner
emoji: π¬
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.30.0
app_file: app.py
short_description: Multi-reference image-grounded video captioning
python_version: '3.10'
startup_duration_timeout: 1h
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Demo of RefCaptioner (paper: 2607.28509), an 8B vision-language model fine-tuned from Qwen3-VL-8B-Instruct for multi-reference image-grounded video captioning.
Given a video and an ordered list of reference images, the model writes a fluent English video caption and places <Image_n> tags immediately after the visual phrases that each reference grounds. References that are distractors (not visible in the video) are omitted.
Inputs
- π₯ one video
- πΌοΈ 1β6 reference images β upload order defines the tag mapping (
<Image_1>,<Image_2>, β¦) - optional: leave a reference slot empty to see the model reject a distractor
Output
- one grounded caption paragraph with phrase-level
<Image_n>bindings (tags highlighted)
The demo follows the authors' released inference recipe exactly (prompt Prompt_1.0, greedy decoding, 2 FPS sampling, up to 20 frames, 602112 max pixels per image/frame, thinking disabled).
Links
- π Paper: RefCaptioner: Multi-Reference Image-Grounded Video Captioning
- π€ Weights: TengfeiLiuCoder/RefCaptioner
- π» Code: pkucs-Ltf/RefCaptioner
- π Benchmark: TengfeiLiuCoder/MRVBench
Example media in this Space comes from the linoyts/repo-to-space-example-videos and linoyts/repo-to-space-example-inputs datasets (free-to-use / CC0-style stock media).