refcaptioner / README.md
multimodalart's picture
multimodalart HF Staff
Upload folder using huggingface_hub
f3fc14a verified
|
Raw History Blame Contribute Delete
2.06 kB
metadata
title: RefCaptioner
emoji: 🎬
colorFrom: indigo
colorTo: green
sdk: gradio
sdk_version: 6.30.0
app_file: app.py
short_description: Multi-reference image-grounded video captioning
python_version: '3.10'
startup_duration_timeout: 1h

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Demo of RefCaptioner (paper: 2607.28509), an 8B vision-language model fine-tuned from Qwen3-VL-8B-Instruct for multi-reference image-grounded video captioning.

Given a video and an ordered list of reference images, the model writes a fluent English video caption and places <Image_n> tags immediately after the visual phrases that each reference grounds. References that are distractors (not visible in the video) are omitted.

Inputs

  • πŸŽ₯ one video
  • πŸ–ΌοΈ 1–6 reference images β€” upload order defines the tag mapping (<Image_1>, <Image_2>, …)
  • optional: leave a reference slot empty to see the model reject a distractor

Output

  • one grounded caption paragraph with phrase-level <Image_n> bindings (tags highlighted)

The demo follows the authors' released inference recipe exactly (prompt Prompt_1.0, greedy decoding, 2 FPS sampling, up to 20 frames, 602112 max pixels per image/frame, thinking disabled).

Links

Example media in this Space comes from the linoyts/repo-to-space-example-videos and linoyts/repo-to-space-example-inputs datasets (free-to-use / CC0-style stock media).