Spaces:
Running on Zero
Running on Zero
|
Download README.md from hugging-apps/refcaptioner: direct link, hf CLI and curl.
- Browser
- Download file 2.06 kB
-
https://huggingface.co/spaces/hugging-apps/refcaptioner/resolve/main/README.md
- Command line
-
hf download hf://spaces/hugging-apps/refcaptioner/README.md
-
curl -L -o README.md https://huggingface.co/spaces/hugging-apps/refcaptioner/resolve/main/README.md
2.06 kB
| title: RefCaptioner | |
| emoji: π¬ | |
| colorFrom: indigo | |
| colorTo: green | |
| sdk: gradio | |
| sdk_version: 6.30.0 | |
| app_file: app.py | |
| short_description: Multi-reference image-grounded video captioning | |
| python_version: "3.10" | |
| startup_duration_timeout: 1h | |
| # RefCaptioner: Multi-Reference Image-Grounded Video Captioning | |
| Demo of [RefCaptioner](https://huggingface.co/TengfeiLiuCoder/RefCaptioner) (paper: [2607.28509](https://arxiv.org/abs/2607.28509)), an 8B vision-language model fine-tuned from Qwen3-VL-8B-Instruct for **multi-reference image-grounded video captioning**. | |
| Given a video and an ordered list of reference images, the model writes a fluent English video caption and places `<Image_n>` tags immediately after the visual phrases that each reference grounds. References that are distractors (not visible in the video) are omitted. | |
| **Inputs** | |
| - π₯ one video | |
| - πΌοΈ 1β6 reference images β upload order defines the tag mapping (`<Image_1>`, `<Image_2>`, β¦) | |
| - optional: leave a reference slot empty to see the model reject a distractor | |
| **Output** | |
| - one grounded caption paragraph with phrase-level `<Image_n>` bindings (tags highlighted) | |
| The demo follows the authors' released inference recipe exactly (prompt `Prompt_1.0`, greedy decoding, 2 FPS sampling, up to 20 frames, 602112 max pixels per image/frame, thinking disabled). | |
| ## Links | |
| - π Paper: [RefCaptioner: Multi-Reference Image-Grounded Video Captioning](https://arxiv.org/abs/2607.28509) | |
| - π€ Weights: [TengfeiLiuCoder/RefCaptioner](https://huggingface.co/TengfeiLiuCoder/RefCaptioner) | |
| - π» Code: [pkucs-Ltf/RefCaptioner](https://github.com/pkucs-Ltf/RefCaptioner) | |
| - π Benchmark: [TengfeiLiuCoder/MRVBench](https://huggingface.co/datasets/TengfeiLiuCoder/MRVBench) | |
| Example media in this Space comes from the [`linoyts/repo-to-space-example-videos`](https://huggingface.co/datasets/linoyts/repo-to-space-example-videos) and [`linoyts/repo-to-space-example-inputs`](https://huggingface.co/datasets/linoyts/repo-to-space-example-inputs) datasets (free-to-use / CC0-style stock media). |