Video-Text-to-Text
Transformers
Safetensors
English
qwen3_vl
image-text-to-text
video
retrieval
reranking
qwen3-vl
Instructions to use hltcoe/RankVideo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hltcoe/RankVideo with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("hltcoe/RankVideo") model = AutoModelForMultimodalLM.from_pretrained("hltcoe/RankVideo", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 1,944 Bytes
c510afa ea365ec c510afa ea365ec c510afa c52c833 c510afa ea365ec c510afa ea365ec 7d4aee5 ea365ec cb132fc ea365ec cb132fc c52c833 7d4aee5 c510afa ea365ec c510afa ea365ec c510afa ea365ec c510afa ea365ec 7d4aee5 ea365ec 7d4aee5 ea365ec | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | ---
base_model: Qwen/Qwen3-VL-8B-Instruct
language:
- en
license: mit
pipeline_tag: video-text-to-text
library_name: transformers
arxiv: 2602.02444
tags:
- video
- retrieval
- reranking
- qwen3-vl
datasets:
- hltcoe/RankVideo-Dataset
---
# RankVideo
RankVideo is a video-native reasoning reranker for text-to-video retrieval, fine-tuned from [Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct).
The model explicitly reasons over query-video pairs using video content to assess relevance. It was introduced in the paper [RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval](https://huggingface.co/papers/2602.02444).
- **Repository:** [https://github.com/tskow99/RANKVIDEO-Reasoning-Reranker](https://github.com/tskow99/RANKVIDEO-Reasoning-Reranker)
- **Paper:** [RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval](https://arxiv.org/abs/2602.02444)
## Training Data
This model was trained using the [MultiVENT 2.0 dataset](https://huggingface.co/datasets/hltcoe/MultiVENT2.0) and [RankVideo-Dataset](https://huggingface.co/datasets/hltcoe/RankVideo-Dataset).
## Usage
You can use the model for scoring query-video pairs via the `rankvideo` library as follows:
```python
from rankvideo import VLMReranker
reranker = VLMReranker(model_path="hltcoe/RankVideo")
# Score query-video pairs for relevance
scores = reranker.score_batch(
queries=["person playing guitar"],
video_paths=["/path/to/video.mp4"],
)
print(f"Relevance score: {scores[0]['logit_delta_yes_minus_no']:.3f}")
```
## BibTeX
```bibtex
@misc{skow2026rankvideoreasoningrerankingtexttovideo,
title={RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval},
author={Tyler Skow and Alexander Martin and Benjamin Van Durme and Rama Chellappa and Reno Kriz},
year={2026},
eprint={2602.02444},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2602.02444},
}
``` |