Instructions to use ngqtrung/Video-HopChain-8B-Standard-RL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ngqtrung/Video-HopChain-8B-Standard-RL with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ngqtrung/Video-HopChain-8B-Standard-RL") model = AutoModelForMultimodalLM.from_pretrained("ngqtrung/Video-HopChain-8B-Standard-RL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Video-HopChain-8B-Standard-RL
This is the stage-1 checkpoint of Video-HopChain. It is Qwen3-VL-8B-Instruct trained with GRPO on a 105,993-row general video QA mixture (LLaVA-Video 72,421, STAR 11,455, CLEVRER 8,220, NExT-QA 7,549, PerceptionTest 6,348) at 24 frames, taken at step 80. It saw no Video-HopChain data and used no Confidence-Gated Exploration (CGE). It is the "+ standard RL" row of Table 1 in the paper. Use it as the starting point of both second-stage runs, including ngqtrung/Video-HopChain-8B, as a baseline, or to reproduce Table 1.
- Paper: Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models (arXiv:2609.25773)
- Final model: ngqtrung/Video-HopChain-8B
- Dataset: ngqtrung/Video-HopChain
- Project page: ngquangtrung57.github.io/video-hopchain-page
- Code: github.com/ngquangtrung57/video-hopchain
- Collection: Video-HopChain
Results
Accuracy in percent, evaluated with lmms-eval at 100 frames per video under one setting for every
row: at most 501,760 pixels per frame, a 33,792-token context, and at most 16,384 generated tokens.
"In-domain" is the 1,000-question held-out Video-HopChain split.
| Model | Video-MME | PerceptionComp | Video-MMMU | Video-Holmes | VCRBench | MMR-V | LongVideo-Reason | VRBench | mean | in-domain |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-8B-Instruct | 64.0 | 28.1 | 63.5 | 40.7 | 32.0 | 42.7 | 72.6 | 74.7 | 52.3 | 13.4 |
| this model (+ standard RL) | 65.8 | 34.3 | 63.0 | 47.4 | 35.9 | 43.4 | 76.1 | 77.6 | 55.4 | 13.4 |
| Video-HopChain-8B (final) | 69.2 | 37.4 | 67.5 | 48.6 | 44.9 | 46.6 | 78.9 | 81.0 | 59.3 | 23.2 |
VCRBench is reported on its multiple-choice subset.
Usage
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
model = Qwen3VLForConditionalGeneration.from_pretrained(
"ngqtrung/Video-HopChain-8B-Standard-RL", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("ngqtrung/Video-HopChain-8B-Standard-RL")
The weights are BF16. The model was trained to reason inside <think>...</think> and to put the final
answer in <answer>\boxed{...}</answer>, so use the same system prompt as training. It is in the
dataset rows.
Citation
@misc{nguyenquang2026videohopchain,
title = {Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models},
author = {Nguyen Quang, Trung and Dong, Yuhao and Sun, Shuo and Liu, Shuai and Tian, Shulin and Yap, Kim-Hui and Liu, Ziwei},
year = {2026},
eprint = {2609.25773},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.25773}
}
- Downloads last month
- 51