Video-HopChain-8B

Qwen3-VL-8B-Instruct trained with GRPO in two stages, first on a general video question answering mixture and then on Video-HopChain with Confidence-Gated Exploration (CGE). This is Video-HopChain-8B, the "+ standard RL + Video-HopChain + CGE" row of Table 1 in the paper, at global_step_10 of the second stage.

Results

Accuracy in percent, evaluated with lmms-eval at 100 frames per video under one setting for every row. "In-domain" is the 1,000-question held-out Video-HopChain split.

Model Video-MME PerceptionComp Video-MMMU Video-Holmes VCRBench MMR-V LongVideo-Reason VRBench mean in-domain
Qwen3-VL-8B-Instruct 64.0 28.1 63.5 40.7 32.0 42.7 72.6 74.7 52.3 13.4
+ standard RL 65.8 34.3 63.0 47.4 35.9 43.4 76.1 77.6 55.4 13.4
+ standard RL + CGE 67.9 34.4 62.5 48.5 35.1 46.4 78.8 79.5 56.6 16.4
+ standard RL + Video-HopChain 68.6 36.2 64.7 48.6 42.1 45.9 77.8 79.4 57.9 21.2
this model 69.2 37.4 67.5 48.6 44.9 46.6 78.9 81.0 59.3 23.2

VCRBench is reported on its multiple-choice subset.

Usage

from transformers import AutoProcessor, Qwen3VLForConditionalGeneration

model = Qwen3VLForConditionalGeneration.from_pretrained(
    "ngqtrung/Video-HopChain-8B", dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained("ngqtrung/Video-HopChain-8B")

The weights are BF16. The model was trained to reason inside <think>...</think> and to put the final answer in <answer>\boxed{...}</answer>, so use the same system prompt as training. It is in the dataset rows and in the repository.

Training

Confidence-Gated Exploration recovers zero-variance groups in GRPO (also called advantage collapse). With 8 rollouts per question, CGE draws the first 4 normally. If those 4 are all correct or all incorrect, the group carries no reward variance and no gradient, so CGE draws the last 4 with the policy's top token masked wherever its probability exceeds tau = 0.95, inside the reasoning span only. The masked positions are dropped from the loss while all 8 rollouts enter the group advantage, so the method spends no extra rollouts.

Base model Qwen3-VL-8B-Instruct
Stage 1 GRPO on a 105,993-row general video QA mixture, 24 frames
Stage 2 GRPO + CGE on Video-HopChain, 22,550 rows, 140 frames
Rollouts per question 8: the first 4 as usual, the last 4 under the mask
Mask threshold 0.95, reasoning span only
Reward exact match on the integer sum, plus a format term
Hardware 4 nodes of 8 H100 80GB, asynchronous verl trainer

Appendix C of the paper gives every hyperparameter.

Limitations

Evaluated at one model scale only. The dataset is synthetic and verified against captions rather than against the frames, so a caption error can reach a label. The paper reports both limitations, together with the exploration settings that we did not ablate.

Citation

@misc{nguyenquang2026videohopchain,
  title         = {Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models},
  author        = {Nguyen Quang, Trung and Dong, Yuhao and Sun, Shuo and Liu, Shuai and Tian, Shulin and Yap, Kim-Hui and Liu, Ziwei},
  year          = {2026},
  eprint        = {2609.25773},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.25773}
}
Downloads last month
174
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ngqtrung/Video-HopChain-8B

Finetuned
(614)
this model
Quantizations
2 models

Dataset used to train ngqtrung/Video-HopChain-8B

Paper for ngqtrung/Video-HopChain-8B