Instructions to use Alibaba-VELLDEPTH/CTVG-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Alibaba-VELLDEPTH/CTVG-4B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Alibaba-VELLDEPTH/CTVG-4B") model = AutoModelForMultimodalLM.from_pretrained("Alibaba-VELLDEPTH/CTVG-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CTVG-4B
Confidence-based Temporal Video Grounding
Generate once. Rank, select, or reject.
CTVG-4B is the final 4B-family model from Grounding with Confidence: Controllable Generative Video Temporal Grounding. Given a video and a text query, it generates candidate time intervals and scores each interval with a lightweight confidence head using decoder states from the same generation pass.
Paper on arXiv · Code on GitHub · Base model
This repository contains a complete merged generator and its matched confidence head. Loading only the generator with Transformers does not apply the head or reproduce confidence-based selection. Use the accompanying Grounding-with-Confidence inference code.
Model details
| Property | Value |
|---|---|
| Model name | CTVG-4B |
| Model repository | Alibaba-VELLDEPTH/CTVG-4B |
| Task | Text-conditioned video temporal grounding with interval-level confidence |
| Base model | MCG-NJU/TimeLens2-4B (Qwen3-VL architecture) |
| Base revision | ddbb6cb944f13ce21e59e85da23c5f356107260e |
| Training | Supervised fine-tuning followed by set-level reinforcement learning |
| Released checkpoint | Final low-learning-rate RL generator, step 200 |
| Confidence head | Matching online head after 150 updates |
| Generator format | Merged BF16 SafeTensors; not a LoRA-only adapter |
| Confidence architecture | LayerNorm(2560) → Linear(2560, 256) → GELU → Dropout(0.1) → Linear(256, 1), followed by sigmoid |
| Confidence features | Mean of semantic interval-token decoder states, with stop-gradient during head training |
The model separates candidate generation from acceptance. Generation-order NMS at temporal IoU 0.3 runs before confidence thresholding; confidence does not reorder NMS. Threshold changes can reuse the saved candidate scores without another decoding pass. No external verifier or additional confidence tokens are needed at inference.
Getting started
Get the inference code
Clone the code repository, then run the following download and inference commands from its root:
git clone https://github.com/Alibaba-VELLDEPTH/Grounding-with-Confidence.git
cd Grounding-with-Confidence
Download the complete model
Use the Hugging Face CLI in a separate download environment if needed:
hf download Alibaba-VELLDEPTH/CTVG-4B \
--local-dir models/CTVG-4B
For reproducible experiments, add --revision COMMIT_SHA using the commit shown in this repository's Files and versions tab. Access to a private repository requires login with an authorized account.
Download the whole repository, not just the generator shards. release.json, span-confidence-head.safetensors, and span-head-config.json are required for model/head pairing validation. No separate base-model download is needed for this merged release.
Run confidence-aware inference
From the root of the accompanying Grounding-with-Confidence codebase, use Python 3.10+ and a CUDA GPU:
python -m pip install -e .
python -m pip install -r requirements.txt
python scripts/infer.py \
--model models/CTVG-4B \
--input requests.jsonl \
--output-dir outputs/ctvg-example \
--threshold 0.5
Create requests.jsonl with one request per line, replacing the example path, query, and duration with your own data:
{"request_id":"example-1","video_path":"videos/example.mp4","query":"A person opens a door.","source_duration_s":30.0}
A relative video_path is resolved against the request JSONL's directory. Add gt_intervals to every request and pass --evaluate to also produce an evaluation report. Add --check-only to validate inputs and artifact hashes without loading the GPU model.
The output predictions.jsonl includes candidate intervals, their span_scores, and selected_intervals. The example cutoff of 0.5 is illustrative, not a calibrated default; select a threshold on separate validation data for your application. Failed requests remain explicit in the output and evaluation denominator.
The released inference configuration uses 2 fps, min_pixels=2048, total_pixels=8388608, a 16,384-token context, and at most 512 output tokens. SDPA is the default attention backend. The full model/head inference path requires the companion scripts rather than a generic text-generation pipeline.
Training data and procedure
Public video-grounding data sources are TimeLens2-93K and OMTG-56K. Obtain the data from the upstream repositories under their respective terms; source videos are not redistributed here.
- SFT: GT-anchored interval targets train the generator with LoRA. A separate fixed candidate replay and verifier-v2 continuous soft labels train the detached confidence head with soft BCE and ranking supervision. The historical run used 60,405 units and 1,888 optimizer steps.
- RL: Set-level rewards train the generator, while online overlap supervision updates the detached head after actor updates. The released checkpoint is actor step 200 with 150 online head updates, not the bootstrap SFT head.
The upstream downloads alone do not reconstruct the paper-specific splits, offline replay, or verifier scores. This weight release does not include original training videos, those additional supervision files, optimizer state, or distributed training checkpoints.
Evaluation
Reported results on all 320 OMTG-Bench queries:
| Metric | Reported value (%) |
|---|---|
| tF1@0.5 | 67.39 |
| Temporal IoU | 63.24 |
| EtF1 | 42.64 |
These are the manuscript's retrospective benchmark results. The operating threshold, 0.18404516577720642, was selected on the test set and must not be treated as an independently calibrated deployment threshold. The default companion evaluator follows the official OMTG metric protocol, including its merge-before-matching behavior; it is distinct from generic interval-set metrics.
Confidence also improves selection from the same fixed post-NMS candidate pool under equal global return budgets. Official query-macro Recall@0.5 (%), with 1,314 candidates across 320 queries:
| Global return budget | Generation order | Token likelihood | External verifier | Confidence |
|---|---|---|---|---|
| 10% | 9.95 | 9.41 | 11.00 | 14.42 |
| 25% | 26.48 | 22.37 | 26.31 | 31.12 |
| 50% | 51.34 | 43.05 | 47.72 | 52.82 |
| 75% | 64.76 | 61.00 | 62.45 | 65.30 |
The budget is shared across queries, not assigned independently to each query. These tables report prior experiments, not a new inference run performed during upload. GPU end-to-end acceptance of the companion release wrapper remains a separate validation step; CPU artifact and metric checks do not establish GPU execution correctness.
Intended use and limitations
CTVG-4B is intended for research on video moment retrieval, repeated-event localization, and controllable selection or rejection of generated temporal intervals, subject to applicable permissions.
- Confidence can only select among generated candidates. It cannot recover missing events, repair boundaries, or reverse NMS suppression.
- Scores do not establish calibrated correctness probabilities. Precision targets and thresholds may not transfer to new datasets or applications.
- Most of the full-system tF1 improvement comes from NMS. The additional confidence gain after NMS is +0.32 percentage points with a 95% source-video bootstrap interval of [−0.14, 0.76], so that incremental improvement is not statistically established.
- Reported rejection experiments use synthetic cross-video query mismatches, not general open-world rejection.
- Do not rely on the model as the sole basis for safety-critical decisions or consequential decisions about people. Review outputs and comply with privacy and video-use permissions.
License and upstream attribution
This repository preserves the pinned TimeLens2 source notice in LICENSE-TIMELENS2.txt, including its academic-use and geographic restrictions. The upstream HF model card labels TimeLens2-4B as Apache-2.0; that model metadata and the pinned source notice are distinct. This release does not resolve their applicability or grant an additional unrestricted MIT/Apache license to CTVG-4B. Review the applicable upstream terms and obtain any required permissions before use or redistribution; public availability alone does not imply permission for commercial use.
We thank the authors of TimeLens2 and OMTG for releasing their code and training datasets. The work also builds on Qwen3-VL, Transformers, PEFT, PyTorch, and OMTG's vendored verl integration. Their resources retain their respective terms.
Citation
@article{chen2026grounding,
title={Grounding with Confidence: Controllable Generative Video Temporal Grounding},
author={Chen, Jinhao and Cui, Benlei and Jia, Ruijian and Wang, Ziheng and Wo, Tianyu and Sun, Pengfei and Huang, Longtao and Xue, Hui and Yang, Yitong and Hong, Haiwen},
journal={arXiv preprint arXiv:2609.39883},
year={2026}
}
- Downloads last month
- -