Video-Text-to-Text
Transformers
Safetensors
English
qwen3_vl
image-text-to-text
video-temporal-grounding
temporal-localization
confidence-estimation
Instructions to use Alibaba-VELLDEPTH/CTVG-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Alibaba-VELLDEPTH/CTVG-4B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Alibaba-VELLDEPTH/CTVG-4B") model = AutoModelForMultimodalLM.from_pretrained("Alibaba-VELLDEPTH/CTVG-4B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download tokenizer.json from Alibaba-VELLDEPTH/CTVG-4B: direct link, hf CLI and curl.
- Browser
- Download file 11.4 MB
-
https://huggingface.co/Alibaba-VELLDEPTH/CTVG-4B/resolve/main/tokenizer.json
- Command line
-
hf download hf://Alibaba-VELLDEPTH/CTVG-4B/tokenizer.json
-
curl -L -o tokenizer.json https://huggingface.co/Alibaba-VELLDEPTH/CTVG-4B/resolve/main/tokenizer.json
11.4 MB
- Xet hash:
- 693ec4b3922b0bd306bf7b4989e115ffbfeb7b0c08b31bc6d956818c6bb07f61
- Size of remote file:
- 11.4 MB
- SHA256:
- aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
·
Xet efficiently stores Large Files inside Git, intelligently splitting files into unique chunks and accelerating uploads and downloads. More info.