Video-Text-to-Text
Transformers
Safetensors
English
Chinese
moss_vl
feature-extraction
MOSS-VL
realtime
streaming
video-understanding
FP8
compressed-tensors
HQQ
quantized
custom_code
Instructions to use OpenMOSS-Team/MOSS-VL-Realtime-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenMOSS-Team/MOSS-VL-Realtime-FP8 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("OpenMOSS-Team/MOSS-VL-Realtime-FP8", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download generation_config.json from OpenMOSS-Team/MOSS-VL-Realtime-FP8: direct link, hf CLI and curl.
- Browser
- Download file 314 Bytes
-
https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-FP8/resolve/main/generation_config.json
- Command line
-
hf download hf://OpenMOSS-Team/MOSS-VL-Realtime-FP8/generation_config.json
-
curl -L -o generation_config.json https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-FP8/resolve/main/generation_config.json
314 Bytes
| { | |
| "_from_model_config": true, | |
| "bos_token_id": 151643, | |
| "eos_token_id": 151645, | |
| "transformers_version": "4.57.1", | |
| "cache_implementation": "quantized", | |
| "cache_config": { | |
| "backend": "hqq", | |
| "nbits": 8, | |
| "axis_key": 0, | |
| "axis_value": 0, | |
| "q_group_size": 64, | |
| "residual_length": 128 | |
| } | |
| } | |