Video-Text-to-Text
Transformers
Safetensors
English
Chinese
moss_vl
feature-extraction
Base
Video-Understanding
Image-Understanding
MOSS-VL
OpenMOSS
multimodal
video
vision-language
custom_code
Instructions to use OpenMOSS-Team/MOSS-VL-Base-0708 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenMOSS-Team/MOSS-VL-Base-0708 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("OpenMOSS-Team/MOSS-VL-Base-0708", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
sync offline system prompts with training data: no_thinking -> no system message, deep_thinking -> <think>/<answer> tags
Browse files- modeling_moss_vl.py +4 -4
modeling_moss_vl.py
CHANGED
|
@@ -55,12 +55,12 @@ from transformers.generation.streamers import TextIteratorStreamer
|
|
| 55 |
|
| 56 |
_OFFLINE_SYSTEM_PROMPTS = {
|
| 57 |
"no_thinking": {
|
| 58 |
-
"text_image":
|
| 59 |
-
"video":
|
| 60 |
},
|
| 61 |
"deep_thinking": {
|
| 62 |
-
"text_image": "A conversation between User and Assistant. The user
|
| 63 |
-
"video": "A conversation between User and Assistant
|
| 64 |
},
|
| 65 |
}
|
| 66 |
|
|
|
|
| 55 |
|
| 56 |
_OFFLINE_SYSTEM_PROMPTS = {
|
| 57 |
"no_thinking": {
|
| 58 |
+
"text_image": None,
|
| 59 |
+
"video": None,
|
| 60 |
},
|
| 61 |
"deep_thinking": {
|
| 62 |
+
"text_image": "A conversation between User and Assistant. The user asks a question, and the Assistant solves it. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer. The reasoning process and answer are enclosed within <think>...</think> and <answer>...</answer> tags, respectively, i.e., <think> reasoning process here </think> <answer> answer here </answer>.",
|
| 63 |
+
"video": "A conversation between User and Assistant. The user asks a question, and the Assistant solves it. The assistant first thinks about the reasoning process in the mind and then provides the user with the answer. The reasoning process and answer are enclosed within <think>...</think> and <answer>...</answer> tags, respectively, i.e., <think> reasoning process here </think> <answer> answer here </answer>.",
|
| 64 |
},
|
| 65 |
}
|
| 66 |
|