Instructions to use MCG-NJU/OneStreamer-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MCG-NJU/OneStreamer-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="MCG-NJU/OneStreamer-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("MCG-NJU/OneStreamer-4B") model = AutoModelForMultimodalLM.from_pretrained("MCG-NJU/OneStreamer-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MCG-NJU/OneStreamer-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MCG-NJU/OneStreamer-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MCG-NJU/OneStreamer-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/MCG-NJU/OneStreamer-4B
- SGLang
How to use MCG-NJU/OneStreamer-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MCG-NJU/OneStreamer-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MCG-NJU/OneStreamer-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MCG-NJU/OneStreamer-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MCG-NJU/OneStreamer-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use MCG-NJU/OneStreamer-4B with Docker Model Runner:
docker model run hf.co/MCG-NJU/OneStreamer-4B
OneStreamer-4B
Perceive the present. Remember the past. Respond at the right time.
Paper · GitHub · Project page · OneStreamer-1M
OneStreamer-4B is a streaming video-language model initialized from Qwen3-VL-4B-Instruct. It interprets incoming video, records time-grounded evidence, and decides when to respond. The model is introduced in OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction.
Performance
OneStreamer-4B achieves the best aggregate results among the methods compared in the paper on all eight benchmarks. Each benchmark uses its own metric; all values below are higher-is-better.
| Benchmark | Metric | Qwen3-VL-4B base | OneStreamer-4B |
|---|---|---|---|
| OVOBench | Overall | 58.8 | 72.1 |
| StreamingBench | Real-Time | 81.8 | 86.9 |
| OVBench | Average | 55.4 | 66.8 |
| ODVBench | Overall | 57.6 | 71.3 |
| ProactiveVideoQA | Average | 34.3 | 48.7 |
| OmniMMI | Average | 29.4 | 36.6 |
| OVO-Timing | Average F1 | 29.4 | 41.6 |
| ViSpeak | Average, 0–5 scale | 2.41 | 2.87 |
The figure and numbers follow the GitHub README. OneStreamer's OmniMMI result uses ASR for AP, SI, MD, and SG; PA is visual-only. See the evaluation settings for full protocols.
Inference
OneStreamer processes video incrementally through a recent visual window, while PHCM retains time-aligned caption memory for later questions. For proactive interaction, the model generates </Silence> to keep observing, </Standby> when more evidence is needed, or </Response> followed by an answer once the observed evidence is sufficient.
See Inference/ on GitHub for the implementation and usage.
Citation
@misc{zeng2026onestreamerunifyingperceptionmemory,
title={OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction},
author={Xiangyu Zeng and Yuandong Yang and Zhiqiu Zhang and Yuhan Zhu and Xinhao Li and Qingyi Si and Dingyu Yao and Changlian Ma and Haoran Chen and Xinyu Chen and Yansong Shi and Junhao Zhou and Yifei Li and Jun Zhang and Chuanyu Qin and Chenxu Yang and Xinlei Yu and Kun Ouyang and Yuchen Shao and Qianshan Wei and Changhai Zhou and Jun Gao and Jiaqi Wang and Limin Wang},
year={2026},
eprint={2610.01762},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.01762}
}
- Downloads last month
- -
Model tree for MCG-NJU/OneStreamer-4B
Base model
Qwen/Qwen3-VL-4B-Instruct