OneStreamer-4B

Perceive the present. Remember the past. Respond at the right time.

OneStreamer-4B performance on eight streaming-video benchmarks

Paper · GitHub · Project page · OneStreamer-1M

OneStreamer-4B is a streaming video-language model initialized from Qwen3-VL-4B-Instruct. It interprets incoming video, records time-grounded evidence, and decides when to respond. The model is introduced in OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction.

Performance

OneStreamer-4B achieves the best aggregate results among the methods compared in the paper on all eight benchmarks. Each benchmark uses its own metric; all values below are higher-is-better.

Benchmark Metric Qwen3-VL-4B base OneStreamer-4B
OVOBench Overall 58.8 72.1
StreamingBench Real-Time 81.8 86.9
OVBench Average 55.4 66.8
ODVBench Overall 57.6 71.3
ProactiveVideoQA Average 34.3 48.7
OmniMMI Average 29.4 36.6
OVO-Timing Average F1 29.4 41.6
ViSpeak Average, 0–5 scale 2.41 2.87

The figure and numbers follow the GitHub README. OneStreamer's OmniMMI result uses ASR for AP, SI, MD, and SG; PA is visual-only. See the evaluation settings for full protocols.

Inference

OneStreamer processes video incrementally through a recent visual window, while PHCM retains time-aligned caption memory for later questions. For proactive interaction, the model generates </Silence> to keep observing, </Standby> when more evidence is needed, or </Response> followed by an answer once the observed evidence is sufficient.

See Inference/ on GitHub for the implementation and usage.

Citation

@misc{zeng2026onestreamerunifyingperceptionmemory,
  title={OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction},
  author={Xiangyu Zeng and Yuandong Yang and Zhiqiu Zhang and Yuhan Zhu and Xinhao Li and Qingyi Si and Dingyu Yao and Changlian Ma and Haoran Chen and Xinyu Chen and Yansong Shi and Junhao Zhou and Yifei Li and Jun Zhang and Chuanyu Qin and Chenxu Yang and Xinlei Yu and Kun Ouyang and Yuchen Shao and Qianshan Wei and Changhai Zhou and Jun Gao and Jiaqi Wang and Limin Wang},
  year={2026},
  eprint={2610.01762},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2610.01762}
}
Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MCG-NJU/OneStreamer-4B

Finetuned
(458)
this model

Space using MCG-NJU/OneStreamer-4B 1

Collection including MCG-NJU/OneStreamer-4B

Paper for MCG-NJU/OneStreamer-4B