Benchmarks, tradeoffs, and live video inference across five Qwen 3.5 models

#43
by younes-ovs - opened

We're Overshoot (YC W26), we do real-time vision inference with VLMs.

The Qwen 3.5 family is the strongest open-weight vision model family available today. We deployed five models (2B, 4B, 9B, 27B, 35B-A3B) and on Overshoot they run inference on continuous live video at under 200ms.

We wrote up our findings: which model wins on which benchmark, where each one falls short, and how they compare to GPT-5-mini and Claude Sonnet 4.5 on vision tasks.

Some highlights:

  • 9B beats last-gen Qwen3-VL-30B on every vision benchmark at 1/3 the size
  • 4B retains 96% of the 27B's video understanding score
  • 27B beats GPT-5-mini on VideoMME, MathVision, and OmniDocBench
  • All five models running on live video streams at under 200ms

Full post with interactive benchmark charts: https://blog.overshoot.ai/blog/qwen3.5-on-overshoot

2 hours of free video inference if you want to try them: https://playground.overshoot.ai

Sign up or log in to comment