Papers
arxiv:2608.03979

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Published on Aug 4
· Submitted by
Zhen Fang
on Aug 5
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

Community

Paper submitter

🎬 We introduce Video-DeepResearch (Video-DR) — a framework that redefines what a multimodal agent looks like when the input is not an image or a document, but the full
temporal stream of a video. The framework rests on three pillars: a decoupled perception→exploration paradigm with stage-wise tool unlocking (👁️ look first, 🌐 search
later) that structurally cures modality bias; a scalable video-grounded data engine producing multi-hop QA that cannot be shortcutted by parametric memory; and a two-stage
SFT → GRPO recipe that lets agents surpass their imitation ceiling and discover tool-use patterns rather than just copy them. The framework generalizes across model
scales and video categories — a shared blueprint for the next generation of vision-native research agents. 🔭 Our vision: streaming DR — know everything in vision. 🚀
code: https://github.com/Osilly/Vision-DeepResearch/tree/main/Video-DeepResearch

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.03979
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.03979 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.03979 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.03979 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.