PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
Abstract
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typically only a few minutes long, many answer pairs can be separated from the transcript alone, and collecting human judgments does not scale to ultra-long videos. We introduce PlaylistEval, an agentic framework that builds video-language judge benchmarks over 100-hour playlist collection without human annotation. It automatically generates questions with paired answers whose differences are controlled by causal degradation, so that every pair demands retrieval across the collection. The resulting benchmark contains 630 pairs across seven domains spanning both static and dynamic knowledge, and on a stratified subset of 152 pairs it agrees with human judgments 93.0% of the time (IAA 0.781). Evaluating 17 omnimodal and multimodal models from eight families reveals that frontier judges reach only 75.4% pairwise accuracy, while open-source judge models perform far behind. We further show that both retrieval and final judgment depend on using multiple modalities, and that judge accuracy degrades as the playlist set grows. We release our pipeline, benchmark, and evaluation code at https://playlisteval.github.io.
Community
We benchmark 17 video-language judges from eight model families on PlaylistEval, covering 630 preference pairs across seven domains, with roughly 100 hours of video per domain. Several findings stood out:
The strongest judge reaches only 75.4% accuracy, compared with 93.0% human agreement.
Judge accuracy consistently drops as the video collection grows from about 1 hour to 100 hours.
Retrieval matters a lot. It improves accuracy by up to 10.5 points, yet the best retriever finds both relevant segments in the top 10 only 37.9% of the time.
Both visual and textual evidence are important. In contrast, simply increasing reasoning budget or visual resolution brings only limited gains.
Video reward models trained on short clips perform around chance in this setting, suggesting that short-video judging does not transfer well to day-scale video.
Get this paper in your agent:
hf papers read 2609.34314 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper