Papers
arxiv:2609.33627

StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

Published on Sep 27
¡ Submitted by
Yeying Jin
on Sep 30
Authors:
,

Abstract

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.

Community

Paper author Paper submitter

🎬 StoryEngine: A State-Driven Agentic Framework for Multi-Shot Video Storytelling
AI can generate beautiful short videos. But keeping a story coherent across multiple shots remains a challenge: characters drift, objects become inconsistent, and errors accumulate.
Our approach: give the story an explicit “world state.”
This state tracks where characters and objects are, and how events change them. It also defines how each shot should begin and end—so generation follows the intended story.
✨ Highlights
🧠 Track the story
Turn narrative events into explicit state changes, keeping track of what happens and what changes from shot to shot.
🖼️ Keep visuals consistent
Create shared visual references for characters, objects, and scenes. Select camera viewpoints to suit the action, and translate story states into clear shot-generation requirements.
🔁 Verify and correct
Use the next shot’s expected state to decide how to use the previous shot’s visuals: reuse them, reference them partially, or discard them.
Check actions and states in every shot, then make targeted corrections within a fixed retry budget. The goal: prevent visual errors from changing what happens next in the story.
📊 Evaluate the whole story—not just individual shots
We also introduce a video evaluation benchmark with 8 metrics across 3 dimensions:
• Storytelling
• Cross-shot coherence
• Visual consistency

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33627
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33627 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33627 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33627 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.