Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Abstract
Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.
Community
Video models can make anything look real. Can they actually do the task?
Ego2Act gives a video model one real egocentric start frame and one everyday goal (e.g. "the round yellow paper is stapled to the white A4 paper") and asks it to generate the whole manipulation.
- 110 real-world tasks, 660 human recordings (3 correct + 3 wrong per task), 1,980 generations from 6 models Ć 3 seeds
- The gap: rated by humans, the best model (Seedance-2.0) averages 64/100 with worst model scoring below 10/100
- Why they fail: 88.7% of failed steps are skipped or left half-done; physics breaks too (objects morph, pass through solids, open on impossible hinges)
- Ego2ActJudge: a reference-free judge that plans subgoals in order and scores Task + Physics. r = 0.69 with human consensus (best baseline 0.40, human rater 0.76)
Everything is open: benchmark, code, and every generated video with its scores.
š Benchmark: https://huggingface.co/datasets/ego2act/ego2act-bench
š¬ Generated videos + scores: https://huggingface.co/datasets/ego2act/ego2act-vidgen
Get this paper in your agent:
hf papers read 2610.01092 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 2
ego2act/ego2act-vidgen
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper