Papers
arxiv:2610.01092

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Authors:
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.

Community

Video models can make anything look real. Can they actually do the task?

Ego2Act gives a video model one real egocentric start frame and one everyday goal (e.g. "the round yellow paper is stapled to the white A4 paper") and asks it to generate the whole manipulation.

  • 110 real-world tasks, 660 human recordings (3 correct + 3 wrong per task), 1,980 generations from 6 models Ɨ 3 seeds
  • The gap: rated by humans, the best model (Seedance-2.0) averages 64/100 with worst model scoring below 10/100
  • Why they fail: 88.7% of failed steps are skipped or left half-done; physics breaks too (objects morph, pass through solids, open on impossible hinges)
  • Ego2ActJudge: a reference-free judge that plans subgoals in order and scores Task + Physics. r = 0.69 with human consensus (best baseline 0.40, human rater 0.76)

Everything is open: benchmark, code, and every generated video with its scores.
šŸ“Š Benchmark: https://huggingface.co/datasets/ego2act/ego2act-bench
šŸŽ¬ Generated videos + scores: https://huggingface.co/datasets/ego2act/ego2act-vidgen

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.01092
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.01092 in a model README.md to link it from this page.

Datasets citing this paper 2

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.01092 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.