Papers
arxiv:2610.03632

World Embedding Benchmark

Published on Oct 2
· Submitted by
Chenghao Xiao
on Oct 5
#1 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,

Abstract

Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.

Community

Paper submitter

We introduce World Embedding Benchmark, a benchmark for studying how video representations capture and expose physical information. Building the benchmark requires substantial simulation engineering: we develop physics simulation pipelines spanning 80 physical families, systematically varying physical parameters and visual conditions to construct 8,000 controlled video-text cases with grounded physical annotations.

  • Benchmarking physical representations. We evaluate physical representations through three complementary tasks—text-video retrieval, physical-property regression, and pair classification—across fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism.

  • What do embeddings actually encode? Pre-trained embeddings preserve recoverable physical signals but show weak cross-modal physical alignment. Physics-aware contrastive adaptation substantially improves retrieval and pair classification, yet degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability.

  • From representation to generation. We further study retrieval-augmented video generation, using different video retrieval models to retrieve physically relevant reference videos for MiniMax-H3. Better physics retrieval leads to more physically faithful generation, with the physics-adapted LCO-Embedding-Omni achieving the strongest retrieval-augmented generation performance, highlighting the value of physically grounded representations for generative world models.

We will release all physics simulation engines, the complete benchmark dataset, evaluation code, and trained models/checkpoints to support future research on physically grounded representations and world models.

Paper submitter

task_overview

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.03632
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.03632 in a model README.md to link it from this page.

Datasets citing this paper 5

Browse 5 datasets citing this paper

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.03632 in a Space README.md to link it from this page.

Collections including this paper 1