Stereo World Model: Camera-Guided Stereo Video Generation
Paper β’ 2603.17375 β’ Published β’ 11
How to use Yang-Tian/StereoWorld with Diffusers:
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("Yang-Tian/StereoWorld", dtype=torch.bfloat16, device_map="cuda")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]Official model weights for StereoWorld: Camera-Guided Stereo Video Generation.
This repository contains two models:
| Directory | Model | Description |
|---|---|---|
StereoWorldModel/ |
Fixed-Baseline Stereo | Generates side-by-side stereo video with a configurable but fixed stereo baseline. Use --use_raymap during inference. |
StereoWorldFlexModel/ |
Flexible Stereo | Independent left/right camera control with multiple right-camera modes (converging, horizontal offset, depth offset, height offset). |
StereoWorldModel/)
A binocular teacher model that produces left-right stereo pairs with a fixed baseline between cameras. Best for standard stereo video generation where consistent disparity is desired.
StereoWorldFlexModel/)
A more general multi-view world model that allows independent camera trajectories for left and right views. Supports four right-camera modes for creative stereo effects.
huggingface-cli download Yang-Tian/StereoWorld --local-dir weights
huggingface-cli download Yang-Tian/StereoWorld \
--include "StereoWorldModel/*" \
--local-dir weights
huggingface-cli download Yang-Tian/StereoWorld \
--include "StereoWorldFlexModel/*" \
--local-dir weights
StereoWorld/
βββ StereoWorldModel/
β βββ transformer/
β βββ vae/
β βββ tokenizer/
β βββ text_encoder/
β βββ scheduler/
βββ StereoWorldFlexModel/
βββ transformer/
βββ vae/
βββ tokenizer/
βββ text_encoder/
βββ scheduler/
See the GitHub repository for installation and inference instructions.
# Fixed-Baseline Stereo
python3 inference.py \
--pipeline_dir weights/StereoWorldModel \
--use_raymap \
--eval_json ExpData/demo_custom_eval.json
# Flexible Stereo
python3 inference_flex.py \
--pipeline_dir weights/StereoWorldFlexModel \
--eval_json ExpData/flex_demo_custom_eval.json
@article{sun2026stereo,
title={Stereo World Model: Camera-Guided Stereo Video Generation},
author={Sun Yang-Tian and Huang Zehuan and Niu Yifan and Ma Lin and Cao Yan-Pei and Ma Yuewen and Qi Xiaojuan},
journal={arXiv preprint arXiv:2603.17375},
year={2026}
}