EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model
From Semantic Understanding through Visual Foresight to Action
Project Page | Paper (arXiv) | Code
EWAM is a unified embodied model that jointly processes vision-language, future-video, and action tokens in a single flow-matching Diffusion Transformer. Without any explicit stage design, the action stream learns a depth-wise division of labor: shallow layers retrieve instruction and current-scene semantics from the vision-language expert, intermediate layers shift toward predicted future-frame representations, and deep layers are dominated by action self-attention for motor refinement.
Results
| Benchmark | Setting | Success rate (%) |
|---|---|---|
| RoboTwin 2.0 | Clean-to-random (C2C / C2R / Avg.) | 82.2 / 72.1 / 77.2 |
| RoboTwin 2.0 | In-domain (Clean / Randomized / Avg.) | 93.0 / 92.8 / 92.9 |
| LIBERO | Spatial / Object / Goal / Long / Avg. | 98.6 / 99.8 / 98.6 / 98.2 / 98.8 |
Real-robot results on Franka, Dobot, and Unitree G1-D are reported in the technical report.
Released Checkpoints
| Checkpoint | Folder | Usage |
|---|---|---|
| Pretrain (multi-source, stage 1) | pretrain/ |
Initialization for stage-2 finetuning |
| RoboTwin 2.0 clean-to-random | robotwin-c2r/ |
RoboTwin 2.0 evaluation |
| RoboTwin 2.0 in-domain | robotwin-indomain/ |
RoboTwin 2.0 evaluation |
| LIBERO | libero/ |
LIBERO evaluation |
Each checkpoint is a DeepSpeed-style directory containing mp_rank_00_model_states.pt.
Download
pip install -U "huggingface_hub"
huggingface-cli download HaoWang00/EWAM --local-dir /path/to/ewam_weights
# Download a single checkpoint only
huggingface-cli download HaoWang00/EWAM --include "libero/*" --local-dir /path/to/ewam_weights
The backbone weights are not included. Download Wan2.2-TI2V-5B and Qwen3-VL-2B-Instruct from their official repositories; they provide the configs, tokenizers, VAE, and T5 encoder used at inference.
License
The EWAM weights and code are released under the Apache License 2.0. The Wan2.2 and Qwen3-VL components retain their original licenses.
Citation
@article{wang2026ewam,
title = {{EWAM}: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action},
author = {Wang, Hao and Wen, Jiajun and Liu, Jingzhi and Xue, Shuoshuo and Chen, Zhiliang and Lin, Min and Chang, Yicheng and Guo, Xiaoyu and Zhuo, Yukang and Chong, Zheng and Nie, Yunshuang and Zhang, Jian and Liufu, Weijia and Wu, Qingman and Xu, Heming and Song, Bingchang and Wu, Dantong and Wang, Zhiyuan and Xu, Hang and Han, Jianhua and Chen, Bokui and Zhao, Shen and Li, Rui and Liang, Xiaodan},
journal = {arXiv preprint arXiv:2609.39973},
year = {2026}
}