EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model

From Semantic Understanding through Visual Foresight to Action

Project Page | Paper (arXiv) | Code

EWAM is a unified embodied model that jointly processes vision-language, future-video, and action tokens in a single flow-matching Diffusion Transformer. Without any explicit stage design, the action stream learns a depth-wise division of labor: shallow layers retrieve instruction and current-scene semantics from the vision-language expert, intermediate layers shift toward predicted future-frame representations, and deep layers are dominated by action self-attention for motor refinement.

Results

Benchmark Setting Success rate (%)
RoboTwin 2.0 Clean-to-random (C2C / C2R / Avg.) 82.2 / 72.1 / 77.2
RoboTwin 2.0 In-domain (Clean / Randomized / Avg.) 93.0 / 92.8 / 92.9
LIBERO Spatial / Object / Goal / Long / Avg. 98.6 / 99.8 / 98.6 / 98.2 / 98.8

Real-robot results on Franka, Dobot, and Unitree G1-D are reported in the technical report.

Released Checkpoints

Checkpoint Folder Usage
Pretrain (multi-source, stage 1) pretrain/ Initialization for stage-2 finetuning
RoboTwin 2.0 clean-to-random robotwin-c2r/ RoboTwin 2.0 evaluation
RoboTwin 2.0 in-domain robotwin-indomain/ RoboTwin 2.0 evaluation
LIBERO libero/ LIBERO evaluation

Each checkpoint is a DeepSpeed-style directory containing mp_rank_00_model_states.pt.

Download

pip install -U "huggingface_hub"
huggingface-cli download HaoWang00/EWAM --local-dir /path/to/ewam_weights

# Download a single checkpoint only
huggingface-cli download HaoWang00/EWAM --include "libero/*" --local-dir /path/to/ewam_weights

The backbone weights are not included. Download Wan2.2-TI2V-5B and Qwen3-VL-2B-Instruct from their official repositories; they provide the configs, tokenizers, VAE, and T5 encoder used at inference.

License

The EWAM weights and code are released under the Apache License 2.0. The Wan2.2 and Qwen3-VL components retain their original licenses.

Citation

@article{wang2026ewam,
  title   = {{EWAM}: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action},
  author  = {Wang, Hao and Wen, Jiajun and Liu, Jingzhi and Xue, Shuoshuo and Chen, Zhiliang and Lin, Min and Chang, Yicheng and Guo, Xiaoyu and Zhuo, Yukang and Chong, Zheng and Nie, Yunshuang and Zhang, Jian and Liufu, Weijia and Wu, Qingman and Xu, Heming and Song, Bingchang and Wu, Dantong and Wang, Zhiyuan and Xu, Hang and Han, Jianhua and Chen, Bokui and Zhao, Shen and Li, Rui and Liang, Xiaodan},
  journal = {arXiv preprint arXiv:2609.39973},
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for HaoWang00/EWAM