ImageWAM on Unitree G1 with Three Image-Editing Backbones
Three ImageWAM world-action models for a dual-arm Unitree G1, fine-tuned with the paper's real-world recipe on the four LGG100 datasets. The three runs share data, split, inputs, outputs, recipe and length. They differ only in the image-editing backbone and the parts bound to it: the editing DiT, its text features, its VAE, and the depth of the action expert.
| Directory | Backbone | Video expert | Action expert | Base license |
|---|---|---|---|---|
g1_lgg100_flux2_klein_4b_base_imagewam/ |
FLUX.2 [klein] base 4B | 3.88 B | 0.64 B (5 double + 20 single blocks) | Apache-2.0 |
g1_lgg100_flux2_klein_9b_base_imagewam/ |
FLUX.2 [klein] base 9B | 9.08 B | 0.95 B (8 double + 24 single blocks) | FLUX Non-Commercial License v2.1 |
g1_lgg100_qwenimage21_imagewam/ |
Qwen-Image-2.1 | 7.12 B | 0.85 B (32 single-stream blocks) | Qwen Research License |
Built with Qwen: the Qwen-Image-2.1 model is fine-tuned from Qwen-Image-2.1.
Repository Layout
<run>/
checkpoints/weights/step_005000.pt … step_025000.pt, step_027950.pt # weights every 5,000 steps and the final step
config.yaml # the full resolved training config (Hydra); `model` builds the network, `data` holds the preprocessing
dataset_stats.json # z-score statistics of the 16-dim state and action, needed to normalize inputs and de-normalize outputs
text_cache/ # text features of the four training task strings, in the training cache format
LICENSE(.md), NOTICE # license of the backbone this run derives from
- Training logs, videos, optimizer states and data are not included.
- A weight file is a
torch.savedict:mot: the mixture-of-transformers state dict, video expert and action expert, bfloat16;proprio_encoder: the state-token encoder;stepandtorch_dtype.
- The 9B files were saved through the LoRA-merged path. Each video-expert block appears under two names (
mixtures.video.transformer.*andmixtures.video.{double,single}_blocks.*), so those files are about twice the size of the parameters. Loading is unaffected. config.yamlkeeps the absolute paths of the training machine (dataset roots, base weights, text caches). Point them to local copies before use.
Inputs and Outputs
All three models take the same inputs and return the same outputs. Control runs at 30 Hz, the frame rate of the data.
Image (one frame, time t). Three RGB cameras of the G1, each 480 × 640, are tiled into one 288 × 256 image in the RoboTwin compact layout:
| Order | Camera key | Tile | Position |
|---|---|---|---|
| 1 | cam_left_high (left half of the head stereo camera) |
192 × 256 | top |
| 2 | cam_left_wrist |
96 × 128 | bottom left |
| 3 | cam_right_wrist |
96 × 128 | bottom right |
Preprocessing, as in training:
- Each view is converted to float in [0, 1] and resized to 240 × 320.
- Each view is resized to its tile size, bilinear with antialiasing.
- The tiles are stacked and the result is scaled to [-1, 1].
The image is encoded by the backbone's VAE: the FLUX.2 VAE (ae.safetensors from FLUX.2-dev) or the Qwen-Image-2.1 VAE.
State (proprio), 16 values in raw units, z-scored with dataset_stats.json → state:
| Index | Meaning |
|---|---|
| 0–6 | left arm joints (rad): shoulder pitch, shoulder roll, shoulder yaw, elbow, wrist roll, wrist pitch, wrist yaw |
| 7–13 | right arm joints (rad), same order |
| 14 | left Dex1 gripper |
| 15 | right Dex1 gripper |
The proprio encoder maps the normalized state to one token, packed right after the valid text tokens.
Text. The task string is wrapped as
A video recorded from a robot's point of view executing the following instruction: {task}
and encoded to 128 positions with a validity mask:
- FLUX.2: Qwen3 hidden states of layers 9, 18 and 27, concatenated. Qwen3-4B (7,680 dims) for the 4B model; Qwen3-8B (12,288 dims) for the 9B model.
- Qwen-Image-2.1: the last Qwen3-VL decoder layer before the final norm (4,096 dims), with the text-to-image chat template without a system message.
text_cache/ holds these features for the four training tasks. The file name is the SHA-256 of the wrapped string, and each file has text_hidden_states [128, D] (bfloat16) and text_attention_mask [128] (bool).
| Dataset | Task string |
|---|---|
| drawer | Use the arm nearer to the drawer to pull it open, then use the other arm to pick up the purple cube and place it inside the drawer, and finally close the drawer. |
| biomanual | Pick up the blue cup with the nearer arm and place it on the blue coaster, then pick up the green cup with the nearer arm and place it on the now-empty green coaster. |
| simple-pick | Pick up the pink cube with the nearer arm and place it into the green bin. |
| Stack-the-cubes | Stack the blocks by color: put the red block in the center, then stack the blue block on the red block, then stack the yellow block on the blue block. |
Output.
- Action chunk [16, 16]: 16 future actions starting at t (0.53 s at 30 Hz), with the same 16-dim layout as the state. The values are absolute joint and gripper targets, z-scored; de-normalize them with
dataset_stats.json→action. - Video branch (optional): the predicted 288 × 256 frame at t + 16, from joint image and action inference.
Timing in the data. In LGG100, the recorded action leads the recorded state by about 3 frames. The camera frames lag the joint state by about 3 frames (100 ms). Deployments should keep this relative timing between the images and the state.
Model and Training Configuration
Shared by the three runs:
| Setting | Value |
|---|---|
| Architecture | ImageWAM mixture of transformers: the image-editing DiT as video expert, a smaller action DiT paired with it layer by layer as action expert |
| Sequence and attention | `[text + proprio token |
| Objective | flow matching; λ_video 0.5, λ_action 1.0; train and inference shift 5.0 for both experts |
| Inference | 10 action denoising steps (as in validation), seed 42 for evaluation, action horizon 16 (max_action_horizon 64) |
| Data | LGG100 drawer, biomanual, simple-pick and Stack-the-cubes, trained jointly. 339 valid episodes (11 invalid ones excluded), 181,403 frames at 30 fps. Episode-level 1% validation split (seed 42): 335 training episodes and 4 held-out episodes, one per dataset. |
| Augmentation | color jitter, gamma, Gaussian noise, random resized crop (scale 0.95–1.0) and rotation ±5°, applied with p = 0.5 |
| Optimizer | AdamW, betas (0.9, 0.95), learning rate 1e-4, weight decay 0.01, 5% linear warmup, then cosine to 0.01× |
| Batch and length | effective batch 64, 27,950 steps (10 epochs) |
| Precision and parallelism | bf16, DeepSpeed ZeRO-2, gradient clipping 1.0 |
| Trainable | the whole editing DiT, the action expert and the proprio encoder. The text encoder and the VAE are frozen. |
| Initialization | video expert from the backbone weights; action expert copied from the backbone's blocks layer by layer and interpolated to its width (no embodied-data pretraining) |
Per run:
| Run | Text features | GPUs | Gradient accumulation | Wall clock | Peak GPU memory |
|---|---|---|---|---|---|
| FLUX.2-4B | Qwen3-4B, 128 × 7,680 | 2 × H100 NVL | 8 | 31.3 h | 60.9 GiB |
| FLUX.2-9B | Qwen3-8B, 128 × 12,288 | 4 × H100 NVL | 4 | 28.1 h | 86.1 GiB |
| Qwen-Image-2.1 | Qwen3-VL-8B, 128 × 4,096 | 2 × H100 NVL | 8 | 51.8 h | 84.6 GiB |
Usage
The weights load into the ImageWAM code (github.com/yuyangalin/ImageWAM, MIT). Two parts of these runs are additions to that code and are not part of the public repository:
- the 3-camera, 16-dim G1 data configuration;
- the Qwen-Image-2.1 backbone (
imagewam.runtime.create_imagewam_qwenimage21).
Building a model needs the base weights the config names:
- FLUX.2: the FLUX.2 [klein] base transformer and the FLUX.2-dev
ae.safetensorsVAE. Qwen3-4B or Qwen3-8B is needed only for prompts outsidetext_cache/. - Qwen-Image-2.1: the Qwen-Image-2.1
transformer/andvae/. New prompts need its Qwen3-VL text encoder, which requires transformers ≥ 5.
import torch
from hydra.utils import instantiate
from omegaconf import OmegaConf
from imagewam.datasets.lerobot.utils.normalizer import load_dataset_stats_from_json
run = "g1_lgg100_flux2_klein_4b_base_imagewam"
cfg = OmegaConf.load(f"{run}/config.yaml") # edit the base-weight paths in cfg.model first
model = instantiate(cfg.model, model_dtype=torch.bfloat16, device="cuda")
model.load_checkpoint(f"{run}/checkpoints/weights/step_027950.pt")
model.eval()
processor = instantiate(cfg.data.val.processor).eval() # state normalization / action de-normalization
processor.set_normalizer_from_stats(load_dataset_stats_from_json(f"{run}/dataset_stats.json"))
cache = torch.load(f"{run}/text_cache/<sha256>.qwen3_flux2_len128.pt")
pred = model.infer_action(
prompt=None,
input_image=image, # [1, 3, 288, 256] in [-1, 1], tiled as above
proprio=proprio, # [16] normalized state
context=cache["text_hidden_states"],
context_mask=cache["text_attention_mask"],
action_horizon=16, num_inference_steps=10, seed=42,
)
action_normalized = pred["action"] # [16, 16]; de-normalize with the processor's action normalizer
Evaluation
Offline only. No physical-robot results are reported yet.
Held-out actions.
- Setup: 316 samples, every 8th frame of the four held-out episodes. The preprocessing is that of training, applied to the raw 480 × 640 frames; 10 denoising steps, seed 42, final weights.
- Metric: mean absolute error against the recorded 16-step action chunk, in raw units.
| Model | Arm joints (rad) | Grippers | drawer | biomanual | simple-pick | Stack-the-cubes |
|---|---|---|---|---|---|---|
| FLUX.2-4B | 0.0209 ± 0.0007 | 0.038 | 0.036 | 0.019 | 0.015 | 0.017 |
| FLUX.2-9B | 0.0209 ± 0.0007 | 0.039 | 0.036 | 0.018 | 0.015 | 0.017 |
| Qwen-Image-2.1 | 0.0206 ± 0.0007 | 0.044 | 0.033 | 0.019 | 0.016 | 0.017 |
- The per-dataset columns are arm-joint errors.
- Sending the frames through a 256 × 320 JPEG (quality 90) camera path changes the predicted actions by 0.001 on average.
- On this held-out set, the three backbones give the same arm-joint error within its standard error.
Training-time validation (2 random held-out samples per evaluation, 4 for the 9B run; mean of the last 10 evaluations, steps 23,000–27,500):
| Model | val_loss |
action_l1 (normalized) |
PSNR, predicted vs. true frame t + 16 | SSIM | PSNR, VAE reconstruction |
|---|---|---|---|---|---|
| FLUX.2-4B | 0.174 | 0.019 | 21.7 | 0.843 | 43.3 |
| FLUX.2-9B | 0.176 | 0.023 | 21.7 | 0.831 | 43.3 |
| Qwen-Image-2.1 | 0.093 | 0.018 | 21.6 | 0.833 | 43.4 |
The Qwen-Image-2.1 val_loss is lower because its latent space has a smaller frame-to-frame change per element. Its predicted frames and actions are not better by these measures.
Inference latency (one H100 NVL, batch 1, 10 denoising steps, prefix cache; from the observation to the action chunk, including preprocessing):
| Model | Model time | Total |
|---|---|---|
| FLUX.2-4B | 156 ms | 170 ms |
| FLUX.2-9B | 202 ms | 215 ms |
| Qwen-Image-2.1 | 237 ms | 258 ms |
Limitations
- Tasks: the models are trained on four tasks with one fixed instruction each. Other instructions are untested.
- Embodiment: the data come from one G1 setup with retrofitted Dex1-1 grippers and non-standard camera mounting. Other setups are untested.
- Evaluation: offline only. Closed-loop success on the robot is not reported.
License
Each directory carries the license of the backbone it derives from. The weights inherit the license of their base model.
g1_lgg100_flux2_klein_4b_base_imagewam/: derived from FLUX.2 [klein] base 4B, Apache License 2.0 (LICENSE.mdin the directory).g1_lgg100_flux2_klein_9b_base_imagewam/: a Derivative of FLUX.2 [klein] base 9B under the FLUX Non-Commercial License v2.1 (LICENSE.mdandNOTICEin the directory).- Non-commercial use only.
- Any rights to use it are granted directly by Black Forest Labs Inc. under that license.
- The FLUX Model was modified (fine-tuned, with an action expert added).
- It is not an official product of, nor endorsed by, Black Forest Labs Inc.
g1_lgg100_qwenimage21_imagewam/: derived from Qwen-Image-2.1 under the Qwen Research License Agreement (LICENSEandNOTICEin the directory). Use is limited to research and evaluation.- Notice: Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved.
- Shared components:
- The FLUX.2 runs also use the FLUX.2-dev VAE, under the FLUX Non-Commercial License, which is not included here.
- The training data are the public LGG100 datasets, which declare no license.
- The ImageWAM code is MIT-licensed.
IN NO EVENT SHALL THE AUTHORS OR THE LICENSORS OF THE BASE MODELS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY ARISING FROM, OUT OF OR IN CONNECTION WITH THE USE OF THESE WEIGHTS. The weights are provided as is, without warranty of any kind.
Citation
@article{zhang2026imagewam,
title = {ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?},
author = {Zhang, Yuyang and Zhang, Wenyao and Qi, Zekun and Zhang, He and Lin, Haitao and Zhang, Jingbo and
Mu, Yao and Yang, Xiaokang and Zeng, Wenjun and Jin, Xin},
journal = {arXiv preprint arXiv:2606.19531},
year = {2026}
}
Model tree for JingwuLuo/ImageWAM3backbones
Base model
Qwen/Qwen-Image-2.1