File size: 2,210 Bytes
bb40549
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
---
license: apache-2.0
library_name: pytorch
base_model: Wan-AI/Wan2.2-TI2V-5B
tags:
- robotics
- world-action-model
- action-chunking
- piper
- fastwam
---

# SpaceDreamer Piper WAM

This repository contains the runtime artifacts for the Piper RGB-state WAM
policy. Given one `cam_high` RGB frame, the current 14D dual-arm qpos, and one
of four supported task IDs, the policy generates a 32-step action chunk.

Inference code and documentation:
[qshou-coder/SpaceDreamer, `piper_inference` branch](https://github.com/qshou-coder/SpaceDreamer/tree/piper_inference)

## Artifacts

| File | Purpose |
|---|---|
| `policy.pt` | Complete RGB DiT, ActionDiT, state encoder, and action normalizer |
| `Wan2.2_VAE.pth` | Wan2.2 VAE used to encode the live RGB frame |
| `prompt_contexts.pt` | Frozen UMT5 contexts for the four supported tasks |
| `inference_config.yaml` | Portable model and input/output contract |
| `SHA256SUMS` | Artifact integrity checksums |

The policy checkpoint already contains the action normalization buffers. It
does not require the Wan base DiT shards or the UMT5 encoder at runtime.

## Input/output contract

- Image: `uint8[H,W,3]`, from `cam_high`, explicitly marked RGB or BGR.
- State: raw, unnormalized qpos with shape `[14]`.
- State/action order: left joints 1-6, left gripper, right joints 1-6,
  right gripper.
- Output: `float32[32,14]` absolute joint-position commands at 25 Hz.
- Sampling: joint RGB/action flow matching, 20 inference steps, shift 5.

Supported task IDs:

- `assemble_battery_long`
- `battery_assemble`
- `pack_3_objects_plus`
- `stack_3_cups_gen`

See the GitHub usage guide for the Python API and offline CLI.

## Safety

The model only predicts actions. It does not enforce robot joint limits,
velocity limits, observation freshness, collision avoidance, or emergency
stop state. A robot-specific safety/control layer must validate every chunk
before execution.

## Upstream attribution

The included VAE is redistributed from
[`Wan-AI/Wan2.2-TI2V-5B`](https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B),
which is licensed under Apache-2.0. The inference implementation also uses
FastWAM components distributed with the accompanying GitHub code under MIT.