WasserMan: paper models
Project website · Code · Paper · Demonstrations
Original final-budget policies accompanying WasserMan. This release contains 69 checkpoints:
54 ACT/DP/chunked-BC models for the six-task wm-open-v2-20260929 primary comparison,
six additional end-effector-interface DP models for PushSlider/PullLever,
and nine SmolVLA models for PressButton/PushSlider/PullLever. Every training seed
(17, 43, 101), including failed and zero-success policies, is retained.
The checkpoint payload is 22.98 GiB before small configuration and evidence files.
These models use the open-procedural-v1 asset profile. Historical CAD-profile policies and later experimental models belong to different versioned studies. Original checkpoint, configuration, completion-record and training-loss bytes are preserved. The release includes recorded result summaries and per-episode SmolVLA result CSVs. Full closed-loop rollout archives are not included in this HF release.
Download one model
pip install huggingface_hub
hf download dancher00/WasserMan-models --include README.md release-index.json 'tools/*' --local-dir WasserMan-models
python WasserMan-models/tools/download_model.py --model PressButton-ACT-17 --output WasserMan-models
The tool downloads only the four files for the requested model and verifies their
SHA-256 values. Use --list to see every model. Checkpoint configurations and
normalization statistics are also embedded in final.pt; inspect safely with:
import torch
checkpoint = torch.load('WasserMan-models/models/PressButton-ACT-17/final.pt', map_location='cpu', weights_only=True)
print(checkpoint['config']['task'], checkpoint['config']['model'])
print(checkpoint['statistics'])
ACT, BC and SmolVLA contain state/action statistics. DP stores its normalizers inside the state dictionary. Preserve the task-specific image transform, clock, history and action interface from the runtime; a weight file alone does not define the correct observation/action contract. These are native PyTorch checkpoint payloads; use the provided runtime loaders rather than a Transformers AutoModel.
Run in the published simulator
Install the public runtime
and its pinned policy dependencies. The public source snapshot used by the release
is commit 1c3b3108d31f46cb7fa2df9c6f5bded496144bda. From that installed checkout:
export WASMAN_ASSET_PROFILE=open-procedural-v1
.venv/bin/python /path/to/WasserMan-models/tools/run_policy.py --runtime . --weights /path/to/WasserMan-models --model PressButton-ACT-17 --purpose development --seeds 31000 --steps 16 --output artifacts/hf-button-check
This bounded example verifies loading and stepping. It does not measure task
success. For a full original primary cohort use --purpose test and omit --steps
and --seeds; the launcher supplies the complete task horizon and all 30 ordered
test resets. Use a new output directory for each run. The six extra EE models are
interface-study policies; evaluate them with --purpose research and explicit
paired reset IDs from the interface protocol, rather than treating them as extra
primary models. --dry-run prints the exact command without starting simulation.
SmolVLA delta reconstruction
The nine SmolVLA files contain trainable parameter deltas, 99,880,992 parameter
elements each. Reconstruct them from the pinned lerobot/smolvla_base revision
d9f33c94a60fb382c90dea2164c96845bd955e28 and tokenizer/config revision
HuggingFaceTB/SmolVLM2-500M-Video-Instruct@7b375e1b73b11138ff12fe22c8f2822d8fe03467.
The setup helper downloads those public dependencies with hash checks, installs
the recorded small text dependency, and copies five original audited rollout/
adapter scripts into the installed runtime, refusing changed existing files.
hf download dancher00/WasserMan-models --include 'runtime-extension/*' 'provenance/pretrained-dependencies.json' --local-dir WasserMan-models
.venv/bin/python /path/to/WasserMan-models/tools/setup_smolvla.py --runtime .
python /path/to/WasserMan-models/tools/download_model.py --model PushSlider-SmolVLA-17 --output /path/to/WasserMan-models
.venv/bin/python /path/to/WasserMan-models/tools/run_policy.py --runtime . --weights /path/to/WasserMan-models --model PushSlider-SmolVLA-17 --purpose development --seeds 31000 --steps 16 --output artifacts/hf-smolvla-check
The original SmolVLA configs inherit several ACT-template bookkeeping fields.
provenance/SmolVLA-metadata-clarification.json, the SmolVLA environment receipts
and executed trainer sources describe the actual parameter counts, worker count
and source identity; the original configuration bytes are retained as provenance.
The original checkpoints do not contain optimizer-resume state.
Recorded primary results
Success percentage, mean ± sample SD across three independent trainings, each evaluated on 30 resets on the primary workstation. These are original scores; simulator/renderer/numerical differences can change outcomes in new executions. Cross-workstation repetitions remain separate from the primary comparison.
| Task | ACT | DP | Chunked BC |
|---|---|---|---|
| PressButton | 45.6 ± 5.1 | 60.0 ± 12.0 | 40.0 ± 5.8 |
| RotateValve | 84.4 ± 24.1 | 84.4 ± 6.9 | 74.4 ± 5.1 |
| OpenHatch | 81.1 ± 16.8 | 100.0 ± 0.0 | 16.7 ± 17.6 |
| CollectShell | 0.0 ± 0.0 | 33.3 ± 12.0 | 0.0 ± 0.0 |
| PushSlider | 10.0 ± 12.0 | 42.2 ± 10.2 | 2.2 ± 1.9 |
| PullLever | 40.0 ± 38.4 | 51.1 ± 6.9 | 24.4 ± 15.0 |
Recorded SmolVLA test results
| Task | Successes for training seeds 17 / 43 / 101 |
|---|---|
| PressButton | 15/30 / 18/30 / 16/30 |
| PushSlider | 5/30 / 1/30 / 2/30 |
| PullLever | 4/30 / 1/30 / 2/30 |
The constant task instructions do not test language generalization. Pretraining,
compute and observation encoders differ from ACT/DP/BC. All 270 test episodes and
72 validation episodes were physically rescored; the result summaries and
per-episode CSVs are included under results/.
Integrity and scope
release-index.json lists each checkpoint and its configuration, completion
record and training metrics with sizes and SHA-256 values.
Each checkpoint hash was verified against its original completion record before
publication. SHA256SUMS also covers the tools, protocols, summaries and source
extension. The demonstration dataset has its own unchanged archive manifest.
Historical absolute paths in scientific receipts identify original inputs; the
published download/setup/launch tools accept paths for the current machine.
Simulation success does not establish transfer to a physical underwater robot.
The retained license and third-party notices apply.
Citation
@misc{belov2026wasserman,
title = {{WasserMan}: Benchmark for Underwater Manipulation Policy Learning},
author = {Belov, Danil and Erkhov, Artem and Parsegov, Sergei and Osinenko, Pavel},
year = {2026},
eprint = {2610.04536},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2610.04536}
}
