|
Download README.md from HorizonRobotics/Ego4WAM: direct link, hf CLI and curl.
- Browser
- Download file 2.29 kB
-
https://huggingface.co/HorizonRobotics/Ego4WAM/resolve/main/README.md
- Command line
-
hf download hf://HorizonRobotics/Ego4WAM/README.md
-
curl -L -o README.md https://huggingface.co/HorizonRobotics/Ego4WAM/resolve/main/README.md
2.29 kB
| license: apache-2.0 | |
| library_name: starvla | |
| tags: | |
| - robotics | |
| - vision-language-action | |
| - world-model | |
| - robodojo | |
| base_model: | |
| - Wan-AI/Wan2.2-TI2V-5B-Diffusers | |
| - google/umt5-xxl | |
| # Ego4WAM-Joint | |
| Released weights for the Joint variant of Ego4WAM. | |
| ## Contents | |
| | File | Size | Description | | |
| | --- | ---: | --- | | |
| | `model.safetensors` | 25.07 GiB | 2,092 BF16 tensors containing the complete framework state dict | | |
| | `inference_config.yaml` | <1 KiB | Minimal model selection: `Ego4WAM` with `interaction_mode: joint` | | |
| | `model_assets/` | 20.46 MiB | VAE/text-encoder configs and the UMT5 tokenizer required for offline construction | | |
| | `LICENSE` | | Apache-2.0 license for the released weights | | |
| The checkpoint contains the Video DiT, VAE, UMT5 text encoder, Action DiT, and | |
| proprioceptive encoder weights. No additional backbone weights are downloaded | |
| when loading this release. | |
| ## Loading | |
| ```python | |
| from huggingface_hub import snapshot_download | |
| from starVLA.model.framework.base_framework import baseframework | |
| root = snapshot_download("HorizonRobotics/Ego4WAM") | |
| model = baseframework.from_pretrained(f"{root}/model.safetensors") | |
| ``` | |
| This constructs the framework from `inference_config.yaml`, resolves the local | |
| files in `model_assets/`, and loads the state dict with `strict=True`. | |
| ## Model geometry | |
| | | | | |
| | --- | --- | | |
| | Tensors / dtype | 2,092 / BF16 | | |
| | Parameters | 13,456,608,092 | | |
| | State-dict prefixes | `backbone.` 1,266; `action_model.` 824; `proprio_encoder.` 2 | | |
| | Interaction mode | Joint synchronous video/action attention | | |
| | State / action dimension | 32 / 32 | | |
| | Action horizon | 32 | | |
| | RoboDojo execution output | First 14 dimensions after inverse normalization | | |
| | MoT | 30 layers; video hidden size 3,072; action hidden size 1,024 | | |
| | Sampling | 20-step Euler flow matching | | |
| | Camera input | Head, left wrist, and right wrist RGB views | | |
| The SHA-256 digest of `model.safetensors` is: | |
| ```text | |
| ae1b4595e1909e4573ad3214e1c5112f1920a34675340f6adfd4f2e9ec6c12c4 | |
| ``` | |
| ## License and attribution | |
| The released weights are provided under Apache-2.0. Ego4WAM source code is | |
| provided under MIT. Ego4WAM builds on StarVLA and uses Wan2.2 TI2V-5B and | |
| UMT5-XXL components; retain the corresponding upstream attribution when | |
| redistributing derivative work. | |