Add LIBERO Cosmos Policy Predict2-2B checkpoint
Browse filesMirror the verified upstream model, configuration, statistics, text embeddings, original model card and license under libero/cosmos-2. Update dataset indexes and document model-specific licenses.
- README.md +18 -6
- checkpoint_manifest.json +30 -0
- libero/README.md +4 -1
- libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt +3 -0
- libero/cosmos-2/LICENSE +35 -0
- libero/cosmos-2/README.md +20 -0
- libero/cosmos-2/SHA256SUMS +8 -0
- libero/cosmos-2/UPSTREAM_README.md +236 -0
- libero/cosmos-2/config.json +67 -0
- libero/cosmos-2/libero_dataset_statistics.json +102 -0
- libero/cosmos-2/libero_t5_embeddings.pkl +3 -0
- libero/cosmos-2/provenance.json +62 -0
README.md
CHANGED
|
@@ -1,10 +1,15 @@
|
|
| 1 |
---
|
| 2 |
pretty_name: RIFT Checkpoints
|
| 3 |
-
license:
|
|
|
|
|
|
|
| 4 |
library_name: pytorch
|
| 5 |
-
base_model:
|
|
|
|
|
|
|
| 6 |
datasets:
|
| 7 |
- yuanty/LIBERO-fastwam
|
|
|
|
| 8 |
tags:
|
| 9 |
- robotics
|
| 10 |
- robot-learning
|
|
@@ -27,16 +32,22 @@ Each dataset has one top-level directory. Model variants live inside that direct
|
|
| 27 |
|
| 28 |
| Dataset | Directory | Models |
|
| 29 |
| --- | --- | --- |
|
| 30 |
-
| LIBERO | [libero](https://huggingface.co/PoopBear/RIFT/tree/main/libero) | FastWAM; RIFT, FastWAM-IDM and FastWAM-Joint (step 21,700) |
|
| 31 |
| RoboCOIN, three tasks | [robocoin_multitask_mm](https://huggingface.co/PoopBear/RIFT/tree/main/robocoin_multitask_mm) | RIFT (final), 10 epochs / step 9,340 |
|
| 32 |
| Galaxea indoor cleaning | [galaxea_indoor_cleaning_3cam224](https://huggingface.co/PoopBear/RIFT/tree/main/galaxea_indoor_cleaning_3cam224) | FastWAM, FastWAM-Joint and RIFT, 10 epochs / step 14,830 |
|
| 33 |
|
| 34 |
-
Shared VAE, text-encoder, and tokenizer files are under `assets/`. Exact checkpoint paths,
|
| 35 |
hashes, and source revisions are listed in [checkpoint_manifest.json](checkpoint_manifest.json).
|
| 36 |
Use each model's own normalization statistics.
|
| 37 |
|
| 38 |
## Download
|
| 39 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
LIBERO FastWAM:
|
| 41 |
|
| 42 |
```bash
|
|
@@ -79,5 +90,6 @@ with `weights_only=True`.
|
|
| 79 |
|
| 80 |
## License
|
| 81 |
|
| 82 |
-
|
| 83 |
-
|
|
|
|
|
|
| 1 |
---
|
| 2 |
pretty_name: RIFT Checkpoints
|
| 3 |
+
license: other
|
| 4 |
+
license_name: model-specific-licenses
|
| 5 |
+
license_link: https://huggingface.co/PoopBear/RIFT/blob/main/README.md#license
|
| 6 |
library_name: pytorch
|
| 7 |
+
base_model:
|
| 8 |
+
- Wan-AI/Wan2.2-TI2V-5B
|
| 9 |
+
- nvidia/Cosmos-Predict2-2B-Video2World
|
| 10 |
datasets:
|
| 11 |
- yuanty/LIBERO-fastwam
|
| 12 |
+
- nvidia/LIBERO-Cosmos-Policy
|
| 13 |
tags:
|
| 14 |
- robotics
|
| 15 |
- robot-learning
|
|
|
|
| 32 |
|
| 33 |
| Dataset | Directory | Models |
|
| 34 |
| --- | --- | --- |
|
| 35 |
+
| LIBERO | [libero](https://huggingface.co/PoopBear/RIFT/tree/main/libero) | FastWAM; RIFT, FastWAM-IDM and FastWAM-Joint (step 21,700); Cosmos Policy Predict2-2B |
|
| 36 |
| RoboCOIN, three tasks | [robocoin_multitask_mm](https://huggingface.co/PoopBear/RIFT/tree/main/robocoin_multitask_mm) | RIFT (final), 10 epochs / step 9,340 |
|
| 37 |
| Galaxea indoor cleaning | [galaxea_indoor_cleaning_3cam224](https://huggingface.co/PoopBear/RIFT/tree/main/galaxea_indoor_cleaning_3cam224) | FastWAM, FastWAM-Joint and RIFT, 10 epochs / step 14,830 |
|
| 38 |
|
| 39 |
+
Shared Wan VAE, text-encoder, and tokenizer files are under `assets/`. Exact checkpoint paths,
|
| 40 |
hashes, and source revisions are listed in [checkpoint_manifest.json](checkpoint_manifest.json).
|
| 41 |
Use each model's own normalization statistics.
|
| 42 |
|
| 43 |
## Download
|
| 44 |
|
| 45 |
+
LIBERO Cosmos-2:
|
| 46 |
+
|
| 47 |
+
```bash
|
| 48 |
+
hf download PoopBear/RIFT --include 'libero/cosmos-2/*' --local-dir ./checkpoints
|
| 49 |
+
```
|
| 50 |
+
|
| 51 |
LIBERO FastWAM:
|
| 52 |
|
| 53 |
```bash
|
|
|
|
| 90 |
|
| 91 |
## License
|
| 92 |
|
| 93 |
+
RIFT checkpoints are distributed under the Apache License 2.0 (see `LICENSE`). RIFT source code is distributed separately under the MIT License. See `NOTICE` for attribution.
|
| 94 |
+
|
| 95 |
+
The third-party Cosmos Policy files in `libero/cosmos-2/` retain their original NVIDIA One-Way Noncommercial License (NSCLv1). See [the included license](libero/cosmos-2/LICENSE) and [original model card](libero/cosmos-2/UPSTREAM_README.md).
|
checkpoint_manifest.json
CHANGED
|
@@ -137,6 +137,36 @@
|
|
| 137 |
"checkpoint": "libero_uncond_2cam224.pt",
|
| 138 |
"dataset_stats": "libero_uncond_2cam224_dataset_stats.json"
|
| 139 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 140 |
}
|
| 141 |
]
|
| 142 |
}
|
|
|
|
| 137 |
"checkpoint": "libero_uncond_2cam224.pt",
|
| 138 |
"dataset_stats": "libero_uncond_2cam224_dataset_stats.json"
|
| 139 |
}
|
| 140 |
+
},
|
| 141 |
+
{
|
| 142 |
+
"dataset": "libero",
|
| 143 |
+
"variant": "cosmos-2",
|
| 144 |
+
"model_name": "Cosmos-Policy-LIBERO-Predict2-2B",
|
| 145 |
+
"checkpoint": "libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt",
|
| 146 |
+
"bytes": 3913017345,
|
| 147 |
+
"sha256": "8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2",
|
| 148 |
+
"step": 40000,
|
| 149 |
+
"step_source": "upstream config.json training.gradient_steps",
|
| 150 |
+
"config": "libero/cosmos-2/config.json",
|
| 151 |
+
"config_kind": "original upstream model metadata",
|
| 152 |
+
"dataset_stats": "libero/cosmos-2/libero_dataset_statistics.json",
|
| 153 |
+
"text_embeddings": "libero/cosmos-2/libero_t5_embeddings.pkl",
|
| 154 |
+
"source": {
|
| 155 |
+
"repo": "nvidia/Cosmos-Policy-LIBERO-Predict2-2B",
|
| 156 |
+
"revision": "cb689ec0e3347c13667d70a78a3447388f5c3bb8",
|
| 157 |
+
"path": "Cosmos-Policy-LIBERO-Predict2-2B.pt"
|
| 158 |
+
},
|
| 159 |
+
"license": {
|
| 160 |
+
"name": "NVIDIA One-Way Noncommercial License (NSCLv1)",
|
| 161 |
+
"path": "libero/cosmos-2/LICENSE",
|
| 162 |
+
"source": {
|
| 163 |
+
"repo": "NVlabs/HMAR",
|
| 164 |
+
"revision": "7e17e31191ee0287c3043e3a8523c23ee3d355dd",
|
| 165 |
+
"url": "https://github.com/NVlabs/HMAR/blob/7e17e31191ee0287c3043e3a8523c23ee3d355dd/LICENSE",
|
| 166 |
+
"bytes": 4060,
|
| 167 |
+
"sha256": "6d1daa9c89ec421ba80fe051ac9f0dd46010d4fa4e8f4af654eb8e62f64e2ad5"
|
| 168 |
+
}
|
| 169 |
+
}
|
| 170 |
}
|
| 171 |
]
|
| 172 |
}
|
libero/README.md
CHANGED
|
@@ -2,9 +2,12 @@
|
|
| 2 |
|
| 3 |
| Model | Directory | Step |
|
| 4 |
| --- | --- | --- |
|
|
|
|
| 5 |
| FastWAM | [fastwam](fastwam/) | Not recorded |
|
| 6 |
| RIFT | [rift](rift/) | 21,700 |
|
| 7 |
| FastWAM-IDM | [idm](idm/) | 21,700 |
|
| 8 |
| FastWAM-Joint | [joint](joint/) | 21,700 |
|
| 9 |
|
| 10 |
-
Each directory contains the original weights, matching normalization statistics,
|
|
|
|
|
|
|
|
|
| 2 |
|
| 3 |
| Model | Directory | Step |
|
| 4 |
| --- | --- | --- |
|
| 5 |
+
| Cosmos Policy Predict2-2B | [cosmos-2](cosmos-2/) | 40,000 (upstream config) |
|
| 6 |
| FastWAM | [fastwam](fastwam/) | Not recorded |
|
| 7 |
| RIFT | [rift](rift/) | 21,700 |
|
| 8 |
| FastWAM-IDM | [idm](idm/) | 21,700 |
|
| 9 |
| FastWAM-Joint | [joint](joint/) | 21,700 |
|
| 10 |
|
| 11 |
+
Each directory contains the original weights, matching normalization statistics, configuration files, provenance, and SHA-256 checksums. RIFT's configuration is a model configuration; the FastWAM, IDM and Joint configs are workspace references.
|
| 12 |
+
|
| 13 |
+
Cosmos-2 includes the original upstream model metadata, cached text embeddings, model card, and NVIDIA license.
|
libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2
|
| 3 |
+
size 3913017345
|
libero/cosmos-2/LICENSE
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
NVIDIA License
|
| 2 |
+
|
| 3 |
+
1. Definitions
|
| 4 |
+
|
| 5 |
+
“Licensor” means any person or entity that distributes its Work.
|
| 6 |
+
“Work” means (a) the original work of authorship made available under this license, which may include software, documentation, or other files, and (b) any additions to or derivative works thereof that are made available under this license.
|
| 7 |
+
The terms “reproduce,” “reproduction,” “derivative works,” and “distribution” have the meaning as provided under U.S. copyright law; provided, however, that for the purposes of this license, derivative works shall not include works that remain separable from, or merely link (or bind by name) to the interfaces of, the Work.
|
| 8 |
+
Works are “made available” under this license by including in or with the Work either (a) a copyright notice referencing the applicability of this license to the Work, or (b) a copy of this license.
|
| 9 |
+
|
| 10 |
+
2. License Grant
|
| 11 |
+
|
| 12 |
+
2.1 Copyright Grant. Subject to the terms and conditions of this license, each Licensor grants to you a perpetual, worldwide, non-exclusive, royalty-free, copyright license to use, reproduce, prepare derivative works of, publicly display, publicly perform, sublicense and distribute its Work and any resulting derivative works in any form.
|
| 13 |
+
|
| 14 |
+
3. Limitations
|
| 15 |
+
|
| 16 |
+
3.1 Redistribution. You may reproduce or distribute the Work only if (a) you do so under this license, (b) you include a complete copy of this license with your distribution, and (c) you retain without modification any copyright, patent, trademark, or attribution notices that are present in the Work.
|
| 17 |
+
|
| 18 |
+
3.2 Derivative Works. You may specify that additional or different terms apply to the use, reproduction, and distribution of your derivative works of the Work (“Your Terms”) only if (a) Your Terms provide that the use limitation in Section 3.3 applies to your derivative works, and (b) you identify the specific derivative works that are subject to Your Terms. Notwithstanding Your Terms, this license (including the redistribution requirements in Section 3.1) will continue to apply to the Work itself.
|
| 19 |
+
|
| 20 |
+
3.3 Use Limitation. The Work and any derivative works thereof only may be used or intended for use non-commercially. Notwithstanding the foregoing, NVIDIA Corporation and its affiliates may use the Work and any derivative works commercially. As used herein, “non-commercially” means for non-commercial research and educational purposes only.
|
| 21 |
+
|
| 22 |
+
3.4 Patent Claims. If you bring or threaten to bring a patent claim against any Licensor (including any claim, cross-claim or counterclaim in a lawsuit) to enforce any patents that you allege are infringed by any Work, then your rights under this license from such Licensor (including the grant in Section 2.1) will terminate immediately.
|
| 23 |
+
|
| 24 |
+
3.5 Trademarks. This license does not grant any rights to use any Licensor’s or its affiliates’ names, logos, or trademarks, except as necessary to reproduce the notices described in this license.
|
| 25 |
+
|
| 26 |
+
3.6 Termination. If you violate any term of this license, then your rights under this license (including the grant in Section 2.1) will terminate immediately.
|
| 27 |
+
|
| 28 |
+
4. Disclaimer of Warranty.
|
| 29 |
+
|
| 30 |
+
THE WORK IS PROVIDED “AS IS” WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING WARRANTIES OR CONDITIONS OF
|
| 31 |
+
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE OR NON-INFRINGEMENT. YOU BEAR THE RISK OF UNDERTAKING ANY ACTIVITIES UNDER THIS LICENSE.
|
| 32 |
+
|
| 33 |
+
5. Limitation of Liability.
|
| 34 |
+
|
| 35 |
+
EXCEPT AS PROHIBITED BY APPLICABLE LAW, IN NO EVENT AND UNDER NO LEGAL THEORY, WHETHER IN TORT (INCLUDING NEGLIGENCE), CONTRACT, OR OTHERWISE SHALL ANY LICENSOR BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES ARISING OUT OF OR RELATED TO THIS LICENSE, THE USE OR INABILITY TO USE THE WORK (INCLUDING BUT NOT LIMITED TO LOSS OF GOODWILL, BUSINESS INTERRUPTION, LOST PROFITS OR DATA, COMPUTER FAILURE OR MALFUNCTION, OR ANY OTHER DAMAGES OR LOSSES), EVEN IF THE LICENSOR HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES.
|
libero/cosmos-2/README.md
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# LIBERO Cosmos-2
|
| 2 |
+
|
| 3 |
+
This directory mirrors **Cosmos-Policy-LIBERO-Predict2-2B**, the NVIDIA Cosmos Policy Predict2-2B checkpoint for LIBERO. All original model files match both the verified server bundle and the pinned upstream revision.
|
| 4 |
+
|
| 5 |
+
- Weights: `Cosmos-Policy-LIBERO-Predict2-2B.pt`.
|
| 6 |
+
- Original model metadata: `config.json`.
|
| 7 |
+
- Normalization: `libero_dataset_statistics.json`.
|
| 8 |
+
- Cached text embeddings: `libero_t5_embeddings.pkl`.
|
| 9 |
+
- Original model card: [UPSTREAM_README.md](UPSTREAM_README.md).
|
| 10 |
+
- File hashes and source revisions: `SHA256SUMS` and `provenance.json`.
|
| 11 |
+
|
| 12 |
+
Upstream: [nvidia/Cosmos-Policy-LIBERO-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-LIBERO-Predict2-2B/tree/cb689ec0e3347c13667d70a78a3447388f5c3bb8). Use the upstream Cosmos Policy code and its required pretrained assets for inference.
|
| 13 |
+
|
| 14 |
+
```bash
|
| 15 |
+
hf download PoopBear/RIFT --include 'libero/cosmos-2/*' --local-dir ./checkpoints
|
| 16 |
+
```
|
| 17 |
+
|
| 18 |
+
## License
|
| 19 |
+
|
| 20 |
+
These Cosmos files retain the NVIDIA One-Way Noncommercial License (NSCLv1), as specified by the upstream model card. A complete copy is included in [LICENSE](LICENSE). The original model card and attribution notices are preserved.
|
libero/cosmos-2/SHA256SUMS
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2 Cosmos-Policy-LIBERO-Predict2-2B.pt
|
| 2 |
+
3bd882474403177ab22e65598d8ab1c1b299804b5bff0019f30c4a01aa405506 UPSTREAM_README.md
|
| 3 |
+
6c246e6588762546fd91c4fac62af570583da1156aa9b2800f2268e6a2c31f43 config.json
|
| 4 |
+
5b119a98ad7824507ddff3c6c7507ca244261ee416dcfa623876702160c580d3 libero_dataset_statistics.json
|
| 5 |
+
8a03499676c6c196127c577144b9fd09bb02c30f9caf058adbcef09bb99ad8f5 libero_t5_embeddings.pkl
|
| 6 |
+
6d1daa9c89ec421ba80fe051ac9f0dd46010d4fa4e8f4af654eb8e62f64e2ad5 LICENSE
|
| 7 |
+
ab9b5f6d5f126b63bcf5f20c24594f36ae2bb490ffaf625845af8f689adb9a5c README.md
|
| 8 |
+
bd742a925ea01fda83403adc90aa915b96fdfed700dd374b8212a455194da70d provenance.json
|
libero/cosmos-2/UPSTREAM_README.md
ADDED
|
@@ -0,0 +1,236 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model:
|
| 3 |
+
- nvidia/Cosmos-Predict2-2B-Video2World
|
| 4 |
+
---
|
| 5 |
+
# **Cosmos-Policy-LIBERO-Predict2-2B**
|
| 6 |
+
|
| 7 |
+
[**Cosmos Policy**](https://huggingface.co/collections/nvidia/cosmos-policy) | [**Code**](http://github.com/NVlabs/cosmos-policy) | [**White Paper**](https://arxiv.org/abs/2601.16163) | [**Website**](https://research.nvidia.com/labs/dir/cosmos-policy/)
|
| 8 |
+
|
| 9 |
+
# Model Overview
|
| 10 |
+
|
| 11 |
+
## Description:
|
| 12 |
+
|
| 13 |
+
Cosmos-Policy-LIBERO-Predict2-2B is a 2B-parameter robot manipulation policy model fine-tuned from the [NVIDIA Cosmos-Predict2-2B-Video2World](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World) video foundation model. This model achieves state-of-the-art performance on the LIBERO simulation benchmark with a 98.5% average success rate across four task suites.
|
| 14 |
+
|
| 15 |
+
Key features:
|
| 16 |
+
|
| 17 |
+
* **Single-stage fine-tuning**: Adapted from pretrained video model with no architectural modifications
|
| 18 |
+
* **Multimodal outputs**: Jointly predicts actions, future states, and values through unified video diffusion
|
| 19 |
+
* **High performance**: 98.5% average success rate on LIBERO (Spatial: 98.1%, Object: 100.0%, Goal: 98.2%, Long: 97.6%)
|
| 20 |
+
|
| 21 |
+
Use cases:
|
| 22 |
+
|
| 23 |
+
* Robotic manipulation and control in simulation environments
|
| 24 |
+
* Imitation learning and policy learning for table-top manipulation tasks
|
| 25 |
+
* Vision-based robot learning with multiple camera viewpoints
|
| 26 |
+
* Long-horizon task planning and execution
|
| 27 |
+
* Lifelong learning and transfer learning in robotics
|
| 28 |
+
|
| 29 |
+
This model is for research and development only.
|
| 30 |
+
|
| 31 |
+
**Model Developer**: NVIDIA
|
| 32 |
+
|
| 33 |
+
## Model Versions
|
| 34 |
+
|
| 35 |
+
Cosmos Policy models include the following:
|
| 36 |
+
|
| 37 |
+
- [Cosmos-Policy-LIBERO-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-LIBERO-Predict2-2B): Given current state observations and a task description, generate action sequences, future state predictions, and value estimates for robot manipulation in simulated LIBERO environments.
|
| 38 |
+
- [Cosmos-Policy-RoboCasa-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-RoboCasa-Predict2-2B): Given current state observations and a task description, generate action sequences, future state predictions, and value estimates for robot manipulation in simulated RoboCasa environments.
|
| 39 |
+
- [Cosmos-Policy-ALOHA-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-ALOHA-Predict2-2B): Given current state observations and a task description, generate action sequences, future state predictions, and value estimates for robot manipulation in real-world ALOHA robot environments.
|
| 40 |
+
- [Cosmos-Policy-ALOHA-Planning-Model-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Policy-ALOHA-Planning-Model-Predict2-2B): Given current state observations, a task description, and action sequences, generate future state predictions and value estimates for robot manipulation in real-world ALOHA robot environments. (This checkpoint is meant to be deployed alongside Cosmos-Policy-ALOHA-Predict2-2B, not independently.)
|
| 41 |
+
|
| 42 |
+
### License:
|
| 43 |
+
|
| 44 |
+
This model is released under the [NVIDIA One-Way Noncommercial License (NSCLv1)](https://github.com/NVlabs/HMAR/blob/main/LICENSE). For a custom license, please contact [cosmos-license@nvidia.com](mailto:cosmos-license@nvidia.com).
|
| 45 |
+
|
| 46 |
+
Under the NVIDIA One-Way Noncommercial License (NSCLv1), NVIDIA confirms:
|
| 47 |
+
|
| 48 |
+
* Models are not for commercial use.
|
| 49 |
+
* NVIDIA does not claim ownership to any outputs generated using the Models or Derivative Models.
|
| 50 |
+
|
| 51 |
+
### Deployment Geography:
|
| 52 |
+
|
| 53 |
+
Global
|
| 54 |
+
|
| 55 |
+
### Use Case:
|
| 56 |
+
|
| 57 |
+
Physical AI: Robot manipulation and control, encompassing tabletop manipulation and imitation learning in simulation environments.
|
| 58 |
+
|
| 59 |
+
### Release Date:
|
| 60 |
+
|
| 61 |
+
GitHub [01/22/2026] via [https://github.com/nvlabs/cosmos-policy](https://github.com/nvlabs/cosmos-policy)
|
| 62 |
+
|
| 63 |
+
Hugging Face [01/22/2026] via [https://huggingface.co/collections/nvidia/cosmos-policy](https://huggingface.co/collections/nvidia/cosmos-policy)
|
| 64 |
+
|
| 65 |
+
## Model Architecture:
|
| 66 |
+
|
| 67 |
+
Architecture Type: A diffusion transformer with latent video diffusion, fine-tuned from Cosmos-Predict2-2B-Video2World.
|
| 68 |
+
|
| 69 |
+
Network Architecture: The model uses the same architecture as the base [Cosmos-Predict2-2B](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World) model (a diffusion transformer with latent video diffusion).
|
| 70 |
+
|
| 71 |
+
**Key adaptation**: Actions, proprioceptive states, and values are encoded as latent frames and injected directly into the video model's latent diffusion sequence, enabling the model to generate these modalities alongside predicted future images.
|
| 72 |
+
|
| 73 |
+
**Number of model parameters:**
|
| 74 |
+
|
| 75 |
+
2B (inherited from base model)
|
| 76 |
+
|
| 77 |
+
## Input
|
| 78 |
+
|
| 79 |
+
**Input Type(s)**: Text + Multi-view Images + Proprioceptive State
|
| 80 |
+
|
| 81 |
+
**Input Format(s)**:
|
| 82 |
+
|
| 83 |
+
* Text: String (natural language task description)
|
| 84 |
+
* Images: RGB images from multiple camera views
|
| 85 |
+
* Proprioception: Numerical array
|
| 86 |
+
|
| 87 |
+
**Input Parameters**:
|
| 88 |
+
|
| 89 |
+
* Text: One-dimensional (1D) - Task description (e.g., "put the black bowl on top of the cabinet")
|
| 90 |
+
* Images: Two-dimensional (2D) - Third-person camera (agentview): 224×224 RGB; Wrist-mounted camera (eye-in-hand): 224×224 RGB
|
| 91 |
+
* Proprioception: One-dimensional (1D) - 9-dimensional state (2 gripper joints + 3 end-effector position + 4 end-effector quaternion)
|
| 92 |
+
|
| 93 |
+
**Other Properties Related to Input**:
|
| 94 |
+
|
| 95 |
+
* Requires specific camera configuration (third-person + wrist views)
|
| 96 |
+
* Images resized to 224×224 pixels from original resolution
|
| 97 |
+
* Trained exclusively for Franka Emika Panda robot arm in LIBERO simulation environments
|
| 98 |
+
|
| 99 |
+
## Output
|
| 100 |
+
|
| 101 |
+
**Output Type(s)**: Action Sequence + Future State Predictions + Value Estimate
|
| 102 |
+
|
| 103 |
+
**Output Format**:
|
| 104 |
+
|
| 105 |
+
* Actions: Numerical array
|
| 106 |
+
* Future states: Images + Proprioception
|
| 107 |
+
* Value: Scalar
|
| 108 |
+
|
| 109 |
+
**Output Parameters**:
|
| 110 |
+
|
| 111 |
+
* Action chunk: 16-timestep sequence of 7-dimensional actions (6-DoF end-effector control + 1 gripper)
|
| 112 |
+
* Future robot proprioception: 9-dimensional state at timestep t+16
|
| 113 |
+
* Future state images: Third-person camera prediction (224×224 RGB) and wrist camera prediction (224×224 RGB) at timestep t+16
|
| 114 |
+
* Future state value: Expected cumulative reward from future state (scalar)
|
| 115 |
+
|
| 116 |
+
**Other Properties Related to Output**:
|
| 117 |
+
|
| 118 |
+
* Action chunk size: 16 timesteps
|
| 119 |
+
* Denoising steps: 5 (configurable without retraining)
|
| 120 |
+
* Noise level range: σ_min = 4.0, σ_max = 80.0
|
| 121 |
+
* Generation mode: Parallel (action, future state, and value generated simultaneously)
|
| 122 |
+
|
| 123 |
+
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
|
| 124 |
+
|
| 125 |
+
## Software Integration
|
| 126 |
+
|
| 127 |
+
**Runtime Engine(s):**
|
| 128 |
+
|
| 129 |
+
* [Transformers](https://github.com/huggingface/transformers)
|
| 130 |
+
|
| 131 |
+
**Supported Hardware Microarchitecture Compatibility:**
|
| 132 |
+
|
| 133 |
+
* NVIDIA Hopper (e.g., H100)
|
| 134 |
+
|
| 135 |
+
**Note**: We have only tested doing inference with BF16 precision.
|
| 136 |
+
|
| 137 |
+
**Operating System(s):**
|
| 138 |
+
|
| 139 |
+
* Linux
|
| 140 |
+
|
| 141 |
+
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
|
| 142 |
+
|
| 143 |
+
# Usage
|
| 144 |
+
|
| 145 |
+
See [Cosmos Policy GitHub](http://github.com/NVlabs/cosmos-policy) for details.
|
| 146 |
+
|
| 147 |
+
## Training and Evaluation Sections:
|
| 148 |
+
|
| 149 |
+
### Training Datasets:
|
| 150 |
+
|
| 151 |
+
**Data Collection Method**:
|
| 152 |
+
|
| 153 |
+
* LIBERO-Cosmos-Policy: Hybrid: Human - Human-teleoperated demonstrations recorded in simulation environment
|
| 154 |
+
|
| 155 |
+
**Labeling Method**:
|
| 156 |
+
|
| 157 |
+
* LIBERO-Cosmos-Policy: Automated - Success/failure labels automatically determined by simulation environment evaluation; task descriptions from benchmark specification
|
| 158 |
+
|
| 159 |
+
##### Properties:
|
| 160 |
+
|
| 161 |
+
**Training Data**: [LIBERO-Cosmos-Policy](https://huggingface.co/datasets/nvidia/LIBERO-Cosmos-Policy) dataset
|
| 162 |
+
|
| 163 |
+
- 4 task suites: LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, LIBERO-Long
|
| 164 |
+
- 500 demonstrations per suite (50 demos × 10 tasks)
|
| 165 |
+
- Successful demonstrations used for policy training
|
| 166 |
+
- All demonstrations (including failures) used for world model and value function training
|
| 167 |
+
|
| 168 |
+
**Training Configuration**:
|
| 169 |
+
|
| 170 |
+
- **Base model**: NVIDIA Cosmos-Predict2-2B-Video2World (`model-480p-16fps.pt`)
|
| 171 |
+
- **Training steps**: 40,000 gradient steps
|
| 172 |
+
- **Batch size**: 1,920 (global)
|
| 173 |
+
- **GPUs**: 64 H100 GPUs
|
| 174 |
+
- **Training time**: ~48 hours
|
| 175 |
+
- **Optimization**: Full model fine-tuning (all weights updated)
|
| 176 |
+
- **Action chunk size**: 16 timesteps
|
| 177 |
+
- **Image resolution**: 224×224 pixels
|
| 178 |
+
|
| 179 |
+
**Training Objective**: The model is trained with a hybrid log-normal-uniform noise distribution (modified from the base model's log-normal distribution; see paper for details) to improve action prediction accuracy. Training batches are split 50/25/25 for policy, world model, and value function objectives, respectively.
|
| 180 |
+
|
| 181 |
+
### Evaluation Datasets:
|
| 182 |
+
|
| 183 |
+
Data Collection Method: Not Applicable
|
| 184 |
+
|
| 185 |
+
Labeling Method: Not Applicable
|
| 186 |
+
|
| 187 |
+
Properties: Not Applicable - we use the LIBERO simulation environments for direct evaluations.
|
| 188 |
+
|
| 189 |
+
## Inference:
|
| 190 |
+
|
| 191 |
+
**Test Hardware:** H100, A100
|
| 192 |
+
|
| 193 |
+
See [Cosmos Policy GitHub](http://github.com/NVlabs/cosmos-policy) for details.
|
| 194 |
+
|
| 195 |
+
#### System Requirements and Performance
|
| 196 |
+
|
| 197 |
+
Inference with base Cosmos Policy only (i.e., no model-based planning):
|
| 198 |
+
|
| 199 |
+
* 1 GPU with 6.8 GB VRAM for LIBERO sim benchmark tasks
|
| 200 |
+
* 1 GPU with 8.9 GB VRAM for RoboCasa sim benchmark tasks
|
| 201 |
+
* 1 GPU with 6.0 GB VRAM for ALOHA robot tasks
|
| 202 |
+
|
| 203 |
+
#### Quality Benchmarks
|
| 204 |
+
|
| 205 |
+
### LIBERO Benchmark Results
|
| 206 |
+
|
| 207 |
+
| Task Suite | Success Rate |
|
| 208 |
+
| ----------------- | --------------- |
|
| 209 |
+
| LIBERO-Spatial | 98.1% |
|
| 210 |
+
| LIBERO-Object | 100.0% |
|
| 211 |
+
| LIBERO-Goal | 98.2% |
|
| 212 |
+
| LIBERO-Long | 97.6% |
|
| 213 |
+
| **Average** | **98.5%** |
|
| 214 |
+
|
| 215 |
+
Success rates are averaged over 500 trials per suite (10 tasks × 50 episodes) across 3 random seeds (6,000 trials total).
|
| 216 |
+
|
| 217 |
+
## Ethical Considerations
|
| 218 |
+
|
| 219 |
+
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
|
| 220 |
+
|
| 221 |
+
Users are responsible for model inputs and outputs. Users are responsible for ensuring safe integration of this model, including implementing guardrails as well as other safety mechanisms, prior to deployment.
|
| 222 |
+
|
| 223 |
+
Please report security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
|
| 224 |
+
|
| 225 |
+
## Related Resources
|
| 226 |
+
|
| 227 |
+
- **Base Model**: [Cosmos-Predict2-2B-Video2World](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Video2World)
|
| 228 |
+
- **Training Dataset**: [LIBERO-Cosmos-Policy](https://huggingface.co/datasets/nvidia/LIBERO-Cosmos-Policy)
|
| 229 |
+
- **Paper**: [Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning](https://arxiv.org/abs/2601.16163)
|
| 230 |
+
- **Original LIBERO**: [LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning](https://arxiv.org/abs/2306.03310)
|
| 231 |
+
|
| 232 |
+
## Citation
|
| 233 |
+
|
| 234 |
+
If you use this model, please cite the Cosmos Policy paper:
|
| 235 |
+
|
| 236 |
+
(Cosmos Policy BibTeX citation coming soon!)
|
libero/cosmos-2/config.json
ADDED
|
@@ -0,0 +1,67 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "cosmos-policy",
|
| 3 |
+
"architecture": "diffusion-transformer",
|
| 4 |
+
"base_model": "nvidia/Cosmos-Predict2-2B-Video2World",
|
| 5 |
+
"num_parameters": "2B",
|
| 6 |
+
|
| 7 |
+
"input_spec": {
|
| 8 |
+
"text": {
|
| 9 |
+
"type": "string",
|
| 10 |
+
"description": "Natural language task description"
|
| 11 |
+
},
|
| 12 |
+
"images": {
|
| 13 |
+
"format": "RGB",
|
| 14 |
+
"resolution": [224, 224],
|
| 15 |
+
"views": ["agentview", "eye_in_hand"]
|
| 16 |
+
},
|
| 17 |
+
"proprioception": {
|
| 18 |
+
"dim": 9,
|
| 19 |
+
"components": ["gripper_joints", "end_effector_position", "quaternion"]
|
| 20 |
+
}
|
| 21 |
+
},
|
| 22 |
+
|
| 23 |
+
"output_spec": {
|
| 24 |
+
"actions": {
|
| 25 |
+
"dim": 7,
|
| 26 |
+
"horizon": 16,
|
| 27 |
+
"components": ["end_effector_6dof", "gripper"]
|
| 28 |
+
},
|
| 29 |
+
"future_proprioception": {
|
| 30 |
+
"dim": 9
|
| 31 |
+
},
|
| 32 |
+
"future_images": {
|
| 33 |
+
"resolution": [224, 224]
|
| 34 |
+
},
|
| 35 |
+
"value": {
|
| 36 |
+
"dim": 1
|
| 37 |
+
}
|
| 38 |
+
},
|
| 39 |
+
|
| 40 |
+
"diffusion_config": {
|
| 41 |
+
"denoising_steps": 5,
|
| 42 |
+
"sigma_min": 4.0,
|
| 43 |
+
"sigma_max": 80.0,
|
| 44 |
+
"generation_mode": "parallel"
|
| 45 |
+
},
|
| 46 |
+
|
| 47 |
+
"training": {
|
| 48 |
+
"dataset": "LIBERO-Cosmos-Policy",
|
| 49 |
+
"gradient_steps": 40000,
|
| 50 |
+
"batch_size": 1920,
|
| 51 |
+
"hardware": "64x H100",
|
| 52 |
+
"action_chunk_size": 16
|
| 53 |
+
},
|
| 54 |
+
|
| 55 |
+
"benchmark_results": {
|
| 56 |
+
"libero_spatial": 0.981,
|
| 57 |
+
"libero_object": 1.0,
|
| 58 |
+
"libero_goal": 0.982,
|
| 59 |
+
"libero_long": 0.976,
|
| 60 |
+
"average": 0.985
|
| 61 |
+
},
|
| 62 |
+
|
| 63 |
+
"inference": {
|
| 64 |
+
"precision": "bf16",
|
| 65 |
+
"vram_gb": 6.8
|
| 66 |
+
}
|
| 67 |
+
}
|
libero/cosmos-2/libero_dataset_statistics.json
ADDED
|
@@ -0,0 +1,102 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"actions_min": [
|
| 3 |
+
-0.9375,
|
| 4 |
+
-0.9375,
|
| 5 |
+
-0.9375,
|
| 6 |
+
-0.2582142949104309,
|
| 7 |
+
-0.375,
|
| 8 |
+
-0.3642857074737549,
|
| 9 |
+
-1.0
|
| 10 |
+
],
|
| 11 |
+
"actions_max": [
|
| 12 |
+
0.9375,
|
| 13 |
+
0.9375,
|
| 14 |
+
0.9375,
|
| 15 |
+
0.3557142913341522,
|
| 16 |
+
0.375,
|
| 17 |
+
0.375,
|
| 18 |
+
1.0
|
| 19 |
+
],
|
| 20 |
+
"actions_mean": [
|
| 21 |
+
0.0627586841583252,
|
| 22 |
+
0.08706478029489517,
|
| 23 |
+
-0.09038306772708893,
|
| 24 |
+
0.0004062088264618069,
|
| 25 |
+
0.005638255272060633,
|
| 26 |
+
-0.004925117362290621,
|
| 27 |
+
-0.0539495050907135
|
| 28 |
+
],
|
| 29 |
+
"actions_std": [
|
| 30 |
+
0.3358899652957916,
|
| 31 |
+
0.3785707950592041,
|
| 32 |
+
0.44441089034080505,
|
| 33 |
+
0.03959346562623978,
|
| 34 |
+
0.06336744129657745,
|
| 35 |
+
0.07815399020910263,
|
| 36 |
+
0.9980894923210144
|
| 37 |
+
],
|
| 38 |
+
"actions_median": [
|
| 39 |
+
0.0,
|
| 40 |
+
0.0,
|
| 41 |
+
-0.05624999850988388,
|
| 42 |
+
0.0,
|
| 43 |
+
0.0,
|
| 44 |
+
0.0,
|
| 45 |
+
-1.0
|
| 46 |
+
],
|
| 47 |
+
"proprio_min": [
|
| 48 |
+
-0.005054043605923653,
|
| 49 |
+
-0.042120561003685,
|
| 50 |
+
-0.48564884066581726,
|
| 51 |
+
-0.33136284351348877,
|
| 52 |
+
0.008128181099891663,
|
| 53 |
+
0.2902947962284088,
|
| 54 |
+
-0.8897988200187683,
|
| 55 |
+
-0.5616180300712585,
|
| 56 |
+
-0.5712438225746155
|
| 57 |
+
],
|
| 58 |
+
"proprio_max": [
|
| 59 |
+
0.04238177835941315,
|
| 60 |
+
0.0013513736193999648,
|
| 61 |
+
0.2103137969970703,
|
| 62 |
+
0.3904264271259308,
|
| 63 |
+
1.3660907745361328,
|
| 64 |
+
0.9999998807907104,
|
| 65 |
+
0.9094555377960205,
|
| 66 |
+
0.34964922070503235,
|
| 67 |
+
0.6364591121673584
|
| 68 |
+
],
|
| 69 |
+
"proprio_mean": [
|
| 70 |
+
0.026981549337506294,
|
| 71 |
+
-0.027262669056653976,
|
| 72 |
+
-0.04627704992890358,
|
| 73 |
+
0.03411973640322685,
|
| 74 |
+
0.7628858089447021,
|
| 75 |
+
0.9422642588615417,
|
| 76 |
+
-0.05906042829155922,
|
| 77 |
+
-0.04012482985854149,
|
| 78 |
+
-0.0029695192351937294
|
| 79 |
+
],
|
| 80 |
+
"proprio_std": [
|
| 81 |
+
0.014144331216812134,
|
| 82 |
+
0.014038173481822014,
|
| 83 |
+
0.10425098240375519,
|
| 84 |
+
0.15156729519367218,
|
| 85 |
+
0.37788110971450806,
|
| 86 |
+
0.11462057381868362,
|
| 87 |
+
0.2741202116012573,
|
| 88 |
+
0.10206353664398193,
|
| 89 |
+
0.0958578810095787
|
| 90 |
+
],
|
| 91 |
+
"proprio_median": [
|
| 92 |
+
0.03365287557244301,
|
| 93 |
+
-0.03424428775906563,
|
| 94 |
+
-0.02911928854882717,
|
| 95 |
+
0.02497534640133381,
|
| 96 |
+
0.9323337078094482,
|
| 97 |
+
0.9929837584495544,
|
| 98 |
+
-0.023691711947321892,
|
| 99 |
+
-0.027851801365613937,
|
| 100 |
+
0.0008408676367253065
|
| 101 |
+
]
|
| 102 |
+
}
|
libero/cosmos-2/libero_t5_embeddings.pkl
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8a03499676c6c196127c577144b9fd09bb02c30f9caf058adbcef09bb99ad8f5
|
| 3 |
+
size 41957938
|
libero/cosmos-2/provenance.json
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"dataset": "libero",
|
| 3 |
+
"variant": "cosmos-2",
|
| 4 |
+
"model_name": "Cosmos-Policy-LIBERO-Predict2-2B",
|
| 5 |
+
"checkpoint": "libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt",
|
| 6 |
+
"bytes": 3913017345,
|
| 7 |
+
"sha256": "8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2",
|
| 8 |
+
"step": 40000,
|
| 9 |
+
"step_source": "upstream config.json training.gradient_steps",
|
| 10 |
+
"config": "libero/cosmos-2/config.json",
|
| 11 |
+
"config_kind": "original upstream model metadata",
|
| 12 |
+
"dataset_stats": "libero/cosmos-2/libero_dataset_statistics.json",
|
| 13 |
+
"text_embeddings": "libero/cosmos-2/libero_t5_embeddings.pkl",
|
| 14 |
+
"source": {
|
| 15 |
+
"repo": "nvidia/Cosmos-Policy-LIBERO-Predict2-2B",
|
| 16 |
+
"revision": "cb689ec0e3347c13667d70a78a3447388f5c3bb8",
|
| 17 |
+
"path": "Cosmos-Policy-LIBERO-Predict2-2B.pt"
|
| 18 |
+
},
|
| 19 |
+
"license": {
|
| 20 |
+
"name": "NVIDIA One-Way Noncommercial License (NSCLv1)",
|
| 21 |
+
"path": "libero/cosmos-2/LICENSE",
|
| 22 |
+
"source": {
|
| 23 |
+
"repo": "NVlabs/HMAR",
|
| 24 |
+
"revision": "7e17e31191ee0287c3043e3a8523c23ee3d355dd",
|
| 25 |
+
"url": "https://github.com/NVlabs/HMAR/blob/7e17e31191ee0287c3043e3a8523c23ee3d355dd/LICENSE",
|
| 26 |
+
"bytes": 4060,
|
| 27 |
+
"sha256": "6d1daa9c89ec421ba80fe051ac9f0dd46010d4fa4e8f4af654eb8e62f64e2ad5"
|
| 28 |
+
}
|
| 29 |
+
},
|
| 30 |
+
"files": [
|
| 31 |
+
{
|
| 32 |
+
"source_name": "Cosmos-Policy-LIBERO-Predict2-2B.pt",
|
| 33 |
+
"path": "libero/cosmos-2/Cosmos-Policy-LIBERO-Predict2-2B.pt",
|
| 34 |
+
"bytes": 3913017345,
|
| 35 |
+
"sha256": "8818528d8c9150cda0ddf8c711b0f221b21dac8ac379bd26d5690235954d33e2"
|
| 36 |
+
},
|
| 37 |
+
{
|
| 38 |
+
"source_name": "README.md",
|
| 39 |
+
"path": "libero/cosmos-2/UPSTREAM_README.md",
|
| 40 |
+
"bytes": 11233,
|
| 41 |
+
"sha256": "3bd882474403177ab22e65598d8ab1c1b299804b5bff0019f30c4a01aa405506"
|
| 42 |
+
},
|
| 43 |
+
{
|
| 44 |
+
"source_name": "config.json",
|
| 45 |
+
"path": "libero/cosmos-2/config.json",
|
| 46 |
+
"bytes": 1354,
|
| 47 |
+
"sha256": "6c246e6588762546fd91c4fac62af570583da1156aa9b2800f2268e6a2c31f43"
|
| 48 |
+
},
|
| 49 |
+
{
|
| 50 |
+
"source_name": "libero_dataset_statistics.json",
|
| 51 |
+
"path": "libero/cosmos-2/libero_dataset_statistics.json",
|
| 52 |
+
"bytes": 2372,
|
| 53 |
+
"sha256": "5b119a98ad7824507ddff3c6c7507ca244261ee416dcfa623876702160c580d3"
|
| 54 |
+
},
|
| 55 |
+
{
|
| 56 |
+
"source_name": "libero_t5_embeddings.pkl",
|
| 57 |
+
"path": "libero/cosmos-2/libero_t5_embeddings.pkl",
|
| 58 |
+
"bytes": 41957938,
|
| 59 |
+
"sha256": "8a03499676c6c196127c577144b9fd09bb02c30f9caf058adbcef09bb99ad8f5"
|
| 60 |
+
}
|
| 61 |
+
]
|
| 62 |
+
}
|