Instructions to use Aether258/pi05_single_insert_01_step2000 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Aether258/pi05_single_insert_01_step2000 with LeRobot:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
|
Download README.md from Aether258/pi05_single_insert_01_step2000: direct link, hf CLI and curl.
- Browser
- Download file 4.12 kB
-
https://huggingface.co/Aether258/pi05_single_insert_01_step2000/resolve/main/README.md
- Command line
-
hf download hf://Aether258/pi05_single_insert_01_step2000/README.md
-
curl -L -o README.md https://huggingface.co/Aether258/pi05_single_insert_01_step2000/resolve/main/README.md
4.12 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| library_name: lerobot | |
| pipeline_tag: robotics | |
| tags: | |
| - lerobot | |
| - robotics | |
| - manipulation | |
| - pi0 | |
| - single-arm | |
| # pi05_single — insert_01, step 2000 | |
| openpi `pi05_single` checkpoint for a **single-arm, vision-only** insertion task. | |
| Step 2000 (1.13 epochs), held-out validation loss **0.1063**. | |
| This is the best checkpoint of the run -- validation loss got monotonically worse afterwards. | |
| ## Configuration | |
| This run differs from the sibling `pi05_bi_*` models in two ways that make their | |
| loss values **not comparable**: | |
| | | `pi05_single` (this) | `pi05_bi` (two_tubes / task2 / bread) | | |
| |---|---|---| | |
| | state / action dim | 7 / 10 | bimanual dims | | |
| | image streams used | **1** (`camera0` only) | 6 (2 RGB + 4 tactile) | | |
| | lr schedule | constant 2e-4 after 1k warmup | cosine 2.5e-5 -> 2.5e-6 | | |
| The dataset ships six image streams (`camera0`, `camera1`, and four tactile | |
| sensors) but `pi05_single` reads only `camera0`; `camera1` and all tactile inputs | |
| are discarded. | |
| ## Data | |
| `KaiyueChen/insert_01`: 500 episodes / 252,795 frames, LeRobot v2.1, 30 fps, | |
| images embedded in the parquet files. | |
| **One frame per episode was dropped during preparation.** The collection pipeline | |
| writes an all-zero action vector on each episode's terminal frame (it has no | |
| successor action). Because the action vector contains two quaternions whose `w` | |
| components sit at ~0.99999 with a q01-q99 spread of only 2e-4, quantile | |
| normalization mapped those zeros to about -10,000, and the step-0 training loss | |
| came out at **193,510** while validation read 1.78 (validation always draws the | |
| same leading batches and happened to miss the terminal frames). Dropping the | |
| terminal frame removes 500 of 252,795 frames (0.2%), all of them invalid, and | |
| brings step-0 training loss to 0.74. Merged frame count: 252,295. | |
| Split, held out per source (10%, seed 42): train 450 / val_seen 50 / val_unseen 50. | |
| Normalization statistics (quantile q01/q99) over `train` only. | |
| ## Training | |
| | | | | |
| |---|---| | |
| | config | `pi05_single` | | |
| | hardware | 2 x A100-80GB, FSDP | | |
| | batch size | 128 | | |
| | this checkpoint | step 2000 (~1.13 epoch; 1 epoch = 1,764 steps) | | |
| | run length | 6,000 steps (3.40 epochs), 4h08m at 2.4 s/it | | |
| | lr | **constant 2e-4** after 1,000 warmup steps | | |
| | | (`pi05_single` sets `peak_lr = decay_lr = 2e-4`, so the cosine schedule does not actually decay) | | |
| | LoRA | rank 16 on the LLM, rank 32 on the action expert | | |
| | vision tower | **fully fine-tuned** -- the freeze filter matches only `.*llm.*` | | |
| ## Validation curve | |
| Flow-matching loss, 20 batches per split, same leading batches every evaluation. | |
| | step | train | val_seen | val_unseen | gap | | |
| |---:|---:|---:|---:|---:| | |
| | 0 | 0.7447 | 0.7284 | 0.7377 | +0.0093 | | |
| | **2000** | 0.0987 | 0.0860 | **0.1063** | +0.0203 | | |
| | 4000 | 0.0881 | 0.0750 | 0.1164 | +0.0414 | | |
| | 5999 | 0.0813 | 0.0673 | 0.1192 | +0.0519 | | |
| `val_unseen` was already at its minimum by the first validation point (step 2000, | |
| 1.13 epochs) and got monotonically worse afterwards, while `val_seen` improved | |
| monotonically. The gap grew from +0.0203 to +0.0519 (+156%) across three | |
| consecutive points -- an unambiguous overfitting signature, not sampling noise. | |
| For contrast, the sibling six-input `pi05_bi` runs on this codebase turned much | |
| later and ended with far smaller gaps: two_tubes at ~2.5 epochs (+0.0220), task2 | |
| at ~2 epochs (+0.0041), bread at ~2.4 epochs (-0.0037). This run, with a single | |
| camera and no tactile input, turned before 1.1 epochs and ended an order of | |
| magnitude wider. The most likely explanation is the reduced observation space: | |
| with one view and no touch, memorizing training trajectories is an easier way to | |
| lower the loss than learning a transferable policy. `pi05_single_vitac` | |
| (`camera0` + two tactile streams) would be the natural comparison, and insertion | |
| is a task where contact feedback plausibly matters. | |
| ## Contents | |
| ``` | |
| checkpoint/ | |
| params/ # inference weights | |
| train_state/ # optimizer state, for resuming | |
| assets/insert_01/ | |
| norm_stats.json # computed over the train split only | |
| ``` | |