You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

pi05_plus checkpoints

pi05_plus is a pi0.5 policy with a progress head and a discrete action head โ€” an autoregressive model of the action chunk's FAST codes, which gives an exact likelihood over actions. Trained on RoboPRO demonstrations.

Every folder is self-contained: params and an assets/roboreal_lerobot/ holding all three files serving reads. The two base checkpoints also carry train_state, so they can be resumed from; the RL folders are serving checkpoints only.

Base checkpoints

Scored on the whole RoboPRO grid โ€” all 80 tasks, each task's clean scene and its ten clutter levels, 3,146 episodes, under the per-episode-keyed protocol.

folder warm start steps clean clutter all (SR / HSR)
pi05_plus_18k5_warm RoboPRO's jax_30000 18,500 71.9 / 66.5 64.8 / 46.4 68.3 / 56.4
pi05_plus_30k_scratch pi0.5 base 29,999 70.2 / 65.8 61.8 / 43.3 66.0 / 54.5

RoboPRO's own jax_30000 scores 61.3 / 50.8 on the same grid. SR is the benchmark's success; HSR is success without a collision.

RL post-training

One round of KTO on the policy's own rollouts, starting from pi05_plus_18k5_warm, 3,000 steps. The three differ only in which rollouts they learned from. Same whole-grid measurement.

folder rollouts it learned from all (SR / HSR)
rl_kto_alltasks_3000 1,735 plain rollouts, every task and clutter level 69.5 / 55.6
rl_ktofans_alltasks_3000 1,753 six-branch fans, same tasks and seeds 68.4 / 55.3
rl_kto_8task_3000 960 fan rollouts over 8 clean tasks 70.0 / 56.0

Each gains about a point of success over its 68.3 starting point and gives back about a point of hard success, the loss concentrated in clutter. The wider pools did not beat the narrow one.

Cherry-picked checkpoints

Read these differently from the rows above. They are the best of 25 checkpoints on a 317-episode probe โ€” all 80 tasks at two clean seeds and two d9 clutter seeds โ€” and were chosen on the very episodes their numbers come from. That selection is optimistic, the probe is a twentieth of the grid, and neither has a whole-grid number. They are here because each is the strongest model found for one half of the benchmark, not because they are better overall.

folder picked for clean d9 clutter probe all
cherry_picked_clean_1250 clean scenes 81.0 / 75.9 62.9 / 47.8 71.9 / 61.8
cherry_picked_clutter_2500 cluttered scenes 72.2 / 69.0 71.1 / 50.3 71.6 / 59.6

The same 317 episodes put pi05_plus_18k5_warm at 70.0 / 58.4, clean 74.7 / 68.4, d9 65.4 / 48.4. The two picks sit at opposite ends: the checkpoints strongest on clean scenes tend to be weakest in clutter, and the reverse.

Assets each folder carries

file read by if it is missing
norm_stats.json the normalize / unnormalize transforms state and actions are scaled wrong
actions_per_timestep.npz PerTimestepActions and its inverse the policy returns normalized numbers as if they were actions
action_correlation.npy noise_cholesky, for the correlated noise the flow head samples model construction raises, since correlated_noise needs it

The config resolves the last two through assets_dirs, which names an assets directory rather than the checkpoint. Running elsewhere, point ROBORESEARCH_ASSETS at a directory holding the folder's own assets/, or copy the two files into the directory the config names.

Verified: pi05_plus_18k5_warm downloaded fresh, with an assets directory built from nothing but its own assets/, reproduces the published evaluation on 151 of 151 episodes โ€” identical verdicts, not merely an identical rate.

Downloading

RoboResearch's scripts/download_checkpoint.py takes the repo, where it lands under $ROBORESEARCH_CHECKPOINTS, and the folder to fetch.

uv run python scripts/download_checkpoint.py mahgoobi/pi05_plus pi05_plus_18k5_warm pi05_plus_18k5_warm

To go on training from a base checkpoint's state, place it as its step of a run and resume:

uv run python scripts/download_checkpoint.py mahgoobi/pi05_plus \
    pi05_plus_robopro_jax30000/pi05_plus_18k5_warm/18500 pi05_plus_18k5_warm
CUDA_VISIBLE_DEVICES=<gpus> uv run python -m roboresearch.policies.pi05_plus.train \
    pi05_plus_robopro_jax30000 --exp-name=pi05_plus_18k5_warm --resume

pi05_plus_30k_scratch the same way, with pi05_plus_robopro and step 29999. A resumed run restores onto whatever cards it has, not only the ones it was saved on.

Code: the pi05_plus policy in RoboResearch, on openpi.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading