Pi0.5 SnapFlow distilled for LIBERO
This repository contains 10k- and 20k-step SnapFlow checkpoints distilled from lerobot/pi05_libero_finetuned, the 10-step Pi0.5 LIBERO teacher.
Results
Each policy was evaluated on 50 episodes for each of the 40 tasks in the four standard LIBERO suites. The 20k student completed 10 more episodes than the teacher. Its median action-chunk latency was 58.9 ms, compared with 175.4 ms for the teacher on an NVIDIA H100 PCIe.
| Checkpoint | Steps | Spatial | Object | Goal | LIBERO-10 | Overall success | Median latency |
|---|---|---|---|---|---|---|---|
| Pi0.5 teacher | 10-NFE | 98.0% | 99.0% | 96.4% | 94.8% | 1,941/2,000 (97.05%) | 175.4 ms/chunk |
| SnapFlow student | 10,000 | 97.4% | 99.8% | 96.4% | 93.6% | 1,936/2,000 (96.80%) | 58.4 ms/chunk |
| SnapFlow student | 20,000 | 98.8% | 99.8% | 96.8% | 94.8% | 1,951/2,000 (97.55%) | 58.9 ms/chunk |
The 20k student measured a 2.98× median latency speedup over the teacher in this Torch pipeline. Its success-rate difference is small and comes from one seeded evaluation, so it does not establish a statistically significant improvement or predict performance on other seeds, robots, or hardware.
Evaluation method
Studio's LiberoBenchmark evaluated Spatial, Object, Goal, and LIBERO-10 with seed 42 and no video recording. The overall score is the arithmetic mean across the four equally sized suites, equivalent to the success count across 2,000 episodes.
Latency was measured separately with Runtime's InferenceLatencyBenchmark on the exported Torch models. It timed preprocessing, model execution, and postprocessing for 100 batch-one action chunks after 5 warmup iterations, using seeded synthetic inputs matching the exported input schema. The timing excludes simulator stepping, rendering, network transport, model loading, and robot I/O. Rollout FPS is not used as an inference-latency measure.
The run used Torch 2.11.0+cu128 and Runtime 0.2.1.dev4+g0ad4548. The Studio source was based on commit 18c4c300894d8ac3f1fa0478e0af68401e520a90. The teacher was resolved at 8e174154ef5f6c60a8da12ae99c303d8963138c1; the dataset was HuggingFaceVLA/libero revision 86958911c0f959db2bbbdb107eb3e17c5f9c798e.
Training targeted 30k steps but stopped at step 27,647, after the 20k evaluation. No 30k checkpoint or result was produced.
Files
checkpoints/snapflow-step010000.ckptandcheckpoints/snapflow-step020000.ckpt: full Lightning checkpoints. The evaluated 20k checkpoint SHA-256 iseec7419542f4b0cb5e2ab57948eb5e03c7ccad141db4d679c6716ee8a050322b.results/steps-10000.jsonandresults/steps-20000.json: per-task LIBERO success rates, counts, latency statistics, and run metadata.benchmark/benchmark_snapflow_libero.py: evaluation and latency script.training/resolved-config.yaml: the resolved training config used for this run.source/library/src/physicalai/benchmark/snapflow_libero.py: checkpoint and evaluation callback.
Loading a checkpoint
Install the Physical AI Studio library with its Pi0.5 and LIBERO extras, then load the checkpoint:
from physicalai.policies import Pi05
policy = Pi05.load_from_checkpoint(
"checkpoints/snapflow-step020000.ckpt",
map_location="cpu",
compile_model=False,
)
policy.eval()
To rerun the benchmark, use the script from a matching Studio library/ checkout:
uv run --no-sync python /path/to/snapflow-libero/benchmark/benchmark_snapflow_libero.py \
--snapflow-checkpoint /path/to/snapflow-libero/checkpoints/snapflow-step020000.ckpt \
--snapflow-steps 20000 \
--episodes 50 --iterations 100 --warmup 5 \
--output results/reproduced-20k.json
License and use
This checkpoint is a Gemma Model Derivative of the Pi0.5 LIBERO teacher. The repository is gated. By requesting access, users agree to Google's Gemma Terms of Use and must follow the Gemma Prohibited Use Policy. A complete copy of the terms is included in LICENSE-GEMMA.html, and NOTICE.txt contains the required Gemma notice. This model is not offered under Apache-2.0.
Model tree for Daankrol/pi05-snapflow-libero
Base model
google/paligemma-3b-pt-224