Instructions to use PruhaNLP/TurboVLA-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use PruhaNLP/TurboVLA-base with LeRobot:
- Notebooks
- Google Colab
- Kaggle
TurboVLA-base (SO-100)
A ready-to-fine-tune TurboVLA checkpoint for the SO-100 / SO-101 arm, in the spirit of lerobot/smolvla_base.
No manual model assembly: one folder holds the config, the full weights (DINOv3 ViT-B + BERT + interaction + action decoder) and the tokenizer.
TurboVLA drops the LLM from the VLA loop. DINOv3 encodes each camera, BERT encodes the instruction, 6 bidirectional vision-language cross-attention layers fuse them, and an ACT-style decoder predicts a 12-step continuous action chunk in one pass (~0.2B params, L1 behavior cloning).
How it was made
Built by model/build_turbovla_base.py from the official LIBERO checkpoint
H-EmbodVis/TurboVLA/checkpoints/libero/turbovla_libero.pth, the same starting point the paper uses for real-robot fine-tuning.
Every tensor is copied as is, except the embodiment-specific ones, which are freshly initialized (seed 0):
action_head.state_projection.net.0.weightaction_head.state_projection.net.0.biasaction_head.state_projection.net.1.weightaction_head.state_projection.net.1.biasaction_head.decoder.action_projection.layers.2.weightaction_head.decoder.action_projection.layers.2.bias
I/O contract
| Cameras | 2 RGB views, order front, wrist, resized + padded to 256x256, ImageNet mean/std |
| Instruction | English text, BERT uncased, padded to 32 tokens |
| State | 6-D SO-100 joint positions, normalized with mean/std |
| Action | 12 x 6 absolute joint targets, tanh output mapped from per-joint dataset min/max |
This base has no normalization stats: they come from your dataset on the first fine-tune and are saved as stats.safetensors.
The paper recipe freezes BERT and trains everything else with lr 5e-5, AdamW (0.9, 0.95), L1 loss.
Use it
In robosim-at-home: switch Mode → TurboVLA, open Train, pick PruhaNLP/TurboVLA-base, tick datasets, start.
Inference and Eval then list the fine-tuned runs.
Standalone, with model/turbovla.py:
from model.turbovla import TurboEngine
engine = TurboEngine(device="cuda")
engine.load("PruhaNLP/TurboVLA-base") # config + weights + tokenizer from this repo
# after fine-tuning (stats present):
chunk = engine.predict_chunk({"front": front_pil, "wrist": wrist_pil}, joints, "pick up the cube and put it in the tray")
Files
| File | Content |
|---|---|
config.json |
type: turbovla, cameras, sizes, interaction / action head, full DINOv3 and BERT configs |
model.safetensors |
fp32 weights; key names match the official turbovla.models.turbovla.TurboVLA module (transformers 4.x DINOv3 layout vision_encoder.backbone.layer.N) |
tokenizer.json, tokenizer_config.json |
google-bert/bert-base-uncased tokenizer |
DINOv3_LICENSE.md, LICENSE |
licenses inherited from H-EmbodVis/TurboVLA |
The weights load strictly into the official TurboVLA class built with action_dim=6, state_dim=6, num_views=2,
image_size=256, text.padding_length=32, and give the same outputs as model/turbovla.py.
License
Weights contain DINOv3-derived parameters and are distributed under the DINOv3 License; TurboVLA code is Apache-2.0.
Citation
@article{xie2026turbovla,
title = {TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM},
author = {Xie, Hengyi and Yao, Chenfei and Wu, Xianjin and Xi, Xuanyang and Tang, Yiping and Xu, Di and
Zhu, Yingying and Liang, Dingkang and Bai, Xiang and Ding, Han},
journal = {arXiv preprint arXiv:2607.27205},
year = {2026}
}
- Downloads last month
- -
Model tree for PruhaNLP/TurboVLA-base
Base model
H-EmbodVis/TurboVLA