TurboVLA-base (SO-100)

A ready-to-fine-tune TurboVLA checkpoint for the SO-100 / SO-101 arm, in the spirit of lerobot/smolvla_base. No manual model assembly: one folder holds the config, the full weights (DINOv3 ViT-B + BERT + interaction + action decoder) and the tokenizer.

TurboVLA drops the LLM from the VLA loop. DINOv3 encodes each camera, BERT encodes the instruction, 6 bidirectional vision-language cross-attention layers fuse them, and an ACT-style decoder predicts a 12-step continuous action chunk in one pass (~0.2B params, L1 behavior cloning).

How it was made

Built by model/build_turbovla_base.py from the official LIBERO checkpoint H-EmbodVis/TurboVLA/checkpoints/libero/turbovla_libero.pth, the same starting point the paper uses for real-robot fine-tuning. Every tensor is copied as is, except the embodiment-specific ones, which are freshly initialized (seed 0):

  • action_head.state_projection.net.0.weight
  • action_head.state_projection.net.0.bias
  • action_head.state_projection.net.1.weight
  • action_head.state_projection.net.1.bias
  • action_head.decoder.action_projection.layers.2.weight
  • action_head.decoder.action_projection.layers.2.bias

I/O contract

Cameras 2 RGB views, order front, wrist, resized + padded to 256x256, ImageNet mean/std
Instruction English text, BERT uncased, padded to 32 tokens
State 6-D SO-100 joint positions, normalized with mean/std
Action 12 x 6 absolute joint targets, tanh output mapped from per-joint dataset min/max

This base has no normalization stats: they come from your dataset on the first fine-tune and are saved as stats.safetensors. The paper recipe freezes BERT and trains everything else with lr 5e-5, AdamW (0.9, 0.95), L1 loss.

Use it

In robosim-at-home: switch Mode → TurboVLA, open Train, pick PruhaNLP/TurboVLA-base, tick datasets, start. Inference and Eval then list the fine-tuned runs.

Standalone, with model/turbovla.py:

from model.turbovla import TurboEngine

engine = TurboEngine(device="cuda")
engine.load("PruhaNLP/TurboVLA-base")          # config + weights + tokenizer from this repo
# after fine-tuning (stats present):
chunk = engine.predict_chunk({"front": front_pil, "wrist": wrist_pil}, joints, "pick up the cube and put it in the tray")

Files

File Content
config.json type: turbovla, cameras, sizes, interaction / action head, full DINOv3 and BERT configs
model.safetensors fp32 weights; key names match the official turbovla.models.turbovla.TurboVLA module (transformers 4.x DINOv3 layout vision_encoder.backbone.layer.N)
tokenizer.json, tokenizer_config.json google-bert/bert-base-uncased tokenizer
DINOv3_LICENSE.md, LICENSE licenses inherited from H-EmbodVis/TurboVLA

The weights load strictly into the official TurboVLA class built with action_dim=6, state_dim=6, num_views=2, image_size=256, text.padding_length=32, and give the same outputs as model/turbovla.py.

License

Weights contain DINOv3-derived parameters and are distributed under the DINOv3 License; TurboVLA code is Apache-2.0.

Citation

@article{xie2026turbovla,
  title   = {TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM},
  author  = {Xie, Hengyi and Yao, Chenfei and Wu, Xianjin and Xi, Xuanyang and Tang, Yiping and Xu, Di and
             Zhu, Yingying and Liang, Dingkang and Bai, Xiang and Ding, Han},
  journal = {arXiv preprint arXiv:2607.27205},
  year    = {2026}
}
Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Video Preview
loading

Model tree for PruhaNLP/TurboVLA-base

Finetuned
(2)
this model

Paper for PruhaNLP/TurboVLA-base