xvla-base-p150

X-VLA-base (the cross-embodiment vision-language-action model, lerobot checkpoint: Florence-2 encoder + soft-prompted flow-matching transformer) running on one Tenstorrent Blackhole p150a via tt-nn: 3 camera views + an instruction + proprio state in, a 30-step x 20-D action chunk out. Weights: lerobot/xvla-base · Paper: arXiv:2510.10274 · Upstream code: 2toinf/X-VLA · Port: changh95/tt-XVLA

Runs on p150 (mesh P150).

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  changh95/xvla-base-p150 --with-weights
tt-model serve changh95/xvla-base-p150
  • Weights lerobot/xvla-base at cdb7964e4fe8 go to your HF cache; the image does not contain them.
  • Serves on port 20000 (or the next free port); ready when the log says Application startup complete.

Run with tt-cli

tt serve changh95/xvla-base-p150
IMG=$(base64 -w0 media/pusht_synthetic.png)
printf '{"images":["%s","%s","%s"],"instruction":"push the T","state":[0,0,0,0,0,0,0,0],"seed":42}' "$IMG" "$IMG" "$IMG" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/xvla-base-p150
  • POST /predict: images (1 or 3 base64 PNG/JPEG views; 1 is copied into all 3 slots), instruction (or task); optional state (1–20 floats, default 8 zeros), domain_id (0), num_denoising_steps (1, max 50), seed.
  • GET /health, GET /info, POST /reset.

Response

{"actions": [[-0.04, -0.11, 0.258, 0.146, -0.197, 0.146, 0.025, -0.019, -0.414, 0.406, ...], ...],
 "chunk_size": 30, "n_action_steps": 30, "action_dim": 20, "action_space": "ee6d", "normalized": false,
 "num_denoising_steps": 1, "seed": 42, "domain_id": 0, "instruction": "push the T", "language_tokens": 32,
 "state_dim": 8, "views_used": ["image", "image2", "image3"], "single_view_replicated": false,
 "input_size": [224, 224], "image_sizes": [{"width": 256, "height": 256}, ...],
 "timing_ms": {"preprocess": 3.82, "inference": 102.21, "total": 106.04}}
  • actions is 30 rows of [x, y, z, r1..r6 (6-D rotation), gripper, 0 x 10] in raw model space (normalized: false): the base checkpoint ships no dataset statistics, so the values are not directly usable on a robot without fine-tuning.
  • Views are ImageNet-normalised and resize_with_pad-ed to 224x224; the instruction is tokenised with the vendored facebook/bart-large tokenizer to a fixed 32 tokens.

Demo

Input (media/pusht_synthetic.png, synthetic push-T scene, sent as all 3 views)

Accuracy and speed

Metric Value
Action-chunk PCC vs fp32 torch (5 seeds, 10 denoising steps) 0.999982 (min 0.999979, max abs err 8.15e-03)
Open-loop action MAE vs fp32 torch (lerobot/pusht_image, 10 samples, 10 denoising steps) 253.78 vs 253.78 (delta +4.73e-04, +0.00%)
Inference, served over HTTP (warm, 3 views, 1 denoising step, 30 requests) ~90–109 ms inference (median 94.4) · ~93–111 ms end-to-end per 30-step chunk
Inference, served over HTTP (warm, 3 views, 10 denoising steps, 30 requests) ~174 ms inference (min 168 / max 181) · ~177 ms end-to-end
Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) 15.5 / 15.1 ms (1 step) and 43.5 / 43.7 ms (10 steps) → GPU 6.3× / 4.0× faster than the p150a's 97.4 / 175.8 ms; fp32-strict 34.6 / 95.4 ms (2.8× / 1.8×); best torch.compile 12.8 / 38.1 ms

Caveats

  • Fixed shapes: 3 views at 224x224, 32 language tokens, batch 1; one request at a time (a lock serialises the device); TT_MESH_SHAPE must be 1x1.
  • Default num_denoising_steps=1 is a speed setting (upstream uses 10); it costs <0.05% open-loop PCC but closed-loop flow-matching policies usually want 4–10 steps. Override per request.
  • This is the BASE checkpoint, meant for fine-tuning: outputs are raw model-space actions, not normalised to any robot or dataset.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • Validated on tt-metal v0.71.0-dev20260509-4 (main 2a6ddd8e572) with lerobot 0.5.0 / transformers 5.4.0, single p150a only.
  • Device path: fused ttnn ops (SDPA, minimal_matmul, matmul+residual, LN+residual, DaViT window permutes on device) with the 24-block transformer and the BART encoder replayed as metal traces (TT_FUSED, default on; TT_FUSED=0 = the 2026-09-12 legacy op chain, ~177 ms).
  • GPU comparison: GPU bf16 is 6.3× (1 step) / 4.0× (10 steps) faster; the p150a path is bound by the DaViT vision tower (48 eager ttnn calls plus ~21 ms of CPU depthwise convs, ~73 of 87 ms), not by the diffusion steps. RTX 5090 rows (2026-09-14): same host, the port's own torch reference (same weights) run eagerly in PyTorch 2.11 cu128 with fp32 weights + autocast unless stated, no TensorRT; medians of 50 iterations after warm-up, H2D/D2H included; the p150a rows are the served bf16 fused path incl. upload/readback. p150a power was not measured, so no efficiency comparison is made. Full table: GPU_COMPARISON.md.

Licensing

  • Weights: lerobot/xvla-base, Apache-2.0; fetched from the Hub at serve time, not redistributed here.
  • Port and serving code (code/): Apache-2.0, from changh95/tt-XVLA; it patches the lerobot 0.5.0 X-VLA policy (Apache-2.0), and code/tt/assets/bart-large-tokenizer/ vendors the facebook/bart-large tokenizer files (Apache-2.0).

Provenance

The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal 2a6ddd8e572bb09b236a2adbd3afab1153e0a17e
code/ digest 57c9edfc340cc2ac (sha256, first 16 hex digits)
built 2026-09-13T15:42:11+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for changh95/xvla-base-p150

Finetuned
(62)
this model

Collection including changh95/xvla-base-p150

Paper for changh95/xvla-base-p150