Robotics
Collection
Robotics models that runs on Tenstorrent AI accelerators • 9 items • Updated
X-VLA-base (the cross-embodiment vision-language-action model, lerobot checkpoint: Florence-2 encoder + soft-prompted flow-matching transformer) running on one Tenstorrent Blackhole p150a via tt-nn: 3 camera views + an instruction + proprio state in, a 30-step x 20-D action chunk out. Weights: lerobot/xvla-base · Paper: arXiv:2510.10274 · Upstream code: 2toinf/X-VLA · Port: changh95/tt-XVLA
Runs on p150 (mesh P150).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
tt-model pull changh95/xvla-base-p150 --with-weights
tt-model serve changh95/xvla-base-p150
lerobot/xvla-base at cdb7964e4fe8 go to your HF cache; the image does not contain them.Application startup complete.tt serve changh95/xvla-base-p150
IMG=$(base64 -w0 media/pusht_synthetic.png)
printf '{"images":["%s","%s","%s"],"instruction":"push the T","state":[0,0,0,0,0,0,0,0],"seed":42}' "$IMG" "$IMG" "$IMG" > req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/xvla-base-p150
POST /predict: images (1 or 3 base64 PNG/JPEG views; 1 is copied into all 3 slots), instruction (or task); optional state (1–20 floats, default 8 zeros), domain_id (0), num_denoising_steps (1, max 50), seed.GET /health, GET /info, POST /reset.{"actions": [[-0.04, -0.11, 0.258, 0.146, -0.197, 0.146, 0.025, -0.019, -0.414, 0.406, ...], ...],
"chunk_size": 30, "n_action_steps": 30, "action_dim": 20, "action_space": "ee6d", "normalized": false,
"num_denoising_steps": 1, "seed": 42, "domain_id": 0, "instruction": "push the T", "language_tokens": 32,
"state_dim": 8, "views_used": ["image", "image2", "image3"], "single_view_replicated": false,
"input_size": [224, 224], "image_sizes": [{"width": 256, "height": 256}, ...],
"timing_ms": {"preprocess": 3.82, "inference": 102.21, "total": 106.04}}
actions is 30 rows of [x, y, z, r1..r6 (6-D rotation), gripper, 0 x 10] in raw model space (normalized: false): the base checkpoint ships no dataset statistics, so the values are not directly usable on a robot without fine-tuning.resize_with_pad-ed to 224x224; the instruction is tokenised with the vendored facebook/bart-large tokenizer to a fixed 32 tokens.| Metric | Value |
|---|---|
| Action-chunk PCC vs fp32 torch (5 seeds, 10 denoising steps) | 0.999982 (min 0.999979, max abs err 8.15e-03) |
| Open-loop action MAE vs fp32 torch (lerobot/pusht_image, 10 samples, 10 denoising steps) | 253.78 vs 253.78 (delta +4.73e-04, +0.00%) |
| Inference, served over HTTP (warm, 3 views, 1 denoising step, 30 requests) | ~90–109 ms inference (median 94.4) · ~93–111 ms end-to-end per 30-step chunk |
| Inference, served over HTTP (warm, 3 views, 10 denoising steps, 30 requests) | ~174 ms inference (min 168 / max 181) · ~177 ms end-to-end |
| Same forward on an RTX 5090 (same host, port's torch reference, eager PyTorch, batch 1, incl. H2D/D2H; bf16 / fp16 autocast) | 15.5 / 15.1 ms (1 step) and 43.5 / 43.7 ms (10 steps) → GPU 6.3× / 4.0× faster than the p150a's 97.4 / 175.8 ms; fp32-strict 34.6 / 95.4 ms (2.8× / 1.8×); best torch.compile 12.8 / 38.1 ms |
TT_MESH_SHAPE must be 1x1.num_denoising_steps=1 is a speed setting (upstream uses 10); it costs <0.05% open-loop PCC but closed-loop flow-matching policies usually want 4–10 steps. Override per request.GET /v1/models is a stub so the tt-model ready card does not 404.v0.71.0-dev20260509-4 (main 2a6ddd8e572) with lerobot 0.5.0 / transformers 5.4.0, single p150a only.TT_FUSED, default on; TT_FUSED=0 = the 2026-09-12 legacy op chain, ~177 ms).GPU_COMPARISON.md.code/): Apache-2.0, from changh95/tt-XVLA; it patches the lerobot 0.5.0 X-VLA policy (Apache-2.0), and code/tt/assets/bart-large-tokenizer/ vendors the facebook/bart-large tokenizer files (Apache-2.0).The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | 2a6ddd8e572bb09b236a2adbd3afab1153e0a17e |
code/ digest |
57c9edfc340cc2ac (sha256, first 16 hex digits) |
| built | 2026-09-13T15:42:11+00:00 by tt-model 0.1.0 |
Base model
lerobot/xvla-base