streampetr-p150
StreamPETR, the camera-only 3D detector of Autoware's autoware_camera_streampetr package (a VoVNet-99-eSE + CPFPN image backbone, the PETR 3D position embedding and a 6-layer transformer decoder over 644 learned and 256 propagated queries, with a 1,024-row object memory carried from frame to frame), ported to one Tenstorrent Blackhole p150 with tt-nn. Five surround-view camera images with their calibration, the frame's time stamp and the ego pose in; 3D boxes with scores, velocities and the Autoware label (CAR, TRUCK, BUS, BICYCLE, PEDESTRIAN) out.
Weights: AutowareFoundation/camera_streampetr v1.0 · Paper: StreamPETR (arXiv:2303.11926) · Autoware package: autoware_camera_streampetr · Training code: tier4/AWML projects/StreamPETR (model zoo t4base v2.5, byte-identical to the v1.0 ONNX files) · Port: code/
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. The whole network runs on the chip as ONE metal trace per frame (frame: the backbone, the head, the top-256 selection and the memory update, with the object memory kept on the chip between frames); the image pre-processing, the position embedding (once per calibration), the stream clock and ego poses, and the box decoding and NMS run on the host. All numbers on this card were measured in this configuration.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/streampetr-p150 --exclude "image/*" --local-dir streampetr-p150 && cd streampetr-p150
pip install -e . # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors; ttnn and torch come from tt-metal
pip install -e ".[server,test]" # optional: the HTTP server and the tests
Run the snippet from the model repo root (the sample path is relative to it).
from tt_streampetr import StreamPETR, load_sample
with StreamPETR.from_pretrained(device_id=0) as model: # weights -> your HF cache, trace captured
out = model(**load_sample("code/tt_streampetr/samples/synthetic_5cam.json")) # five cameras + calibration + stream
for d in out.to_dicts():
print(d["label"], d["score"], d["center"], d["size"], d["yaw"], d["velocity"])
from_pretraineddownloadsAutowareFoundation/camera_streampetrat the pinned commit90f37d64186(tagv1.0; the three ONNX files, 399 MB) to your HF cache, opens the chip, builds the graph and captures the metal trace. The first load on a machine compiles the kernels (minutes: 155-165 s of warm-up measured for a cold dispatch configuration on the host, 224 s in the container's first boot); later loads take 12-13 s.- The first call is as fast as the later calls: the load runs the whole graph eagerly before it captures the trace.
- The shipped sample is a synthetic test pattern with one object of each class (see
code/tt_streampetr/samples/README.md); feed your own rig's five images, calibration, time stamps and ego poses for real use. - The
withblock releases the trace and closes the chip. Withoutwith, callmodel.close().
| Input | images: the five cameras CAM_FRONT, CAM_FRONT_LEFT, CAM_BACK_LEFT, CAM_FRONT_RIGHT, CAM_BACK_RIGHT, any order, any size (path, bytes, PIL, uint8 array; resized and cropped to 640×480 like the node). calibration: per camera the intrinsics (CameraInfo P or K of the image as sent) and T_ref_from_camera (camera optical frame -> base_link), or a preset. stream: id, reset, timestamp_s (the CAM_FRONT stamp), T_world_from_ego (map <- base_link) for the temporal object memory; without stream, one frame from zero memory. |
| Options | score_threshold=None (the node's per-class thresholds .36 / .39 / .38 / .41 / .43), iou_nms_threshold=0.5, iou_nms_search_distance=10.0, circle_nms_distance=0.0, decode_layers=6, awml_decode=False, max_detections=500. from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None, normalization="autoware_main", use_temporal=True, max_streams=1). |
| Output | Detections3D in base_link: boxes float32 [N, 7] (x, y, z gravity centre, length, width, height, yaw), scores, label_ids, labels (the Autoware labels), velocities [N, 2], timing_ms, meta (the decoder layer and query of every box, the stream state). |
| Methods | out.to_dict() gives the /predict JSON; out.to_dicts() one dict per box (label, label_id, score, center, size, yaw, velocity). |
- Temporal memory. Each
stream["id"]keeps its object memory on the chip (one stream at a time by default,max_streams; another id evicts the least recently used one). Send the frames of a stream in time order, each with its CAM_FRONTtimestamp_sand itsT_world_from_ego: the memory is aligned with the ego motion. A new id orreset=Truestarts like a fresh node. As in the node, a camera-sync failure (stamps more than 0.15 s apart) resets the stream, and a frame without an ego pose is skipped (no objects,meta["skipped"]). - The API gives the same output as the HTTP server
/predict: both share the decoders, the device trace and the host post-processing. One model uses one chip; calls from several threads are serialised. - Full reference:
code/PYTHON.md. Runnable example:examples/quickstart.py(writes the/predictJSON).
Serving (HTTP)
tt-model pull changh95/streampetr-p150 --with-weights
tt-model serve changh95/streampetr-p150 # or with tt-cli: tt serve changh95/streampetr-p150
S=code/tt_streampetr/samples/synthetic_5cam
python3 code/tt_streampetr/server/client.py --image CAM_FRONT=$S/CAM_FRONT.png --image CAM_FRONT_LEFT=$S/CAM_FRONT_LEFT.png \
--image CAM_BACK_LEFT=$S/CAM_BACK_LEFT.png --image CAM_FRONT_RIGHT=$S/CAM_FRONT_RIGHT.png \
--image CAM_BACK_RIGHT=$S/CAM_BACK_RIGHT.png --calib-preset synthetic_5cam --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/streampetr-p150
- The image does not contain the weights.
--with-weightsputs them in your HF cache. - The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete. POST /predict:images(the five cameras, base64 JPEG / PNG /.npy, rgb8),calibration(per cameraintrinsics= CameraInfo P or K of the image as sent, andT_ref_from_camera; or{"preset": name}); optionalstream(id,reset,timestamp_s,T_world_from_ego),params(score_threshold,iou_nms_threshold,decode_layers, ...),output_format. AlsoGET /health,GET /info,GET /v1/models(stub). Contract:SERVING.mdsection 3.
The shipped sample (first request after the warm-up, host server, 2026-10-09):
{"model": "streampetr-p150", "frame_id": "base_link", "num_detections": 5,
"detections": [
{"label": "PEDESTRIAN", "label_id": 4, "score": 0.9715, "center": [9.334, -2.159, 0.65], "size": [0.803, 0.749, 1.813], "yaw": 2.3695, "velocity": [-0.003, 0.003]},
{"label": "BUS", "label_id": 2, "score": 0.9697, "center": [8.977, 13.014, 1.606], "size": [11.762, 2.718, 3.189], "yaw": -2.8072, "velocity": [-0.001, -0.002]},
{"label": "CAR", "label_id": 0, "score": 0.9669, "center": [11.932, 1.191, 0.554], "size": [4.186, 1.731, 1.701], "yaw": 2.9098, "velocity": [-0.001, 0.001]},
...],
"meta": {"class_names": ["CAR", "TRUCK", "BUS", "BICYCLE", "PEDESTRIAN"], "num_pre_nms": 30, "decoder_layer": [2, 5, 4, 5, 5], "query": [451, 299, 335, 144, 487], "stream": "synthetic_5cam", "fresh": true, "skipped": null, "ts": 0.0, "frame_index": 1},
"timing_ms": {"host_in": 15.045, "trace": 149.158, "preprocess": 316.876, "device": 168.683, "postprocess": 9.794, "total": 621.573, "decode": 125.161, "model_call": 495.405}}
- Boxes are in
base_linkmetres / radians, as the Autoware node publishes them:centeris the gravity centre,sizeis[length along the heading, width, height](the node'sshape.dimensions) andyawis counter-clockwise from +x; detections are sorted by score. labelis the class name, which camera_streampetr v1.0 also publishes as the Autoware label;label_idis the network class.velocityis the head's[vx, vy](m/s); the node predicts it but does not publish it (has_twist = false).metacarries the decoder layer and query of every object and the stream state (fresh,skipped,ts,frame_index);/inforeports the device (ETH, 12×10), the labels, the calibration presets and the load-time knobs.
Demo
| PandaSet 019, frame 40 (San Francisco; CC BY 4.0): TT output | nuScenes scene-0103, key-frame 20 (Boston; CC BY-NC-SA 4.0, non-commercial): TT output |
|---|---|
![]() |
![]() |
| TT vs the fp32 CPU reference: PandaSet 019, 80 frames free-running | The shipped synthetic sample: TT vs the stored CPU reference |
|---|---|
![]() |
![]() |
More: PandaSet 019 frames 0-78 as a GIF (every second frame, real-time playback), the bird's-eye views with LiDAR context of PandaSet 019 frame 40 and of nuScenes scene-0103 key-frame 20 (non-commercial). Every image shows outputs of this port on the p150 (the boxes the node publishes), some next to the fp32 CPU reference; grey outlines are the dataset annotations, and head and licence-plate regions of annotated people and vehicles are blurred. Sources, licences and changes: media/ATTRIBUTION.md. The weights were trained on TIER IV's T4 data only, so both datasets are out of domain: the detections show the domain gap, not the port (CPU reference, one-to-one matches within 2 m and 40 m: vehicles recall / precision 0.37 / 0.40 on PandaSet 019, 0.80 / 0.68 on nuScenes scene-0103).
Demo & Performances
Warm, batch 1, five cameras. Accuracy: the TT output against the port's fp32 CPU reference of the same network (same weights, same pre- and post-processing), which matches ONNX Runtime on the deployed ONNX files (PCC ≥ 0.9999; same objects). Speed: code/scripts/bench.py, p50 (p99) of 100 iterations, on a shared 8-core host (OPT_BASELINE.md).
| Metric | Performance |
|---|---|
| Per-stage PCC vs the CPU reference (gates) | backbone img_feats 0.999988 (gate 0.999); head, teacher-forced on PandaSet 019 frames 1 / 40: classes 0.99953 / 0.99970 (0.99), boxes 0.99997 / 0.999994 (0.999), decoder output 0.99970 / 0.99979 (0.995); top-256 memory selection overlap 0.992 / 0.996 (0.9); box centre error max 0.169 / 0.130 m (0.2 m) |
| Box agreement, PandaSet 019 frames 0-79, free-running (each side with its own memory) | mean F1 0.921 (min 0.816) vs the CPU run (same class, 1 m; gate 0.85); the scores of matched boxes drift apart as the two memories diverge (max |difference| per frame: median 0.16, up to 0.46) |
| Box agreement, nuScenes scene-0103 / scene-0916, free-running (229 / 240 frames) | mean F1 0.956 / 0.974 vs the CPU run |
| nuScenes mini_val 5-class mAP (81 key-frames; cross-domain, indicative) | TT 0.1933, CPU 0.1942 (the node's published boxes; the weights never saw nuScenes) |
| Box agreement, PandaSet 019 frame 40 (cold start) | recall 0.957, precision 1.000, max |score difference| 0.032 (gates 0.95 / 0.95 / 0.05) |
| Box agreement, the shipped synthetic sample (the container smoke's check) | 5 / 5 boxes (recall 1.000, precision 1.000), max |score difference| 0.0021 vs the stored CPU reference (gates 0.95 / 0.95 / 0.05) |
Device, one frame (frame replay) |
74.97 ms; back to back 74.91 ms: 13.3 frames/s device-bound |
Python model(), five 1920×1080 images (synthetic) |
601.0 ms (726.7): preprocess 474.1, image upload 38.6, device 74.98, postprocess 0.9 |
Python model(), PandaSet 1920×1080 JPEGs |
706.1 ms (966.5): preprocess 542.3, image upload 40.5, device 75.00, postprocess 68.3 |
Served /predict on the host, shipped sample (uvicorn, loopback) |
one request after the warm-up: timing_ms.total 621.6 ms (decode of the 4.0 MB base64 PNG body 125.2, preprocess 316.9, device incl. upload 168.7, postprocess 9.8); client round trip 668.9 ms |
All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, AICLK 1350 MHz. The device is kernel-bound (74.48 ms of kernels and 0.91 ms of op-to-op gaps for 1,484 programs); the backbone takes 54.8 ms of it, and 54 % of the replay is layout, data movement and precision glue. The end-to-end latency is dominated by the host pre-processing (the node's anti-aliased resize of five 1920×1080 images, 70-77 % of a call), which moves by ±100 ms with the load of the shared host. This is the first release: optimization has not started (OPT_REPORT.md ranks the plan; OPT_BASELINE.md has every measurement). Verification: VERIFICATION_2026-10-09.md.
Box agreement: a reference box counts as recalled when a TT box of the same class lies within the stated distance, and the reverse for precision; free-running means each side carries its own object memory through the whole sequence. The nuScenes mAP uses the nuScenes devkit functions (detection_cvpr_2019 settings) on the 5 Autoware classes (motorcycle annotations scored as bicycle), on the boxes the node publishes, free-running at the CAM_FRONT rate and scored at the key-frames; it is a cross-domain sanity check, not comparable to the paper. nuScenes dataset © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use); non-commercial: the data were used for local validation only and are not distributed here; Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020.
No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). The Autoware package README lists 22.13 ms of inference (26.04 ms in total, 3.25 ms of pre-processing per image) per frame on an RTX 3090 with TensorRT; its example configuration uses trt_precision: fp16, and the README does not state the precision of the measurement (not like-for-like). p150 power was not measured, so no efficiency comparison is made.
Caveats
- Deployment status.
autoware_camera_streampetris not launched by the default Autoware stack: the camera-only BEV detector of autoware_launch (0.52.0 and main) is BEVDet, and StreamPETR runs only through its own package launch file. This port is a drop-in for the network and the node's pre- and post-processing, not a ROS node, and not a certified Autoware component: do not use it for safety-critical driving decisions. - First release: optimization pending. This is the functional baseline port; no optimization round has run yet (
OPT_REPORT.md). - Precision. The node runs fp16 TensorRT engines by default; this port follows the fp32 graphs. On the chip: fp32 weights and biases, HiFi4, fp32 and packer L1 accumulation; bf16 backbone activations; fp32 head activations with the attention keys split into two bf16 terms; the temporal positional chain as exact-product sliced fp32 maps (the precision policy of the first release). Outputs differ slightly from the fp32 reference (agreement figures above); objects near a class threshold can appear or disappear.
- Documented behaviour and deviations from the node (
SERVING.mdsection 4; PORT_LOG.md section 6):- Normalization: Autoware universe main (
autoware_main, RGB planes and RGB statistics) by default;STREAMPETR_NORMALIZATION=autoware_0.52reproduces an rgb8 camera on the pinned universe 0.52.1,awml_trainingthe training pipeline. - The memory ages follow the node exactly, including its quirks after a camera-sync reset and at the 600 s clock-origin reset (
STREAMPETR_AUTOWARE_COMPAT=0gives clean ages instead). - A stream frame without an ego pose is skipped, as the node skips a cycle whose TF lookup fails; skipped frames return an empty object list with
meta.skippedwhere the node publishes nothing. - The position embedding is recomputed when a request's calibration changes (the node computes it once; equal for a fixed rig).
velocityis reported although the node does not publish it.- The resize kernel is emulated with IEEE single rounding; nvcc's contraction (
STREAMPETR_PREPROCESS_FMA=1) differs by at most 3e-5 in the normalised input.
- Normalization: Autoware universe main (
- Domain gap. The weights were trained on TIER IV's T4 five-camera data only; no public dataset is in domain. Other rigs work (the network takes the calibration) but lose accuracy (vehicles recall / precision 0.37 / 0.40 on PandaSet 019 with the CPU reference).
- Fixed shapes. Five cameras in the node's slots, a 480×640 network input, 644 + 256 queries, a 1,024-row memory, 5 classes; batch 1, one frame per request; requests are serialised on the chip; one stream memory by default (
STREAMPETR_MAX_STREAMS, PLAN D16). - Validation scope. Agreement with the fp32 CPU reference on public data (PandaSet, nuScenes) and on the synthetic sample; the nuScenes mAP is a cross-domain, indicative number. The shipped sample is synthetic; a PandaSet sample is not shipped yet.
- Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (
patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(server:STREAMPETR_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (80.8 ms per frame), and its numerics are not validated: it fails the frozen PandaSet sample gate (recall 0.913). If ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404.
Licensing
- Weights: AutowareFoundation/camera_streampetr at tag
v1.0(commit90f37d64186686ebc458509d69d3d4ca16d185c7), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The weights were trained by TIER IV on their T4 dataset (AWML StreamPETR t4base v2.5, 123,708 frames); the upstream card is Apache-2.0 and documents no further data restriction. - Pre- and post-processing ported from autoware_universe
perception/autoware_camera_streampetr(Apache-2.0). - Port and serving code (
code/): Apache-2.0. - Sample data:
code/tt_streampetr/samples/synthetic_5cam*and the presetcalib/synthetic_5cam.jsonwere generated by this repository (code/scripts/make_sample.py; no third-party data), Apache-2.0. No dataset sample is shipped; nuScenes never ships, and a PandaSet sample (CC BY 4.0) is prepared but not yet included. - Demo media (
media/; per-file sources, frames and changes:media/ATTRIBUTION.md):- PandaSet renders (
media/streampetr_pandaset_*): Contains data from PandaSet (Scale AI and Hesai), https://pandaset.org, licensed under CC BY 4.0 and the PandaSet Dataset Terms. Changes: the 4:3 network crop of each camera image, downscaled, head and licence-plate regions blurred; predicted 3D boxes, annotation outlines and LiDAR points drawn; bird's-eye-view plots. Scale AI and Hesai do not endorse this work. Cite: P. Xiao et al., PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving, ITSC 2021. - nuScenes renders (
media/*_NC.*), non-commercial, CC BY-NC-SA 4.0: Rendered from the nuScenes dataset, © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., nuScenes: A Multimodal Dataset for Autonomous Driving, CVPR 2020. - The synthetic-sample render (
media/streampetr_synthetic_5cam_tt.png): generated by this repository, Apache-2.0.
- PandaSet renders (
Provenance
These are the exact sources the container image was built from:
| component | built from |
|---|---|
| tt-metal | 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (dirty tree: the image includes the patch) |
| weights | AutowareFoundation/camera_streampetr@90f37d64186686ebc458509d69d3d4ca16d185c7 (tag v1.0), files simplify_extract_img_feat.onnx (sha256 322c2f42…3efd), simplify_position_embedding.onnx (629f1482…7163), simplify_pts_head_memory.onnx (49b038bd…4015), ml_package_camera_streampetr.param.yaml (f539407f…5218) |
| Autoware reference | autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (package 0.53.0; the autoware_0.52 normalization preset reproduces the pinned universe 0.52.1) |
| shared port code | ttaw 0.20.0 @ common 89dec49 (vendored in code/tt_streampetr/ttaw, VENDORED.json) |
code/ digest (image) |
95413849922807ea (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json) |
| image | tt-model/streampetr-p150:ecba3ee8ed81 (sha256:ecba3ee8ed81e0a53703d8acc43289d4797fcc1a9f057cdea72f69e28319c016) |
| base images | build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json) |
| built | 2026-10-09T17:24:03+00:00 by tt-model 0.1.0 |
Model tree for changh95/streampetr-p150
Base model
AutowareFoundation/camera_streampetr


