bevdet-p150
BEVDet, the camera-only 3D detector Autoware deploys in autoware_tensorrt_bevdet (BEVDet-R50-4DLongterm-Depth: ResNet-50, depth-aware lift-splat into a 128×128 bird's-eye view, an 8-frame BEV history), ported to one Tenstorrent Blackhole p150 with tt-nn. Six surround-view camera images with their calibration in, 3D boxes with scores, velocities, the network class and the Autoware label out.
Weights: AutowareFoundation/tensorrt_bevdet v1.0 · Papers: BEVDet (arXiv:2112.11790), BEVDet4D (arXiv:2203.17054) · Autoware package: autoware_tensorrt_bevdet · Training code: LCH1238/BEVDet (export branch of HuangJunJie2017/BEVDet dev2.1) · Port: code/
Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. The image encoder and the BEV decoder run on the chip as two metal traces; the BEVPool lift-splat between them and the box decoding run on the host. All numbers on this card were measured in this configuration.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart (Python)
Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.
hf download changh95/bevdet-p150 --exclude "image/*" --local-dir bevdet-p150 && cd bevdet-p150
pip install -e . # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors, opencv-python-headless; ttnn and torch come from tt-metal
pip install -e ".[server,test]" # optional: the HTTP server and the tests
Run the snippet from the model repo root (the sample path is relative to it).
from tt_bevdet import BEVDet, load_sample
with BEVDet.from_pretrained(device_id=0) as model: # weights -> your HF cache, traces captured
out = model(**load_sample("code/tt_bevdet/samples/synthetic_6cam.json")) # six cameras + calibration
for d in out.to_dicts():
print(d["label"], d["autoware_label"], d["score"], d["center"], d["size"], d["yaw"], d["velocity"])
from_pretraineddownloadsAutowareFoundation/tensorrt_bevdetat the pinned commit391466349a1(tagv1.0, 305 MB) to your HF cache, opens the chip, builds the graph and captures the four metal traces (encode,decodeand the two small history-ring updatespush/fill). The first load compiles the kernels (about 4 minutes on the measurement host); later loads take 15-17 s.- The first call is as fast as the later calls: the load ends with one synthetic frame through the whole path.
- The shipped sample is a synthetic test pattern with eight objects (see
code/tt_bevdet/samples/README.md); feed your own rig's six images and calibration for real use. - The
withblock releases the traces and closes the chip. Withoutwith, callmodel.close().
| Input | images: the six cameras CAM_FRONT_LEFT, CAM_FRONT, CAM_FRONT_RIGHT, CAM_BACK_LEFT, CAM_BACK, CAM_BACK_RIGHT, any order, any size (path, bytes, PIL, uint8 array; images that are not 1600×900 are resized like the node). calibration: per camera the intrinsics of the image as sent and T_ref_from_camera (camera optical frame -> base_link), or a preset. stream: id, reset, T_world_from_ego for the 8-frame BEV history. |
| Options | score_threshold=0.2 (the node's filter), decode_score_threshold=0.1, nms_iou_threshold=0.2, max_pre_nms=500, max_detections=500, nms_rescale_factor (per class). from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None, max_streams=4). |
| Output | BEVDetDetections: boxes float32 [N, 7] (x, y, z, length, width, height, yaw in base_link), scores, label_ids / labels (the 10 network classes), autoware_labels, velocities [N, 2], twists(), timing_ms, meta. |
| Methods | out.to_dict() gives the /predict JSON; out.to_dicts() one dict per box (label, label_id, autoware_label, score, center, size, yaw, velocity, twist). |
- Ego motion. Without
stream["T_world_from_ego"]the history is fused without motion compensation, exactly like the deployed node, which never sets the pose. Pass the ego pose (4×4,base_linkin a world frame) whenever you have a localisation: on nuScenes mini_val it lifts mAP from 0.253 to 0.367 and NDS from 0.274 to 0.436 (below). - Streams. Each
stream["id"]has its own history on the chip; up to 4 at once (max_streams), a fifth evicts the least recently used one, which restarts like a freshly launched node.reset=Trueis a scene change. - The API gives the same output as the HTTP server
/predict: both share the decoders, the device traces and the host post-processing. One model uses one chip; calls from several threads are serialised. - Full reference:
code/PYTHON.md. Runnable example:examples/quickstart.py(also writes a bird's-eye view PNG).
Serving (HTTP)
tt-model pull changh95/bevdet-p150 --with-weights
tt-model serve changh95/bevdet-p150 # or with tt-cli: tt serve changh95/bevdet-p150
S=code/tt_bevdet/samples/synthetic_6cam
python3 code/tt_bevdet/server/client.py --image CAM_FRONT_LEFT=$S/CAM_FRONT_LEFT.png --image CAM_FRONT=$S/CAM_FRONT.png \
--image CAM_FRONT_RIGHT=$S/CAM_FRONT_RIGHT.png --image CAM_BACK_LEFT=$S/CAM_BACK_LEFT.png \
--image CAM_BACK=$S/CAM_BACK.png --image CAM_BACK_RIGHT=$S/CAM_BACK_RIGHT.png --calib-preset synthetic_6cam --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/bevdet-p150
- The image does not contain the weights.
--with-weightsputs them in your HF cache. - The server uses port 20000 (or the next free port). It is ready when the log shows
Application startup complete. POST /predict:images(the six cameras, base64 JPEG / PNG /.npy),calibration(per cameraintrinsics+T_ref_from_camera, or{"preset": name}); optionalstream(id,reset,T_world_from_ego),params(score_threshold,nms_iou_threshold,nms_rescale_factor, ...),output_format. AlsoGET /health,GET /info,GET /v1/models(stub). Contract:SERVING.mdsection 3.
{"model": "bevdet-p150", "frame_id": "base_link", "num_detections": 10, "detections": [
{"label": "traffic_cone", "label_id": 9, "score": 0.9492, "center": [7.68, -3.513, 0.367], "size": [0.385, 0.348, 0.733], "yaw": 0.134, "velocity": [0.0, 0.0], "autoware_label": "UNKNOWN", "twist": {"linear_x": 0.0, "angular_z": 2.1597}},
{"label": "car", "label_id": 0, "score": 0.8457, "center": [-8.51, -5.941, 0.845], "size": [4.606, 1.92, 1.699], "yaw": 0.156, "velocity": [-0.0, 0.0], "autoware_label": "CAR", "twist": {"linear_x": 0.0, "angular_z": 4.8626}}],
"timing_ms": {"encode": 60.0, "bevpool": 25.7, "bev_decode": 15.1, "preprocess": 60.6, "device": 101.4, "postprocess": 18.6, "total": 584.6, "decode": 403.6, "model_call": 180.6},
"meta": {"class_names": ["car", "truck", ...], "autoware_labels": ["CAR", "TRUCK", ...], "score_threshold": 0.2, "stream": {"id": "timing", "frame": 42, "flag": 1}}}
center(box centre, z = gravity centre) /size([length, width, height]) /yaw(counter-clockwise from +x) are inbase_linkmetres / radians, as the Autoware node publishes them; detections are sorted by score.labelis the network class (one of the 10 nuScenes classes of the Autoware config,label_idits index);autoware_labelis the label the node publishes (construction_vehicle, barrier and traffic_cone become UNKNOWN).velocityis[vx, vy]in m/s (absolute,base_linkaxes);twistis the node'stwist.linear.x/twist.angular.z, reproduced as deployed./inforeports the device (ETH, 12×10), the labels, the stream table (4 streams, least-recently-used eviction) and the load-time knobs.
Demo
| PandaSet 019, frame 40 (San Francisco; CC BY 4.0): TT output, ego-pose mode | nuScenes scene-0103, key-frame 20 (Boston; CC BY-NC-SA 4.0, non-commercial): TT output |
|---|---|
![]() |
![]() |
| TT vs the fp32 CPU reference: PandaSet 019 at 2 Hz | The shipped synthetic sample: TT vs the stored CPU reference |
|---|---|
![]() |
![]() |
More: nuScenes scene-0916 key-frame 20 (Singapore) and the 40 key-frames of scene-0103 as a GIF, both non-commercial. Every image shows outputs of this port on the p150, some next to the fp32 CPU reference; the camera mosaics and the mode / K-quirk panels draw the boxes with score >= 0.4 (the node publishes >= 0.2, too dense to read at this size), the TT-vs-CPU panel and the synthetic sample every box >= 0.2. Sources, licences and changes: media/ATTRIBUTION.md. The weights were trained on nuScenes; PandaSet is another rig (taller cameras, other lenses), and its ranges come out biased (front objects about 15 % too close).
Demo & Performances
Warm, batch 1, six cameras. Accuracy: the TT output against the port's fp32 CPU reference of the same Autoware network (same weights, same pre- and post-processing), checked against ONNX Runtime on the plugin-free sub-graphs of the deployed ONNX (PCC 1.0000000000). Speed: code/scripts/bench.py, p50 (p99) of 100 iterations, host load average 7-9 on the shared 8-core measurement host (OPT_BASELINE.md).
| Metric | Performance |
|---|---|
| Per-stage PCC vs the CPU reference, PandaSet frames (gates) | depth 0.99994, features 0.99994 (gate 0.999); BEV map after the host BEVPool 0.99995 (0.999); six head maps >= 0.99952 (0.998) |
| Box agreement, nuScenes mini_val (scene-0103 / scene-0916, 81 key-frames, ego poses) | recall 0.9897 / 0.9897, precision 0.9899 / 0.9902, p99 |score difference| 0.0083 / 0.0089 (gates 0.95 / 0.95 / 0.05) |
| Box agreement, PandaSet 019 frames 38-40 (Autoware semantics) | recall 1.000 / 0.996 / 0.995, precision 1.000 / 0.996 / 0.986, max |score difference| 0.0059 / 0.0080 / 0.0079 |
| Box agreement, the shipped synthetic sample (the container smoke's check) | recall 1.000, precision 1.000 (all 10 reference boxes), max |score difference| 0.0056, centres within 0.005 m (gates 0.95 / 0.95 / 0.05) |
| nuScenes mini_val mAP / NDS, ego-pose mode (indicative: 81 frames, 3 of 10 classes absent) | TT 0.3670 / 0.4362, CPU 0.3673 / 0.4367 (BEVDet's own full-val result for this config: 0.394 / 0.515) |
| nuScenes mini_val mAP / NDS, Autoware node semantics (no ego motion, camera rate) | TT 0.2530 / 0.2739, CPU 0.2535 / 0.2746 |
Device, one steady frame (encode + decode + ring push, back to back) |
34.15 ms (34.45 with ego poses): 29.3 frames/s device-bound |
Trace replays: encode / decode / push |
24.24 / 8.68 / 1.26 ms back to back (24.30 / 8.73 / 1.29 ms one blocking replay) |
Python model(), 1600×900 images, Autoware mode (synthetic input) |
173.3 ms (217.5); 131.8-155.2 ms p50 in the other runs on a less loaded host |
Python model(), nuScenes 1600×900 with ego poses |
220.3 ms (262.5) |
Python model(), PandaSet 1920×1080 (host resize to 1600×900) |
341.9 ms (370.5) |
Served /predict on the host, shipped sample (uvicorn, loopback) |
client round trip 623.9 ms p50 (p90 703.7; a 9.6 MB JSON request of six base64 PNGs); timing_ms.model_call 156.7 ms; timing_ms.total 538.2 ms, most of it the request decode (base64 + PNG, 404 ms in the response shown in SERVING.md) |
All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150; AICLK 1343-1350 MHz. The device is kernel-bound (34.0 of 34.3 ms are kernels; 265 µs of op-to-op gaps for 457 programs), and the end-to-end latency is set by host work: the image preprocessing (35-43 ms at 1600×900, 209 ms at 1920×1080), the 6.5 MB image upload (27 ms, sent as 6-byte pages), the host BEVPool (15-36 ms) and the Scale-NMS (about 50 ms on real scenes). This is the first release: optimization has not started (OPT_REPORT.md ranks the plan; OPT_BASELINE.md has every measurement). Verification: VERIFICATION_2026-10-08.md.
The nuScenes metrics use the official devkit (detection_cvpr_2019, mini_val) on every box the vendor decode keeps; nuScenes is CC BY-NC-SA 4.0 (non-commercial; the data were used for local validation only and are not distributed here). Box agreement: a reference box with score >= 0.2 counts as recalled when a same-class TT box (any score above the 0.1 decode threshold) lies within 0.5 m, and the reverse for precision; scores that straddle 0.2 within bf16 noise are not counted as disagreements.
No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). The Autoware README claims up to 35 FPS (fp16) / 17 FPS (fp32) on an unspecified GPU (not like-for-like). p150 power was not measured (OPT_BASELINE.md lists only the board telemetry range during the timed loops), so no efficiency comparison is made.
Caveats
- Deployment status. BEVDet is not Autoware's default perception:
autoware.launch.xmldefaults toperception_mode:=lidar, andautoware_tensorrt_bevdetis the detector of the camera-only mode (perception_mode:=camera). This port is a drop-in for the network and the node's pre- and post-processing, not a ROS node, and not a certified Autoware component: do not use it for safety-critical driving decisions. - First release: optimization pending. This is the functional baseline port; no optimization round has run yet, and the end-to-end latency is set by host stages (
OPT_REPORT.mdranks the plan). - Precision. The node defaults to its fp16 TensorRT engine; this port implements the fp32 path. By code reading (not run here: no GPU), the fp16 engine also mis-reads the AlignBEV transforms (the history branch is undefined in the default mode), keeps BEVPool sums in fp16 and leaves unreached BEV cells uninitialised; the port follows the intended fp32 semantics (zero-initialised cells). On the chip: fp32 weights and biases, HiFi4, fp32 and packer L1 accumulation, bf16 activations (the precision policy of the first release; bf16 weights fail the accuracy gates). Outputs differ slightly from the fp32 reference (agreement figures above), with a small systematic score offset of about -0.0008.
- Temporal semantics. Default = the deployed node: no ego motion (identity alignment), the history advanced once per request, a fresh stream fuses an empty history on its first frame. The ego-pose input (
stream.T_world_from_ego) is a documented extension;BEVDET_HISTORY_STRIDE(load time) spaces the history like training's 2 Hz key-frames. - Intrinsics. The API takes the intrinsics of the image actually sent and scales K with the node's resize to 1600×900; the node itself keeps the raw
CameraInfo.kwhen it resizes (calibration.scale_intrinsics: falsereproduces it: far fewer and misplaced boxes on a 1920×1080 rig, see the K-quirk image). Rectified rigs should pass P[0:3, 0:3]. - Domain gap. The weights were trained on nuScenes (Boston, Singapore; camera heights ~1.5 m). On other rigs the monocular depth prior is biased: on PandaSet 019 front objects come out ~15 % too close and side objects 10-20 % too far.
- Host stages. Image preprocessing, the BEVPool lift-splat between the two traces and the vendor Scale-NMS run on the host (the documented fallbacks of this first release); the end-to-end latency is host-bound (above).
- Fixed shapes. Six cameras, 256×704 network input, 118 depth bins, a 128×128 BEV grid (±51.2 m), 8 history frames, 10 classes; rigs whose BEVPool tables exceed M 370,000 / N 14,000 are refused, as with TensorRT. Batch 1, one frame per request; requests are serialised on the chip.
- Streams. At most 4 stream ids keep a history at once (
BEVDET_MAX_STREAMS); a fifth evicts the least recently used one, which restarts with an empty history. - Validation scope. Agreement with the fp32 CPU reference on public data (nuScenes mini_val, PandaSet) and on the synthetic sample; the mAP / NDS on 81 mini_val frames are indicative, not a full-val reproduction. The shipped sample is synthetic; the PandaSet sample is not shipped yet.
- Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (
patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication. dispatch="worker"(server:BEVDET_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (34.83 ms per frame, the same outputs); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.- Not an OpenAI-compatible API;
GET /v1/modelsis a stub so the tt-model ready card does not 404. - p150 power was not measured (only the board telemetry was logged, OPT_BASELINE.md), so no efficiency comparison is made.
Licensing
- Weights: AutowareFoundation/tensorrt_bevdet at tag
v1.0(commit391466349a1b39c93b9865364a8088110ca8d97b), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card's Legal Notice: the training data is nuScenes (CC BY-NC-SA 4.0, non-commercial); read it before commercial use. - Pre- and post-processing ported from autoware_universe
perception/autoware_tensorrt_bevdet(Apache-2.0). The node's inference library,bevdet_vendor, carries no licence: its behaviour was re-implemented clean-room from a written description, and no code was copied. - Port and serving code (
code/): Apache-2.0. - Sample data:
code/tt_bevdet/samples/synthetic_6cam*and the presetcalib/synthetic_6cam.jsonwere generated by this repository (code/scripts/make_sample.py; no third-party data), Apache-2.0. No dataset sample is shipped; nuScenes, Argoverse 2 and KITTI never ship, and a PandaSet sample (CC BY 4.0) is prepared but not yet included. - Demo media (
media/): renders of PandaSet (Scale AI and Hesai, CC BY 4.0 and the PandaSet Dataset Terms; Scale AI and Hesai do not endorse this work) and of nuScenes (© Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use: those renders are non-commercial, ShareAlike, labelled_NC; Motional does not endorse this work). Attribution texts and changes:media/ATTRIBUTION.md.
Provenance
These are the exact sources the container image was built from:
| component | built from |
|---|---|
| tt-metal | 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (dirty tree: the image includes the patch) |
| weights | AutowareFoundation/tensorrt_bevdet@391466349a1b39c93b9865364a8088110ca8d97b (tag v1.0), files bevdet_one_lt_d.onnx (sha256 47f156cb…0d77) |
| Autoware reference | autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (package 0.53.0); bevdet_vendor 0.2.1 semantics |
| shared port code | ttaw 0.20.0 @ common 89dec49 (vendored in code/tt_bevdet/ttaw, VENDORED.json) |
code/ digest (image) |
a67de74da4365b53 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json) |
| image | tt-model/bevdet-p150:673a2b830216 (sha256:673a2b8302161c0c92000591efc289d1c4aaa34b6464818144ee28fc59d942a4) |
| base images | build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json) |
| built | 2026-10-09T05:04:41+00:00 by tt-model 0.1.0 |
Model tree for changh95/bevdet-p150
Base model
AutowareFoundation/tensorrt_bevdet




