bevdet-p150

BEVDet, the camera-only 3D detector Autoware deploys in autoware_tensorrt_bevdet (BEVDet-R50-4DLongterm-Depth: ResNet-50, depth-aware lift-splat into a 128×128 bird's-eye view, an 8-frame BEV history), ported to one Tenstorrent Blackhole p150 with tt-nn. Six surround-view camera images with their calibration in, 3D boxes with scores, velocities, the network class and the Autoware label out. Weights: AutowareFoundation/tensorrt_bevdet v1.0 · Papers: BEVDet (arXiv:2112.11790), BEVDet4D (arXiv:2203.17054) · Autoware package: autoware_tensorrt_bevdet · Training code: LCH1238/BEVDet (export branch of HuangJunJie2017/BEVDet dev2.1) · Port: code/

Runs on p150 (mesh P150). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. The image encoder and the BEV decoder run on the chip as two metal traces; the BEVPool lift-splat between them and the box decoding run on the host. All numbers on this card were measured in this configuration.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart (Python)

Prerequisite: a tt-metal / ttnn environment at tt-metal 44d66500520 with patches/tt-metal-eth-dispatch.patch applied. ttnn is not on PyPI.

hf download changh95/bevdet-p150 --exclude "image/*" --local-dir bevdet-p150 && cd bevdet-p150
pip install -e .                        # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors, opencv-python-headless; ttnn and torch come from tt-metal
pip install -e ".[server,test]"         # optional: the HTTP server and the tests

Run the snippet from the model repo root (the sample path is relative to it).

from tt_bevdet import BEVDet, load_sample

with BEVDet.from_pretrained(device_id=0) as model:    # weights -> your HF cache, traces captured
    out = model(**load_sample("code/tt_bevdet/samples/synthetic_6cam.json"))   # six cameras + calibration

for d in out.to_dicts():
    print(d["label"], d["autoware_label"], d["score"], d["center"], d["size"], d["yaw"], d["velocity"])
  • from_pretrained downloads AutowareFoundation/tensorrt_bevdet at the pinned commit 391466349a1 (tag v1.0, 305 MB) to your HF cache, opens the chip, builds the graph and captures the four metal traces (encode, decode and the two small history-ring updates push / fill). The first load compiles the kernels (about 4 minutes on the measurement host); later loads take 15-17 s.
  • The first call is as fast as the later calls: the load ends with one synthetic frame through the whole path.
  • The shipped sample is a synthetic test pattern with eight objects (see code/tt_bevdet/samples/README.md); feed your own rig's six images and calibration for real use.
  • The with block releases the traces and closes the chip. Without with, call model.close().
Input images: the six cameras CAM_FRONT_LEFT, CAM_FRONT, CAM_FRONT_RIGHT, CAM_BACK_LEFT, CAM_BACK, CAM_BACK_RIGHT, any order, any size (path, bytes, PIL, uint8 array; images that are not 1600×900 are resized like the node). calibration: per camera the intrinsics of the image as sent and T_ref_from_camera (camera optical frame -> base_link), or a preset. stream: id, reset, T_world_from_ego for the 8-frame BEV history.
Options score_threshold=0.2 (the node's filter), decode_score_threshold=0.1, nms_iou_threshold=0.2, max_pre_nms=500, max_detections=500, nms_rescale_factor (per class). from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None, max_streams=4).
Output BEVDetDetections: boxes float32 [N, 7] (x, y, z, length, width, height, yaw in base_link), scores, label_ids / labels (the 10 network classes), autoware_labels, velocities [N, 2], twists(), timing_ms, meta.
Methods out.to_dict() gives the /predict JSON; out.to_dicts() one dict per box (label, label_id, autoware_label, score, center, size, yaw, velocity, twist).
  • Ego motion. Without stream["T_world_from_ego"] the history is fused without motion compensation, exactly like the deployed node, which never sets the pose. Pass the ego pose (4×4, base_link in a world frame) whenever you have a localisation: on nuScenes mini_val it lifts mAP from 0.253 to 0.367 and NDS from 0.274 to 0.436 (below).
  • Streams. Each stream["id"] has its own history on the chip; up to 4 at once (max_streams), a fifth evicts the least recently used one, which restarts like a freshly launched node. reset=True is a scene change.
  • The API gives the same output as the HTTP server /predict: both share the decoders, the device traces and the host post-processing. One model uses one chip; calls from several threads are serialised.
  • Full reference: code/PYTHON.md. Runnable example: examples/quickstart.py (also writes a bird's-eye view PNG).

Serving (HTTP)

tt-model pull  changh95/bevdet-p150 --with-weights
tt-model serve changh95/bevdet-p150       # or with tt-cli: tt serve changh95/bevdet-p150
S=code/tt_bevdet/samples/synthetic_6cam
python3 code/tt_bevdet/server/client.py --image CAM_FRONT_LEFT=$S/CAM_FRONT_LEFT.png --image CAM_FRONT=$S/CAM_FRONT.png \
    --image CAM_FRONT_RIGHT=$S/CAM_FRONT_RIGHT.png --image CAM_BACK_LEFT=$S/CAM_BACK_LEFT.png \
    --image CAM_BACK=$S/CAM_BACK.png --image CAM_BACK_RIGHT=$S/CAM_BACK_RIGHT.png --calib-preset synthetic_6cam --out req.json
curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json
tt model stop changh95/bevdet-p150
  • The image does not contain the weights. --with-weights puts them in your HF cache.
  • The server uses port 20000 (or the next free port). It is ready when the log shows Application startup complete.
  • POST /predict: images (the six cameras, base64 JPEG / PNG / .npy), calibration (per camera intrinsics + T_ref_from_camera, or {"preset": name}); optional stream (id, reset, T_world_from_ego), params (score_threshold, nms_iou_threshold, nms_rescale_factor, ...), output_format. Also GET /health, GET /info, GET /v1/models (stub). Contract: SERVING.md section 3.
{"model": "bevdet-p150", "frame_id": "base_link", "num_detections": 10, "detections": [
   {"label": "traffic_cone", "label_id": 9, "score": 0.9492, "center": [7.68, -3.513, 0.367], "size": [0.385, 0.348, 0.733], "yaw": 0.134, "velocity": [0.0, 0.0], "autoware_label": "UNKNOWN", "twist": {"linear_x": 0.0, "angular_z": 2.1597}},
   {"label": "car", "label_id": 0, "score": 0.8457, "center": [-8.51, -5.941, 0.845], "size": [4.606, 1.92, 1.699], "yaw": 0.156, "velocity": [-0.0, 0.0], "autoware_label": "CAR", "twist": {"linear_x": 0.0, "angular_z": 4.8626}}],
 "timing_ms": {"encode": 60.0, "bevpool": 25.7, "bev_decode": 15.1, "preprocess": 60.6, "device": 101.4, "postprocess": 18.6, "total": 584.6, "decode": 403.6, "model_call": 180.6},
 "meta": {"class_names": ["car", "truck", ...], "autoware_labels": ["CAR", "TRUCK", ...], "score_threshold": 0.2, "stream": {"id": "timing", "frame": 42, "flag": 1}}}
  • center (box centre, z = gravity centre) / size ([length, width, height]) / yaw (counter-clockwise from +x) are in base_link metres / radians, as the Autoware node publishes them; detections are sorted by score.
  • label is the network class (one of the 10 nuScenes classes of the Autoware config, label_id its index); autoware_label is the label the node publishes (construction_vehicle, barrier and traffic_cone become UNKNOWN).
  • velocity is [vx, vy] in m/s (absolute, base_link axes); twist is the node's twist.linear.x / twist.angular.z, reproduced as deployed.
  • /info reports the device (ETH, 12×10), the labels, the stream table (4 streams, least-recently-used eviction) and the load-time knobs.

Demo

PandaSet 019, frame 40 (San Francisco; CC BY 4.0): TT output, ego-pose mode nuScenes scene-0103, key-frame 20 (Boston; CC BY-NC-SA 4.0, non-commercial): TT output
PandaSet 019 frame 40, six cameras and bird's-eye view nuScenes scene-0103 key-frame 20, non-commercial
TT vs the fp32 CPU reference: PandaSet 019 at 2 Hz The shipped synthetic sample: TT vs the stored CPU reference
TT vs CPU on PandaSet 019 The synthetic sample
Temporal modes on PandaSet 019 frame 40 (11 m/s) The Autoware K quirk on a 1920×1080 rig
Temporal modes K quirk

More: nuScenes scene-0916 key-frame 20 (Singapore) and the 40 key-frames of scene-0103 as a GIF, both non-commercial. Every image shows outputs of this port on the p150, some next to the fp32 CPU reference; the camera mosaics and the mode / K-quirk panels draw the boxes with score >= 0.4 (the node publishes >= 0.2, too dense to read at this size), the TT-vs-CPU panel and the synthetic sample every box >= 0.2. Sources, licences and changes: media/ATTRIBUTION.md. The weights were trained on nuScenes; PandaSet is another rig (taller cameras, other lenses), and its ranges come out biased (front objects about 15 % too close).

Demo & Performances

Warm, batch 1, six cameras. Accuracy: the TT output against the port's fp32 CPU reference of the same Autoware network (same weights, same pre- and post-processing), checked against ONNX Runtime on the plugin-free sub-graphs of the deployed ONNX (PCC 1.0000000000). Speed: code/scripts/bench.py, p50 (p99) of 100 iterations, host load average 7-9 on the shared 8-core measurement host (OPT_BASELINE.md).

Metric Performance
Per-stage PCC vs the CPU reference, PandaSet frames (gates) depth 0.99994, features 0.99994 (gate 0.999); BEV map after the host BEVPool 0.99995 (0.999); six head maps >= 0.99952 (0.998)
Box agreement, nuScenes mini_val (scene-0103 / scene-0916, 81 key-frames, ego poses) recall 0.9897 / 0.9897, precision 0.9899 / 0.9902, p99 |score difference| 0.0083 / 0.0089 (gates 0.95 / 0.95 / 0.05)
Box agreement, PandaSet 019 frames 38-40 (Autoware semantics) recall 1.000 / 0.996 / 0.995, precision 1.000 / 0.996 / 0.986, max |score difference| 0.0059 / 0.0080 / 0.0079
Box agreement, the shipped synthetic sample (the container smoke's check) recall 1.000, precision 1.000 (all 10 reference boxes), max |score difference| 0.0056, centres within 0.005 m (gates 0.95 / 0.95 / 0.05)
nuScenes mini_val mAP / NDS, ego-pose mode (indicative: 81 frames, 3 of 10 classes absent) TT 0.3670 / 0.4362, CPU 0.3673 / 0.4367 (BEVDet's own full-val result for this config: 0.394 / 0.515)
nuScenes mini_val mAP / NDS, Autoware node semantics (no ego motion, camera rate) TT 0.2530 / 0.2739, CPU 0.2535 / 0.2746
Device, one steady frame (encode + decode + ring push, back to back) 34.15 ms (34.45 with ego poses): 29.3 frames/s device-bound
Trace replays: encode / decode / push 24.24 / 8.68 / 1.26 ms back to back (24.30 / 8.73 / 1.29 ms one blocking replay)
Python model(), 1600×900 images, Autoware mode (synthetic input) 173.3 ms (217.5); 131.8-155.2 ms p50 in the other runs on a less loaded host
Python model(), nuScenes 1600×900 with ego poses 220.3 ms (262.5)
Python model(), PandaSet 1920×1080 (host resize to 1600×900) 341.9 ms (370.5)
Served /predict on the host, shipped sample (uvicorn, loopback) client round trip 623.9 ms p50 (p90 703.7; a 9.6 MB JSON request of six base64 PNGs); timing_ms.model_call 156.7 ms; timing_ms.total 538.2 ms, most of it the request decode (base64 + PNG, 404 ms in the response shown in SERVING.md)

All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150; AICLK 1343-1350 MHz. The device is kernel-bound (34.0 of 34.3 ms are kernels; 265 µs of op-to-op gaps for 457 programs), and the end-to-end latency is set by host work: the image preprocessing (35-43 ms at 1600×900, 209 ms at 1920×1080), the 6.5 MB image upload (27 ms, sent as 6-byte pages), the host BEVPool (15-36 ms) and the Scale-NMS (about 50 ms on real scenes). This is the first release: optimization has not started (OPT_REPORT.md ranks the plan; OPT_BASELINE.md has every measurement). Verification: VERIFICATION_2026-10-08.md.

The nuScenes metrics use the official devkit (detection_cvpr_2019, mini_val) on every box the vendor decode keeps; nuScenes is CC BY-NC-SA 4.0 (non-commercial; the data were used for local validation only and are not distributed here). Box agreement: a reference box with score >= 0.2 counts as recalled when a same-class TT box (any score above the 0.1 decode threshold) lies within 0.5 m, and the reverse for precision; scores that straddle 0.2 within bf16 noise are not counted as disagreements.

No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). The Autoware README claims up to 35 FPS (fp16) / 17 FPS (fp32) on an unspecified GPU (not like-for-like). p150 power was not measured (OPT_BASELINE.md lists only the board telemetry range during the timed loops), so no efficiency comparison is made.

Caveats

  • Deployment status. BEVDet is not Autoware's default perception: autoware.launch.xml defaults to perception_mode:=lidar, and autoware_tensorrt_bevdet is the detector of the camera-only mode (perception_mode:=camera). This port is a drop-in for the network and the node's pre- and post-processing, not a ROS node, and not a certified Autoware component: do not use it for safety-critical driving decisions.
  • First release: optimization pending. This is the functional baseline port; no optimization round has run yet, and the end-to-end latency is set by host stages (OPT_REPORT.md ranks the plan).
  • Precision. The node defaults to its fp16 TensorRT engine; this port implements the fp32 path. By code reading (not run here: no GPU), the fp16 engine also mis-reads the AlignBEV transforms (the history branch is undefined in the default mode), keeps BEVPool sums in fp16 and leaves unreached BEV cells uninitialised; the port follows the intended fp32 semantics (zero-initialised cells). On the chip: fp32 weights and biases, HiFi4, fp32 and packer L1 accumulation, bf16 activations (the precision policy of the first release; bf16 weights fail the accuracy gates). Outputs differ slightly from the fp32 reference (agreement figures above), with a small systematic score offset of about -0.0008.
  • Temporal semantics. Default = the deployed node: no ego motion (identity alignment), the history advanced once per request, a fresh stream fuses an empty history on its first frame. The ego-pose input (stream.T_world_from_ego) is a documented extension; BEVDET_HISTORY_STRIDE (load time) spaces the history like training's 2 Hz key-frames.
  • Intrinsics. The API takes the intrinsics of the image actually sent and scales K with the node's resize to 1600×900; the node itself keeps the raw CameraInfo.k when it resizes (calibration.scale_intrinsics: false reproduces it: far fewer and misplaced boxes on a 1920×1080 rig, see the K-quirk image). Rectified rigs should pass P[0:3, 0:3].
  • Domain gap. The weights were trained on nuScenes (Boston, Singapore; camera heights ~1.5 m). On other rigs the monocular depth prior is biased: on PandaSet 019 front objects come out ~15 % too close and side objects 10-20 % too far.
  • Host stages. Image preprocessing, the BEVPool lift-splat between the two traces and the vendor Scale-NMS run on the host (the documented fallbacks of this first release); the end-to-end latency is host-bound (above).
  • Fixed shapes. Six cameras, 256×704 network input, 118 depth bins, a 128×128 BEV grid (±51.2 m), 8 history frames, 10 classes; rigs whose BEVPool tables exceed M 370,000 / N 14,000 are refused, as with TensorRT. Batch 1, one frame per request; requests are serialised on the chip.
  • Streams. At most 4 stream ids keep a history at once (BEVDET_MAX_STREAMS); a fifth evicts the least recently used one, which restarts with an empty history.
  • Validation scope. Agreement with the fp32 CPU reference on public data (nuScenes mini_val, PandaSet) and on the synthetic sample; the mAP / NDS on 81 mini_val frames are indicative, not a full-val reproduction. The shipped sample is synthetic; the PandaSet sample is not shipped yet.
  • Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (patches/tt-metal-eth-dispatch.patch), so this build assumes that you do not need chip-to-chip ethernet communication.
  • dispatch="worker" (server: BEVDET_DISPATCH=worker) is an A/B opt-in. On a p150 it gives an 11×10 grid (34.83 ms per frame, the same outputs); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode.
  • Not an OpenAI-compatible API; GET /v1/models is a stub so the tt-model ready card does not 404.
  • p150 power was not measured (only the board telemetry was logged, OPT_BASELINE.md), so no efficiency comparison is made.

Licensing

  • Weights: AutowareFoundation/tensorrt_bevdet at tag v1.0 (commit 391466349a1b39c93b9865364a8088110ca8d97b), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card's Legal Notice: the training data is nuScenes (CC BY-NC-SA 4.0, non-commercial); read it before commercial use.
  • Pre- and post-processing ported from autoware_universe perception/autoware_tensorrt_bevdet (Apache-2.0). The node's inference library, bevdet_vendor, carries no licence: its behaviour was re-implemented clean-room from a written description, and no code was copied.
  • Port and serving code (code/): Apache-2.0.
  • Sample data: code/tt_bevdet/samples/synthetic_6cam* and the preset calib/synthetic_6cam.json were generated by this repository (code/scripts/make_sample.py; no third-party data), Apache-2.0. No dataset sample is shipped; nuScenes, Argoverse 2 and KITTI never ship, and a PandaSet sample (CC BY 4.0) is prepared but not yet included.
  • Demo media (media/): renders of PandaSet (Scale AI and Hesai, CC BY 4.0 and the PandaSet Dataset Terms; Scale AI and Hesai do not endorse this work) and of nuScenes (© Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use: those renders are non-commercial, ShareAlike, labelled _NC; Motional does not endorse this work). Attribution texts and changes: media/ATTRIBUTION.md.

Provenance

These are the exact sources the container image was built from:

component built from
tt-metal 44d66500520fda9f2c7060c0f6b41ec48f7ab37e + patches/tt-metal-eth-dispatch.patch (dirty tree: the image includes the patch)
weights AutowareFoundation/tensorrt_bevdet@391466349a1b39c93b9865364a8088110ca8d97b (tag v1.0), files bevdet_one_lt_d.onnx (sha256 47f156cb…0d77)
Autoware reference autoware_universe 9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd (package 0.53.0); bevdet_vendor 0.2.1 semantics
shared port code ttaw 0.20.0 @ common 89dec49 (vendored in code/tt_bevdet/ttaw, VENDORED.json)
code/ digest (image) a67de74da4365b53 (sha256, first 16 hex digits; built.code_sha256 of tt_kernel_manifest.json)
image tt-model/bevdet-p150:673a2b830216 (sha256:673a2b8302161c0c92000591efc289d1c4aaa34b6464818144ee28fc59d942a4)
base images build stage ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest @ sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc; runtime stage docker.io/library/ubuntu:22.04 @ sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401 (tt-model's FROM tags float; these are the digests this build resolved, see build_info.json)
built 2026-10-09T05:04:41+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for changh95/bevdet-p150

Finetuned
(1)
this model

Collection including changh95/bevdet-p150

Papers for changh95/bevdet-p150